A monocular panoramic vision 3D reconstruction method for forest scenes

By using a monocular panoramic camera and deep learning technology to extract shallow details and deep semantic features from panoramic images of forest areas, the high cost and complexity of traditional forest 3D reconstruction methods are solved, and efficient and accurate 3D reconstruction is achieved, supporting forestry resource monitoring and intelligent robot navigation.

CN119027593BActive Publication Date: 2025-09-16FOSHAN ZHONGKE AGRI ROBOT & INTELLIGENT AGRI INNOVATION RES INST
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202411188074.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-28
Publication Date
2025-09-16
Estimated Expiration
2044-08-28

AI Technical Summary

Technical Problem

Traditional forest 3D reconstruction methods rely on multi-camera vision systems or lidar technology, which have limitations such as high equipment cost, complex deployment, and large data processing volume, making them difficult to apply in remote or inaccessible forest areas.

Method used

A monocular panoramic camera is used to collect panoramic images of the forest area. Through deep learning-based image processing technology, the shallow detail features and deep semantic features of the panoramic images of the forest area are extracted, and their bidirectional dependency is optimized and fused to generate depth information for three-dimensional reconstruction.

Benefits of technology

It effectively reduces the cost of 3D reconstruction equipment, simplifies system deployment, improves the efficiency and accuracy of 3D reconstruction in complex forest environments, and provides support for forestry resource monitoring and intelligent robot navigation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119027593B_ABST
    Figure CN119027593B_ABST
Patent Text Reader

Abstract

The present application relates to the field of three-dimensional reconstruction of forest areas, and specifically to a monocular panoramic vision three-dimensional reconstruction method for forest scenes. It uses a monocular panoramic camera to collect panoramic images of the forest area, and uses image processing technology based on deep learning to perform image analysis on the panoramic images of the forest area, respectively extracting the shallow detail features and deep semantic features of the panoramic images of the forest area, and then based on the bidirectional dependency relationship between its shallow features and deep features, optimizes and fuses its shallow and deep features to achieve multi-level joint perception of the panoramic images of the forest area, thereby intelligently predicting the depth information of each position of the panoramic image of the forest area, and reconstructing it in three dimensions. In this way, the equipment cost of three-dimensional reconstruction can be effectively reduced, the system deployment can be simplified, and the efficiency and accuracy of three-dimensional reconstruction in complex forest environments can be improved, providing strong support for the precise monitoring of forestry resources, intelligent robot navigation and operations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of three-dimensional reconstruction of forest areas, and specifically to a monocular panoramic vision three-dimensional reconstruction method for forest scenes. Background Art

[0002] my country currently ranks first in the world in terms of planted forest area. Actively developing forest informatization theory is a key approach to implementing this development philosophy. Improving robots' ability to perceive and understand forest environments, and based on this, collecting forest resource information and enabling intelligent robot navigation, positioning, and target recognition, are key research directions for forest informatization theory under the new normal. Three-dimensional reconstruction is a key technology for representing physical objects in computers using relevant sensors. Its value lies in implicitly modeling the relative positions of objects in space. For forestry robots, acquiring three-dimensional structural information about forest scenes is key to improving their perception and understanding of the forest environment.

[0003] Traditional 3D reconstruction methods for forest areas mostly rely on multi-camera vision systems or laser radar (LiDAR) technology. Although these methods can provide high-precision 3D models, they often have limitations such as high equipment costs, complex deployment, and large data processing volumes. Their application is particularly restricted in remote or difficult-to-reach forest areas.

[0004] Therefore, a monocular panoramic vision 3D reconstruction method for forest scenes is expected. Summary of the Invention

[0005] This application is made in consideration of the above problems. One purpose of this application is to provide a monocular panoramic vision 3D reconstruction method for forest scenes.

[0006] The embodiments of the present application provide a monocular panoramic vision 3D reconstruction method for forest scenes, which includes:

[0007] Obtain a panoramic image of the forest area captured by a monocular panoramic camera;

[0008] Performing multi-level feature extraction on the forest area panoramic image to obtain a forest area panoramic shallow feature map and a forest area panoramic deep semantic feature map;

[0009] Inputting the forest area panoramic shallow feature map into a feature distribution gradient mask salient device to obtain a forest area panoramic shallow salient feature map;

[0010] generating a depth information decoding value of the forest area image based on the joint perception information of the forest area panoramic deep semantic feature map and the forest area panoramic shallow salient feature map;

[0011] Based on the depth information decoding value, the panoramic image of the forest area is projected into a three-dimensional space to obtain a three-dimensional model of the forest area scene.

[0012] For example, according to an embodiment of the present application, a monocular panoramic vision 3D reconstruction method for a forest scene, wherein multi-level feature extraction is performed on the forest panoramic image to obtain a forest panoramic shallow feature map and a forest panoramic deep semantic feature map, including:

[0013] The forest area panoramic image is input into a panoramic image multi-scale feature extractor based on a pyramid network to obtain the forest area panoramic shallow feature map and the forest area panoramic deep semantic feature map.

[0014] For example, according to an embodiment of the present application, a monocular panoramic vision 3D reconstruction method for a forest scene, wherein the forest panoramic shallow feature map is input into a feature distribution gradient mask salient device to obtain a forest panoramic shallow salient feature map, includes:

[0015] Calculating the multidirectional gradient value distribution of each position in the forest area panoramic shallow feature map, and determining the gradient amplitude value of each position in the forest area panoramic shallow feature map based on the multidirectional gradient value distribution of each position to obtain a forest area panoramic shallow feature gradient amplitude distribution map;

[0016] Calculating the gradient amplitude local description operator of each position in the forest area panoramic shallow feature gradient amplitude distribution map to obtain the forest area panoramic shallow feature gradient amplitude local significant distribution map;

[0017] Inputting the forest area panoramic shallow feature gradient amplitude local significant distribution map into a gated masker based on a GELU function to obtain a gradient amplitude local significant gated mask map;

[0018] The forest area panoramic shallow salient feature map is obtained by calculating the multiplication of the gradient amplitude local salient gated mask map and the forest area panoramic shallow feature map by position points.

[0019] For example, according to an embodiment of the present application, a monocular panoramic vision 3D reconstruction method for a forest scene is provided, wherein a local description operator of the gradient amplitude at each position in the forest panoramic shallow feature gradient amplitude distribution map is calculated to obtain a local significant distribution map of the forest panoramic shallow feature gradient amplitude, including:

[0020] Determine the scale of the local neighborhood, calculate the average of the differences between the gradient amplitude value of a predetermined position in the forest area panoramic shallow feature gradient amplitude distribution map and the gradient amplitude values ​​of other positions in the local neighborhood to obtain a gradient amplitude local description operator corresponding to the predetermined position.

[0021] For example, according to an embodiment of the present application, a monocular panoramic vision 3D reconstruction method for a forest scene, wherein a depth information decoding value of a forest image is generated based on the joint perception information of the forest panoramic deep semantic feature map and the forest panoramic shallow salient feature map, including:

[0022] Performing joint perception on the forest area panoramic deep semantic feature map and the forest area panoramic shallow salient feature map to obtain a forest area panoramic shallow-deep salient joint perception feature map;

[0023] The forest area panoramic shallow-deep significant joint perceptual feature map is input into a decoder-based depth estimation module to obtain the depth information decoding value.

[0024] For example, according to an embodiment of the present application, a monocular panoramic vision 3D reconstruction method for a forest scene, wherein the forest panoramic deep semantic feature map and the forest panoramic shallow salient feature map are jointly perceived to obtain a forest panoramic shallow-deep salient joint perception feature map, including:

[0025] The forest area panoramic deep semantic feature map and the forest area panoramic shallow salient feature map are input into a salient joint perception module based on the attention mechanism to obtain the forest area panoramic shallow-deep salient joint perception feature map.

[0026] For example, according to an embodiment of the present application, a monocular panoramic vision 3D reconstruction method for a forest scene, wherein the forest panoramic deep semantic feature map and the forest panoramic shallow salient feature map are input into a salient joint perception module based on an attention mechanism to obtain the forest panoramic shallow-deep salient joint perception feature map, includes:

[0027] Performing feature shape reshaping on the forest area panoramic shallow layer salient feature map and the forest area panoramic deep layer semantic feature map to obtain a forest area panoramic shallow layer feature shape reshaping matrix and a forest area panoramic deep layer semantic feature shape reshaping matrix;

[0028] Input the forest area panoramic shallow feature shape reshaping matrix and the forest area panoramic deep semantic feature shape reshaping matrix into a feature channel-by-channel interactive perception module to obtain a forest area panoramic shallow detail-deep semantic dependency matrix and a forest area panoramic deep semantic-shallow detail dependency matrix;

[0029] Inputting the forest area panoramic shallow detail-deep semantic dependency matrix and the forest area panoramic deep semantic-shallow detail dependency matrix into a random dropout module to obtain a pruned forest area panoramic shallow detail-deep semantic dependency matrix and a pruned forest area panoramic deep semantic-shallow detail dependency matrix;

[0030] Based on the pruned forest area panoramic shallow detail-deep semantic dependency matrix and the pruned forest area panoramic deep semantic-shallow detail dependency matrix, feature optimization is performed on the forest area panoramic deep semantic feature shape reshaping matrix and the forest area panoramic shallow feature shape reshaping matrix to obtain a dependency-optimized forest area panoramic deep semantic feature matrix and a dependency-optimized forest area panoramic shallow feature matrix;

[0031] Reshaping the dependency-optimized forest area panoramic shallow feature matrix and the dependency-optimized forest area panoramic deep semantic feature matrix to obtain an optimized forest area panoramic shallow salient feature map and an optimized forest area panoramic deep semantic feature map;

[0032] The weighted sum of the optimized forest area panoramic shallow salient feature map and the optimized forest area panoramic deep semantic feature map is calculated to obtain the forest area panoramic shallow-deep salient joint perceptual feature map.

[0033] For example, according to an embodiment of the present application, a monocular panoramic vision 3D reconstruction method for a forest scene, wherein the forest panoramic shallow feature shape reconstruction matrix and the forest panoramic deep semantic feature shape reconstruction matrix are input into a feature channel-by-channel interactive perception module to obtain a forest panoramic shallow detail-deep semantic dependency matrix and a forest panoramic deep semantic-shallow detail dependency matrix, including:

[0034] Calculating the forest area panoramic shallow feature shape reconstruction matrix multiplied by the transposed matrix of the forest area panoramic deep semantic feature shape reconstruction matrix to obtain a forest area panoramic shallow detail-deep semantic association representation matrix;

[0035] The forest panorama shallow detail-deep semantic association representation matrix is ​​divided by the scale of the forest panorama shallow feature shape reshaping matrix and then input into the softmax function to obtain the forest panorama shallow detail-deep semantic dependency relationship matrix;

[0036] Calculating the transposed matrix of the forest area panoramic deep semantic feature shape reconstruction matrix and multiplying it by the forest area panoramic shallow feature shape reconstruction matrix to obtain a forest area panoramic deep semantic-shallow detail association representation matrix;

[0037] The forest area panoramic deep semantics-shallow detail association representation matrix is ​​divided by the scale of the forest area panoramic shallow feature shape reshaping matrix and then input into the softmax function to obtain the forest area panoramic deep semantics-shallow detail dependency matrix.

[0038] For example, according to an embodiment of the present application, a monocular panoramic vision 3D reconstruction method for forest scenes is provided, wherein, based on the pruned forest panoramic shallow detail-deep semantic dependency matrix and the pruned forest panoramic deep semantic-shallow detail dependency matrix, the forest panoramic deep semantic feature shape reshaping matrix and the forest panoramic shallow feature shape reshaping matrix are feature optimized to obtain a dependency-optimized forest panoramic deep semantic feature matrix and a dependency-optimized forest panoramic shallow feature matrix, including:

[0039] Performing matrix multiplication on the pruned forest area panoramic shallow detail-deep semantic dependency matrix and the forest area panoramic deep semantic feature shape reshaping matrix to obtain the dependency-optimized forest area panoramic deep semantic feature matrix;

[0040] The pruned forest area panoramic deep semantic-shallow detail dependency matrix is ​​matrix multiplied with the forest area panoramic shallow feature shape reshaping matrix to obtain the dependency optimized forest area panoramic shallow feature matrix.

[0041] According to the monocular panoramic vision 3D reconstruction method for forest scenes in the embodiment of the present application, a monocular panoramic camera is used to collect panoramic images of the forest area, and image processing technology based on deep learning is used to analyze the panoramic images of the forest area, respectively extracting shallow detail features and deep semantic features of the panoramic images of the forest area, and then optimizing and fusing the shallow and deep features based on the bidirectional dependency between the shallow and deep features to achieve multi-level joint perception of the panoramic images of the forest area, thereby intelligently predicting the depth information of each position of the panoramic image of the forest area, and thus reconstructing the three-dimensional image. In this way, the equipment cost of the three-dimensional reconstruction can be effectively reduced, the system deployment can be simplified, and the efficiency and accuracy of the three-dimensional reconstruction in complex forest environments can be improved, providing strong support for the precise monitoring of forestry resources, intelligent robot navigation and operations. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings of the embodiments of the present application. Obviously, the drawings described below only relate to some embodiments of the present application, and are not intended to limit the present application.

[0043] Figure 1 A schematic diagram of the application architecture of a monocular panoramic vision 3D reconstruction method for forest scenes in an embodiment of the present application is shown;

[0044] Figure 2 A flowchart of a monocular panoramic vision 3D reconstruction method for a forest scene in an embodiment of the present application is shown;

[0045] Figure 3A flowchart of sub-step S530 of the monocular panoramic vision 3D reconstruction method for forest scenes in an embodiment of the present application is shown;

[0046] Figure 4 A flowchart of sub-step S540 of the monocular panoramic vision 3D reconstruction method for forest scenes in an embodiment of the present application is shown;

[0047] Figure 5 A schematic diagram of the structure of a monocular panoramic vision 3D reconstruction system for forest scenes in an embodiment of the present application is shown;

[0048] Figure 6 A diagram showing an application scenario of a monocular panoramic vision 3D reconstruction method for a forest scene in an embodiment of the present application is shown;

[0049] Figure 7 A flowchart showing a monocular panoramic vision 3D reconstruction method for a forest scene in another embodiment of the present application is shown; and

[0050] Figure 8 A schematic diagram of panoramic depth map acquisition in a monocular panoramic vision 3D reconstruction method for forest scenes in another embodiment of the present application is shown. DETAILED DESCRIPTION

[0051] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present application, not all embodiments. Based on the embodiments of the present application, all other embodiments obtained by ordinary technicians in this field without making creative efforts also fall within the scope of protection of this application.

[0052] The terms used in this specification are those commonly used in the art currently in consideration of the functions of the present application, but these terms may vary according to the intentions of those skilled in the art, precedents, or new technologies in the art. In addition, specific terms may be selected, and in such cases, their detailed meanings will be described in the detailed description of the present application. Therefore, the terms used in the specification should not be understood as simple names, but rather as the meaning of the terms and the overall description of the present application.

[0053] Although the present application makes various references to certain modules in the system according to embodiments of the present application, any number of different modules can be used and run on the user terminal and / or server. The modules are illustrative only, and different aspects of the system and method can use different modules.

[0054] Flowcharts are used throughout this application to illustrate the operations performed by the systems of the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, the various steps may be processed in reverse order or simultaneously, as needed. Furthermore, other operations may be added to these processes, or one or more operations may be removed from these processes.

[0055] Figure 1 A schematic diagram of the application architecture of a monocular panoramic vision 3D reconstruction method for forest scenes in an embodiment of the present application is shown, including a server 100 and a terminal device 200.

[0056] The terminal device 200 and the server 100 can be connected via the Internet to enable communication between them. Optionally, the Internet utilizes standard communication technologies and / or protocols. The Internet is typically the Internet, but may also be any network, including but not limited to a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired, or wireless network, a private network, or any combination of a virtual private network. In some embodiments, technologies and / or formats such as Hypertext Markup Language (HTML) and Extensible Markup Language (XML) are used to represent data exchanged over the network. Conventional encryption technologies such as Secure Sockets Layer (SSL), Transport Layer Security (TLS), Virtual Private Networks (VPN), and Internet Protocol Security (IPsec) may also be used to encrypt all or some links. In other embodiments, customized and / or dedicated data communication technologies may be used in place of or in addition to the aforementioned data communication technologies.

[0057] Server 100 can provide various network services to terminal device 200. Server 100 can be a single server, a server cluster consisting of multiple servers, or a cloud computing center. Specifically, server 100 may include a processor 110 (Center Processing Unit, CPU), memory 120, input devices 130, and output devices 140. Input devices 130 may include a keyboard, mouse, touch screen, etc. Output devices 140 may include display devices such as a liquid crystal display (LCD) or a cathode ray tube (CRT).

[0058] The memory 120 may include a read-only memory (ROM) and a random access memory (RAM), and provides program instructions and data stored in the memory 120 to the processor 110. In an embodiment of the present application, the memory 120 may be used to store the program for the monocular panoramic vision 3D reconstruction method for forest scenes in an embodiment of the present application.

[0059] The processor 110 calls the program instructions stored in the memory 120, and the processor 110 is used to execute the steps of any one of the monocular panoramic vision three-dimensional reconstruction methods for forest scenes in the embodiments of the present application according to the obtained program instructions.

[0060] In addition, the application architecture diagram in the embodiment of the present application is intended to more clearly illustrate the technical solution in the embodiment of the present application, and does not constitute a limitation on the technical solution provided in the embodiment of the present application. Of course, for other application architectures and business applications, the technical solution provided in the embodiment of the present application is also applicable to similar problems.

[0061] The following is a non-restrictive description of the monocular panoramic vision three-dimensional reconstruction method for forest scenes provided by at least one embodiment of the present application through several examples or embodiments. As described below, different features in these specific examples or embodiments can be combined with each other without conflicting with each other to obtain new examples or embodiments, and these new examples or embodiments also fall within the scope of protection of this application.

[0062] In response to the above technical problems, the technical concept of this application is to collect panoramic images of the forest area through a monocular panoramic camera, and use deep learning-based image processing technology to analyze the panoramic images of the forest area, respectively extract the shallow detail features and deep semantic features of the panoramic images of the forest area, and then optimize and fuse the shallow and deep features based on the bidirectional dependency between the shallow and deep features to achieve multi-level joint perception of the panoramic images of the forest area, thereby intelligently predicting the depth information of each position of the panoramic image of the forest area, and thus reconstructing it in three dimensions. In this way, the equipment cost of three-dimensional reconstruction can be effectively reduced, the system deployment can be simplified, and the efficiency and accuracy of three-dimensional reconstruction in complex forest environments can be improved, providing strong support for the precise monitoring of forestry resources, intelligent robot navigation and operations.

[0063] Based on this, Figure 2 The flowchart of the monocular panoramic vision 3D reconstruction method for forest scenes in the embodiment of the present application is shown. For example, the monocular panoramic vision 3D reconstruction method for forest scenes can be executed by a server, which can be Figure 1 The server 100 shown in FIG. Figure 2 As shown, the monocular panoramic vision three-dimensional reconstruction method for forest scenes according to the embodiment of the present application includes the following steps: S510, obtaining a panoramic image of the forest area collected by a monocular panoramic camera; S520, performing multi-level feature extraction on the panoramic image of the forest area to obtain a panoramic shallow feature map of the forest area and a panoramic deep semantic feature map of the forest area; S530, inputting the panoramic shallow feature map of the forest area into a feature distribution gradient mask salient device to obtain a panoramic shallow salient feature map of the forest area; S540, generating a depth information decoding value of the forest area image based on the joint perception information of the panoramic deep semantic feature map of the forest area and the panoramic shallow salient feature map of the forest area; S550, projecting the panoramic image of the forest area into three-dimensional space based on the depth information decoding value to obtain a three-dimensional model of the forest scene.

[0064] Specifically, in the technical solution of the present application, a panoramic image of a forest area is first acquired using a monocular panoramic camera. It should be understood that traditional monocular or binocular cameras can only capture images with a limited viewing angle and require stitching multiple captured images, which not only increases computational complexity but may also introduce stitching errors due to differences in lighting and viewing angles. A monocular panoramic camera, on the other hand, can capture 360-degree horizontal and vertical images at a certain angle (e.g., 180 degrees), achieving panoramic coverage of the complex environment of the forest area and directly providing a continuous, seamless view of the scene. This avoids the image stitching required after using multiple cameras or multiple shots, thereby reducing data redundancy and errors and facilitating the subsequent three-dimensional reconstruction process.

[0065] Taking into account the large scale differences of objects such as trees and shrubs in the forest area, everything from tiny branches and leaves to the global spatial distribution of trees plays an important role in the three-dimensional reconstruction of the forest area. Therefore, in order to simultaneously consider the depth and breadth of feature extraction, in the technical solution of the present application, a panoramic image multi-scale feature extractor based on a pyramid network is used to process the panoramic image of the forest area. By utilizing the multi-level feature extraction capability of the pyramid network, feature maps of different scales are constructed by downsampling layer by layer, effectively capturing the shallow detail information and deep semantic information of the panoramic image of the forest area, and generating a shallow feature map of the panoramic view of the forest area and a deep semantic feature map of the panoramic view of the forest area. Among them, the shallow feature map of the panoramic view of the forest area mainly focuses on the detailed information such as local texture and edges in the image, while the deep semantic feature map of the panoramic view of the forest area focuses more on high-level semantic information such as the global structure of the image and object categories, providing a rich information basis for the prediction of forest depth information.

[0066] Accordingly, in step S520, multi-level feature extraction is performed on the panoramic image of the forest area to obtain a panoramic shallow feature map of the forest area and a panoramic deep semantic feature map of the forest area, including: inputting the panoramic image of the forest area into a panoramic image multi-scale feature extractor based on a pyramid network to obtain the panoramic shallow feature map of the forest area and the panoramic deep semantic feature map of the forest area.

[0067] Next, in order to further highlight the key information such as the main objects, edge contours, texture details, etc. in the image, the present application further introduces a feature distribution gradient mask salient device to perform feature enhancement processing on the shallow feature map of the panoramic forest area. By analyzing the pixel change intensity in the shallow feature map of the panoramic forest area, the expression of edge contours and detail areas in the feature map is enhanced, and the local details of the forest area are more clearly highlighted. Specifically, the feature distribution gradient mask salient device first quantifies the local change intensity of the pixels at each position by calculating the gradient amplitude value at each position in the shallow feature map of the panoramic forest area. Then, the gradient amplitude local description operator is calculated by measuring the significance of the gradient amplitude value of each position relative to the gradient amplitude values ​​of other positions in the local neighborhood, and the feature importance of each position is evaluated. Then, the GELU function is used to perform nonlinear activation on the gradient amplitude local description operator of each position to generate a gated mask to guide the enhancement of key information in the shallow feature map of the panoramic forest area. Finally, the shallow feature map of the forest area panorama is weighted element by element based on the generated gated mask, thereby focusing on important visual clues in the image, removing background noise, irrelevant details and other information, and generating a shallow salient feature map of the forest area panorama.

[0068] Accordingly, in step S530, if Figure 3As shown, the forest area panoramic shallow feature map is input into the feature distribution gradient mask salient device to obtain the forest area panoramic shallow salient feature map, including: S531, calculating the multi-directional gradient value distribution of each position in the forest area panoramic shallow feature map, and determining the gradient amplitude value of each position in the forest area panoramic shallow feature map based on the multi-directional gradient value distribution of each position to obtain the forest area panoramic shallow feature gradient amplitude distribution map; S532, calculating the gradient amplitude local description operator of each position in the forest area panoramic shallow feature gradient amplitude distribution map to obtain the forest area panoramic shallow feature gradient amplitude local salient distribution map; S533, inputting the forest area panoramic shallow feature gradient amplitude local salient distribution map into the gated masker based on the GELU function to obtain the gradient amplitude local salient gated mask map; S534, calculating the position point multiplication between the gradient amplitude local salient gated mask map and the forest area panoramic shallow feature map to obtain the forest area panoramic shallow salient feature map.

[0069] Among them, in step S532, the gradient amplitude local description operator of each position in the forest area panoramic shallow feature gradient amplitude distribution map is calculated to obtain the forest area panoramic shallow feature gradient amplitude local significant distribution map, including: determining the scale of the local neighborhood, calculating the mean of the difference between the gradient amplitude value of the predetermined position in the forest area panoramic shallow feature gradient amplitude distribution map and the gradient amplitude values ​​of other positions in the local neighborhood to obtain the gradient amplitude local description operator corresponding to the predetermined position.

[0070] In one example, the forest area panoramic shallow feature map is input into a feature distribution gradient mask salient device to obtain a forest area panoramic shallow salient feature map, including: processing the forest area panoramic shallow feature map using the following feature distribution gradient mask formula to obtain the forest area panoramic shallow salient feature map, wherein the feature distribution gradient mask formula is:

[0071]

[0072]

[0073]

[0074]

[0075]

[0076]

[0077]

[0078] in, 、 、 、 、 and Respectively represent the grayscale values ​​of the corresponding positions of the shallow feature map of the forest panorama, Indicates the shallow feature map of the forest area The gradient value in the horizontal direction of the pixel point, Indicates the shallow feature map of the forest area The gradient value in the vertical direction of the pixel point, The shallow characteristic map of the forest area is shown in The channel direction gradient value of the position pixel point, The first The gradient magnitude value at the position, The shallow characteristic gradient amplitude distribution map of the forest area is represented by The set of gradient magnitude values ​​in the local neighborhood centered at position, represents the number of gradient magnitude values ​​in the set of gradient magnitude values ​​within the local neighborhood, 、 and Respectively represent the offset in the horizontal coordinate direction, vertical coordinate direction and channel direction, The first The gradient magnitude local description operator of the position, represents the GELU function, The first part of the gradient magnitude local significant gated mask map is represented by The gradient magnitude local description operator of the position, represents the gradient magnitude local significant gated mask map, Indicates point multiplication by position, A panoramic shallow-layer significant feature map of the forest area is shown.

[0079] Furthermore, since the forest area panoramic deep semantic feature map and the forest area panoramic shallow salient feature map are feature representations of different levels and focuses of the forest area panoramic image, there is a natural dependency and information complementarity between the two. Therefore, in order to fully understand the overall content and spatial contextual relationship in the forest area panoramic image, the present application introduces a significant joint perception module based on the attention mechanism to process the forest area panoramic deep semantic feature map and the forest area panoramic shallow salient feature map, and models the bidirectional dependency between the two through a dual interactive attention mechanism to achieve optimized fusion of deep and shallow features. Specifically, first, the module reshapes the forest area panoramic deep semantic feature map and the forest area panoramic shallow salient feature map into a feature matrix so that its feature dimension and shape adapt to the input requirements of the attention mechanism. Then, the bidirectional dependency between the two sets of feature matrices is calculated through the feature channel-by-channel interactive perception module, the interaction between the features is quantified, and the shallow detail-deep semantic and deep semantic-shallow detail dependency matrices are generated, which serve as a guiding basis for feature attention optimization fusion. At the same time, to reduce sensitivity to image noise and prevent overfitting, the two dependency matrices are further input into a random dropout module for pruning. This selectively ignores or suppresses some feature interactions to enhance the model's generalization ability. Furthermore, based on the pruned dependency matrices, the reshaped feature matrices are weighted optimized. Through reshaping and weighted fusion, the original feature map is restored to form, generating a shallow-deep joint salient perceptual feature map of the forest panorama.

[0080] Then, the forest area panoramic shallow-deep significant joint perceptual feature map is input into a decoder-based depth estimation module to obtain a depth information decoding value. In the technical solution of the present application, the decoder uses a trained and optimized network structure and parameters to perform multi-layer feature parsing and information recovery on the forest area panoramic shallow-deep significant joint perceptual feature map, gradually restoring the feature size of the forest area panoramic shallow-deep significant joint perceptual feature map to the same size as the original image, and generating a corresponding depth prediction value at each pixel position, thereby accurately representing the three-dimensional depth information of each position in the forest area panoramic image.

[0081] Accordingly, in step S540, if Figure 4 As shown, based on the joint perception information of the forest area panoramic deep semantic feature map and the forest area panoramic shallow salient feature map, a depth information decoding value of the forest area image is generated, including: S541, jointly perceiving the forest area panoramic deep semantic feature map and the forest area panoramic shallow salient feature map to obtain a forest area panoramic shallow-deep salient joint perception feature map; S542, inputting the forest area panoramic shallow-deep salient joint perception feature map into a decoder-based depth estimation module to obtain the depth information decoding value.

[0082] Among them, in step S541, the forest area panoramic deep semantic feature map and the forest area panoramic shallow salient feature map are jointly perceived to obtain the forest area panoramic shallow-deep salient joint perception feature map, including: inputting the forest area panoramic deep semantic feature map and the forest area panoramic shallow salient feature map into a salient joint perception module based on the attention mechanism to obtain the forest area panoramic shallow-deep salient joint perception feature map.

[0083] Specifically, the forest area panoramic deep semantic feature map and the forest area panoramic shallow salient feature map are input into a salient joint perception module based on an attention mechanism to obtain the forest area panoramic shallow-deep salient joint perception feature map, including: performing feature shape reshaping on the forest area panoramic shallow salient feature map and the forest area panoramic deep semantic feature map to obtain a forest area panoramic shallow feature shape reshaping matrix and a forest area panoramic deep semantic feature shape reshaping matrix; inputting the forest area panoramic shallow feature shape reshaping matrix and the forest area panoramic deep semantic feature shape reshaping matrix into a feature channel-by-channel interactive perception module to obtain a forest area panoramic shallow detail-deep semantic dependency matrix and a forest area panoramic deep semantic-shallow detail dependency matrix; inputting the forest area panoramic shallow detail-deep semantic dependency matrix and the forest area panoramic deep semantic-shallow detail dependency matrix into a random dropout module to obtain a pruned forest area panoramic shallow detail- A deep semantic dependency matrix and a pruned forest area panoramic deep semantic-shallow detail dependency matrix; based on the pruned forest area panoramic shallow detail-deep semantic dependency matrix and the pruned forest area panoramic deep semantic-shallow detail dependency matrix, feature optimization is performed on the forest area panoramic deep semantic feature shape reshaping matrix and the forest area panoramic shallow feature shape reshaping matrix to obtain a dependency optimized forest area panoramic deep semantic feature matrix and a dependency optimized forest area panoramic shallow feature matrix; feature shape reshaping is performed on the dependency optimized forest area panoramic shallow feature matrix and the dependency optimized forest area panoramic deep semantic feature matrix to obtain an optimized forest area panoramic shallow salient feature map and an optimized forest area panoramic deep semantic feature map; the weighted sum of the optimized forest area panoramic shallow salient feature map and the optimized forest area panoramic deep semantic feature map is calculated to obtain the forest area panoramic shallow-deep salient joint perception feature map.

[0084] Furthermore, the forest area panoramic shallow feature shape reshaping matrix and the forest area panoramic deep semantic feature shape reshaping matrix are input into the feature channel-by-channel interactive perception module to obtain the forest area panoramic shallow detail-deep semantic dependency matrix and the forest area panoramic deep semantic-shallow detail dependency matrix, including: calculating the forest area panoramic shallow feature shape reshaping matrix multiplied by the transposed matrix of the forest area panoramic deep semantic feature shape reshaping matrix to obtain the forest area panoramic shallow detail-deep semantic association representation matrix; dividing the forest area panoramic shallow detail-deep semantic association representation matrix by the forest area panoramic The scale of the shallow feature shape reshaping matrix is ​​input into the softmax function to obtain the forest area panoramic shallow detail-deep semantic dependency matrix; the transposed matrix of the forest area panoramic deep semantic feature shape reshaping matrix is ​​calculated and multiplied by the forest area panoramic shallow feature shape reshaping matrix to obtain the forest area panoramic deep semantic-shallow detail association representation matrix; the forest area panoramic deep semantic-shallow detail association representation matrix is ​​divided by the scale of the forest area panoramic shallow feature shape reshaping matrix and then input into the softmax function to obtain the forest area panoramic deep semantic-shallow detail dependency matrix.

[0085] Furthermore, based on the pruned forest area panoramic shallow detail-deep semantic dependency matrix and the pruned forest area panoramic deep semantic-shallow detail dependency matrix, the forest area panoramic deep semantic feature shape reshaping matrix and the forest area panoramic shallow feature shape reshaping matrix are feature optimized to obtain a dependency optimized forest area panoramic deep semantic feature matrix and a dependency optimized forest area panoramic shallow feature matrix, including: matrix multiplying the pruned forest area panoramic shallow detail-deep semantic dependency matrix with the forest area panoramic deep semantic feature shape reshaping matrix to obtain the dependency optimized forest area panoramic deep semantic feature matrix; matrix multiplying the pruned forest area panoramic deep semantic-shallow detail dependency matrix with the forest area panoramic shallow feature shape reshaping matrix to obtain the dependency optimized forest area panoramic shallow feature matrix.

[0086] In a specific example, the forest area panoramic deep semantic feature map and the forest area panoramic shallow salient feature map are input into a salient joint perception module based on an attention mechanism to obtain the forest area panoramic shallow-deep salient joint perception feature map, including: processing the forest area panoramic shallow salient feature map and the forest area panoramic deep semantic feature map using the following joint perception formula to obtain the forest area panoramic shallow-deep salient joint perception feature map, wherein the joint perception formula is:

[0087]

[0088]

[0089]

[0090]

[0091]

[0092]

[0093]

[0094]

[0095]

[0096]

[0097]

[0098] in, A panoramic shallow layer significant feature map of the forest area is shown. represents the panoramic deep semantic feature map of the forest area, represents the feature shape reshaping, Represents the shallow feature shape reshaping matrix of the forest panorama, Represents the shape reshaping matrix of the deep semantic features of the forest panorama, represents matrix multiplication, represents the transpose of the matrix, represents the scale of the forest area panoramic shallow feature shape reshaping matrix, that is, the width multiplied by the height of the forest area panoramic shallow feature shape reshaping matrix, and the scale of the forest area panoramic shallow feature shape reshaping matrix and the forest area panoramic deep semantic feature shape reshaping matrix are the same, is the normalized exponential function, represents the shallow detail-deep semantic dependency matrix of the forest panorama, represents the forest panorama deep semantics-shallow detail dependency matrix, represents random dropout processing, Represents the shallow detail-deep semantic dependency matrix of the pruned forest panorama, Represents the deep semantic-shallow detail dependency matrix of the pruned forest panorama, Represents dependency relationship optimization of the shallow feature matrix of the forest panorama, Representing dependency relationship optimization of forest panorama deep semantic feature matrix, Indicates the optimized shallow feature map of the forest panorama, represents the optimized forest area panoramic deep semantic feature map, and Represent different weight coefficients, Represents the shallow-deep joint significant perceptual feature map of the panoramic view of the forest area.

[0099] Then, based on the decoded depth information values, the panoramic forest image is projected into three-dimensional space to obtain a three-dimensional model of the forest scene. Specifically, by combining the decoded depth information values ​​with the two-dimensional pixel coordinates of the panoramic forest image, a depth mapping algorithm is used to map the two-dimensional image pixels to the corresponding decoded depth information values ​​to construct a three-dimensional point cloud model of the forest scene. This effectively restores the three-dimensional coordinates of each point in the forest environment, including trees, shrubs, ground, and other terrain features, achieving detailed three-dimensional reconstruction of complex forest areas.

[0100] In a preferred example, the forest area panoramic shallow-deep significant joint perception feature map is input into a decoder-based depth estimation module to obtain the depth information decoding value, including the steps of: clustering all eigenvalues ​​of the forest area panoramic shallow-deep significant joint perception feature map based on the L2 distance between the eigenvalues, and arranging the cluster features into a forest area panoramic shallow-deep significant joint perception cluster vector; determining the cluster ratio of the number of eigenvalues ​​of the forest area panoramic shallow-deep significant joint perception cluster vector to the number of eigenvalues ​​of the forest area panoramic shallow-deep significant joint perception feature map value; dividing the two norms of the shallow-deep significant joint perception clustering vector of the forest panorama by the two norms of the shallow-deep significant joint perception feature vector of the forest panorama obtained after the shallow-deep significant joint perception feature map of the forest panorama is expanded to obtain the forest panorama shallow-deep significant joint perception conflict representation value; dividing the first power value of the one norm of the shallow-deep significant joint perception clustering vector of the forest panorama with the cluster ratio value as the exponent by the second power value of the one norm of the shallow-deep significant joint perception feature vector of the forest panorama with the cluster ratio value as the exponent to obtain the forest panorama shallow-deep significant joint perception conflict representation value. Shallow-deep significant joint perception confrontation representation value; for each eigenvalue of the forest area panoramic shallow-deep significant joint perception clustering vector, multiply it by the inverse of the difference between the forest area panoramic shallow-deep significant joint perception conflict representation value and the forest area panoramic shallow-deep significant joint perception confrontation representation value to obtain the optimized eigenvalue of the forest area panoramic shallow-deep significant joint perception clustering vector; for each eigenvalue outside the cluster in the forest area panoramic shallow-deep significant joint perception feature map, multiply it by the difference between the forest area panoramic shallow-deep significant joint perception conflict representation value and the forest area panoramic shallow-deep significant joint perception confrontation representation value The inverse of the sum of the forest area panoramic shallow-deep significant joint perception adversarial representation values ​​is obtained to obtain the optimized out-of-class eigenvalue of the forest area panoramic shallow-deep significant joint perception feature map; the optimized eigenvalue of the forest area panoramic shallow-deep significant joint perception clustering vector and the optimized out-of-class eigenvalue of the forest area panoramic shallow-deep significant joint perception feature map are composed of an optimized forest area panoramic shallow-deep significant joint perception feature map; the optimized forest area panoramic shallow-deep significant joint perception feature map is input into the decoder-based depth estimation module to obtain the depth information decoding value.

[0101] The optimization process of the shallow-deep joint salient perceptual feature map of the forest panorama is expressed as follows:

[0102]

[0103]

[0104] in, is the shallow-deep joint significant perceptual feature vector of the forest panorama, is the number of eigenvalues ​​of the shallow-deep significant joint perceptual eigenvector of the forest panorama, is the shallow-deep joint significant perceptual clustering vector of the forest panorama, is the number of eigenvalues ​​of the shallow-deep significant joint perception clustering vector of the forest panorama, is the clustering ratio value, represents the cluster feature set corresponding to the shallow-deep significant joint perception clustering vector of the forest panorama, and Represents the vector's two-norm and one-norm respectively Power, is each eigenvalue of the shallow-deep significant joint perceptual feature map of the forest panorama, It is each eigenvalue of the optimized forest area panoramic shallow-deep significant joint perceptual feature map.

[0105] That is, considering that the forest area panoramic deep semantic feature map and the forest area panoramic shallow salient feature map respectively represent the image shallow features and image deep semantic features of the forest area panoramic image based on feature distribution gradient mask saliency, when further performing salient joint perception based on the attention mechanism, the image semantic salient joint alignment conflict will be caused due to the difference in image shallow-deep fusion attention weights, thereby causing the loss of aggregated key information, affecting the expression effect of the forest area panoramic shallow-deep salient joint perception feature map.

[0106] Therefore, in order to avoid the loss of key suffix semantic information relative to the original feature set as a whole due to aggregation conflict when the forest area panoramic shallow-deep significant joint perception feature map is based on aggregated features, the clustering ratio of the number of eigenvalues ​​of the forest area panoramic shallow-deep significant joint perception clustering vector relative to the number of eigenvalues ​​of the forest area panoramic shallow-deep significant joint perception feature map is used as a judgment function to perform an adversarial judgment of the absolute representation of the set of the forest area panoramic shallow-deep significant joint perception clustering vector and the norm of the forest area panoramic shallow-deep significant joint perception feature vector, and respectively compare it with the forest area panoramic shallow-deep significant joint perception clustering vector and the forest area panoramic shallow-deep significant joint perception feature vector. The clustering intrinsic conflict representation of the two-norm of the forest area panoramic shallow-deep significant joint perception feature vector is positively and negatively interacted to construct the optimized forest area panoramic shallow-deep significant joint perception feature map based on the aggregated features and the solid alignment guardrail with the original feature set as a whole, thereby achieving the intentional mitigation of the harmful information loss of the optimized forest area panoramic shallow-deep significant joint perception feature map based on the aggregation risk transferability, improving the expression effect of the optimized forest area panoramic shallow-deep significant joint perception feature map, and thus improving the accuracy of the depth information decoding value obtained by the forest area panoramic shallow-deep significant joint perception feature map input into the decoder-based depth estimation module.

[0107] Based on the above embodiments, see Figure 5 The figure shows a schematic diagram of the structure of a monocular panoramic vision 3D reconstruction system 800 for forest scenes according to an embodiment of the present application. The monocular panoramic vision 3D reconstruction system 800 for forest scenes includes: an image acquisition module 810 for acquiring a panoramic image of the forest area captured by a monocular panoramic camera; a multi-level feature extraction module 820 for performing multi-level feature extraction on the panoramic image of the forest area to obtain a panoramic shallow feature map of the forest area and a panoramic deep semantic feature map of the forest area; a mask calculation module 830 for inputting the panoramic shallow feature map of the forest area into a feature distribution gradient mask salient detector to obtain a panoramic shallow salient feature map of the forest area; a decoding value generation module 840 for generating a depth information decoding value of the forest area image based on the joint perception information of the panoramic deep semantic feature map of the forest area and the panoramic shallow salient feature map of the forest area; and a 3D model generation module 850 for projecting the panoramic image of the forest area into a 3D space based on the depth information decoding value to obtain a 3D model of the forest area scene.

[0108] Here, those skilled in the art will appreciate that the specific functions and operations of the various modules in the above-mentioned monocular panoramic vision 3D reconstruction system 800 for forest scenes have been described in detail in the above reference. Figures 2 to 4 It has been introduced in detail in the description of the monocular panoramic vision 3D reconstruction method for forest scenes, and therefore, its repeated description will be omitted.

[0109] Figure 6 FIG is an application scenario diagram of a monocular panoramic vision 3D reconstruction method for forest scenes according to an embodiment of the present application. Figure 6 As shown, in this application scenario, first, a panoramic image of the forest area captured by a monocular panoramic camera is obtained (for example, Figure 6 Then, the forest area panoramic image is input to a server that is equipped with a monocular panoramic vision 3D reconstruction algorithm for forest area scenes (for example, Figure 6 In S) as shown in , the server is capable of using the monocular panoramic vision three-dimensional reconstruction algorithm for forest scenes to process the forest panoramic image to generate a depth information decoding value of the forest image, and then, based on the depth information decoding value, project the forest panoramic image into a three-dimensional space to obtain a three-dimensional model of the forest scene.

[0110] Furthermore, in another example of the present application, a monocular panoramic vision 3D reconstruction method for forest scenes is provided, referring to Figure 7As shown, it includes: S610, obtaining panoramic image data of a typical forest area under different environmental factors such as light and climate collected by a panoramic camera; S620, counting the relevant features of the panoramic image of the forest area, and exploring the benefits of relevant clues such as light on the accuracy of panoramic image depth estimation; S630, using a monocular panoramic vision-depth clue fusion method to establish a depth estimation model for forest area images; S640, using the required depth information to project the panoramic image into three-dimensional space and perform three-dimensional reconstruction of the forest area scene. Further, the process of obtaining the panoramic depth map refers to Figure 8 As shown in the figure, it encodes and decodes the implicit depth cues from the panoramic image using a convolutional neural network, and performs multi-scale encoding of the explicit depth cues from the copied image using a wavelet transform to obtain a multi-scale depth map. This is then fused through an attention fusion module to obtain a panoramic depth map. This process also involves spectral-spatial domain transformation and multi-scale mask calculation.

[0111] Based on the above embodiment, the present application also provides another exemplary embodiment of an electronic device. In some possible implementations, the electronic device in the present application may include a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, can implement the steps of the monocular panoramic vision 3D reconstruction method for forest scenes in the above embodiment.

[0112] For example, in the case of electronic equipment Figure 1 Taking the server 100 in the example for explanation, the processor in the electronic device is the processor 110 in the server 100, and the memory in the electronic device is the memory 120 in the server 100.

[0113] Embodiments of the present application also provide a computer-readable storage medium having computer-executable instructions stored thereon. When the computer-executable instructions are executed by a processor, the monocular panoramic vision 3D reconstruction method for forest scenes according to the embodiments of the present application described with reference to the above figures can be executed. The computer-readable storage medium includes, but is not limited to, volatile memory and / or non-volatile memory. The volatile memory may, for example, include random access memory (RAM) and / or cache memory. The non-volatile memory may, for example, include read-only memory (ROM), a hard disk, flash memory, etc.

[0114] Embodiments of the present application also provide a computer program product or computer program, which includes computer-executable instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the computer device to perform the monocular panoramic vision 3D reconstruction method for forest scenes according to an embodiment of the present application.

[0115] Those skilled in the art will appreciate that the contents disclosed in this application may be subject to various modifications and improvements. For example, the various devices or components described above may be implemented through hardware, software, firmware, or a combination of some or all of the three.

[0116] Furthermore, although this application makes various references to certain units in the system according to embodiments of the present application, any number of different units may be used and run on the client and / or server. The units are illustrative only, and different aspects of the system and method may use different units.

[0117] Those skilled in the art will appreciate that all or part of the steps in the above method can be performed by instructing the relevant hardware through a program, and the program can be stored in a computer-readable storage medium, such as a read-only memory, a magnetic disk, or an optical disk. Alternatively, all or part of the steps in the above embodiment can also be implemented using one or more integrated circuits. Accordingly, each module / unit in the above embodiment can be implemented in the form of hardware or in the form of software functional modules. This application is not limited to any particular form of combination of hardware and software.

[0118] Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. It should also be understood that terms such as those defined in common dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology and should not be interpreted in an idealized or highly formal sense, unless expressly defined as such herein.

[0119] The above is an explanation of the present application and should not be considered as limiting thereof. Although several exemplary embodiments of the present application have been described, those skilled in the art will readily appreciate that many modifications may be made to the exemplary embodiments without departing from the novel teachings and advantages of the present application.

Claims

1. A monocular panoramic vision 3D reconstruction method for forest scenes, characterized by: include: Obtain a panoramic image of the forest area captured by a monocular panoramic camera; Performing multi-level feature extraction on the forest area panoramic image to obtain a forest area panoramic shallow feature map and a forest area panoramic deep semantic feature map; Inputting the forest area panoramic shallow feature map into a feature distribution gradient mask salient device to obtain a forest area panoramic shallow salient feature map; Generating a depth information decoding value of the forest area image based on the joint perception information of the forest area panoramic deep semantic feature map and the forest area panoramic shallow salient feature map, specifically comprising: The forest area panoramic deep semantic feature map and the forest area panoramic shallow salient feature map are jointly perceived to obtain a forest area panoramic shallow-deep salient joint perception feature map, wherein the forest area panoramic deep semantic feature map and the forest area panoramic shallow salient feature map are input into a salient joint perception module based on an attention mechanism to obtain the forest area panoramic shallow-deep salient joint perception feature map, specifically including: Performing feature shape reshaping on the forest area panoramic shallow layer salient feature map and the forest area panoramic deep layer semantic feature map to obtain a forest area panoramic shallow layer feature shape reshaping matrix and a forest area panoramic deep layer semantic feature shape reshaping matrix; Input the forest area panoramic shallow feature shape reshaping matrix and the forest area panoramic deep semantic feature shape reshaping matrix into a feature channel-by-channel interactive perception module to obtain a forest area panoramic shallow detail-deep semantic dependency matrix and a forest area panoramic deep semantic-shallow detail dependency matrix; Inputting the forest area panoramic shallow detail-deep semantic dependency matrix and the forest area panoramic deep semantic-shallow detail dependency matrix into a random dropout module to obtain a pruned forest area panoramic shallow detail-deep semantic dependency matrix and a pruned forest area panoramic deep semantic-shallow detail dependency matrix; Based on the pruned forest area panoramic shallow detail-deep semantic dependency matrix and the pruned forest area panoramic deep semantic-shallow detail dependency matrix, feature optimization is performed on the forest area panoramic deep semantic feature shape reshaping matrix and the forest area panoramic shallow feature shape reshaping matrix to obtain a dependency-optimized forest area panoramic deep semantic feature matrix and a dependency-optimized forest area panoramic shallow feature matrix; Reshaping the dependency-optimized forest area panoramic shallow feature matrix and the dependency-optimized forest area panoramic deep semantic feature matrix to obtain an optimized forest area panoramic shallow salient feature map and an optimized forest area panoramic deep semantic feature map; Calculating a weighted sum of the optimized forest area panoramic shallow salient feature map and the optimized forest area panoramic deep semantic feature map to obtain the forest area panoramic shallow-deep salient joint perceptual feature map; Inputting the forest area panoramic shallow-deep joint significant perceptual feature map into a decoder-based depth estimation module to obtain the depth information decoding value; Based on the depth information decoded value, projecting the panoramic image of the forest area into a three-dimensional space to obtain a three-dimensional model of the forest area scene; The forest area panoramic shallow-deep significant joint perception feature map is input into a depth estimation module based on a decoder to obtain the depth information decoding value, including the steps of: clustering all eigenvalues ​​of the forest area panoramic shallow-deep significant joint perception feature map based on the L2 distance between the eigenvalues, and arranging the cluster features into a forest area panoramic shallow-deep significant joint perception cluster vector; determining the cluster ratio value of the number of eigenvalues ​​of the forest area panoramic shallow-deep significant joint perception cluster vector and the number of eigenvalues ​​of the forest area panoramic shallow-deep significant joint perception feature map; clustering the forest area panoramic shallow-deep significant joint perception cluster vector into a forest area panoramic shallow-deep significant joint perception cluster vector; and clustering the forest area panoramic shallow-deep significant joint perception feature map into a forest area panoramic shallow-deep significant joint perception feature map. The forest area panoramic shallow-deep significant joint perception conflict representation value is obtained by dividing the two-norm of the forest area panoramic shallow-deep significant joint perception cluster vector by the two-norm of the forest area panoramic shallow-deep significant joint perception feature vector obtained after the forest area panoramic shallow-deep significant joint perception feature map is expanded; the forest area panoramic shallow-deep significant joint perception conflict representation value is obtained by dividing the first power value of the one-norm of the forest area panoramic shallow-deep significant joint perception cluster vector with the cluster ratio value as the exponent by the second power value of the one-norm of the forest area panoramic shallow-deep significant joint perception feature vector with the cluster ratio value as the exponent Layer significant joint perception confrontation representation value; for each eigenvalue of the forest area panoramic shallow-deep significant joint perception clustering vector, multiply it by the inverse of the difference between the forest area panoramic shallow-deep significant joint perception conflict representation value and the forest area panoramic shallow-deep significant joint perception confrontation representation value to obtain the optimized eigenvalue of the forest area panoramic shallow-deep significant joint perception clustering vector; for each eigenvalue outside the cluster in the forest area panoramic shallow-deep significant joint perception feature map, multiply it by the difference between the forest area panoramic shallow-deep significant joint perception conflict representation value and the forest area panoramic shallow-deep significant joint perception confrontation representation value The optimized out-of-class eigenvalue of the forest area panoramic shallow-deep significant joint perception feature map is obtained by calculating the reciprocal of the sum of the forest area panoramic shallow-deep significant joint perception adversarial representation values; the optimized eigenvalue of the forest area panoramic shallow-deep significant joint perception clustering vector and the optimized out-of-class eigenvalue of the forest area panoramic shallow-deep significant joint perception feature map are combined into an optimized forest area panoramic shallow-deep significant joint perception feature map; the optimized forest area panoramic shallow-deep significant joint perception feature map is input into the decoder-based depth estimation module to obtain the depth information decoding value; By taking the clustering ratio of the number of eigenvalues ​​of the forest area panoramic shallow-deep significant joint perception clustering vector relative to the number of eigenvalues ​​of the forest area panoramic shallow-deep significant joint perception feature map as a judgment function, an adversarial judgment is made on the set absolute representation of the one-norm of the forest area panoramic shallow-deep significant joint perception clustering vector and the forest area panoramic shallow-deep significant joint perception feature vector, and positive and negative interactions are performed with the cluster intrinsic conflict representation of the two-norm of the forest area panoramic shallow-deep significant joint perception clustering vector and the forest area panoramic shallow-deep significant joint perception feature vector respectively, to construct a solid alignment guardrail of the optimized forest area panoramic shallow-deep significant joint perception feature map with the original feature set as a whole based on the aggregated features, thereby achieving the mitigation of the harmful information loss intention of the optimized forest area panoramic shallow-deep significant joint perception feature map based on the aggregated risk transferability, and improving the expression effect of the optimized forest area panoramic shallow-deep significant joint perception feature map.

2. The monocular panoramic vision 3D reconstruction method for forest scenes according to claim 1 is characterized in that: Performing multi-level feature extraction on the forest area panoramic image to obtain a forest area panoramic shallow feature map and a forest area panoramic deep semantic feature map, including: The forest area panoramic image is input into a panoramic image multi-scale feature extractor based on a pyramid network to obtain the forest area panoramic shallow feature map and the forest area panoramic deep semantic feature map.

3. The monocular panoramic vision 3D reconstruction method for forest scenes according to claim 2, characterized in that: Inputting the forest area panoramic shallow feature map into a feature distribution gradient mask salient device to obtain a forest area panoramic shallow salient feature map, including: Calculating the multidirectional gradient value distribution of each position in the forest area panoramic shallow feature map, and determining the gradient amplitude value of each position in the forest area panoramic shallow feature map based on the multidirectional gradient value distribution of each position to obtain a forest area panoramic shallow feature gradient amplitude distribution map; Calculating the gradient amplitude local description operator of each position in the forest area panoramic shallow feature gradient amplitude distribution map to obtain the forest area panoramic shallow feature gradient amplitude local significant distribution map; Inputting the forest area panoramic shallow feature gradient amplitude local significant distribution map into a gated masker based on a GELU function to obtain a gradient amplitude local significant gated mask map; The forest area panoramic shallow salient feature map is obtained by calculating the multiplication of the gradient amplitude local salient gated mask map and the forest area panoramic shallow feature map by position points.

4. The monocular panoramic vision 3D reconstruction method for forest scenes according to claim 3 is characterized in that: Calculating the gradient amplitude local description operator of each position in the forest area panoramic shallow feature gradient amplitude distribution map to obtain the forest area panoramic shallow feature gradient amplitude local significant distribution map, including: Determine the scale of the local neighborhood, calculate the average of the differences between the gradient amplitude value of a predetermined position in the forest area panoramic shallow feature gradient amplitude distribution map and the gradient amplitude values ​​of other positions in the local neighborhood to obtain a gradient amplitude local description operator corresponding to the predetermined position.

5. The monocular panoramic vision 3D reconstruction method for forest scenes according to claim 4, characterized in that: Inputting the forest area panoramic shallow feature shape reshaping matrix and the forest area panoramic deep semantic feature shape reshaping matrix into a feature channel-by-channel interactive perception module to obtain a forest area panoramic shallow detail-deep semantic dependency matrix and a forest area panoramic deep semantic-shallow detail dependency matrix, including: Calculating the forest area panoramic shallow feature shape reconstruction matrix multiplied by the transposed matrix of the forest area panoramic deep semantic feature shape reconstruction matrix to obtain a forest area panoramic shallow detail-deep semantic association representation matrix; The forest panorama shallow detail-deep semantic association representation matrix is ​​divided by the scale of the forest panorama shallow feature shape reshaping matrix and then input into the softmax function to obtain the forest panorama shallow detail-deep semantic dependency relationship matrix; Calculating the transposed matrix of the forest area panoramic deep semantic feature shape reconstruction matrix and multiplying it by the forest area panoramic shallow feature shape reconstruction matrix to obtain a forest area panoramic deep semantic-shallow detail association representation matrix; The forest area panoramic deep semantics-shallow detail association representation matrix is ​​divided by the scale of the forest area panoramic shallow feature shape reshaping matrix and then input into the softmax function to obtain the forest area panoramic deep semantics-shallow detail dependency matrix.

6. The monocular panoramic vision 3D reconstruction method for forest scenes according to claim 5, characterized in that: Based on the pruned forest area panoramic shallow detail-deep semantic dependency matrix and the pruned forest area panoramic deep semantic-shallow detail dependency matrix, feature optimization is performed on the forest area panoramic deep semantic feature shape reshaping matrix and the forest area panoramic shallow feature shape reshaping matrix to obtain a dependency-optimized forest area panoramic deep semantic feature matrix and a dependency-optimized forest area panoramic shallow feature matrix, including: Performing matrix multiplication on the pruned forest area panoramic shallow detail-deep semantic dependency matrix and the forest area panoramic deep semantic feature shape reshaping matrix to obtain the dependency-optimized forest area panoramic deep semantic feature matrix; The pruned forest area panoramic deep semantic-shallow detail dependency matrix is ​​matrix multiplied with the forest area panoramic shallow feature shape reshaping matrix to obtain the dependency optimized forest area panoramic shallow feature matrix.

Citation Information

Patent Citations

  • Image generation method, device and equipment and storage medium

    CN112270745A

  • Forest region positioning and three-dimensional reconstruction method and system based on multi-sensor fusion

    CN116228969A

  • Stereo matching method based on multi-feature aggregation

    CN117475182A

  • Wetland species classification method based on aerospace remote sensing fusion image

    CN118762237A

  • Preservation parameter adaptive optimization control system based on fruit state detection

    CN118778457A