Monocular depth estimation-based winter wheat heading stage growth vigor monitoring method
By applying the MDEHNet network based on monocular depth estimation in crop plant height measurement, the problems of low efficiency and poor adaptability in traditional methods are solved, and more accurate and efficient crop growth monitoring is achieved.
Patent Information
- Application Number
- CN202510637166.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-05-19
AI Technical Summary
The prior art has problems such as low efficiency, cumbersome operation, complex system construction and high cost in the measurement of crop heights, and the monocular depth estimation method has insufficient ability to extract feature and pay attention to detail.
A method for monitoring the growth of winter wheat heading period based on monocular depth estimation is proposed. Through the established MDEHNet network, multi-level feature extraction and refinement of the winter wheat test field acquisition images, enhance the accuracy of depth estimation, and generate a depth map with the same height as the initial image through the depth information acquisition subnet.
The accuracy and adaptability of crop height measurement are improved, the problems of low efficiency and poor adaptability of traditional methods are overcome, and more accurate and efficient crop growth monitoring is achieved.
Smart Images

Figure CN120198688A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and specifically relates to a method for monitoring the growth of winter wheat at the heading stage based on monocular depth estimation. Background Technique
[0002] Plant height is an important indicator reflecting the growth status of crops. Obtaining accurate plant height data is of great significance for agricultural production management, agricultural monitoring and other fields. How to quickly, conveniently and accurately obtain crop plant height data has always been an important issue concerned in the agricultural field.
[0003] Traditional methods for measuring crop height mainly rely on manual measurement, but there are problems such as low efficiency and cumbersome operation. With the development of information technology, non-contact measurement methods have attracted people's attention. Such methods extract plant height information without damaging crop growth, and have the advantages of non-contact and high efficiency, which is an ideal technical path. Based on parametric calibration space calculation, multi-view three-dimensional reconstruction, and depth cameras are the current main measurement methods. The method based on parametric calibration space calculation calculates the actual size of the object to be measured through a standard part with real size information or spatial calibration at a known position. For the three-dimensional reconstruction method, researchers take images of rice crops with a smartphone parallel to the measurement platform to calculate the height of the rice crops, and obtain the left and right views of the lawn by building a binocular stereo vision platform to calculate the overall height of the lawn. The depth camera obtains the three-dimensional structure information of the crop through technologies such as structured light and time of flight, so as to estimate the plant height.
[0004] Although the above methods have high detection accuracy, they usually need to build special data acquisition equipment and environments according to specific plant types, with complex system construction and high costs, and it is difficult to achieve large-scale applications. At the same time, under certain conditions, they may also interfere with or damage the plants, and there are certain limitations. With the continuous development of computer vision technology, monocular vision technology has shown great potential and broad application prospects in many fields. Some research has proposed a depth estimation method for the analysis of single-frame images of apple fruit trees, and introduced a feature fusion module to improve the network's perception ability of the structural details of fruit trees. Other researchers have established a depth estimation model, conducted experiments on tomato plant images, and extracted the deep features of the images through an encoder-decoder structure to effectively model the structure of tomato plants.
[0005] In the monocular depth estimation task, many CNN-based methods are widely used in feature extraction tasks, such as backbone networks like ResNet, HRNet, DenseNet, and EfficientNet. These methods have achieved good results in many image analysis tasks. However, in crop images, there are problems such as poor feature extraction ability and weak attention to fine details of crops. Commonly used feature refinement methods usually use multi-scale convolution, decoder architectures, and upsampling strategies to gradually restore the high-resolution details of the image. Many depth estimation methods, such as FCRN and DORN, use this strategy to enhance the details of depth information. However, these methods also have certain limitations. They are often vulnerable to noise interference. Especially in crop images, background noise and different planting methods may lead to insufficiently fine refinement results and prone to incorrect depth estimation results. Summary of the Invention
[0006] To solve the above technical problems, the present application proposes the following technical solutions: In a first aspect, an embodiment of the present application provides a method for monitoring the growth of winter wheat at the heading stage based on monocular depth estimation, including: Performing multi-level feature extraction on the images collected from the winter wheat experimental field through the established winter wheat heading stage growth monitoring network MDEHNet and adding important feature information to obtain image feature information; The MDEHNet network restores the image details of the image feature information and improves the accuracy of depth estimation to obtain refined image feature information; The refined image feature information is upsampled through the MDEHNet network to achieve size restoration and converted into a depth map with the same height as the initial image, and the depth map reflects the depth information of each pixel in the initial image.
[0007] In a possible implementation, the MDEHNet includes a feature extraction sub-network, a feature refinement module FRM, and a depth information acquisition sub-network, where: the feature extraction sub-network includes a plurality of feature extraction modules FEM connected in sequence, and the FEM sequentially compresses the size of the images collected from the winter wheat experimental field; the FRM receives the image feature information output by the last FEM and refines the image features; the depth information acquisition sub-network includes a plurality of upsampling modules and a depth prediction module DPM connected in sequence, and the refined image feature information output by the FRM sequentially restores the image size through the plurality of upsampling modules, and the information output by the last upsampling module is input to the DPM to obtain a depth map with the same height as the initial image.
[0008] In a possible implementation, the FEM includes a feature extraction sub-network and a CBAM attention mechanism. The feature extraction sub-network consists of multiple convolutional layers, and the ReLU activation function is used between adjacent convolutional layers. The image collected from the winter wheat experimental field is input into the convolutional layer of the first FEM, and is output from the last convolutional layer of the first FEM to the CBAM attention mechanism. After enhancing the feature representation through the CBAM attention mechanism, it is output to the next FEM. Different FEMs process the input image step by step to extract feature maps at different levels.
[0009] In a possible implementation, the CBAM attention mechanism enhances the feature representation of the network through a channel attention mechanism and a spatial attention mechanism; the channel attention mechanism performs global average pooling and max pooling on the channel dimension of each feature map to generate statistical information about the channel importance, and the statistical information is processed through a shared multi-layer perceptron and an activation function to generate a channel attention weight; the spatial attention mechanism obtains spatial information using average pooling and max pooling after receiving the channel attention weight, and generates a spatial attention map through a convolution operation.
[0010] In a possible implementation, the FRM includes multiple parallel processing branches. Each processing branch includes a convolutional kernel and a ReLU activation function. The convolutional kernels of different processing branches have different sizes. The convolutional kernels perform convolutional processing on the feature map output by the last FEM of the feature extraction sub-network, and splice or weighted sum these convolutional outputs to generate a feature map containing multi-scale information; then use consecutive convolutional layers to deeply process the feature map to strengthen the non-linear expression ability of the features; after depth convolution, the ECA attention mechanism is introduced to calculate the importance of each channel and suppress irrelevant features.
[0011] In a possible implementation, the ECA performs global average pooling on each channel of the input feature map through global average pooling (GAP) to obtain a vector and acquire the global information of each channel. Then, the size of the convolutional kernel is selected according to the input feature map. Each time the number of channels for feature extraction transformation changes, the size of the convolutional kernel changes accordingly. After convolution, it passes through an activation function to output the channel attention weight, and the obtained channel attention weight is multiplied element-wise with the original input feature map to output the refined feature map.
[0012] In a possible implementation, the DPM includes two consecutive first convolutional layers and a second convolutional layer. The two convolutional layers extract detailed information to generate a depth map with spatial details. The output of the second convolutional layer is connected to a Batch Normalization (BN), followed by a ReLU activation function. BN enhances the stability of the model, and the ReLU activation function is used to remove negative values to improve the non-linear expression ability of the model. Finally, through a third convolutional layer and a depth-to-height conversion module, the pixel values in the predicted depth map are converted into height values in the physical space.
[0013] In a possible implementation, the depth-to-height conversion module converts the pixel values in the predicted depth map into height values in the physical space, including: Obtaining the internal parameters according to camera calibration, and using the camera parameters to obtain the vertical field of view angle and the horizontal field of view angle : , , where: and are the height and width of the image, and are the calibrated horizontal and vertical focal lengths; After monocular depth estimation, the depth value is obtained on the result map, and the crop pixel height is obtained based on it. Then the depth value is converted into a real value , and finally the real height is calculated: , where: is the parameter field of view angle of the camera, is the vertical resolution of the depth map, is the crop pixel height in the depth map.
[0014] In a possible implementation, multiple loss functions are selected and combined to improve the details and effects of depth prediction, including: Using pixel-level regression loss to measure the absolute error between the predicted depth and the real depth. The formula is: Using scale-invariant loss to eliminate the influence of global scale factors and encourage the model to focus on the relative relationship of depth. The formula is: The combined loss function of the and is the formula: where, is the total number of pixels, is the predicted value of the th pixel, is the true value of the th pixel, and represent the gradients in the horizontal and vertical directions, is the predicted depth value at the pixel ( , ), is the true depth value at the pixel ( , ), is the weight of the is the weight of the
[0015] In the embodiments of the present application, MDEHNet extracts semantic features from images, enhances the attention to the crop area, suppresses redundant background information, and improves the perception ability of the model. Then, refined image feature information with more crop information is obtained, which can improve the accuracy and robustness of depth estimation. Finally, the refined image feature information is processed to output a pixel-level depth map. And a geometric transformation strategy is incorporated into the structure to accurately map the predicted depth value and generate the true height of the crop. Through the above process, the measurement accuracy and adaptability are improved, and the problems of low efficiency and poor adaptability of traditional agricultural measurement methods are overcome. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is a schematic flowchart of a method for monitoring the growth of winter wheat at the heading stage based on monocular depth estimation provided by an embodiment of the present application; Figure 2 is a sample image of winter wheat provided by an embodiment of the present application; Figure 3 is a structural diagram of the MDEHNet model provided by an embodiment of the present application; Figure 4 is a structural diagram of the FEM module provided by an embodiment of the present application; Figure 5 is a structural diagram of the CBAM provided by an embodiment of the present application; Figure 6 is a structural diagram of the FRM module provided by an embodiment of the present application; Figure 7 is a structural diagram of the ECA provided by an embodiment of the present application; Figure 8 is a structural diagram of the DPM module provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] The following describes the present solution in conjunction with the drawings and specific embodiments.
[0018] See Figure 1 , the method for monitoring the growth of winter wheat at the heading stage based on monocular depth estimation in this embodiment includes: S101, extracting multi-level features from the images collected from the winter wheat experimental field through the established winter wheat heading stage growth monitoring network MDEHNet and adding important feature information to obtain image feature information.
[0019] In this embodiment, images of the winter wheat experimental field are collected first. The resolution of the original images is 4032×3024, and the aspect ratio is 4:3. During the shooting process, a parallel bracket is used to fix the mobile phone, making the lens face directly in front of the crop, ensuring that the lowest and highest points of the crop are completely covered in each image, and avoiding the influence of local crop loss on depth estimation. A total of 900 images at the heading stage are taken. See Figure 2 , which is an image sample of part of the winter wheat experimental field.
[0020] After the collection is completed, the images are preprocessed, including removing images with severe overexposure or occlusion, uniformly cropping and scaling the images to a resolution of 640×480. Median filtering is applied to denoise the images, histogram equalization is used to enhance the image contrast, and data augmentation (such as random flipping and slight rotation) is performed.
[0021] The winter wheat heading stage growth monitoring network MDEHNet constructed in this embodiment adopts a design based on the encoder-decoder architecture, uses RGB images to predict depth maps, and measures the height of winter wheat. After MDEHNet is constructed, the random gradient descent algorithm is used to train the model. During the training process, the Adam optimizer is used for optimization. The hyperparameters used for training are: the batch size is 4, the number of iterations is 400, and the learning rate is set to 0.001.
[0022] The steps of each round of training are as follows: 1) Input the label data pair into the model. 2) Use the model to perform a forward calculation on the current training data. 3) Calculate the loss using the loss function. 4) Use the random gradient descent algorithm to update the parameters of the model to complete one training process. 5) Evaluate the performance of the current model on the validation set, record the loss value and other evaluation metrics to select the best model for testing. Repeat steps 3) to 5) until the loss function is less than the specified expected value or the loss value no longer changes.
[0023] After each round of training, save the current model parameters and generate a parameter file for testing. By selecting the model parameters with the smallest validation loss as the best model, it is used to evaluate the test data. After training, select the smallest validation loss value as the optimal parameter to test the model performance.
[0024] In this embodiment, multiple loss functions are selected and used in combination to improve the details and effects of depth prediction, including: using pixel-level The regression loss measures the absolute error between the predicted depth and the true depth. The formula is: Using scale-invariant loss To eliminate the influence of the global scale factor and encourage the model to focus on the relative relationship of depth. The formula is: The and The combined loss function is in the form of: Where Is the total number of pixels, Is the predicted value of the th pixel, Is the th pixel's true value, and Represent the gradients in the horizontal and vertical directions, Is the value of the predicted depth at the pixel ( , ), Is the value of the true depth at the pixel ( , ), Is The weight of the loss, Is The weight of the loss.
[0025] See Figure 3 , the MDEHNet in this embodiment includes a feature extraction sub-network, an FRM (Feature Refinement Module), and a depth information acquisition sub-network. Among them: the feature extraction sub-network includes a plurality of FEMs (Feature Extraction Modules) connected in sequence, and the FEMs sequentially perform size compression on the images collected from the winter wheat test field. The FRM receives the image feature information output by the last FEM and refines the image features. The depth information acquisition sub-network includes a plurality of upsampling modules and a DPM (Depth Prediction Module) connected in sequence. The refined image feature information output by the FRM sequentially restores the image size through a plurality of the upsampling modules, and the information output by the last upsampling module is input into the DPM to obtain a depth map with the same height as the initial image.
[0026] To obtain image feature information, the FEM in this embodiment consists of a feature extraction sub-network and a CBAM attention mechanism. The feature extraction sub-network is combined with the CBAM attention mechanism to enhance the feature extraction effect. Refer to Figure 4 , the feature extraction sub-network consists of multiple convolutional layers, and the ReLU activation function is used between adjacent convolutional layers. The design is based on grouped convolution, which can reduce the computational complexity, maintain the network's expressive ability, and at the same time reduce the number of parameters. The images collected from the winter wheat experimental field are input into the convolutional layer of the first FEM, output from the last convolutional layer of the first FEM to the CBAM attention mechanism, and after enhancing the feature representation through the CBAM attention mechanism, it is output to the next FEM. Different FEMs process the input image step by step to extract feature maps at different levels.
[0027] Refer to Figure 5 , the CBAM attention mechanism enhances the network's feature representation through a channel attention mechanism and a spatial attention mechanism; the channel attention mechanism performs global average pooling and max pooling on the channel dimension of each feature map to generate statistical information about the channel importance, and the statistical information is processed through a shared multi-layer perceptron and an activation function to generate a channel attention weight; the spatial attention mechanism uses average pooling and max pooling to obtain spatial information after receiving the channel attention weight, and generates a spatial attention map through a convolutional operation. The dual attention mechanism enables the model to more accurately focus on key features and effectively improve the robustness and accuracy of depth estimation.
[0028] S102, the MDEHNet network restores the image details of the image feature information and improves the accuracy of depth estimation to obtain refined image feature information.
[0029] The FRM is used to enhance the features extracted by the encoder, restore image details and improve the accuracy of depth estimation, and consists of fine convolutional operations and ECA (Efficient Channel Attention, attention mechanism). Refer to Figure 6 , the FRM includes multiple parallel processing branches, each processing branch includes a convolutional kernel and a ReLU activation function, the convolutional kernels of different processing branches have different sizes, the convolutional kernels perform convolutional processing on the feature map output by the last FEM of the feature extraction sub-network, and these convolutional outputs are concatenated or weighted and summed to generate a feature map containing multi-scale information; then continuous convolutional layers are used to deeply process the feature map to strengthen the non-linear expression ability of the features; after depth convolution, the ECA attention mechanism is introduced to calculate the importance of each channel and suppress irrelevant features.
[0030] The convolution part in the FRM of this embodiment specifically includes multi-scale convolution fusion and depth convolution processing. The input of the feature refinement module comes from the feature extraction sub-network, including the extracted low-resolution feature maps, which have undergone convolution operations and downsampling processing through the encoder layer. The size of the feature map is H×W×C, where H and W are the height and width of the image, and C is the number of channels.
[0031] Specifically, in this embodiment, FRM first processes the feature maps from different levels through multi-scale fusion. Convolution kernels of 3×3, 5×5, and 7×7 are respectively used to perform convolution processing on the downsampled image feature information, and these convolution outputs are concatenated or weighted and summed to generate a feature map containing multi-scale information. Next is depth convolution, which uses three consecutive convolution layers to deeply process the feature map to strengthen the non-linear expression ability of the features. Each convolution layer is equipped with a ReLU activation function to enhance the learning ability of the network and avoid the problem of gradient disappearance. After depth convolution, the ECA attention mechanism is introduced. Compared with CBAM, ECA does not use global pooling operations and directly calculates channel attention through convolution operations. It calculates the importance of each channel in a more lightweight manner and suppresses irrelevant features. The structure of ECA is as Figure 7 shown.
[0032] ECA performs global average pooling on each channel of the input feature map through GAP (Global Average Pooling) to obtain a 1×1×C vector and obtain the global information of each channel. Then, the size of the convolution kernel is selected according to the input feature map (marked in red in the figure). Each time the number of channels for feature extraction transformation changes, the size of the convolution kernel K changes accordingly. Next, the selected kernel size is used to perform convolution operations on the result after GAP. Here, two different convolution kernels are used (marked in yellow and blue in the figure, representing kernels of different sizes respectively). After convolution, a 1×1×C vector is output, and σ is the activation function used to obtain the channel weights. Finally, the obtained channel attention weights are multiplied element-wise with the original input feature map to achieve channel attention weighting, and the output feature map is X'.
[0033] S103, the refined image feature information is upsampled through the MDEHNet network to achieve size restoration and is converted into a depth map with the same height as the initial image. The depth map reflects the depth information of each pixel in the initial image.
[0034] DPM converts the feature map after feature extraction and refinement processing into a depth map and generates a real height through height conversion. Its structure is as Figure 8As shown in the figure. The DPM includes two consecutive first convolutional layers and a second convolutional layer. The two convolutional layers extract detailed information, maintain the spatial consistency of the feature map, and generate a depth map with spatial details. The output of the second convolutional layer is connected to a normalization strategy BN, followed by a ReLU activation function. BN improves the stability of the model, and the ReLU activation function is used to remove negative values to enhance the non-linear expression ability of the model. Finally, through the third convolutional layer and the depth-to-height conversion module, the pixel values in the predicted depth map are converted into height values in the physical space to ensure that the generated depth map can accurately reflect the depth information of each pixel.
[0035] To measure the true height of crops, the DPM module also integrates a depth-to-actual height conversion mechanism. This mechanism converts the pixel values in the predicted depth map into height values in the physical space based on the internal parameters of the camera and the shooting distance.
[0036] Specifically, the internal parameters are obtained through camera calibration, and the vertical field of view angle and the horizontal field of view angle are obtained using the camera parameters: and , where: and are the height and width of the image, and are the calibrated horizontal and vertical focal lengths. After monocular depth estimation, the depth value is obtained on the result map, and the crop pixel height is obtained based on it. Then the depth value is converted into a real value , and finally the true height is calculated: , where: is the parameter field of view angle of the camera, is the vertical resolution of the depth map, is the crop pixel height in the depth map.
[0037] Next, in this embodiment, the method proposed in the above embodiment is evaluated in combination with experimental results.
[0038] The experiment is implemented based on the PyTorch framework of Python, and the running environment is an i5-12400F CPU and an NVIDIA GeForce RTX 3060 GPU. The code is written with the help of open source libraries such as GDAL (Geospatial Data Abstraction Library) and OpenCV as development components.
[0039] Considering that the BTS, FCRN, and DORN network structures are similar to MDEHNet, and the DPT model uses VisionTransformer as the network architecture, it can be used to compare the performance of Transformer and CNN in depth estimation tasks. Therefore, BTS, FCRN, DORN, and DPT are selected as comparison models.
[0040] The main metrics for measuring the performance of models in depth estimation problems are: Root Mean Square Error ( ), Log Root Mean Square Error ( ), Relative Error ( ), Squared Relative Error ( ), and Accuracy ( , , ). The calculation formulas are shown in Table 1.
[0041] Table 1 Evaluation Metrics and Calculation Formulas To further verify the effectiveness of the improved model, we evaluated the improved monocular depth estimation model and compared the depth prediction results of the BTS, FCRN, DORN, and DPT models.
[0042] Table 2 Evaluation Criteria Table for Comparison Models The experimental results are shown in Table 2. It can be seen that the improved model proposed in this study is superior to the existing methods in multiple metrics. Compared with the BTS model, this experiment reduced by 0.381, 0.046, 0.005, and 0.417 in the RMSE, LogRMSE, AbsRel, and SqRel metrics respectively; in the and metrics, it increased by 0.046, 0.018, and 0.002 respectively; compared with the FCRN model, this experiment reduced by 0.591, 0.071, 0.071, and 0.613 in the RMSE, LogRMSE, AbsRel, and SqRel metrics respectively; in the and metrics, it increased by 0.060, 0.019, and 0.011 respectively; compared with the DORN model, this experiment reduced by 0.333, 0.060, 0.051, and 0.386 in the RMSE, LogRMSE, AbsRel, and SqRel metrics respectively; in the and The indicators were improved by 0.033, 0.014, and 0.004 respectively; compared with the DPT model, in this experiment, the RMSE, LogRMSE, AbsRel, and SqRel indicators were reduced by 0.062, 0.015, 0.026, and 0.163 respectively; in and the indicators were improved by 0.028, 0.010, and 0.002 respectively. The RMSE and AbsRel errors are lower, indicating that the overall prediction error of the model is smaller and the depth estimation is more accurate. The lower SqRel error indicates that the improved model has a smaller error in long-distance depth estimation; The higher the
[0043] indicator means that the proportion of prediction samples within a smaller error range of the model is higher, and the accuracy is improved. The decrease in LogRMSE means that the model has a better depth estimation for different scales and reduces the error of distant targets.
[0044] Table 3 Generated height result table As can be seen from Table 3, the method proposed in this paper performs best in the height prediction of randomly selected winter wheat samples, with an average error of 2.93 cm, which is better than other comparison models. The average errors of DPT and DORN are 3.42 cm and 4.07 cm respectively, and their performances are also relatively stable. The errors of BTS and FCRN are relatively large, 5.28 cm and 4.78 cm respectively, showing certain deviations in the prediction task of crop height. Generally speaking, the method in this paper has a good effect in improving the prediction accuracy and reducing the error, verifying the good adaptability and reliability of the model in the actual farmland environment.
[0045] To verify the improvement effect of the CBAM attention mechanism in the FEM module on the model performance, we designed three groups of comparative experiments: in experiment a, only ResNeXt101 was used as the feature extraction backbone network without the CBAM module; in experiment b, the CBAM module was added on this basis. To verify the improvement effect of the FRM module on the model performance, experiment c was continued. Experiment c added the FRM module on the basis of experiment b. Table 4 lists the depth estimation evaluation indicators of the three groups of experiments.
[0046] Table 4 Depth estimation performance of three groups of experiments As can be seen from the above results, Experiment b has certain improvements in all indicators compared to Experiment a, especially in the SqRel and RMSE indicators, indicating that CBAM in FEM can effectively enhance the model's attention to key regions and contribute to the improvement of depth estimation accuracy. After introducing FRM, all evaluation indicators of Experiment c are further optimized, indicating that FRM further reduces the depth prediction error and improves the robustness and generalization ability of the model.
[0047] In addition, to further verify the improvement effect on the actual prediction results, we compared the predicted depth and generated height of the three groups with the camera distance of 100 cm as the benchmark, and the results are shown in Table 5.
[0048] Table 5 Comparison of depth prediction accuracy among three groups As can be seen from the results in Table 5, Experiment b has improved in the predicted depth compared to Experiment a, with an increase of 3.4 cm. This result indicates that after introducing CBAM, the model has become more accurate in depth estimation and can better predict the actual depth. The predicted depth of Experiment c has a 2.2 cm error reduction compared to Experiment b. This indicates that after introducing the FRM module, the FRM module effectively reduces the error and improves the accuracy of the model.
[0049] To verify the accuracy of the generated height, this paper further selected winter wheat samples, generated heights based on the predicted depth, and compared them with the true heights. The results show that the average error of Experiment a is 4.94 cm, Experiment b is reduced to 3.42 cm, and the average error of Experiment c is 1.49 cm, showing the best performance. Experiment b significantly reduces the prediction error compared to Group a, indicating that CBAM enhances the performance of the model; while Experiment c further reduces the error, indicating that the FRM module has an effect in refining features, thus improving the accuracy of the generated height.
[0050] In the depth estimation task, the camera shooting distance is an important factor affecting the depth prediction accuracy. Especially in application scenarios such as agricultural monitoring, changes in the camera distance may lead to significant fluctuations in crop height measurement errors. Different distances between the camera and the target object will result in changes in the captured image features and the quality of depth information, thus affecting the results of depth estimation. Therefore, this paper explores the impact of the camera distance on the depth estimation accuracy. By deeply analyzing the depth prediction results at different camera distances, it can provide a basis for model improvement and help design depth estimation methods more adaptable to different shooting distances. Next, based on the experimental results, the specific impact of the camera distance on the depth estimation accuracy will be analyzed. The experimental results are shown in Table 6.
[0051] Table 6 Depth prediction and true distance results at different camera distances As can be seen from Table 6, the camera shooting distance has a significant impact on the accuracy of depth estimation. For relatively short shooting distances, at 50 cm, 75 cm, and 100 cm, the difference between the predicted depth of the model and the true value is relatively small. Specifically, at 50 cm, the predicted difference of the model is 0.5 cm; at 75 cm, the difference is 0.2 cm; at 100 cm, the difference is 0.7 cm. This indicates that at shorter camera distances, the depth estimation model can perform depth prediction more accurately. As the camera shooting distance increases, the error of the predicted depth gradually increases. When the camera distance is 125 cm, the predicted difference is 4.5 cm; at 150 cm, the predicted difference is 16.7 cm. This trend is more obvious at 175 cm, with a predicted difference of 22.3 cm. This shows that as the camera distance increases, the error of depth estimation gradually increases, especially when shooting at a long distance, and the depth prediction accuracy of the model decreases.
[0052] In the embodiments of the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent the situation where A exists alone, A and B exist simultaneously, or B exists alone. Where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one of the following" and its similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, and c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0053] As described above, the above are only specific embodiments of the present application. Any person skilled in the art within the technical scope disclosed in the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. The protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A method for monitoring the growth of winter wheat at the heading stage based on monocular depth estimation, characterized in that: include: The winter wheat heading period growth monitoring network MDEHNet was established to extract multi-level features from the images collected from the winter wheat experimental field and add important feature information to obtain image feature information. The MDEHNet network restores the image details of the image feature information and improves the accuracy of depth estimation to obtain refined image feature information; The refined image feature information is adopted through the MDEHNet network to achieve size restoration and converted into a depth map with the same height as the initial image, and the depth map reflects the depth information of each pixel in the initial image.
2. The method for monitoring growth of winter wheat at heading stage based on monocular depth estimation according to claim 1, characterized in that: The MDEHNet includes a feature extraction subnetwork, a feature refinement module FRM, and a depth information acquisition subnetwork, wherein: the feature extraction subnetwork includes a plurality of feature extraction modules FEM connected in sequence, and the FEMs sequentially realize size compression of images collected from the winter wheat experimental field; the FRM receives the image feature information output by the last FEM and performs image feature refinement; the depth information acquisition subnetwork includes a plurality of upsampling modules and a depth prediction module DPM connected in sequence, and the refined image feature information output by the FRM sequentially realizes image size restoration through a plurality of the upsampling modules, and the information output by the last upsampling module is input into the DPM to obtain a depth map with the same height as the initial image.
3. The method for monitoring growth of winter wheat at heading stage based on monocular depth estimation according to claim 2, characterized in that: The FEM includes a feature extraction subnetwork and a CBAM attention mechanism. The feature extraction subnetwork is composed of multiple convolutional layers. Adjacent convolutional layers are connected by ReLU activation functions. The images collected from the winter wheat experimental field are input into the convolutional layer in the first FEM, and are output to the CBAM attention mechanism by the last convolutional layer of the first FEM. The feature representation is enhanced by the CBAM attention mechanism and then output to the next FEM. Different FEMs process the input image step by step to extract feature maps at different levels.
4. The method for monitoring growth of winter wheat at heading stage based on monocular depth estimation according to claim 3, characterized in that: The CBAM attention mechanism enhances the feature representation of the network through the channel attention mechanism and the spatial attention mechanism; the channel attention mechanism performs global average pooling and maximum pooling on the channel dimension of each feature map to generate statistical information about the importance of the channel, and the statistical information is processed through a shared multi-layer perceptron and an activation function to generate a channel attention weight; After receiving the channel attention weights, the spatial attention mechanism obtains spatial information using average pooling and maximum pooling, and generates a spatial attention map through convolution operations.
5. The method for monitoring growth of winter wheat at heading stage based on monocular depth estimation according to claim 2, characterized in that: The FRM includes multiple parallel processing branches, each processing branch includes a convolution kernel and a ReLU activation function. The convolution kernels of different processing branches have different sizes. The convolution kernel performs convolution processing on the feature map output by the last FEM of the feature extraction subnetwork, and concatenates or weighted sums these convolution outputs to generate a feature map containing multi-scale information; then, continuous convolution layers are used to perform in-depth processing on the feature map to enhance the nonlinear expression ability of the features; after deep convolution, an ECA attention mechanism is introduced to calculate the importance of each channel and suppress irrelevant features.
6. The method for monitoring growth of winter wheat at heading stage based on monocular depth estimation according to claim 5, characterized in that: The ECA performs global average pooling GAP on each channel of the input feature map to obtain a vector, obtains the global information of each channel, and then selects the size of the convolution kernel according to the input feature map. Each time the number of feature extraction channels is changed, the size of the convolution kernel changes accordingly. After convolution, an activation function is used to output the channel attention weight, and the obtained channel attention weight is multiplied element by element with the original input feature map to output the refined feature map.
7. The method for monitoring growth of winter wheat at heading stage based on monocular depth estimation according to claim 2, characterized in that: The DPM includes two consecutive first convolutional layers and a second convolutional layer. The two convolutional layers extract detail information and generate a depth map with spatial details. The output of the second convolutional layer is connected to the normalization strategy BN, and the BN is followed by the ReLU activation function. BN improves the stability of the model. The ReLU activation function is used to remove negative values and improve the nonlinear expression ability of the model. Finally, the pixel values in the predicted depth map are converted into height values in the physical space through the third convolutional layer and the depth-to-height conversion module.
8. The method for monitoring growth of winter wheat at heading stage based on monocular depth estimation according to claim 7, characterized in that: The depth-to-height conversion module converts pixel values in the predicted depth map into height values in the physical space, including: The internal parameters are obtained according to the camera calibration, and the vertical field of view angle is obtained using the camera parameters. and horizontal field of view : , ,in: and are the height and width of the image, and is the horizontal and vertical focal length obtained by calibration; After monocular depth estimation, the depth value is obtained on the result map, and the crop pixel height is obtained based on it, and then the depth value is converted into the real value , and finally calculate the true height : ,in: is the camera's field of view, is the vertical resolution of the depth map, is the crop pixel height in the depth map.
9. The method for monitoring growth of winter wheat at heading stage based on monocular depth estimation according to any one of claims 1 to 8, characterized in that: Select multiple loss functions to improve the details and effect of depth prediction, including: The regression loss measures the absolute error between the predicted depth and the true depth, and the formula is: Using scale-invariant loss Eliminating the influence of global scale factors and encouraging the model to focus on the relative relationship of depth, the formula is: The and The combined loss function is: in, is the total number of pixels, For the The predicted value of pixels, For the The true value of pixels, and Represents the gradient in the horizontal and vertical directions, To predict the depth in pixels ( , ), is the true depth in pixels ( , ), for The weight of the loss, for The weight of the loss.
Citation Information
Patent Citations
Remote sensing image change detection method based on hierarchical cross-scale global feature fusion deep network
CN117853897A
Depth estimation method based on multi-view self-supervised learning
CN118552596A
Lightweight monocular image depth estimation method and device
CN119941817A
Sparse depth estimation from plant traits
US20230136563A1
System, and method for estimating an object size in an image
WO2024144451A1