Plant terminal bud real-time positioning method and system based on double-flow convolutional network, terminal and storage medium
The binocular camera image is processed through the dual-flow convolution network, the optical flow field is calculated and the characteristics are fusion, which solves the problem of detection of plant apical bud position changes in dynamic environments, realizes high-precision real-time positioning, and supports precise agricultural operations.
Patent Information
- Application Number
- CN202510473809.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-08-01
AI Technical Summary
In the prior art, the identification of plant apical buds in dynamic scenarios is difficult to overcome the influence of environmental factors, resulting in difficulty in detecting location changes and high cost of model training, making it difficult to meet the real-time positioning needs of precision agriculture.
Using a method based on a dual-stream convolution network, the left and right images are obtained through a binocular camera, the optical flow algorithm is used to calculate the optical flow field, and the three-dimensional coordinates of the plant are calculated by combining feature extraction and fusion layer processing.
It improves the accuracy and positioning accuracy of plant apical bud recognition, and provides real-time positioning data support required for precise agricultural operations.
Smart Images

Figure CN120411224A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of object positioning, and in particular, to a real-time plant apical bud positioning method, system, terminal, and computer-readable storage medium based on a two-stream convolutional network. Background Art
[0002] With the progress of computer vision technology, 3D positioning and object tracking have become core tasks in many applications. Binocular cameras are widely used in depth estimation and target positioning, and the 3D coordinates of the target can be deduced by calculating the disparity between the left and right cameras. However, in dynamic scenes, due to factors such as the rapid movement, occlusion, and illumination changes of the target, traditional vision methods face many challenges. To improve the accuracy of target positioning, especially in dynamic scenes, optical flow estimation methods can help extract the motion information of the target.
[0003] However, there are still some technical problems and challenges in this field. For example, cotton fields are a complex dynamic environment, and environmental factors (such as illumination changes, shadows, meteorological conditions, etc.) have a great impact on apical bud recognition; cotton apical buds change continuously during growth and are easily affected by factors such as wind and pests, resulting in changes in the target position; to train a high-precision deep learning model, a large amount of high-quality image datasets are required. The apical bud recognition in dynamic scenes requires high real-time performance. Especially when using automated equipment for precision agriculture, the system needs to provide real-time feedback on the position and coordinates of the target, which increases the cost of model training.
[0004] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention
[0005] The main purpose of the present invention is to provide a real-time plant apical bud positioning method, system, terminal, and computer-readable storage medium based on a two-stream convolutional network, aiming to solve the problems in the existing technology that it is very difficult to detect the position change of plant apical buds due to environmental influence, and the model training cost is too high to improve the accuracy.
[0006] To achieve the above object, the present invention provides a real-time plant apical bud positioning method based on a two-stream convolutional network. The real-time plant apical bud positioning method based on a two-stream convolutional network includes the following steps:
[0007] Obtain the left and right images of the target plant collected by the binocular camera, input the left and right images into the constructed two-stream convolutional network, and perform optical flow calculation on all image frames of the left and right images through the optical flow algorithm to obtain the optical flow field of the left and right images;
[0008] Extract features from the left and right images through the two-stream convolutional network to obtain the static information of the target plant;
[0009] Through the feature fusion layer of the dual-stream convolutional network, fuse the optical flow field and the static information, and output the global features of the target plant;
[0010] Obtain the projection matrix of the binocular camera, and calculate the three-dimensional spatial coordinates of the target plant according to the projection matrix and the global features.
[0011] Optionally, in the method for real-time localization of plant apical buds based on a dual-stream convolutional network, wherein, obtaining the left and right images of the target plant collected by the binocular camera, inputting the left and right images into the pre-constructed dual-stream convolutional network, and performing optical flow calculation on all image frames of the left and right images through an optical flow algorithm to obtain the optical flow field of the left and right images, specifically including:
[0012] Obtain the initial left and right images of the target plant collected by the binocular camera, and preprocess the initial left and right images to obtain the left and right images;
[0013] Input the left and right images into the pre-constructed dual-stream convolutional network, and the dual-stream convolutional network performs optical flow prediction on all pixels of the left and right images through an optical flow algorithm to obtain the optical flow information of each pixel, and perform matching learning on each optical flow information to obtain a matching result;
[0014] Calculate the similarity between each pixel according to the matching result, and obtain the initial optical flow field of the left and right images according to all the similarities;
[0015] Update the initial optical flow field to obtain the optical flow field of the left and right images, wherein the optical flow field represents the dynamic information of the target plant.
[0016] Optionally, in the method for real-time localization of plant apical buds based on a dual-stream convolutional network, wherein, calculating the similarity between each pixel according to the matching result, and obtaining the initial optical flow field of the left and right images according to all the similarities, specifically including:
[0017] Extract the high-dimensional features of each pixel through the pooling layer of the dual-stream convolutional network:
[0018]
[0019] where F i represents the high-dimensional feature of the i-th pixel, CNN represents the dual-stream convolutional network, and I i represents the feature map of the i-th pixel extracted, represents a real number, H and W both represent the spatial dimensions of the feature map of the extracted pixels, and C represents the number of channels of the dual-stream convolutional network;
[0020] Calculate the similarity between each of the high-dimensional features:
[0021]
[0022] where S(i, j) represents the similarity between the i-th pixel and the j-th pixel, and ||F i || represents the norm of F i and F j represents the high-dimensional feature of the j-th pixel, and ||F j || represents the norm of F j ;
[0023] Store all the similarities as a four-dimensional tensor, and construct the initial optical flow field of the left and right images according to the four-dimensional tensor:
[0024]
[0025] where CV represents the four-dimensional tensor.
[0026] Optionally, in the real-time plant apical bud localization method based on a two-stream convolutional network, where updating the initial optical flow field to obtain the optical flow field of the left and right images specifically includes:
[0027] Predict the update amount of each optical flow information through the gated recurrent unit in the two-stream convolutional network according to the four-dimensional tensor;
[0028] Predict the optical flow residual according to the initial optical flow field and each update amount:
[0029]
[0030] where ΔFlow k represents the optical flow residual of the k-th update, h t represents the initial optical flow field state at the t-th time step, and ConvNet represents the two-stream convolutional network;
[0031] Add the optical flow residual to the initial optical flow field for updating to obtain the optical flow field of the left and right images:
[0032] Flow k+1 = Flow k + ΔFlow k ;
[0033] where Flow k+1 represents the optical flow field of the (k + 1)-th update, and Flow k represents the initial optical flow field.
[0034] Optionally, in the real-time plant apical bud localization method based on a two-stream convolutional network, the step of extracting features from the left and right images through the two-stream convolutional network to obtain static information of the target plant specifically includes:
[0035] Classify all pixels in the left and right images through the fully connected layer of the two-stream convolutional network to obtain a class probability distribution;
[0036] Extract features from the left and right images according to the class probability distribution to obtain multiple groups of feature maps;
[0037] Activate all the feature maps through the activation function of the two-stream convolutional network, and perform a pooling operation on all the processed feature maps through the pooling layer to obtain the static information of the target plant.
[0038] Optionally, in the real-time plant apical bud localization method based on a two-stream convolutional network, the step of fusing the optical flow field and the static information through the feature fusion layer of the two-stream convolutional network to output the global feature of the target plant specifically includes:
[0039] Extract the complementary information between the optical flow field and the static information through the feature fusion layer of the two-stream convolutional network;
[0040] Fuse the optical flow field and the static information according to the complementary information through a fusion strategy to obtain a fused feature;
[0041] Perform a convolutional process on the fused feature and output the optimized global feature of the target plant.
[0042] Optionally, in the real-time plant apical bud localization method based on a two-stream convolutional network, the step of obtaining the projection matrix of the binocular camera and calculating the three-dimensional spatial coordinates of the target plant according to the projection matrix and the global feature specifically includes:
[0043] Obtain the projection matrix of the binocular camera, and construct a stereo matching map, a depth map, and a disparity map of the left and right images according to the global feature;
[0044] Calculate the disparity of each pixel according to the stereo matching map, the depth map, and the disparity map:
[0045]
[0046] where z represents depth information, f represents the focus of the binocular camera, B represents the baseline length of the two cameras, and d represents the disparity;
[0047] Calculate the two-dimensional coordinates of the target plant in the left and right images, and convert the two-dimensional coordinates into three-dimensional spatial coordinates according to the depth information and the projection matrix.
[0048] In addition, to achieve the above object, the present invention further provides a real-time plant apical bud localization system based on a two-stream convolutional network, wherein the real-time plant apical bud localization system based on a two-stream convolutional network includes:
[0049] A dynamic information acquisition module, configured to acquire the left and right images of the target plant collected by a binocular camera, input the left and right images into a pre-constructed two-stream convolutional network, and perform optical flow calculation on all image frames of the left and right images through an optical flow algorithm to obtain the optical flow field of the left and right images;
[0050] A static information acquisition module, configured to extract features from the left and right images through the two-stream convolutional network to obtain the static information of the target plant;
[0051] A feature fusion module, configured to fuse the optical flow field and the static information through the feature fusion layer of the two-stream convolutional network, and output the global features of the target plant;
[0052] A localization module, configured to obtain the projection matrix of the binocular camera, and calculate the three-dimensional spatial coordinates of the target plant according to the projection matrix and the global features.
[0053] In addition, to achieve the above object, the present invention further provides a terminal, wherein the terminal includes: a memory, a processor, and a real-time plant apical bud localization program based on a two-stream convolutional network stored on the memory and executable on the processor. When the real-time plant apical bud localization program based on a two-stream convolutional network is executed by the processor, the steps of the real-time plant apical bud localization method based on a two-stream convolutional network as described above are implemented.
[0054] In addition, to achieve the above object, the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a real-time plant apical bud localization program based on a two-stream convolutional network. When the real-time plant apical bud localization program based on a two-stream convolutional network is executed by a processor, the steps of the real-time plant apical bud localization method based on a two-stream convolutional network as described above are implemented.
[0055] In the present invention, left and right images of a target plant collected by a binocular camera are obtained. The left and right images are input into a pre-constructed two-stream convolutional network. The optical flow of all image frames of the left and right images is calculated by an optical flow algorithm to obtain the optical flow field of the left and right images. Feature extraction is performed on the left and right images through the two-stream convolutional network to obtain the static information of the target plant. Through the feature fusion layer of the two-stream convolutional network, the optical flow field and the static information are fused to output the global features of the target plant. The projection matrix of the binocular camera is obtained, and based on the projection matrix and the global features, the three-dimensional spatial coordinates of the target plant are calculated. The present invention can improve the accuracy of plant apical bud recognition and the precision of positioning, thereby providing accurate data support for agricultural operations such as precise fertilization and topping of plants. Description of the Drawings
[0056] Figure 1 is a flowchart of a preferred embodiment of the method for real-time positioning of plant apical buds based on a two-stream convolutional network according to the present invention;
[0057] Figure 2 is a specific step flowchart of a preferred embodiment of the method for real-time positioning of plant apical buds based on a two-stream convolutional network according to the present invention;
[0058] Figure 3 is a schematic diagram of the composition of the network time flow and spatial flow of a preferred embodiment of the method for real-time positioning of plant apical buds based on a two-stream convolutional network according to the present invention;
[0059] Figure 4 is a schematic diagram of the composition of the optical flow estimation algorithm of a preferred embodiment of the method for real-time positioning of plant apical buds based on a two-stream convolutional network according to the present invention;
[0060] Figure 5 is a schematic diagram of the composition of the feature extraction network of a preferred embodiment of the method for real-time positioning of plant apical buds based on a two-stream convolutional network according to the present invention;
[0061] Figure 6 is a schematic diagram of the composition of the fully connected layer of a preferred embodiment of the method for real-time positioning of plant apical buds based on a two-stream convolutional network according to the present invention;
[0062] Figure 7 is a schematic diagram of the composition of the feature fusion layer of a preferred embodiment of the method for real-time positioning of plant apical buds based on a two-stream convolutional network according to the present invention;
[0063] Figure 8 is a schematic diagram of the calculation process of depth estimation of a preferred embodiment of the method for real-time positioning of plant apical buds based on a two-stream convolutional network according to the present invention;
[0064] Figure 9 is a structural diagram of a preferred embodiment of the system for real-time positioning of plant apical buds based on a two-stream convolutional network according to the present invention;
[0065] Figure 10 This is a structural diagram of a preferred embodiment of the terminal of the present invention. Detailed implementation manners
[0066] To make the objectives, technical solutions and advantages of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0067] The real-time plant apical bud localization method based on a two-stream convolutional network according to a preferred embodiment of the present invention, as Figure 1 shown, the real-time plant apical bud localization method based on a two-stream convolutional network includes the following steps:
[0068] Step S10: Obtain the left and right images of a target plant collected by a binocular camera, input the left and right images into a pre-constructed two-stream convolutional network, and perform optical flow calculation on all image frames of the left and right images through an optical flow algorithm to obtain the optical flow field of the left and right images.
[0069] Among them, as Figure 2 shown, the RGB images or video frames of cotton apical buds are collected in real time through a binocular camera. The RGB images or video frames can be images within the range of 0 - 1M taken by a 4-million-pixel dual 1080P binocular synchronous camera, and are generated by inputting the RGB images or video frames into the temporal flow channel using an optical flow estimation algorithm; and it is necessary to ensure the synchronization of the binocular camera to ensure the symmetry of the left and right images, that is, to ensure the corresponding relationship of the objects in the images from two perspectives. Then, the RAFT (Recurrent All-Pairs Field Transforms for Optical Flow, a deep learning-based optical flow estimation method) optical flow algorithm is used to perform optical flow calculation on consecutive left and right image frames, calculate the optical flow vector of each pixel point, and this optical flow field represents the motion information of the target between consecutive frames.
[0070] Specifically, obtain the initial left and right images of the target plant collected by the binocular camera, preprocess the initial left and right images to obtain the left and right images; input the left and right images into the pre-constructed two-stream convolutional network. The two-stream convolutional network performs optical flow prediction on all pixels of the left and right images through an optical flow algorithm to obtain the optical flow information of each pixel, and perform matching learning on each piece of optical flow information to obtain a matching result; calculate the similarity between each pixel according to the matching result, and obtain the initial optical flow field of the left and right images according to all the similarities; update the initial optical flow field to obtain the optical flow field of the left and right images, where the optical flow field represents the dynamic information of the target plant.
[0071] Among them, before inputting the image pair (initial left and right images) into the two-stream convolutional network, the image needs to be preprocessed. For each acquired image with a size of W×H (width is W and height is H), the preprocessing of the image includes steps such as grayscale conversion, denoising, and image enhancement to improve the quality of subsequent processing. Then, as Figure 3 shown, when inputting the image pair into the two-stream convolutional network and obtaining the optical flow image through the RAFT optical flow estimation algorithm, RAFT first extracts the depth features of the input image pair through a convolutional neural network (CNN). These features can effectively represent the visual information in the image (such as edges, textures, etc.) and provide rich information for subsequent optical flow estimation. The two-stream convolutional network usually uses a standard convolutional neural network (such as ResNet, Residual Network, VGG, Visual Geometry Group) as the feature extractor. As Figure 4 shown, the full-pixel matching of the RAFT optical flow estimation algorithm considers the optical flow information between individual pixel pairs, calculates the matching between each pair of pixels, and more accurately estimates the optical flow of the entire image by learning the combination of all pixel pairs. The process uses a Field Transforms module to calculate the pixel-level matching between each pair of pixels in the two images. This method does not rely on traditional local window optimization. Among them, the Field Transforms module is a method used in computer vision and deep learning to estimate the optical flow field between images. By calculating the similarity between each pixel pair, a dense optical flow field is generated.
[0072] Furthermore, through the pooling layer of the two-stream convolutional network, the high-dimensional features of each pixel are extracted:
[0073]
[0074] Among them, F i represents the high-dimensional feature of the i-th pixel, CNN represents the two-stream convolutional network, I i represents the feature map of the i-th pixel extracted, represents a real number, H and W both represent the spatial dimensions of the feature map of the extracted pixel, and C represents the number of channels of the two-stream convolutional network; calculate the similarity between each of the high-dimensional features:
[0075]
[0076] Among them, S(i,j) represents the similarity between the i-th pixel and the j-th pixel, ||F i || represents the modulus of F i , Fj represents the high-dimensional feature of the j-th pixel, ||F j || represents the norm of F j ; store all the similarities as a four-dimensional tensor, and construct the initial optical flow field of the left and right images according to the four-dimensional tensor:
[0077]
[0078] where CV represents the four-dimensional tensor.
[0079] where F i and F j are pixel pairs after mutual matching, as Figure 5 shown. For the extraction of dynamic information, it is executed through the time flow branch, and its structure is composed of multiple convolutional layers. Each convolutional layer is followed by a ReLU (Rectified Linear Unit) activation function. The motion information of the optical flow is extracted through the convolutional layer. The first convolutional layer is a layer with a convolutional kernel size of 3×3, stride = 1, and outputs 64 feature maps; the second convolutional layer is a layer with a convolutional kernel size of 3×3, stride = 1, and outputs 128 feature maps; the third convolutional layer is a layer with a convolutional kernel size of 3×3, stride = 1, and outputs 256 feature maps, and so on to obtain the final feature maps after the convolutional layer.
[0080] Furthermore, through the full-pixel matching of the RAFT optical flow estimation algorithm, considering the optical flow information between individual pixel pairs, calculate the matching between each pair of pixels, and learn through the combination of all pixel pairs to more accurately estimate the optical flow of the entire image. The process uses a Field Transforms module to calculate the pixel-level matching between each pair of pixels in the two images. This method does not rely on traditional local window optimization. The Field Transforms module is a method used in computer vision and deep learning to estimate the optical flow field between images. By calculating the similarity between each pair of pixels, a dense optical flow field is generated.
[0081] Furthermore, according to the four-dimensional tensor, predict the update amount of each optical flow information through the gated recurrent unit in the two-stream convolutional network; according to the initial optical flow field and each update amount, predict the optical flow residual:
[0082]
[0083] where ΔFlow k represents the optical flow residual of the k-th update, h tdenotes the initial optical flow field state at the \(t\)-th time step, and ConvNet denotes the two-stream convolutional network; the optical flow residual is added to the initial optical flow field for updating to obtain the optical flow field of the left and right images:
[0084] Flow k+1 = Flow k + ΔFlow k ;
[0085] where Flow k+1 denotes the optical flow field updated at the \((k + 1)\)-th time, and Flow l denotes the initial optical flow field.
[0086] Among them, the recurrent network of the RAFT optical flow estimation algorithm (based on GRU, Gated Recurrent Unit, a kind of recurrent neural network) is used to gradually optimize the optical flow prediction (iterative refinement). In each iteration, the current optical flow field is used to sample the relevant volume, and then the GRU (gated recurrent unit) is used to predict the update amount of the optical flow by combining this information. At each time step, the GRU updates the optical flow estimation according to the current features and optical flow field, gradually approaching the correct optical flow result. This recursive structure can effectively iterate and update the optical flow estimation, and continuously refine the optical flow result during the training process. Among them, for the optical flow field update of the RAFT optical flow estimation algorithm, in each recursive iteration, RAFT uses the optical flow field update mechanism to gradually optimize the optical flow of each pixel according to the result of full-pixel matching and the currently estimated optical flow field, effectively avoiding the problem of local minimum and making the optical flow estimation more accurate.
[0087] Furthermore, the network structure for performing the above feature extraction specifically includes a first convolutional layer (Conv1). The first convolutional layer is a convolutional layer with a convolutional kernel size of 3×3, stride = 1, input channel number of 3, and output channel number of 64; the first convolutional layer (Conv1) is sequentially connected to a first ReLU activation function layer (non-linear), a first pooling layer (Pool1), and a second convolutional layer (Conv2). The first pooling layer (Pool1) performs max pooling of 2x2 and outputs 112x112x64; the second convolutional layer is a convolutional layer with a convolutional kernel size of 3×3, stride = 1, input channel number of 64, and output channel number of 128; the second convolutional layer (Conv2) is sequentially connected to a second ReLU activation function layer (non-linear) and a second pooling layer (Pool2). The second pooling layer (Pool2) performs max pooling of 2x2 and outputs 56x56x128; the second pooling layer (Pool2) is sequentially connected to a fully connected layer (Full) and a Softmax output layer.
[0088] Step S20: Extract features from the left and right images through the two-stream convolutional network to obtain the static information of the target plant.
[0089] Among them, the extraction of static information is mainly performed through the spatial stream branch. The structure of the spatial stream branch consists of multiple convolutional layers. After each convolutional layer, a ReLU activation function is connected. The role of the convolutional layer is to extract low-level features (such as edges, textures, etc.) in the image, and at the same time extract more advanced image information as the number of layers increases. The first convolutional layer is a layer with a convolutional kernel size of 3×3, stride = 1, and outputs 64 feature maps; the second convolutional layer is a layer with a convolutional kernel size of 3×3, stride = 1, and outputs 128 feature maps; the third convolutional layer is a layer with a convolutional kernel size of 3×3, stride = 1, and outputs 256 feature maps, and so on to obtain the final feature maps after passing through the convolutional layers.
[0090] Specifically, through the fully connected layer of the two-stream convolutional network, classify all pixels in the left and right images to obtain a class probability distribution; according to the class probability distribution, extract the features of the left and right images to obtain multiple groups of feature maps; through the activation function of the two-stream convolutional network, perform activation processing on all the feature maps, and perform pooling operations on all the processed feature maps through the pooling layer to obtain the static information of the target plant.
[0091] Among them, as Figure 6 shown, the fully connected layer (Full) flattens and maps the high-dimensional features extracted by the convolutional and pooling layers to the output space; the class probability distribution finally output by the Softmax output layer is used for classification tasks. The convolutional layer mainly extracts features in the image. The output of each convolutional layer will become a group of feature maps. The ReLU activation function layer (non-linear) performs activation processing on the output of each convolutional layer, introducing non-linearity. The pooling layer performs pooling operations on the output of the convolutional layer to reduce the spatial size of the feature maps while retaining important information.
[0092] Specifically, the input layer (Input X) module is used to input the vector x, which usually comes from the output of the previous layer; the weight matrix (Weights W) module has a corresponding weight connection for each input and output. The dimension of the weight matrix is usually n×m, where n is the number of input neurons and m is the number of output neurons; the weighted sum (Weighted Sum) module is used to calculate the weighted sum of each neuron, that is, to perform a dot product of the input and the weights and add the bias term; the bias (Bias b) module has a bias term for each output neuron, which is used to adjust the output of the neuron. The role of the bias is to make the output of the activation function more flexible; the ReLU activation function module is used to convert the weighted sum into an output, determine whether the neuron is activated, and provide a non-linear transformation for the network; the output layer (Output Y) module is used to output the vector y, that is, the result obtained through the calculation of the fully connected layer, and the dimension of the output is m, corresponding to the number of neurons in the output layer.
[0093] Further, the basic working principle of the fully connected layer described in this embodiment is: given the input vector x = (x1, x2, x3,..., x n ), the fully connected layer outputs a new vector y through weighted summation, adding bias, and passing through the ReLU activation function. The weighted summation is to multiply each input neuron by the corresponding weight and then add the bias term. The formula is as follows:
[0094]
[0095] Among them, Z i represents the i-th weighted summation term, x i represents the i-th input neuron, w i represents the i-th weight, and b represents the bias term.
[0096] The ReLU activation function processes the weighted sum through the activation function f(z) to obtain the output y i :
[0097] yi = f(Z i ).
[0098] Step S30: Through the feature fusion layer of the two-stream convolutional network, fuse the optical flow field and the static information, and output the global features of the target plant.
[0099] Among them, as Figure 7As shown, through the Feature Fusion Layer module, the output feature maps from two input ports (i.e., the optical flow field and the feature map of static information) are received and combined together through a fusion strategy, enabling the network to utilize complementary information from different inputs.
[0100] Specifically, through the feature fusion layer of the two-stream convolutional network, complementary information between the optical flow field and the static information is extracted; according to the complementary information, the optical flow field and the static information are fused through a fusion strategy to obtain fused features; the fused features are subjected to convolutional processing to output the global features of the optimized target plant.
[0101] Among them, through the feature extraction layer (Conv+ReLU) module, the fused feature map is further processed through a convolutional layer and an activation function to optimize the fusion result and reduce possible dimensional mismatches; through the network output layer, the fused feature map is subjected to final processing (fully connected layer, Softmax, etc.) to generate the final output (regression value) of the network; by separately processing the spatial information and temporal information of the image, the accuracy of target localization and tracking can be effectively improved. By combining optical flow estimation with TS-CNN, the network can not only extract spatial features in static images but also effectively capture the motion information of the target, thereby providing more accurate three-dimensional localization in a dynamic environment.
[0102] Step S40: Obtain the projection matrix of the binocular camera, and calculate the three-dimensional spatial coordinates of the target plant according to the projection matrix and the global features.
[0103] Specifically, obtain the projection matrix of the binocular camera, and construct a stereo matching map, a depth map, and a disparity map of the left and right images according to the global features; calculate the disparity of each pixel according to the stereo matching map, the depth map, and the disparity map:
[0104]
[0105] Among them, z represents depth information, f represents the focus of the binocular camera, B represents the maximum distance between the two cameras, d represents the disparity; calculate the two-dimensional coordinates of the target plant in the left and right images, and convert the two-dimensional coordinates into three-dimensional spatial coordinates according to the depth information and the projection matrix.
[0106] Among them, the two-dimensional coordinates are converted into three-dimensional spatial coordinates through the projection matrix of the camera:
[0107]
[0108] Wherein, X represents the abscissa of the three-dimensional spatial coordinates, Y represents the ordinate of the three-dimensional spatial coordinates, Fx and Fy respectively represent the focal lengths of the camera in the horizontal x and vertical y directions, and Cx and Cy respectively represent the principal point coordinates of the left and right images (for example, the center of the image).
[0109] As Figure 8 shown, the present invention synchronously acquires the left and right images of the target through a binocular camera system, and uses an improved two-stream convolutional neural network module to extract the RGB images or video frames captured by the binocular camera, so as to obtain the results after initial spatial and temporal feature extraction. The temporal and spatial feature extraction results of the RGB images or video frames are subjected to feature fusion to obtain the results of the spatial and temporal stream feature fusion of the RGB images or video frames. The feature fusion results of the RGB images or video frames are further subjected to disparity map and depth calculation to obtain the specific three-dimensional coordinates of the features, and the three-dimensional coordinate positions of the target object can be detected in real time, improving the accuracy of plant apical bud recognition and the precision of positioning, thereby providing accurate data support for agricultural operations such as precise fertilization and topping of plants.
[0110] Furthermore, as Figure 9 shown, based on the above-mentioned method for real-time positioning of plant apical buds based on a two-stream convolutional network, the present invention also correspondingly provides a system for real-time positioning of plant apical buds based on a two-stream convolutional network, wherein the system for real-time positioning of plant apical buds based on a two-stream convolutional network includes:
[0111] A dynamic information acquisition module 51, configured to acquire the left and right images of the target plant collected by the binocular camera, input the left and right images into the constructed two-stream convolutional network, and perform optical flow calculation on all image frames of the left and right images through an optical flow algorithm to obtain the optical flow field of the left and right images;
[0112] A static information acquisition module 52, configured to extract features from the left and right images through the two-stream convolutional network to obtain the static information of the target plant;
[0113] A feature fusion module 53, configured to fuse and process the optical flow field and the static information through the feature fusion layer of the two-stream convolutional network, and output the global features of the target plant;
[0114] A positioning module 54, configured to obtain the projection matrix of the binocular camera, and calculate the three-dimensional spatial coordinates of the target plant according to the projection matrix and the global features.
[0115] Furthermore, as Figure 10 shown, based on the above-mentioned method and system for real-time positioning of plant apical buds based on a two-stream convolutional network, the present invention also correspondingly provides a terminal, and the terminal includes a processor 10, a memory 20, and a display 30. Figure 10Only some components of the terminal are shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively.
[0116] The memory 20 may be an internal storage unit of the terminal in some embodiments, such as a hard disk or memory of the terminal. The memory 20 may also be an external storage device of the terminal in other embodiments, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the terminal. Further, the memory 20 may also include both the internal storage unit and the external storage device of the terminal. The memory 20 is used to store application software installed on the terminal and various types of data, such as program codes for installing the terminal. The memory 20 may also be used to temporarily store data that has been output or will be output. In one embodiment, a plant apical bud real-time positioning program 40 based on a two-stream convolutional network is stored on the memory 20, and the plant apical bud real-time positioning program 40 based on the two-stream convolutional network can be executed by the processor 10, so as to implement the plant apical bud real-time positioning method based on the two-stream convolutional network in this application.
[0117] The processor 10 may be a central processing unit (CPU), a microprocessor or other data processing chips in some embodiments, and is used to run the program codes stored in the memory 20 or process data, such as executing the plant apical bud real-time positioning method based on the two-stream convolutional network, etc.
[0118] The display 30 may be an LED display, a liquid crystal display, a touch liquid crystal display, an OLED (Organic Light-Emitting Diode) toucher, etc. in some embodiments. The display 30 is used to display information on the terminal and to display a visual user interface. Components of the terminal communicate with each other through a system bus.
[0119] In one embodiment, when the processor 10 executes the plant apical bud real-time positioning program 40 in the memory 20, the following steps are implemented:
[0120] Obtain the left and right images of the target plant collected by the binocular camera, input the left and right images into the constructed two-stream convolutional network, perform optical flow calculation on all image frames of the left and right images through an optical flow algorithm, and obtain the optical flow field of the left and right images;
[0121] Extract features from the left and right images through the two-stream convolutional network to obtain the static information of the target plant;
[0122] Through the feature fusion layer of the two-stream convolutional network, fuse the optical flow field and the static information, and output the global features of the target plant;
[0123] Obtain the projection matrix of the binocular camera, and calculate the three-dimensional spatial coordinates of the target plant according to the projection matrix and the global features.
[0124] Among them, the steps of acquiring the left and right images of the target plant collected by the binocular camera, inputting the left and right images into the constructed two-stream convolutional network, and calculating the optical flow of all image frames of the left and right images through the optical flow algorithm to obtain the optical flow field of the left and right images specifically include:
[0125] Acquire the initial left and right images of the target plant collected by the binocular camera, and preprocess the initial left and right images to obtain the left and right images;
[0126] Input the left and right images into the constructed two-stream convolutional network. The two-stream convolutional network predicts the optical flow of all pixels of the left and right images through the optical flow algorithm, obtains the optical flow information of each pixel, and performs matching learning on each optical flow information to obtain a matching result;
[0127] Calculate the similarity between each pixel according to the matching result, and obtain the initial optical flow field of the left and right images according to all the similarities;
[0128] Update the initial optical flow field to obtain the optical flow field of the left and right images, where the optical flow field represents the dynamic information of the target plant.
[0129] Among them, the steps of calculating the similarity between each pixel according to the matching result and obtaining the initial optical flow field of the left and right images according to all the similarities specifically include:
[0130] Extract the high-dimensional features of each pixel through the pooling layer of the two-stream convolutional network:
[0131]
[0132] where F i represents the high-dimensional feature of the i-th pixel, CNN represents the two-stream convolutional network, and I i represents the feature map of the i-th pixel extracted, represents a real number, H and W both represent the spatial dimensions of the feature map of the extracted pixel, and C represents the number of channels of the two-stream convolutional network;
[0133] Calculate the similarity between each of the high-dimensional features:
[0134]
[0135] where S(i,j) represents the similarity between the i-th pixel and the j-th pixel, ||F i || represents the modulus of F i and F j represents the high-dimensional feature of the j-th pixel, ||F j || represents the modulus of F j ;
[0136] Store all the similarities as a four-dimensional tensor, and construct the initial optical flow field of the left and right images according to the four-dimensional tensor:
[0137]
[0138] where CV represents the four-dimensional tensor.
[0139] Among them, updating the initial optical flow field to obtain the optical flow field of the left and right images specifically includes:
[0140] Predict the update amount of each optical flow information according to the four-dimensional tensor through the gated recurrent unit in the two-stream convolutional network;
[0141] Predict the optical flow residual according to the initial optical flow field and each update amount:
[0142]
[0143] where ΔFlow k represents the optical flow residual of the k-th update, h t represents the state of the initial optical flow field at the t-th time step, and ConvNet represents the two-stream convolutional network;
[0144] Add the optical flow residual to the initial optical flow field for updating to obtain the optical flow field of the left and right images:
[0145] Flow k+1 = Flow k + ΔFlow k ;
[0146] where Flow k+1 represents the optical flow field of the (k + 1)-th update, and Flow k represents the initial optical flow field.
[0147] Among them, extracting the features of the left and right images through the two-stream convolutional network to obtain the static information of the target plant specifically includes:
[0148] Classify all pixels in the left and right images through the fully connected layer of the two-stream convolutional network to obtain a class probability distribution;
[0149] Extract features of the left and right images according to the class probability distribution to obtain multiple groups of feature maps;
[0150] Activate all the feature maps through the activation function of the two-stream convolutional network, and perform pooling operations on all the processed feature maps through the pooling layer to obtain the static information of the target plant.
[0151] Among them, the feature fusion layer of the two-stream convolutional network fuses the optical flow field and the static information and outputs the global feature of the target plant, specifically including:
[0152] Extract the complementary information between the optical flow field and the static information through the feature fusion layer of the two-stream convolutional network;
[0153] According to the complementary information, fuse the optical flow field and the static information through a fusion strategy to obtain a fusion feature;
[0154] Perform convolutional processing on the fusion feature and output the optimized global feature of the target plant.
[0155] Among them, obtaining the projection matrix of the binocular camera, and calculating the three-dimensional spatial coordinates of the target plant according to the projection matrix and the global feature, specifically including:
[0156] Obtain the projection matrix of the binocular camera, and construct a stereo matching map, a depth map, and a disparity map of the left and right images according to the global feature;
[0157] Calculate the disparity of each pixel according to the stereo matching map, the depth map, and the disparity map:
[0158]
[0159] Where z represents depth information, f represents the focus of the binocular camera, B represents the maximum length of the two cameras, and d represents the disparity;
[0160] Calculate the two-dimensional coordinates of the target plant in the left and right images, and convert the two-dimensional coordinates into three-dimensional spatial coordinates according to the depth information and the projection matrix.
[0161] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a real-time plant apical bud localization program based on a two-stream convolutional network, and when the real-time plant apical bud localization program based on the two-stream convolutional network is executed by a processor, the steps of the real-time plant apical bud localization method based on the two-stream convolutional network as described above are implemented.
[0162] In summary, the present invention provides a real-time plant apical bud localization method and related devices based on a two-stream convolutional network. The method includes: acquiring left and right images of a target plant collected by a binocular camera, inputting the left and right images into a pre-constructed two-stream convolutional network, performing optical flow calculation on all image frames of the left and right images through an optical flow algorithm to obtain the optical flow field of the left and right images; extracting features from the left and right images through the two-stream convolutional network to obtain static information of the target plant; fusing the optical flow field and the static information through a feature fusion layer of the two-stream convolutional network to output global features of the target plant; acquiring a projection matrix of the binocular camera, and calculating spatial three-dimensional coordinates of the target plant according to the projection matrix and the global features. The present invention can improve the accuracy of plant apical bud recognition and the precision of localization, thereby providing accurate data support for agricultural operations such as precise fertilization and topping of plants.
[0163] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or terminal. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or terminal including the element.
[0164] Of course, those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program. The program can be stored in a computer-readable storage medium readable by a computer, and when the program is executed, it can include the processes of the above method embodiments. The computer-readable storage medium can be a memory, a magnetic disk, an optical disc, etc.
[0165] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description, and all such improvements and transformations should fall within the protection scope of the appended claims of the present invention.
Claims
1. A real-time plant apical bud localization method based on a two-stream convolutional network, characterized in that, The real-time plant apical bud localization method based on a two-stream convolutional network includes: Obtain the left and right images of the target plant collected by a binocular camera, input the left and right images into the constructed two-stream convolutional network, perform optical flow calculation on all image frames of the left and right images through an optical flow algorithm, and obtain the optical flow field of the left and right images; Extract features from the left and right images through the two-stream convolutional network to obtain the static information of the target plant; Through the feature fusion layer of the two-stream convolutional network, fuse the optical flow field and the static information, and output the global features of the target plant; Obtain the projection matrix of the binocular camera, and calculate the three-dimensional spatial coordinates of the target plant according to the projection matrix and the global features.
2. The real-time plant apical bud localization method based on a two-stream convolutional network according to claim 1, wherein The step of obtaining the left and right images of the target plant collected by the binocular camera, inputting the left and right images into the constructed two-stream convolutional network, and performing optical flow calculation on all image frames of the left and right images through an optical flow algorithm to obtain the optical flow field of the left and right images specifically includes: Obtain the initial left and right images of the target plant collected by the binocular camera, preprocess the initial left and right images to obtain the left and right images; Input the left and right images into the constructed two-stream convolutional network. The two-stream convolutional network performs optical flow prediction on all pixels of the left and right images through an optical flow algorithm, obtains the optical flow information of each pixel, and performs matching learning on each optical flow information to obtain a matching result; Calculate the similarity between each pixel according to the matching result, and obtain the initial optical flow field of the left and right images according to all the similarities; Update the initial optical flow field to obtain the optical flow field of the left and right images, where the optical flow field represents the dynamic information of the target plant.
3. The real-time plant apical bud localization method based on a two-stream convolutional network according to claim 2, characterized in that The step of calculating the similarity between each pixel according to the matching result and obtaining the initial optical flow field of the left and right images according to all the similarities specifically includes: Extract the high-dimensional features of each pixel through the pooling layer of the two-stream convolutional network: Among them, F i represents the high-dimensional feature of the i-th pixel, CNN represents the two-stream convolutional network, and I i represents the feature map of the i-th pixel extracted, represents a real number, both H and W represent the spatial dimensions of the feature map of the extracted pixel, and C represents the number of channels of the two-stream convolutional network; Calculate the similarity between each high-dimensional feature: Among them, S(i,j) represents the similarity between the i-th pixel and the j-th pixel, ||F i || represents the modulus of F i , F j represents the high-dimensional feature of the j-th pixel, ||F j || represents the modulus of F j ; Store all the similarities as a four-dimensional tensor, and construct the initial optical flow field of the left and right images according to the four-dimensional tensor: Where CV represents the four-dimensional tensor.
4. The real-time plant apical bud localization method based on a two-stream convolutional network according to claim 3, characterized in that The step of updating the initial optical flow field to obtain the optical flow field of the left and right images specifically includes: According to the four-dimensional tensor, predict the update amount of each optical flow information through the gated recurrent unit in the two-stream convolutional network; Predict the optical flow residual according to the initial optical flow field and each update amount; where, ΔFlow k represents the optical flow residual of the k-th update, and h t represents the initial optical flow field state at the t-th time step, and ConvNet represents the two-stream convolutional network; Add the optical flow residual to the initial optical flow field for updating to obtain the optical flow field of the left and right images: Flow k+1 = Flow k + ΔFlow k ; Among them, Flow k+1 represents the optical flow field of the (k + 1)-th update, and Flow k represents the initial optical flow field.
5. The real-time plant apical bud localization method based on a two-stream convolutional network according to claim 1, characterized in that The step of extracting features from the left and right images through the two-stream convolutional network to obtain the static information of the target plant specifically includes: Classify all pixels in the left and right images through the fully connected layer of the two-stream convolutional network to obtain a class probability distribution; Extract the features of the left and right images according to the class probability distribution to obtain multiple groups of feature maps; Through the activation function of the two-stream convolutional network, all the feature maps are activated, and through the pooling layer, pooling operations are performed on all the processed feature maps to obtain the static information of the target plant.
6. The real-time plant apical bud localization method based on a two-stream convolutional network according to claim 1, characterized in that, The fusion process of the optical flow field and the static information through the feature fusion layer of the two-stream convolutional network outputs the global features of the target plant, specifically including: Through the feature fusion layer of the two-stream convolutional network, complementary information between the optical flow field and the static information is extracted; According to the complementary information, the optical flow field and the static information are fused through a fusion strategy to obtain fused features; Convolution processing is performed on the fused features to output the optimized global features of the target plant.
7. The real-time plant apical bud localization method based on a two-stream convolutional network according to claim 2, wherein The projection matrix of the binocular camera is obtained, and according to the projection matrix and the global features, the three-dimensional spatial coordinates of the target plant are calculated, specifically including: The projection matrix of the binocular camera is obtained, and according to the global features, a stereo matching map, a depth map, and a disparity map of the left and right images are constructed; According to the stereo matching map, the depth map, and the disparity map, the disparity of each pixel is calculated: where z represents depth information, f represents the focus of the binocular camera, B represents the maximum distance between the two cameras, and d represents disparity; The two-dimensional coordinates of the target plant in the left and right images are calculated, and according to the depth information and the projection matrix, the two-dimensional coordinates are converted into three-dimensional spatial coordinates.
8. A real-time plant apical bud localization system based on a two-stream convolutional network, characterized in that, The real-time plant apical bud localization system based on a two-stream convolutional network includes: A dynamic information acquisition module, configured to acquire the left and right images of the target plant collected by the binocular camera, input the left and right images into the constructed two-stream convolutional network, and perform optical flow calculation on all the image frames of the left and right images through the optical flow algorithm to obtain the optical flow field of the left and right images; A static information acquisition module, configured to extract features from the left and right images through the two-stream convolutional network to obtain the static information of the target plant; A feature fusion module, configured to fuse the optical flow field and the static information through the feature fusion layer of the two-stream convolutional network to output the global features of the target plant; A localization module, configured to obtain the projection matrix of the binocular camera, and calculate the three-dimensional spatial coordinates of the target plant according to the projection matrix and the global features.
9. A terminal, characterized in that, The terminal includes: a memory, a processor, and a real-time plant apical bud localization program based on a two-stream convolutional network stored on the memory and executable on the processor. When the real-time plant apical bud localization program based on the two-stream convolutional network is executed by the processor, the steps of the real-time plant apical bud localization method based on the two-stream convolutional network according to any one of claims 1-7 are implemented.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a real-time plant apical bud localization program based on a two-stream convolutional network. When the real-time plant apical bud localization program based on the two-stream convolutional network is executed by a processor, the steps of the real-time plant apical bud localization method based on the two-stream convolutional network according to any one of claims 1-7 are implemented.