Fruit tree field ridge identification method and electronic device
By combining the time-space visual attention mechanism and the semantic label model of the fully convolutional neural network, and the deep semantic perception model of the convolutional neural network, the accuracy and efficiency problems of fruit tree and field ridge recognition in the existing technology are solved, and real-time recognition effect with high accuracy is achieved.
Patent Information
- Application Number
- CN202210711029.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-22
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-06-22
AI Technical Summary
The existing fruit tree and field ridge recognition technology has problems such as overfitting, reduced computing efficiency and limitations in core function use, making it difficult to achieve real-time recognition with high accuracy.
The semantic label model based on the time-space visual attention mechanism and a fully convolutional neural network is adopted, combined with the deep semantic perception model of the convolutional neural network, and the fruit tree and field ridge areas are identified through a classifier, and the conditional random field model and motion optical flow are used for online contour reasoning and target bounding box repositioning. Finally, the kernel-related filtering algorithm is used to track and update the target of interest.
It significantly improves the real-time identification accuracy of fruit trees and field ridges, and can identify fruit trees and field ridges 100% when the equipment is running normally, reducing application costs.
Smart Images

Figure CN115063724B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the field of image recognition, and in particular to a method and electronic equipment for recognizing fruit tree field ridges. Background Art
[0002] Image recognition technology is a key module in harvesting robots, and the development of this field plays an indispensable role in the research of robots. In recent years, a large number of researchers at home and abroad have devoted themselves to research in this field, which has led to the rapid development of the field of image recognition. There are also many existing studies on image segmentation algorithms. For example, BAI et al. adopted a rice canopy segmentation method based on SVM classifier, and used the automatic learning characteristics of convolutional neural network to reduce the misclassification rate. This type of algorithm has a good segmentation effect, but the learning of the segmentation model depends on a large number of sample annotations, which has high requirements for computer hardware, so the application cost is high.
[0003] There are also many studies on the recognition and perception of fruit trees and tea ridges in wild orchards and tea gardens, but the existing recognition and perception technologies for fruit trees and tea ridges generally use deep learning technology or support vector machine (SVM) technology. Although support vector machines have advantages in solving nonlinear model recognition, they still have some disadvantages in practical applications. For example, as the training sample set gradually increases, the support vectors of the support vector machine (SVM) will also increase significantly. When a certain limit is exceeded, it may cause overfitting and reduced computational efficiency. In addition, the core function of SVM is relatively restrictive to use and must meet certain conditions. Summary of the invention
[0004] The purpose of the present invention is to provide a method and electronic equipment for recognizing fruit tree ridges, which is not limited by the core function of SVM and greatly improves the accuracy of real-time recognition of interesting targets in fruit tree ridges.
[0005] The technical solution to achieve the purpose of the present invention is:
[0006] A method for identifying fruit tree field ridges comprises the following steps:
[0007] Get the video sequence captured by the camera;
[0008] The semantic labels of fruit trees and field ridges are generated online through a semantic labeling model based on temporal and spatial visual attention mechanism and fully convolutional neural network.
[0009] The fruit tree and field ridge depth semantic perception model based on convolutional neural network is used for fusion semantic perception;
[0010] Extract the feature values of fruit trees and ridges, and identify the fruit trees and ridges through classifiers;
[0011] By estimating the optical flow of the video frames, online contour inference and target bounding box relocation are performed for the fruit trees and ridges of interest based on the conditional random field model and motion optical flow;
[0012] The targets of interest are tracked based on the kernel correlation filtering algorithm, and the depth semantic perception models of fruit trees and field ridges are updated.
[0013] Furthermore, the video sequence captured by the camera specifically includes:
[0014] Step 1.1: Move the mobile robot equipped with a camera in the orchard to take pictures of fruit trees and field ridges;
[0015] Step 1.2: Obtain a video sequence that captures the output of the target information of interest.
[0016] Furthermore, the semantic label model is obtained through offline training, which specifically includes:
[0017] Step 2.1: Based on the image dataset containing two types of semantic labels, fruit trees and ridges, train the fully convolutional neural network offline.
[0018] Step 2.2: Connect the gated recurrent unit after the fully convolutional neural network to capture the temporal information of the video, improve the GRU to a convolutional GRU layer, improve the efficiency and performance of the algorithm, and obtain the semantic label models of the fruit trees and ridges in the image;
[0019] Step 2.3: In the semantic segmentation process of the semantic label model, a temporal and spatial selective attention mechanism is introduced to collect two adjacent frames of the video sequence, and the corresponding semantic labels are generated online through the semantic label model.
[0020] Furthermore, the method for acquiring the image data set is:
[0021] After acquiring the video image sequence, the multi-video sequence is detected frame by frame, and each frame of the acquired image is converted into grayscale, a digital grayscale image mathematical model is established, and an image with enhanced grayscale value is obtained;
[0022] Performing the first filtering, the second filtering and the noise reduction processing on the image after the gray value enhancement;
[0023] Detect the image frame by frame. When a fruit tree or a ridge is detected to appear suddenly in the image, the frame is updated to the initial frame. The fruit tree or ridge in the image is the target of interest, and the target area of interest is locked.
[0024] Obtain multiple sets of images containing objects of interest as image datasets.
[0025] Furthermore, the deep semantic perception model is obtained through offline training, including:
[0026] Get the i-th frame of the video sequence image;
[0027] Obtain target tracking confidence map based on Gaussian perturbation model;
[0028] The semantic labels of the generated fruit trees and ridges are semantically selected and semantically filtered based on the kernelized correlation filter to obtain the semantic dense confidence map of the objects of interest.
[0029] The target tracking confidence map and the semantic dense confidence map are used as inputs of the deep perception network, and the deep perception network is trained offline to generate parameters of the deep perception network;
[0030] A multi-scale recurrent convolutional network is used to deeply fuse spatiotemporal features at multiple levels, and a gated recurrent network is used as the recurrent unit to generate a deep semantic perception model.
[0031] Furthermore, the extracting of the feature values of the fruit trees and the ridges and identifying the fruit trees and the ridges through the classifier specifically includes:
[0032] Step 4.1: Obtain an image containing fruit trees and field ridges;
[0033] Step 4.2: The images of fruit trees and ridges are subjected to denoising through a denoising network;
[0034] Step 4.3: Extract feature values through the deep residual shrinkage network. The fully connected output layer of the deep residual shrinkage network is a classifier to classify and identify fruit trees and ridges.
[0035] Furthermore, the online contour reasoning and target bounding box relocation of the fruit trees and field ridges of interest based on the conditional random field model and motion optical flow specifically include:
[0036] Step 5.1: Obtain a color image in a certain frame of the video, and obtain the image pixel intensity and feature map through pixel enhancement processing;
[0037] Step 5.2: Based on the deep semantic perception model, obtain the semantic perception confidence map of the target of interest in a certain frame of the video;
[0038] Step 5.3: Obtain an inter-frame optical flow motion estimation map of the target of interest based on the video frames;
[0039] Step 5.4: Based on the random conditional field model constructed offline, the image pixel intensity and feature map, semantic perception confidence map and optical flow motion estimation map are used as inputs for non-sub-module target contour reasoning of the conditional random field model to obtain the target contour and locate the target bounding box.
[0040] Furthermore, the tracking of the target of interest based on the kernel correlation filtering algorithm and updating of the fruit tree and field ridge depth semantic perception model specifically include:
[0041] Step 6.1: Based on the contour reasoning of the object of interest and the positioning of the object bounding box, update the Gaussian perturbation model of the object of interest;
[0042] Step 6.2: Obtain the target tracking confidence map based on the Gaussian perturbation of the kernelized correlation filter, and update the depth semantic perception model of fruit trees and ridges.
[0043] Furthermore, the fully convolutional neural network adopts AlexNet, VGG or GoogleNet network architecture.
[0044] An electronic device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the method for identifying fruit tree field ridges is implemented when the processor executes the program.
[0045] Compared with the prior art, the present invention has the following significant effects: the offline method of the time and space attention mechanism based on the video sequence combined with the online method of neural network training proposed by the present invention can accurately segment the region of interest and the non-region of interest in the image; the present invention accurately obtains the semantic label of the region of interest through the training of the neural network, and then generates a depth perception model of the target of interest based on the generated semantic label and the training based on the depth perception semantic network, and finally performs the contour reasoning of the target and the positioning of the target boundary box; the present invention combines the reduction of the region of interest and the training of the offline semantic model based on the neural network to greatly improve the accuracy of real-time recognition of the target of interest in the fruit tree field ridge, and can 100% recognize the fruit tree and field ridge area when the equipment is operating normally. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 This is a block diagram of the fruit tree tea ridge identification module provided by the present invention.
[0047] Figure 2 It is a schematic diagram of semantic generation of interesting objects in online videos.
[0048] Figure 3 Schematic diagram of the deep residual shrinkage module and classifier unit.
[0049] Figure 4 It is a schematic diagram of the joint recognition of the contour of the target of interest by integrating semantic perception and motion optical flow.
[0050] Figure 5 This is a picture of the robot working in the orchard. DETAILED DESCRIPTION
[0051] In order to better understand the steps, advantages and implementation process of the present invention, the present invention is further described below in conjunction with the accompanying drawings.
[0052] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited to the specific embodiments disclosed below.
[0053] The present invention relates to the generation of online video semantic labels of interested targets of fruit trees and tea ridges in a fruit and tea garden, the construction of a target depth semantic perception model, the contour reasoning of the interested target based on semantic perception and motion optical flow, and the positioning of the target boundary box. The present invention realizes a method for identifying fruit trees and field ridges in a wild orchard by combining a visual attention mechanism with an online semantic generation module of interested targets, a deep semantic perception model module of interested targets, and an interested target contour recognition and target frame positioning module.
[0054] like Figure 5 , now taking the fruit tea picking robot as an example, a camera is installed on the robot's head so that it can detect the environment around the robot in real time, and identify and detect fruit trees, tea ridges, etc. based on the camera. The specific steps are as follows Figure 1 As shown:
[0055] Step 1: Get the video sequence captured by the camera and go to step 2;
[0056] Step 1.1: Obtain a video sequence captured by the robot walking along the ridges in the orchard;
[0057] Step 2: Generate semantic labels of the objects of interest in the fruit tree field online based on the temporal and spatial visual attention mechanism and the fully convolutional neural network, and then proceed to step 3;
[0058] Step 2.1: After acquiring the video image sequence in the first step, the multi-video sequence is detected frame by frame, and each frame of the acquired image is converted into grayscale, a digital grayscale image mathematical model is established, and an image with enhanced grayscale value is obtained;
[0059] Step 2.2: Perform the first filtering and the second filtering as well as the noise reduction process on the enhanced grayscale image;
[0060] Table 1 Image background filtering processing
[0061]
[0062] Step 2.3: Based on the above steps, the image is detected frame by frame. When a fruit tree or a ridge is detected to appear suddenly in the image, the frame is updated to the initial frame. The fruit tree or ridge in the image is the target of interest, and the target area of interest is locked;
[0063] Step 2.4: Obtain multiple sets of images containing objects of interest (fruit trees, field ridges) as training samples and test samples;
[0064] Step 2.5: Offline training: Train the training samples based on the fully convolutional neural network to obtain the semantic model of the image region of interest;
[0065] Step 2.6: Online training: Combine the gated recurrent unit (GRU) and the fully convolutional neural network for forward propagation to capture the temporal information of the video and generate online semantic labels;
[0066] The fully convolutional neural network can select commonly used network architectures such as AlexNet, VGG, GoogleNet, etc.
[0067] Step 3: Combine Figure 2 , Construction and training of fruit tree and field ridge depth semantic perception model based on convolutional neural network;
[0068] Step 3.1: Get the i-th frame image containing the target of interest (fruit tree, field ridge);
[0069] Step 3.2: Online method: The image is processed in step 2 to obtain the semantic label of the object of interest, and the semantic dense confidence map of the object of interest is obtained after semantic selection and semantic filtering based on the kernelized correlation filter.
[0070] Step 3.3: Offline method: Obtain a dense target tracking confidence map based on the Gaussian perturbation model for the i-th frame image;
[0071] Step 3.4: Use the semantic confidence maps in step 3.2 and step 3.3 as inputs of the deep perception network, train the deep perception network and generate the parameters of the perception network;
[0072] Step 3.5: Use a multi-scale recurrent convolutional network (RCN) to deeply fuse spatiotemporal features at multiple levels, and use a gated recurrent network (GRU) as a recurrent unit to quickly capture the temporal features of the video at each spatial resolution and generate a semantic perception model of the target of interest;
[0073] Step 4: Extraction of feature values of fruit tree and tea ridges, using classifiers to identify fruit trees and ridge areas, such as Figure 3 As shown;
[0074] Step 4.1: Obtain an image containing fruit trees and field ridges;
[0075] Step 4.2: The image is denoised through a denoising network;
[0076] Step 4.3: Extract feature values through the currently improved deep residual shrinkage network (DRSN), and the fully connected output layer of the network is used for classification and recognition by the classifier;
[0077] Step 5: By estimating the optical flow of the video frame, the online contour inference and target bounding box positioning of the fruit trees and field ridges of interest based on the fusion semantic perception of the conditional random field and the motion optical flow are performed, refer to Figure 4 ;
[0078] Step 5.1: Based on the video frames, the inter-frame optical flow motion estimation map of the target of interest (fruit trees, field ridges) is obtained as the first input of the non-sub-module target contour inference method based on the conditional random field model (CRF);
[0079] Step 5.2: Obtain the semantic perception confidence map of the target of interest (fruit tree, field ridge) in a certain frame of the video as the second input of the non-submodular target contour inference method based on the conditional random field model (CRF);
[0080] Step 5.3: Obtain a color image in a certain frame of the video, and obtain a pixel color intensity map and a feature map through pixel enhancement processing, which are used as the third input of the non-sub-module target contour inference method based on the conditional random field model (CRF);
[0081] Step 5.4: The output results of the first three paths are fused to obtain the accurate contour mask of the target of interest based on the conditional random field model (CRF) non-sub-module target contour inference method;
[0082] Based on the input of the method in step 5, the conditional random field model (CRF) can be obtained offline in combination with the conditional random field model network structure known in the art, which will not be repeated here; the CRF to be constructed in this study is different from the CRF of traditional video segmentation in two aspects: first, due to the convolutional characteristics of the semantic perception network of the target of interest, there are no holes in the confidence map, so the CRF of this project is used to refine the confidence map, instead of the traditional CRF for smoothing the segmentation results; secondly, this project uses moving optical flow to distinguish targets, instead of using optical flow in traditional methods to force the consistency of the movement of the target of interest, which may be destroyed during the movement, thereby affecting the segmentation effect.
[0083] Step 6: Update the general data model (Gaussian perturbation model) of the object of interest by locating the bounding box of the object of interest in step 5;
[0084] Step 7: Track the target of interest based on the kernel correlation filter algorithm and update the deep semantic perception model of the target of interest in step 3;
[0085] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for identifying fruit tree field ridges, characterized in that: Includes steps: Get the video sequence captured by the camera; The semantic labels of fruit trees and field ridges are generated online through a semantic labeling model based on temporal and spatial visual attention mechanism and fully convolutional neural network. The fruit tree and field ridge depth semantic perception model based on convolutional neural network is used for fusion semantic perception; Extract the feature values of fruit trees and ridges, and identify the fruit trees and ridges through classifiers; By estimating the optical flow of video frames, online contour inference and target bounding box relocation of fruit trees and field ridges are performed based on the conditional random field model and motion optical flow. Track the target of interest based on the kernel correlation filter algorithm and update the depth semantic perception model of fruit trees and field ridges; The deep semantic perception model is obtained through offline training, including: Get the i-th frame of the video sequence image; Obtain target tracking confidence map based on Gaussian perturbation model; The semantic labels of the generated fruit trees and ridges are semantically selected and semantically filtered based on the kernelized correlation filter to obtain the semantic dense confidence map of the objects of interest. The target tracking confidence map and the semantic dense confidence map are used as inputs of the deep perception network, and the deep perception network is trained offline to generate parameters of the deep perception network; A multi-scale recurrent convolutional network is used to deeply fuse spatiotemporal features at multiple levels, and a gated recurrent network is used as the recurrent unit to determine the deep semantic perception model. The online contour reasoning and target bounding box relocation of the fruit trees and field ridges based on the conditional random field model and motion optical flow specifically include: Step 5.1: Obtain a color image in a certain frame of the video, and obtain the image pixel intensity and feature map through pixel enhancement processing; Step 5.2: Based on the deep semantic perception model, obtain the semantic perception confidence map of the target of interest in a certain frame of the video; Step 5.3: Obtain an inter-frame optical flow motion estimation map of the target of interest based on the video frames; Step 5.4: Based on the random conditional field model constructed offline, the image pixel intensity and feature map, the semantic perception confidence map and the optical flow motion estimation map are used as the input of the non-sub-module target contour reasoning of the conditional random field model to obtain the target contour and then locate the target bounding box; The tracking of the target of interest based on the kernel correlation filtering algorithm and updating of the fruit tree and field ridge depth semantic perception model specifically include: Step 6.1: Based on the contour reasoning of the object of interest and the positioning of the object bounding box, update the Gaussian perturbation model of the object of interest; Step 6.2: Obtain the target tracking confidence map based on the Gaussian perturbation of the kernelized correlation filter, and update the depth semantic perception model of fruit trees and ridges.
2. The method for identifying fruit tree field ridges according to claim 1, characterized in that: The video sequence captured by the camera includes: Step 1.1: Move the mobile robot equipped with a camera in the orchard to take pictures of fruit trees and field ridges; Step 1.2: Obtain a video sequence that captures the output of the target information of interest.
3. The method for identifying fruit tree field ridges according to claim 1, characterized in that: The semantic label model is obtained through offline training, which specifically includes: Step 2.1: Based on the image dataset containing two types of semantic labels, fruit trees and ridges, train the fully convolutional neural network offline. Step 2.2: Connect the gated recurrent unit to the fully convolutional neural network, improve the GRU to a convolutional GRU layer, and obtain the semantic label models of the fruit trees and ridges in the image; Step 2.3: In the semantic segmentation process of the semantic label model, a temporal and spatial selective attention mechanism is introduced to collect two adjacent frames of the video sequence, and the corresponding semantic labels are generated online through the semantic label model.
4. The method for identifying fruit tree field ridges according to claim 3, characterized in that: The method for obtaining the image data set is: After acquiring the video sequence, the multi-video sequence is detected frame by frame, and each frame of the acquired image is converted into grayscale, a digital grayscale image mathematical model is established, and an image with enhanced grayscale value is obtained; Perform two filtering and noise reduction processes on the image after enhancing the gray value; Detect the image frame by frame. When a fruit tree or a ridge is detected to appear suddenly in the image, the frame is updated to the initial frame. The fruit tree or ridge in the image is the target of interest, and the target area of interest is locked. Obtain multiple sets of images containing objects of interest as image datasets.
5. The method for identifying fruit tree field ridges according to claim 1, characterized in that: The extracting of the feature values of the fruit trees and the ridges and identifying the fruit trees and the ridges by the classifier specifically includes: Step 4.1: Obtain an image containing fruit trees and field ridges; Step 4.2: The images of fruit trees and ridges are subjected to denoising through a denoising network; Step 4.3: Extract feature values through the deep residual shrinkage network. The fully connected output layer of the deep residual shrinkage network is a classifier to classify and identify fruit trees and ridges.
6. The method for identifying fruit tree field ridges according to claim 1, characterized in that: The fully convolutional neural network adopts AlexNet, VGG or GoogleNet network architecture.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method for identifying fruit tree field ridges as described in any one of claims 1-6 is implemented.
Citation Information
Patent Citations
Image Semantic Segmentation Method Based on Deep Full Convolutional Network and Conditional Random Field
AU2020103901A4
Target detection and identification method and system for real-time video
CN110147702A