Method, device, electronic device and storage medium for detecting crop tillering

Through deep processing of rice tillering images and model detection based on deep learning, the problems of low efficiency and high cost of manual operation in rice tiller counting are solved, and efficient and accurate automatic detection is achieved.

CN120125502BActive Publication Date: 2025-10-10INST OF GENETICS & DEVELOPMENTAL BIOLOGY CHINESE ACAD OF SCI
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510057308.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-10-10
Estimated Expiration
2045-01-14

AI Technical Summary

Technical Problem

Existing technologies for counting rice tillers have problems with low efficiency, high cost, and inaccurate precision due to reliance on manual operations. Accurate detection is particularly difficult to achieve in large-scale genetic material screening and new variety development.

Method used

By performing deep processing, normalization, edge detection, and threshold segmentation on images of harvested crop residues, a deep learning-based model is used for tiller detection, reducing computing resource consumption and improving detection efficiency and accuracy.

Benefits of technology

It significantly improves the efficiency and accuracy of rice tillering detection, reduces labor time and costs, and supports automated processes in multiple scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125502B_ABST
    Figure CN120125502B_ABST
Patent Text Reader

Abstract

The application provides a method and device for detecting crop tillering, an electronic device and a storage medium. The method comprises the following steps: performing deep processing on an original image to obtain a first depth image, performing normalization processing on the first depth image to obtain a second depth image, performing edge detection and threshold segmentation on the second depth image to obtain a to-be-detected crop image. In this way, the original image is processed based on the depth information of the image to obtain the to-be-detected crop image, which reduces the calculation resource consumption in the inference process of the first model, thereby improving the inference efficiency while ensuring the detection accuracy. The tillering in the original image is detected through the first model, which significantly improves the rice tillering detection efficiency, reduces the time consumption and cost of manual work, and ensures the detection accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of agricultural automation detection technology, and in particular to a method, device, electronic device and storage medium for detecting crop tillering. Background Art

[0002] In rice cultivation research, tiller count, as a key indicator for assessing plant growth and yield potential, is crucial for understanding the genetic characteristics of different varieties and optimizing field management strategies. By accurately monitoring rice tiller counts, researchers can effectively evaluate rice growth responses under specific conditions, such as varying fertilizer application levels or planting densities, providing a scientific basis for selecting high-yield and stable varieties. This data also provides an important reference for developing appropriate planting density plans, helping to increase yield per unit area and promote sustainable agricultural development.

[0003] Current methods for assessing viable rice ears in rice fields rely heavily on manual labor. However, manual labor is slow and labor-intensive, and is prone to bias. Furthermore, identifying viable rice tillers requires experience, hindering large-scale, accurate genetic screening and new variety development. Summary of the Invention

[0004] In view of this, the purpose of the present application is to propose a method, device, electronic device and storage medium for detecting crop tillering, so as to solve or partially solve the above-mentioned problems.

[0005] Based on the above objectives, in a first aspect, the present application provides a method for detecting crop tillering, comprising:

[0006] Acquiring an original image, wherein the original image includes the stubble of the crop after being harvested;

[0007] Performing depth processing on the original image to obtain a first depth image;

[0008] Normalizing the first depth image to obtain a second depth image;

[0009] performing edge detection on the second depth image to obtain an edge image;

[0010] performing threshold segmentation on the edge image to obtain a crop image to be detected;

[0011] The tillers in the crop image to be detected are detected based on the first model to obtain a tiller detection result.

[0012] In a second aspect of the present application, a device for detecting crop tillering is provided, comprising:

[0013] An acquisition module is configured to acquire an original image, the original image comprising stubble after the crop is harvested;

[0014] A first processing module is configured to perform depth processing on the original image to obtain a first depth image;

[0015] A second processing module is configured to perform normalization processing on the first depth image to obtain a second depth image;

[0016] A third processing module is configured to perform edge detection on the second depth image to obtain an edge image;

[0017] A fourth processing module is configured to perform threshold segmentation on the edge image to obtain a to-be-detected crop image;

[0018] A detection module is configured to detect tillers in the to-be-detected crop image based on a first model to obtain a tiller detection result.

[0019] In a third aspect, an electronic device is provided, which includes a memory, a processor, and a computer program stored in the memory and capable of running on the processor, and the processor implements the method of the first aspect when executing the program.

[0020] In a fourth aspect, a non-transitory computer-readable storage medium is provided, which stores computer instructions for causing a computer to execute the method of the first aspect.

[0021] As can be seen from the above, the method, device, electronic device, and storage medium for detecting tillers of crops provided by the present application obtain a first depth image by performing depth processing on an original image, obtain a second depth image by performing normalization processing on the first depth image, and obtain a to-be-detected crop image by performing edge detection and threshold segmentation on the second depth image. In this way, the original image is processed based on the depth information of the image to obtain the to-be-detected crop image, which reduces the consumption of computing resources in the inference process of the first model, thereby improving the inference efficiency while ensuring the detection accuracy. By detecting tillers in the original image through the first model, the efficiency of detecting tillers of rice is significantly improved, the time consumption and cost of manual work are reduced, and the detection accuracy is ensured. BRIEF DESCRIPTION OF DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the present application or related art, the drawings needed to be used in the embodiments or related art description will be briefly introduced. Obviously, the drawings in the following description are only embodiments of the present application, and for those skilled in the art, other drawings can be obtained without creative labor.

[0023] Figure 1A A schematic diagram of an exemplary training data sampling process according to an embodiment of the present application is shown.

[0024] Figure 1B A flow chart of an exemplary depth cropping module according to an embodiment of the present application is shown.

[0025] Figure 1C A schematic diagram of an exemplary first depth image according to an embodiment of the present application is shown.

[0026] Figure 1D A schematic diagram of an exemplary second depth image according to an embodiment of the present application is shown.

[0027] Figure 1E A schematic diagram of an exemplary edge image according to an embodiment of the present application is shown.

[0028] Figure 1F A schematic diagram showing exemplary crop segmentation results according to an embodiment of the present application is shown.

[0029] Figure 2 A schematic diagram of the training process of an exemplary first model according to an embodiment of the present application is shown.

[0030] Figure 3A A schematic diagram of an exemplary process for detecting crop tillers according to an embodiment of the present application is shown.

[0031] Figure 3B A schematic diagram of an exemplary original image according to an embodiment of the present application is shown.

[0032] Figure 3C A schematic diagram showing exemplary tillering detection results according to an embodiment of the present application is shown.

[0033] Figure 3D A schematic diagram of the detection process of an exemplary first model according to an embodiment of the present application is shown.

[0034] Figures 4A to 4D A schematic diagram of an exemplary depth cropping process according to an embodiment of the present application is shown.

[0035] Figures 5A to 5H A schematic diagram showing exemplary tillering detection results according to an embodiment of the present application is shown.

[0036] Figure 6 A schematic flow chart of an exemplary method for detecting crop tillering according to an embodiment of the present application is shown.

[0037] Figure 7 A schematic diagram of an exemplary device for detecting crop tillers according to an embodiment of the present application is shown.

[0038] Figure 8 A schematic diagram of an exemplary electronic device according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0039] In order to make the objectives, technical solutions and advantages of this application more clear, this application is further described in detail below in combination with specific embodiments and with reference to the accompanying drawings.

[0040] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present application should have the usual meanings understood by people with ordinary skills in the field to which this application belongs. The "first", "second" and similar words used in the embodiments of the present application do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0041] In rice cultivation research, tiller count, as a key indicator for assessing plant growth and yield potential, is crucial for understanding the genetic characteristics of different varieties and optimizing field management strategies. By accurately monitoring rice tiller counts, researchers can effectively evaluate rice growth responses under specific conditions, such as varying fertilizer application levels or planting densities, providing a scientific basis for selecting high-yield and stable varieties. This data also provides an important reference for developing appropriate planting density plans, helping to increase yield per unit area and promote sustainable agricultural development.

[0042] Current methods for assessing effective rice ears in rice fields rely heavily on manual labor. When rice reaches maturity in its growth cycle, researchers typically record the number of effective ears on each plant directly in the field, or harvest the crops and bring them to a laboratory to count them. While feasible, these methods have significant drawbacks, such as slow execution, high labor intensity, and the potential for bias due to manual labor. Identifying effective rice tillers requires experience, which hinders large-scale, accurate genetic material screening and new variety development.

[0043] Automation technology for rice tiller counting tasks has been under exploration. One related technology developed a high-throughput device based on a traditional X-ray computed tomography system to measure rice tillers on a conveyor belt. However, this method exposes testers to severe radiation exposure. Another related technology uses magnetic resonance imaging to obtain cross-sectional images of rice stalks and determines the number of tillers through separation processing, thus avoiding this adverse effect. Another related technology has developed a high-throughput micro-CT-RGB imaging system (Computed Tomography) to more easily automate the process of rice phenotyping. However, these early methods often involve expensive equipment, and because the plants to be tested need to be transported to the testing area, there are still problems of long time and huge costs.

[0044] In recent years, deep learning-based computer vision has made significant progress in agricultural counting, and work has also been conducted on analyzing rice tillers based on this technology. However, during the rice maturation period, there is severe occlusion between plants, making it very difficult to accurately identify all rice ears. However, there is almost no occlusion for the remaining tillers after harvest, making the target clearer. In related technologies, a rice tiller counting method based on deep convolutional neural networks has been proposed. This method automatically detects and counts the original rice tillers by detecting the remaining tillers after harvest. However, this detection method has high annotation costs, its dataset is collected at very low altitudes from the ground, its model is applicable to fewer scenarios, and its performance is poor in complex environments and densely populated scenes.

[0045] While some methods have made progress in automating rice tiller counting, they are often difficult to implement in real-world agricultural scenarios due to high costs, difficult preconditions, or poor model performance. Therefore, there is an urgent need to develop a low-cost, hardware-reliant rice tiller counting method that also performs well and is robust enough to automate rice tiller counting in multiple scenarios.

[0046] In view of this, the present application provides a method, device, electronic device and storage medium for detecting crop tillers, which performs depth processing on the original image to obtain a first depth image, normalizes the first depth image to obtain a second depth image, performs edge detection and threshold segmentation on the second depth image to obtain the crop image to be detected. In this way, processing the original image based on the depth information of the image to obtain the crop image to be detected reduces the consumption of computing resources in the inference process of the first model, thereby improving the inference efficiency and ensuring the detection accuracy. By detecting tillers in the original image through the first model, the efficiency of rice tiller detection is significantly improved, and the time consumption and cost of manual labor are reduced, while ensuring the detection accuracy.

[0047] Figure 1A FIG. 1 is a schematic diagram of an exemplary training data sampling process 100 according to an embodiment of the present application.

[0048] like Figure 1A As shown, during the model training process, the original image 102 is first randomly sampled and cropped to obtain training data 104, which is then fed into the constructed first model for training. Because the training data used to detect rice tillering has a high degree of differentiation between dense and sparse scenes, and the manually labeled true labels are often densely distributed in the center of the image, to ensure training quality, in some embodiments, a global sampling coefficient θ = 0.1 can be set to help the model avoid overfitting. Simultaneously, a random number generator is used in the training data loader to generate a random number ξ and compare it with θ each time the sampled and cropped data is fed into the model to select different sampling and cropping strategies. In some embodiments, when ξ ≤ θ, global cropping is performed, sampling the entire original image 102, allowing the model to learn features of negative samples in the background; when ξ > θ, local cropping is performed, and the sampled cropped block must contain at least one true labeled point, allowing the first model to learn sufficient features of positive samples. Finally, the training data loader inputs the training data 104 obtained by the above cropping strategy into the first model. For example, the training data 104 may be a square image with a side length of patch_size=256 and a feature channel number of channal=3, such as Figure 1A As shown, the training data 104 obtained through different cropping strategies may include training data of negative sample features in the background and training data of positive sample features.

[0049] In some embodiments, the training data 104 may be obtained based on a depth cropping module.

[0050] Figure 1B FIG. 1 is a flow chart of an exemplary depth cropping module 120 according to an embodiment of the present application. For example, the depth cropping module 120 may crop an image based on depth information.

[0051] The inventors of the present application discovered that, based on the prior knowledge that the plant is approximately located at the center of the original image 102 and the good performance of the monocular depth estimation model, in some embodiments, the training data 104 can be cropped based on the depth information of the image, which can reduce the time and computing resources used in model training.

[0052] like Figure 1BAs shown, in some embodiments, the original image 102 can be input into the second model 122. In some embodiments, the second model 122 can be a monocular depth estimation model (DAM V2) based on ViT-Large to obtain a first depth image 124. The depth information in the first depth image 124 is then normalized. In some embodiments, the normalization process can be performed by subtracting the minimum value of the corresponding row in the first depth image 124 from the row, and then subtracting the minimum value of the corresponding column in the first depth image 124 from the column, thereby obtaining a normalized second depth image 126.

[0053] Figure 1C FIG. 1 is a schematic diagram showing an exemplary first depth image 124 according to an embodiment of the present application. Figure 1D FIG. 1 is a schematic diagram showing an exemplary second depth image 126 according to an embodiment of the present application.

[0054] Combine Figure 1C and Figure 1D The first depth image 124 is normalized to obtain a second depth image 126. The background depth in the second depth image 126 remains substantially consistent, and the normalization process can reduce the impact of different shooting angles.

[0055] Back to Figure 1B , edge detection is performed on the normalized second depth image 126. In some embodiments, the Canny operator may be used to perform edge detection, thereby obtaining the edge image 128.

[0056] Figure 1E FIG. 1 is a schematic diagram showing an exemplary edge image 128 according to an embodiment of the present application.

[0057] like Figure 1E As shown, in the edge image 128, the edge point is 1 and the non-edge point is 0. E(x, y) can represent the pixel value of the edge image at the coordinate (x, y). In some embodiments, the center point of the edge image 128 can be defined as (x c ,y c )=(w / 2,h / 2), where w and h are the length and width of the input image, respectively. In the application scenario of the embodiments of the present application, based on the fact that the plant is approximately located in the center of the image, in some embodiments, the weight W(x,y) of each pixel to the center point can be calculated based on the Euclidean distance. Using the Euclidean distance to calculate the center point weight is more intuitive and can be specifically expressed as:

[0058]

[0059] In some embodiments, a weighted center point calculation may be performed on the edge image 128 to obtain the weighted center of all edge points, which can be specifically expressed as:

[0060]

[0061] Among them, (x c ′,y c ′) represents the coordinates of the weighted center point.

[0062] Compared with performing OTSU threshold segmentation on the entire original image 102, a more accurate local OTSU threshold segmentation image can be obtained by intercepting an image (2w / 3, 2h / 3) with the obtained weighted center as the image center and then performing OTSU threshold segmentation.

[0063] Figure 1F FIG. 1 is a schematic diagram showing an exemplary crop segmentation result 130 according to an embodiment of the present application.

[0064] Combine Figure 1B and Figure 1F In some embodiments, a series of morphological operations are performed on the obtained OTSU threshold segmentation image, such as dilation followed by erosion and foreground concatenation, to obtain a crop segmentation result 130 based on depth information. Using depth information for cropping reduces computing resource consumption during model training and inference, thereby improving inference efficiency while ensuring detection accuracy.

[0065] Back to Figure 1B In some embodiments, the weighted center point of the crop segmentation image 130 is calculated again using the formula in the above embodiment. Finally, the minimum bounding rectangle of the connected domain closest to the weighted center point is taken to accurately frame the crop region 132 to be detected. The crop region 132 to be detected is cropped out from the original image 102 to obtain the training data 104.

[0066] Figure 2 FIG. 2 is a schematic diagram showing a training process 200 of an exemplary first model 202 according to an embodiment of the present application.

[0067] like Figure 2 As shown, in some embodiments, the training data 104 can be input into the constructed first model 202 to train the first model 202. The first model 202 can be a model constructed based on a point query transformer. In some embodiments, the structure of the first model 202 can include a feature extraction module, a target query module, a quadtree splitter, and a target prediction branch module.

[0068] In the training process 200, the final N predicted point query information 202 obtained by the first model 202 after training with the training data 104 will be matched with the M real labels 206 for one-to-one bipartite matching to calculate the loss (for the predicted point query information 202 that does not have a point 208 corresponding to the real label 206, the corresponding empty set ), and combined with the loss of splitting sparse point queries in the quadtree splitter in the first model 202 to obtain the final model training loss.

[0069] The N prediction point query information 202 can be expressed as: The M true labels 206 can be expressed as: The process of bipartite matching is to find a matching that minimizes the cost. Expressed as:

[0070]

[0071] Where i and j represent the index of each point query, i∈{1,...,N},j∈{1,...,M}; c i represents the probability of category prediction and is 1 only when the category prediction is correct, and 0 when it is wrong; ‖·‖2 represents the Euclidean distance; α is the balance factor, and there is always N>M; argmin represents finding the parameter value that minimizes the following expression; q i Indicates the query information of the prediction point; y j represents the true label.

[0072] The loss of point query consists of category and position (or offset), which can be expressed as:

[0073]

[0074] Among them, λ1 represents the hyperparameter; p σ(i) Denotes the output q from the model through bipartite matching σ i The obtained and true label y j One-to-one matching; class represents the classification loss, c i is the category predicted by the model, c i * is the true category; l localization represents the positioning loss.

[0075] The loss of point query splitting can be expressed as:

[0076] l split =1-max(M s )+min(M s )

[0077] Among them, M sQuery split information for sparse points divided into blocks with a fixed step size of k; max indicates maximum; min indicates minimum.

[0078] Finally, the training loss of the first model 202 is given by:

[0079] l total =l pq +λ2l split

[0080] Among them, λ2 is a hyperparameter.

[0081] The first model 202 performs back propagation learning on the resulting loss to reduce the prediction error of the model, thereby further improving the accuracy and robustness of the model.

[0082] Figure 3A FIG. 3 is a schematic diagram showing an exemplary process 300 for detecting crop tillers according to an embodiment of the present application. Figure 3B FIG. 3 is a schematic diagram showing an exemplary original image 302 according to an embodiment of the present application. Figure 3C A schematic diagram showing an exemplary tillering detection result 306 according to an embodiment of the present application is shown.

[0083] like Figure 3A As shown, in some embodiments, the collected original image 302 can be input into the depth cropping module 120 to perform depth processing and threshold segmentation on the original image 302 to obtain the crop image 304 to be detected. Figure 3B As shown, in some embodiments, original image 302 may include harvested crop residue 3022, which may be located at the center of original image 302. Crop residue refers to the residue left in or on the soil after the harvest of mature crops, including all or most of the crop roots and some of the aboveground parts. After decomposition in the soil, these residues can increase the nutrient content and produce a certain amount of crop secretions, which can promote or inhibit crop growth. In reduced tillage or no-till farming, crop residue can be fully utilized to protect the soil.

[0084] like Figure 3A As shown, in some embodiments, the crop image 304 to be detected can be input into the first model 202 for tillering detection, thereby obtaining a tillering detection result 306. Figure 3C As shown, the red dots are true labels, and the green dots are tillers detected using the method of the present embodiment. As can be seen, the tiller detection result 306 is relatively accurate. Therefore, the method provided by the present embodiment can automatically detect crop tillers in images, avoiding reliance on manual labor for effective ear counting, significantly reducing time and labor costs, while achieving accurate detection performance.

[0085] Figure 3D A schematic diagram of a detection process of an exemplary first model 202 according to an embodiment of the present application is shown.

[0086] like Figure 3D As shown, in some embodiments, the structure of the first model 202 may include a feature extraction module 222, a target query module 224, a quadtree splitter 226, and a target prediction branch module 228. The improvement of the model in the embodiment of the present application not only reduces the number of parameters of the model, but also improves the performance of the model. Its modules specifically include: the backbone network of the feature extraction module is a sliding converter series model; the encoder module in the original model is removed, and in the embodiment of the present application, the output of the feature extraction module is directly used as the encoding feature for the model to perform target query, and its target query module is a stack of two decoders; the quadtree splitter is a neural network composed of a pooling layer, a convolutional layer, and a Sigmund activation function in sequence; the target prediction branch is composed of two fully connected layers: a category prediction branch and a deviation prediction branch.

[0087] First model 202 first performs a sparse point query 308 on the input crop image 304 to be detected. It then divides the image into blocks according to a fixed step size k and inputs a sparse point query into each block, thereby obtaining a sparse point query result 314. Simultaneously, crop image 304 to be detected is fed into the backbone network within feature extraction module 222 for feature extraction. Specifically, in some embodiments, a sliding transformer within the backbone network can be used to perform attention calculations within a sliding window on the image. Dense features 312 and sparse features 310 are then output at the feature fusion layer, respectively. The resolution of dense features 312 and sparse features 310 are determined by the fixed step size k and the resolution of the input crop image 304 to be detected.

[0088] The sparse point query result 314 is combined with the sparse features 310 obtained by the feature extraction module 222 and fed into the quadtree splitter 226 to determine whether a sparse point query needs to be split (for example, a sparse point query can be divided into a point query 320 that needs to be split and a point query 318 that does not need to be split) into four dense point queries to obtain a complete point query 316. The processing of sparse point queries by the quadtree splitter 226 makes the model more efficient in processing sparse and dense features, further improving the accuracy and robustness of the model.

[0089] The quadtree splitter 226 is a key component used to construct the point query quadtree, allowing the model to adaptively handle both sparse and dense regions in the image. The quadtree splitter 226 dynamically handles regions of varying density by recursively partitioning the image into four quadrants, or nodes of the quadtree. This splitting process follows a sparse-to-dense principle, initially setting sparse query points uniformly across the image. These points are then adaptively split into denser query points in crowded scenes to improve prediction accuracy.

[0090] The quadtree splitter 226 can determine whether a split is necessary by examining the density of a local area. If an area is deemed dense, the query point in that area is split into more query points, forming the next level of the quadtree. This process is repeated until a predetermined maximum number of splits is reached or other termination conditions are met, thereby constructing a multi-layer quadtree structure. The quadtree splitter 226 outputs a split map indicating which areas of the image need to be split to obtain a complete point query 316.

[0091] The complete point query 316, combined with the dense features 312 obtained by the feature extraction module 222, is fed into the decoder 312 in the target query module 224. The decoder decodes the combined encoded information to obtain a final unified decoded output 322 of the point query category and offset. The unified decoded output 322 of the point query category and offset is fed into the prediction head in the target and prediction branch 228. For each point query, the prediction head uses a linear layer to obtain the category of each point query and a fully connected layer to obtain the offset of each point query. Ultimately, N point query information 324 is obtained, i.e., the detection result of crop tillers in the input image.

[0092] Figures 4A to 4D A schematic diagram of an exemplary depth cropping process according to an embodiment of the present application is shown.

[0093] like Figures 4A to 4D As shown in the figure, the large red box and red center point represent the weighted center point and local OTSU region obtained through edge detection, while the small yellow box and yellow center point represent the final weighted center point and final detection region obtained after local OTSU threshold segmentation and morphological operations. Cropping this region from the original image and feeding it into the trained point query converter allows for lightweight rice tiller number detection.

[0094] Figures 5A to 5H A schematic diagram showing exemplary tillering detection results according to an embodiment of the present application is shown.

[0095] like Figures 5A to 5HAs shown, the number of true labels is 26, 19, 26, 44, 40, 54, 67, and 89, and the tillering detection results are 29, 22, 26, 45, 41, 56, 70, and 86. This shows that the tillering detection results obtained by the method provided in the embodiment of the present application are relatively accurate.

[0096] Figure 6 FIG. 6 is a flow chart of an exemplary method 600 for detecting crop tillering according to an embodiment of the present application. Figure 6 As shown, method 600 may include the following steps.

[0097] In step 602 , an original image is acquired, where the original image includes the stubble of the harvested crops.

[0098] In step 604, depth processing is performed on the original image to obtain a first depth image.

[0099] In some embodiments, performing depth processing on the original image to obtain a first depth image further includes: determining the depth information of the original image based on the original image and a second model, the second model including a monocular depth estimation model; performing depth processing on the original image based on the depth information to obtain the first depth image.

[0100] In step 606 , the first depth image is normalized to obtain a second depth image.

[0101] In step 608 , edge detection is performed on the second depth image to obtain an edge image.

[0102] In step 610 , threshold segmentation is performed on the edge image to obtain a crop image to be detected.

[0103] In some embodiments, performing threshold segmentation on the edge image to obtain the crop image to be detected further includes: determining a first center point of the edge image; calculating the weight of each pixel in the edge image with respect to the first center point to obtain a first center point weight; calculating a first weighted center of the edge image based on the first center point weight; and performing threshold segmentation and morphological operations on the edge image with the first weighted center as the center of the edge image to obtain the crop image to be detected.

[0104] In some embodiments, calculating the weight of each pixel in the edge image to the center point to obtain the center point weight further includes: calculating the weight of each pixel in the edge image to the center point according to the Euclidean distance to obtain the center point weight.

[0105] In some embodiments, calculating the weighted center according to the center point weight further includes: calculating the weighted center according to the center point weight and each pixel in the edge image.

[0106] In some embodiments, taking the weighted center as the center of the edge image, performing threshold segmentation and morphological operations on the edge image to obtain the crop image to be detected further includes: taking the weighted center as the center of the edge image, performing threshold segmentation and morphological operations on the edge image to obtain a crop segmentation image; determining a second center point of the crop segmentation image; calculating the weight of each pixel in the crop segmentation image with respect to the second center point to obtain a second center point weight; calculating the weighted center of the edge image according to the second center point weight; taking the weighted center as the second center of the crop segmentation image; and cropping the crop segmentation image based on the second center to obtain the crop image to be detected.

[0107] In step 612, tillers in the crop image to be detected are detected based on the first model to obtain a tiller detection result.

[0108] In some embodiments, the first model includes a feature extraction module, a target query module, a quadtree splitter and a target prediction branch module, the feature extraction module includes a backbone network, and the backbone network is a sliding converter series model; the quadtree splitter includes a pooling layer, a convolutional layer and a Sigmod activation function; the target prediction branch module includes a category prediction branch and a deviation prediction branch.

[0109] In some embodiments, the training loss of the first model is obtained by combining the first loss with the second loss; the first loss is the loss calculated by performing a one-to-one binary matching between the predicted point query information and the true label obtained after the first model is trained on the training data; the second loss is the loss of splitting the sparse point query in the quadtree splitter.

[0110] In some embodiments, the detecting tillers in the crop image to be detected based on the first model to obtain the tiller detection result further includes: performing a sparse point query on the crop image to be detected to obtain a sparse point query result; extracting dense features and sparse features of the crop image to be detected; obtaining a complete point query based on the sparse point query result and the sparse features; decoding the complete point query and the dense features to obtain a unified decoding output of the point query category and the offset; obtaining the tiller detection result based on the unified decoding output of the point query category and the offset.

[0111] The present application provides a method, device, electronic device, and storage medium for detecting crop tillers. The method performs depth processing on the original image to obtain a first depth image, normalizes the first depth image to obtain a second depth image, and performs edge detection and threshold segmentation on the second depth image to obtain the crop image to be detected. Processing the original image based on the depth information of the image to obtain the crop image to be detected reduces the consumption of computing resources in the first model inference process, thereby improving the inference efficiency and ensuring the detection accuracy. Detecting tillers in the original image using the first model significantly improves the efficiency of rice tiller detection, reduces manual time consumption and costs, and ensures detection accuracy.

[0112] It should be noted that the method of the embodiment of the present application can be performed by a single device, such as a computer or server. The method of this embodiment can also be applied in a distributed scenario and performed by multiple devices working together. In such a distributed scenario, one of the multiple devices may only perform one or more steps of the method of the embodiment of the present application, and the multiple devices will interact with each other to complete the method.

[0113] It should be noted that the above description is limited to some embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in an order different from that described in the above embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0114] Based on the same technical concept, corresponding to any of the above-mentioned embodiments and methods, the present application also provides a device for detecting crop tillering.

[0115] refer to Figure 7 , the device for detecting crop tillering comprises:

[0116] The acquisition module 701 is configured to acquire an original image, where the original image includes the stubble of the harvested crops.

[0117] The first processing module 702 is configured to perform depth processing on the original image to obtain a first depth image.

[0118] The first processing module 702 is further configured to determine depth information of the original image based on the original image and a second model, where the second model includes a monocular depth estimation model; and perform depth processing on the original image based on the depth information to obtain the first depth image.

[0119] The second processing module 703 is configured to perform normalization processing on the first depth image to obtain a second depth image.

[0120] The third processing module 704 is configured to perform edge detection on the second depth image to obtain an edge image.

[0121] The fourth processing module 705 is configured to perform threshold segmentation on the edge image to obtain a to-be-detected crop image.

[0122] The fourth processing module 705 is further configured to determine a first center point of the edge image, calculate a weight of each pixel in the edge image to the first center point to obtain a first center point weight, and calculate a first weighted center of the edge image according to the first center point weight.

[0123] The fourth processing module 705 is further configured to calculate the weight of each pixel in the edge image to the center point according to the Euclidean distance to obtain the center point weight.

[0124] The fourth processing module 705 is further configured to calculate the weighted center according to the center point weight and each pixel in the edge image.

[0125] The fourth processing module 705 is further configured to perform threshold segmentation and morphological operation on the edge image with the weighted center as the center of the edge image to obtain a crop segmentation image, determine a second center point of the crop segmentation image, calculate a weight of each pixel in the crop segmentation image to the second center point to obtain a second center point weight, calculate a weighted center of the edge image according to the second center point weight, and crop the crop segmentation image based on the second center to obtain the to-be-detected crop image.

[0126] The detection module 706 is configured to detect tillers in the to-be-detected crop image based on a first model to obtain a tiller detection result.

[0127] The first model includes a feature extraction module, a target query module, a quadtree splitter, and a target prediction branch module. The feature extraction module includes a backbone network, and the backbone network is a sliding transformer series model. The quadtree splitter includes a pooling layer, a convolution layer, and a Sigmod activation function. The target prediction branch module includes a class prediction branch and a bias prediction branch.

[0128] Among them, the training loss of the first model is obtained by combining the first loss with the second loss; the first loss is the loss of one-to-one binary matching calculation between the predicted point query information and the true label obtained after the first model is trained on the training data; the second loss is the loss of splitting the sparse point query in the quadtree splitter.

[0129] The detection module 706 is further configured to perform a sparse point query on the crop image to be detected to obtain a sparse point query result; extract dense features and sparse features of the crop image to be detected; obtain a complete point query based on the sparse point query result and the sparse features; decode the complete point query and the dense features to obtain a unified decoding output of the point query category and the offset; and obtain the tillering detection result based on the unified decoding output of the point query category and the offset.

[0130] For the convenience of description, the above devices are described as being divided into various modules according to their functions. Of course, when implementing this application, the functions of each module can be implemented in the same or multiple software and / or hardware.

[0131] The apparatus of the above embodiment is used to implement the corresponding method 600 in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.

[0132] Based on the same technical concept, corresponding to any of the above-mentioned embodiments, the present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method 600 described in any of the above embodiments when executing the program.

[0133] Figure 8 A schematic diagram of an exemplary electronic device according to an embodiment of the present application is shown, and the device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other within the device via the bus 1050.

[0134] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0135] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage devices, dynamic storage devices, etc. The memory 1020 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.

[0136] The input / output interface 1030 is used to connect an input / output module to implement information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Input devices may include a keyboard, mouse, touch screen, microphone, various sensors, etc., and output devices may include a display, speaker, vibrator, indicator light, etc.

[0137] The communication interface 1040 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via a wired method (such as USB, network cable, etc.) or a wireless method (such as mobile network, WiFi, Bluetooth, etc.).

[0138] The bus 1050 comprises a path for transmitting information between the various components of the device (eg, the processor 1010 , the memory 1020 , the input / output interface 1030 , and the communication interface 1040 ).

[0139] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in a specific implementation, the device may also include other components necessary for normal operation. In addition, it will be understood by those skilled in the art that the above device may only include the components necessary to implement the embodiments of this specification, and does not necessarily include all the components shown in the figure.

[0140] The electronic device of the above embodiment is used to implement the corresponding method 600 in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.

[0141] Based on the same technical concept, corresponding to any of the above-mentioned embodiments, the present application also provides a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable the computer to execute method 600 described in any of the above embodiments.

[0142] The computer-readable media of this embodiment include permanent and non-permanent, removable and non-removable media, and information storage can be achieved by any method or technology. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.

[0143] The computer instructions stored in the storage medium of the above embodiment are used to enable the computer to execute method 600 as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0144] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present application (including the claims) is limited to these examples. Within the scope of the present application, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present application as described above, which are not provided in detail for the sake of simplicity.

[0145] In addition, for simplicity of description and discussion, and in order not to make the embodiment of the application difficult to understand, the known power supply / ground connection with integrated circuit (IC) chip and other components may or may not be shown in the accompanying drawings provided. In addition, the device can be shown in the form of a block diagram to avoid making the embodiment of the application difficult to understand, and this also takes into account the following fact, that is, the details of the embodiment of these block diagram devices are highly dependent on the platform to be implemented in the embodiment of the application (that is, these details should be fully within the scope of understanding of those skilled in the art). When specific details (for example, circuit) are set forth to describe exemplary embodiments of the application, it will be apparent to those skilled in the art that the embodiment of the application can be implemented without these specific details or when these specific details are changed. Therefore, these descriptions should be considered to be illustrative rather than restrictive.

[0146] Although the present invention has been described in conjunction with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those skilled in the art based on the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may utilize the embodiments discussed.

[0147] The embodiments of the present application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present application should be included in the scope of protection of this application.

Claims

1. A method for detecting crop tillering, comprising: Acquiring an original image, wherein the original image includes crop residues after the crop is harvested; Performing depth processing on the original image to obtain a first depth image; Normalizing the first depth image to obtain a second depth image; performing edge detection on the second depth image to obtain an edge image; performing threshold segmentation on the edge image to obtain a crop image to be detected, comprising: determining a first center point of the edge image; calculating a weight of each pixel in the edge image with respect to the first center point to obtain a first center point weight; calculating a first weighted center of the edge image based on the first center point weight; and performing threshold segmentation and morphological operations on the edge image with the first weighted center as the center of the edge image to obtain the crop image to be detected; The tillers in the crop image to be detected are detected based on the first model to obtain a tiller detection result, including: performing a sparse point query on the crop image to be detected to obtain a sparse point query result; extracting dense features and sparse features of the crop image to be detected; obtaining a complete point query based on the sparse point query result and the sparse features; decoding the complete point query and the dense features to obtain a unified decoding output of a point query category and an offset; and obtaining the tiller detection result based on the unified decoding output of the point query category and the offset.

2. The method according to claim 1, wherein The first model includes a feature extraction module, a target query module, a quadtree splitter and a target prediction branch module; the feature extraction module includes a backbone network, which is a sliding converter series model; the quadtree splitter includes a pooling layer, a convolutional layer and a Sigmod activation function; the target prediction branch module includes a category prediction branch and a deviation prediction branch.

3. The method according to claim 1, wherein Calculating the weight of each pixel in the edge image to the center point to obtain the center point weight further includes: The weight of each pixel in the edge image to the center point is calculated according to the Euclidean distance to obtain the center point weight.

4. The method according to claim 1, wherein The step of calculating the weighted center according to the center point weight further includes: The weighted center is calculated according to the center point weight and each pixel in the edge image.

5. The method according to claim 1, wherein Taking the weighted center as the center of the edge image, performing threshold segmentation and morphological operations on the edge image to obtain the crop image to be detected further includes: Taking the weighted center as the center of the edge image, performing threshold segmentation and morphological operations on the edge image to obtain a crop segmentation image; determining a second center point of the crop segmentation image; calculating a weight of each pixel in the crop segmentation image with respect to the second center point to obtain a second center point weight; Calculating a weighted center of the edge image according to the second center point weight; Taking the weighted center as the second center of the crop segmentation image; The crop segmentation image is cropped based on the second center to obtain the crop image to be detected.

6. The method of claim 2, wherein: The training loss of the first model is obtained by combining the first loss with the second loss; The first loss is a loss calculated by performing a one-to-one bipartite matching between the predicted point query information and the true label obtained after the first model is trained on the training data; The second loss is the loss of splitting sparse point queries in the quadtree splitter.

7. The method of claim 1, wherein: The performing depth processing on the original image to obtain a first depth image further includes: Determining depth information of the original image based on the original image and a second model, wherein the second model includes a monocular depth estimation model; Perform depth processing on the original image based on the depth information to obtain the first depth image.

8. A device for detecting crop tillering, comprising: an acquisition module configured to acquire an original image, wherein the original image includes the stubble of the crop after being harvested; A first processing module is configured to perform depth processing on the original image to obtain a first depth image; a second processing module, configured to perform normalization processing on the first depth image to obtain a second depth image; a third processing module, configured to perform edge detection on the second depth image to obtain an edge image; a fourth processing module configured to perform threshold segmentation on the edge image to obtain a crop image to be detected, comprising: determining a first center point of the edge image; calculating a weight of each pixel in the edge image with respect to the first center point to obtain a first center point weight; calculating a first weighted center of the edge image based on the first center point weight; and performing threshold segmentation and morphological operations on the edge image with the first weighted center as the center of the edge image to obtain the crop image to be detected; The detection module is configured to detect tillers in the crop image to be detected based on a first model to obtain a tiller detection result, including: performing a sparse point query on the crop image to be detected to obtain a sparse point query result; extracting dense features and sparse features of the crop image to be detected; obtaining a complete point query based on the sparse point query result and the sparse features; decoding the complete point query and the dense features to obtain a unified decoding output of a point query category and an offset; and obtaining the tiller detection result based on the unified decoding output of the point query category and the offset.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Plant segmentation method, device and equipment based on depth information

    CN118799343A