Method for identifying leaf uprightness and maximum leaf expansion of mechanically-organized cabbage plug seedling

Through the Swin Transformer model based on RFP structure, the image feature extraction and regression prediction of cabbage tray seedlings was solved, and the problem of lack of discrimination standards in the existing technology was solved, and the accurate identification of the erectness of the leaves and the maximum leaf spread was achieved, which improved the efficiency of mechanized transplantation.

CN119919795AActive Publication Date: 2025-05-02BEIJING RES CENT FOR INFORMATION TECH & AGRI
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411849505.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-16
Publication Date
2025-05-02
Estimated Expiration
2044-12-16

AI Technical Summary

Technical Problem

The prior art lacks the criteria and identification methods for the vernality and maximum leaf spread of the leaves of the kale sapling seedlings suitable for mechanized transplantation, resulting in low transplantation efficiency and frequent seedling damage.

Method used

The binocular images of cabbage tray seedlings were extracted using the Swin Transformer model based on RFP structure. The stem bounding box and leaf bounding box were obtained through regression prediction, and the key points were matched, the three-dimensional geometric position coordinates were determined, and the leaf uprightness and maximum leaf spread were calculated.

Benefits of technology

The detection accuracy of cabbage seedlings has been improved, and a mechanized judgment standard has been formed, which can identify the verticality of the leaves and the maximum leaf spread in real time, improving the efficiency and quality of mechanized transplantation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919795A_ABST
    Figure CN119919795A_ABST
Patent Text Reader

Abstract

The invention provides a method for identifying the leaf uprightness and the maximum leaf expansion of a cabbage plug seedling suitable for mechanization, and the method comprises the steps: calling a Swin Transform model based on an RFP structure to carry out the feature extraction of a collected binocular image of the cabbage plug seedling, and obtaining a feature map of the binocular image; performing regression prediction on the feature map of the binocular image to obtain a stem bounding box and a leaf bounding box of the binocular image; and performing key point matching on the stem bounding box and the leaf bounding box of the binocular image, and determining the three-dimensional geometric position coordinates of the cabbage plug seedling according to the image coordinates of the key points obtained by matching, thereby determining the leaf uprightness and the maximum leaf spread of the cabbage plug seedling. By means of the method, the defects that an existing seedling detection method mainly focuses on seedling condition detection and recognition in the plug cultivation process, a plug seedling mechanical judgment standard is not formed, the detection accuracy is limited, and the method is difficult to apply to actual production are overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target detection, and in particular to a method for identifying the uprightness and maximum leaf span of mechanized cabbage plug seedling leaves. Background Art

[0002] The plant structure of cabbage has an important influence on the quality of mechanized transplanting. Different varieties and different cultivation strategies may lead to different plant structures. Cabbage seedlings with poor uprightness and large leaf spread are easy to hang on the transplanting seedling cup during transplanting, resulting in damage to the seedlings or problems such as missing seedlings and replanting. In addition, the uprightness and leaf spread directly affect the photosynthetic utilization rate of the plant. Plants with good uprightness can make full use of light and have better metabolism. Plants with good leaf clustering grow more stably after transplanting, which is conducive to heading. The degree of leaf transplant damage is small and it is not easy to spread infectious bacteria. Cabbage is a long-day plant, and its yield is also closely related to the plant shape. A good plant shape can better utilize space and sunlight and produce more dry matter. However, the current mechanized cabbage plug seedling cultivation lacks the criteria and identification methods for uprightness and leaf spread, making it difficult for cabbage seedlings to adapt to mechanized transplanting. Problems such as seedling cup blockage, seedling injury, and skewed seedling transplanting during transplanting make it difficult to improve the efficiency of mechanized transplanting.

[0003] In addition, the manual measurement of the absolute and relative height of seedlings requires a lot of labor, and the judgment standard is highly subjective, which cannot be quantified and applied to large-scale plug tray seedlings. Existing research on seedling detection methods mainly focuses on seedling detection and identification during plug tray cultivation, and rarely considers the characteristics of mechanized transplanting agronomy. A judgment standard for mechanization of plug tray seedlings has not been formed, and the detection accuracy is limited, making it difficult to apply to actual production. Summary of the invention

[0004] The present invention provides a method for identifying the leaf uprightness and maximum leaf extension of cabbage plug tray seedlings suitable for mechanization, so as to solve the defects that the existing seedling detection methods mainly focus on the detection and identification of seedling conditions during the plug tray cultivation process, and no standard for judging the suitability of plug tray seedlings for mechanization is formed, the detection accuracy is limited, and it is difficult to be applied in actual production.

[0005] The present invention provides a method for identifying the leaf uprightness and maximum leaf span of mechanized cabbage plug seedlings, characterized in that the method comprises the following steps: Obtain binocular images of collected cabbage plug seedlings; Calling a Swin Transformer model based on the RFP structure to perform feature extraction on the binocular image to obtain a feature map of the binocular image; Performing regression prediction on the feature map of the binocular image to obtain a stem bounding box and a leaf bounding box of the binocular image; Key point matching is performed on the stem boundary frame and the leaf boundary frame of the binocular image, and the three-dimensional geometric position coordinates of the cabbage plug seedlings are determined according to the image coordinates of the key points obtained by matching; The leaf uprightness and maximum leaf span of the cabbage plug seedlings are determined according to the three-dimensional geometric position coordinates.

[0006] In some embodiments, the calling of the Swin Transformer model based on the RFP structure to extract features from the binocular image to obtain a feature map of the binocular image includes: Divide the image to be processed into a plurality of patches, input the patches into a first composite stage module based on a three-layer RFP structure in a first Swin Transformer model, extract features of the patches through the first composite stage module, and obtain a first feature map corresponding to the patches, wherein the image to be processed is a left image or a right image in the binocular image; Input the patch into a second composite stage module based on a 3-layer RFP structure in a second Swin Transformer model; Inputting the first feature map into a three-layer second composite stage module based on the RFP structure in a second Swin Transformer model respectively, and extracting features from the patch and the first feature map through the second composite stage module to obtain a second feature map; The second feature map is used as the feature map of the binocular image.

[0007] In some embodiments, extracting features from the patch by the first composite stage module to obtain a first feature map corresponding to the patch includes: Embedding the patch to obtain a corresponding embedded patch sequence, and reshaping the embedded patch sequence into a two-dimensional feature map through the W-MSA module in the first composite stage module; The two-dimensional feature map is divided by non-overlapping detection windows, and the two-dimensional feature map divided in the detection window is grouped and self-attention is calculated by the W-MSA module and the SW-MSA module in the first composite stage module to obtain a window representation corresponding to the detection window; Perform global downsampling self-attention calculation on the window representation corresponding to each detection window, and input the calculated result into the Token fusion module for feature fusion to obtain the fused representation; Inputting the fused representation into an MLP module to obtain a high-level feature map of the patch; The high-level feature map is input into a feature pyramid network included in the first composite stage module for feature extraction to obtain a first feature map corresponding to the patch.

[0008] In some embodiments, the step of inputting the high-level feature map into a feature pyramid network included in the first composite stage module for feature extraction to obtain a first feature map corresponding to the patch includes: Performing edge-adaptive upsampling on the high-level feature map to obtain multiple upsampled feature maps, and fusing the upsampled feature maps; The fused feature map is input into a feature pyramid network for multiple iterative recursive calculations to obtain a first feature map corresponding to the patch.

[0009] In some embodiments, the edge-adaptively upsampling the high-level feature map to obtain a plurality of upsampled feature maps includes: Traversing each pixel point in the high-level feature map, and calling the Sobel operator to calculate the edge gradient value of each pixel point along the four directions of up, down, left and right, and determining the maximum edge gradient value of each pixel point; Determine a pixel point whose maximum edge gradient value is less than or equal to a preset edge gradient threshold as a first pixel point to be interpolated in the flat area, and upsample the first pixel point to be interpolated in both horizontal and vertical directions using a bilinear interpolation algorithm to obtain a corresponding upsampled feature map; Determine a pixel point whose maximum edge gradient value is greater than a preset edge gradient threshold as a second pixel point to be interpolated in the edge area; Taking the direction of the maximum edge gradient value in the second pixel to be interpolated as the gradient direction, and taking the direction perpendicular to the gradient direction as the edge line direction; For the second pixel point to be interpolated that is not in the gradient direction and not in the edge line direction, upsampling is performed using a bilinear interpolation algorithm to obtain a corresponding upsampled feature map; For the second pixel to be interpolated located in the gradient direction, up-sampling is performed using a gradient value-weighted Lanczos window function to obtain a corresponding up-sampled feature map; For the second pixel point to be interpolated located in the direction of the edge line, up-sampling is performed using a high-order Lanczos window function to obtain a corresponding up-sampled feature map.

[0010] In some embodiments, performing regression prediction on the feature map of the binocular image to obtain a stem bounding box and a leaf bounding box of the binocular image includes: Calling a convolutional neural network to predict the feature map of the binocular image, obtaining the preliminary position and preliminary offset of the representative points of the stems and leaves of the cabbage seedlings in plug trays, and correcting the preliminary position and preliminary offset of the representative points of the stems and leaves through a small regression network to obtain the corrected offset of the representative points of the stems and leaves; Correcting the initial position of the representative point of the stem and leaf portion by the corrected offset to obtain the key points of the stem and leaf portion of the cabbage seedling in the plug tray; The minimum circumscribed rectangle of the key point is determined as the stem bounding box and the leaf bounding box of the binocular image.

[0011] In some embodiments, performing key point matching on the stem bounding box and the leaf bounding box of the binocular image includes: For the left image and the right image in the binocular image, respectively extract candidate points in the stem bounding box and the leaf bounding box; Performing rough matching on the candidate points of the left image and the candidate points of the right image based on binocular vision information to obtain matching points; Determine the Euclidean distance between the matching points of the left image and the matching points of the right image; The matching points whose Euclidean distance is less than a distance threshold are determined as key points in the binocular image.

[0012] In some embodiments, determining the three-dimensional geometric position coordinates of the cabbage seedlings in the plug tray according to the image coordinates of the key points obtained by matching includes: Determine the disparity value of the key point in the binocular image according to the image coordinates of the key point obtained by matching; The three-dimensional geometric position coordinates of the cabbage plug seedlings are determined according to the image coordinates and the parallax value.

[0013] In some embodiments, after acquiring the binocular image of the collected cabbage plug seedlings, the method further comprises: Calling the target detection model to detect the binocular image to obtain a stem bounding box and a leaf bounding box of the binocular image; Determine the leaf uprightness and maximum leaf span of the cabbage plug seedlings based on the stem bounding box and the leaf bounding box of the binocular image; The target detection model includes a feature extraction network and a small multi-task network. The target detection model is trained by binocular image samples of cabbage seedlings in plug trays. The binocular image samples carry real stem and leaf boundary boxes. The training process of the target detection model includes: Inputting the binocular image sample of the cabbage seedlings in plug trays into the feature extraction network for feature extraction processing to obtain an output feature map of the binocular image sample; Inputting the output feature map into the small multi-task network for regression prediction processing to obtain predicted stem and leaf boundary boxes and classification information of the stem and leaf parts of cabbage seedlings in plug trays; Constructing a bounding box regression loss according to the true stem-and-leaf bounding box and the predicted stem-and-leaf bounding box; Extracting real key points within the real stem and leaf boundary box in the binocular image sample, and constructing stem and leaf information classification loss based on the real key points and the classification information of the cabbage plug seedling stem and leaf parts; The final loss of the target detection model is constructed according to the bounding box regression loss and the stem-leaf information classification loss, and the final loss is back-propagated in the feature extraction network and the small multi-task network to update the parameters of the target detection model.

[0014] The present invention also provides a mechanized cabbage plug tray seedling leaf uprightness and maximum leaf extension identification device, the device comprising the following modules: An acquisition module, used for acquiring binocular images of the collected cabbage seedlings in plug trays; An extraction module, used for calling a Swin Transformer model based on an RFP structure to perform feature extraction on the binocular image to obtain a feature map of the binocular image; A prediction module, used for performing regression prediction on the feature map of the binocular image to obtain a stem bounding box and a leaf bounding box of the binocular image; A matching module, used for matching key points of the stem boundary box and the leaf boundary box of the binocular image, and determining the three-dimensional geometric position coordinates of the cabbage plug seedling according to the image coordinates of the key points obtained by matching; The determination module is used to determine the leaf uprightness and maximum leaf span of the cabbage plug seedlings according to the three-dimensional geometric position coordinates.

[0015] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method for identifying the leaf uprightness and maximum leaf span of mechanized cabbage plug tray seedlings as described in any one of the above methods is implemented.

[0016] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for identifying the leaf uprightness and maximum leaf span of cabbage seedlings in a mechanized plug tray as described in any one of the above is implemented.

[0017] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned methods for identifying the uprightness and maximum leaf span of leaves of mechanized cabbage plug tray seedlings.

[0018] The method for identifying the leaf uprightness and maximum leaf extension of cabbage seedlings suitable for mechanization provided by the present invention is based on the binocular image of the cabbage seedlings, and the Swin Transformer model based on the RFP structure is called to extract the feature map, and regression prediction is performed to obtain the stem boundary box and the leaf boundary box of the binocular image. Next, the boundary box is matched with key points, so as to determine the three-dimensional geometric position coordinates of the cabbage seedlings according to the image coordinates of the key points obtained by matching, and finally the leaf uprightness and maximum leaf extension of the cabbage seedlings are calculated. Therefore, on the one hand, through the regression prediction of the model, the stem and leaf parts of the cabbage seedlings are detected, and the detection accuracy is improved. On the other hand, by shooting the image of the seedlings, the leaf uprightness and the maximum leaf extension can be detected and determined in real time, forming a discrimination standard for the mechanization of the seedlings, which can be applied to the actual production of mechanized transplanting agronomy. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced one by one below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0020] Figure 1 It is a schematic diagram of the process of the mechanized cabbage plug tray seedling leaf uprightness and maximum leaf span identification method provided by the present invention.

[0021] Figure 2 It is a schematic diagram of the backbone network of the Swin Transformer model provided by the present invention.

[0022] Figure 3 It is a schematic diagram of extracting features of the stage model provided by the present invention.

[0023] Figure 4 It is a schematic diagram of binocular camera imaging provided by the present invention.

[0024] Figure 5 It is a schematic diagram of determining the uprightness and maximum leaf span of a leaf provided by the present invention.

[0025] Figure 6 It is a structural schematic diagram of a mechanized cabbage plug tray seedling leaf uprightness and maximum leaf span identification device provided by the present invention.

[0026] Figure 7 It is a schematic diagram of the physical structure of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0027] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0028] The following describes the mechanized cabbage plug seedling leaf uprightness and maximum leaf span identification method of the present invention in conjunction with the accompanying drawings. Figure 1 FIG. 1 is a flow chart of a method for identifying leaf uprightness and maximum leaf span of mechanized cabbage plug seedlings provided by the present invention, as shown in FIG. Figure 1 As shown, the method includes the following steps 101 to 105.

[0029] Step 101: Acquire a binocular image of the collected cabbage seedlings in a plug tray.

[0030] In an actual production environment where cabbage seedlings are mechanized for transplanting, some cabbage seedlings are randomly selected from the same batch of cultivated seedlings. Then, the embodiment of the present invention calculates the leaf uprightness and maximum leaf span of the cabbage seedlings by collecting binocular images of the cabbage seedlings in real time.

[0031] In the light box, the cabbage seedlings are clamped with a cabbage transplanting device. When the cabbage seedlings are relatively stable (the substrate of the cabbage seedlings is parallel to the ground), the binocular image data of the cabbage seedlings is collected by a binocular camera and a vernier caliper. After being calibrated using the Zhang calibration method, the binocular camera is horizontally placed 50 cm away from the measured cabbage seedlings, and the cabbage seedlings clamped by the cabbage transplanting device are used as the image center to collect the binocular image. A computer is then connected to the ZED 2i binocular camera to obtain the binocular image of the clamped cabbage seedlings at this time, and the binocular image is stored in jpg format. The binocular image generally includes a left image as the target image and a right image used for matching.

[0032] In order to relieve the hardware pressure of subsequent model processing, the collected binocular images were scaled to 550 pixels × 550 pixels three-channel images using the principle of bilinear interpolation, and data expansion was performed using image augmentation. In addition, combined with the irregular growth characteristics of plants such as cabbage and the characteristics of changes in sunlight intensity in the greenhouse, the binocular images were adjusted by mirror rotation and adjusting the brightness and saturation of the picture to facilitate subsequent model processing.

[0033] Step 102: Call the Swin Transformer model based on the RFP structure to extract features from the binocular image to obtain a feature map of the binocular image.

[0034] The traditional visual Transformer model has achieved good results in the fields of defect detection and face recognition. However, when applied to actual vegetable seedling production, it is affected by complex background, uneven lighting and small data size, and there is a problem of incompatibility between the application scenario and the model. Based on this, the embodiment of the present invention uses the Swin Transformer model based on the Feature Pyramid Networks (RFP) structure to extract features from the binocular image and obtain the feature map of the binocular image. Figure 2 As shown, the backbone network of the Swin Transformer model based on the RFP structure in the embodiment of the present invention is composed of two Swin Transformer models (i.e., the first Swin Transformer model and the second Swin Transformer model), each of which includes 3 layers of top-down stage modules, corresponding to the three stages of feature extraction, and the three stages extract features of different scales of the input image. Each stage module includes a linear embedding layer (Linear Embedding) and a Swin Transformer block (Swin Transformer Block). The Swin Transformer block also includes a W-MSA module, a SW-MSA module, and a multi-layer perceptron (MLP), and a Token Merging (TM) module is inserted between the SW-MSA module and the multi-layer perceptron. In addition, the reason why the Swin Transformer model based on the RFP structure is called here is that a feature pyramid network (RFP) is also designed in the stage module, so the stage module is also called a composite stage module, and the feature pyramid network included in the composite stage module can further extract features from the feature map processed by the stage module. Using the SwinTransformer model based on the RFP structure to extract the feature map of the binocular image can not only reduce the amount of feature map and model calculation, making the network more efficient and faster, but also effectively improve the accuracy of feature extraction. The following is a detailed introduction to the process of extracting feature maps.

[0035] First, the image to be processed is divided into multiple patches. The image to be processed here is the left image or the right image in the binocular image. Both images need to be feature extracted separately. The left image is taken as an example for explanation below.

[0036] like Figure 2As shown in the figure, after the image to be processed (left image Image) is input into the Swin Transformer model based on the RFP structure, the Patch Partition module in the first Swin Transformer model divides the input image to be processed (size is H*W) into non-overlapping patches (patch), and the width, height and channel size are 4*4*3 respectively.

[0037] Next, the patch is input into the three-layer first composite stage module based on the RFP structure in the first Swin Transformer model, and features of the patch are extracted through the first composite stage module to obtain the first feature map corresponding to the patch.

[0038] Specifically, if Figure 3 As shown in the figure, after the patch is input into the first composite stage module, the patch is embedded through the linear embedding layer (LN) therein to obtain the corresponding embedded patch sequence. The embedding process is to directly process the patch into a single-channel embedding feature at one time, thereby forming an embedded patch sequence, namely, an embedded patch, where the single-channel embedding features are all one-dimensional feature vectors.

[0039] After the embedding process is completed, the one-dimensional embedded patch sequence is input into the W-MSA module and the SW-MSA module (corresponding to Figure 3 Improved SW-MSA in ). The embedded patch sequence is reshaped into a two-dimensional feature map of size H*W by the W-MSA module in the first composite stage module to facilitate feature extraction. The W-MSA module and the SW-MSA module adopt a window processing mechanism. The two-dimensional feature map is first divided by a non-overlapping detection window (window), and then the W-MSA module and the SW-MSA module in the first composite stage module perform grouped self-attention calculations on the two-dimensional feature maps divided in the detection window to obtain the window representation corresponding to the detection window.

[0040] In the detection window, the corresponding two-dimensional feature map is further divided into sub-windows, and the sub-window size is recorded as M*M. Then, the grouped self-attention calculation is performed for the pixels of the two-dimensional feature map in each sub-window to obtain the window representation corresponding to the detection window. This window representation only represents the local features in the detection window. The above process is repeated until the self-attention calculation of each sub-window is completed, and then the window representation of each sub-window is downsampled to obtain the feature information of the detection window.

[0041] like Figure 3As shown in the figure, in the SW-MSA module, the window representation corresponding to each detection window is globally downsampled and self-attention is calculated, and the calculated result is input into the Token fusion module (i.e., the Token merging module) for feature fusion to obtain the fused representation. Finally, the fused representation is input into the MLP module to obtain the high-level feature map of the patch.

[0042] Here Figure 3 As shown in the figure, the fused representation output by the Token fusion module is input into the linear embedding layer (LN) of the MLP module, and then input into the MLP layer for mapping processing, and finally the high-level feature map of the patch is obtained. This high-level feature map is the output result of the stage module in the first composite stage module, which contains the high-dimensional feature information of the image to be processed.

[0043] Considering that the stems of cabbage seedlings are slender and the leaves of mature seedlings block each other, the conventional feature pyramid structure is difficult to guarantee the detection effect in complex scenes. Therefore, in the embodiment of the present invention, feature pyramid networks (FPN) are added to the backbone network of the Swin Transformer model, and the feature representation ability of the network is enhanced by introducing a recursive feature pyramid structure (RFP).

[0044] The feature pyramid network is used to process the high-level feature map output by the stage module in the first composite stage module. Therefore, after the high-level feature map of the patch is output in the MLP module, the high-level feature map is input into the feature pyramid network included in the first composite stage module for feature extraction to obtain the first feature map corresponding to the patch. However, since the backbone network of the SwinTransformer model is top-down, the method in the prior art is to upsample the high-level feature map and then perform feature fusion. The upsampling method often uses a bilinear interpolation algorithm, which can only uniformly amplify the whole, which is easy to cause loss of high-frequency components of the image and blur the edges of the seedlings to a certain extent. Therefore, an edge-adaptive upsampling method is proposed in the embodiment of the present invention. First, edge-adaptive upsampling is performed on the high-level feature map to obtain multiple upsampled feature maps, and then the upsampled feature maps are fused, which is described in detail below.

[0045] For edge adaptive upsampling, we first traverse each pixel in the high-level feature map and call the Sobel operator to calculate the edge gradient value of each pixel along the four directions of up, down, left, and right (i.e., 0, 45, 90, and 135 degrees), denoted as , , , , and determine the maximum edge gradient value of each pixel, recorded as .

[0046] Next, an edge gradient threshold T can be preset, and the pixel point whose maximum edge gradient value is less than or equal to the preset edge gradient threshold T is determined as the first pixel point to be interpolated in the flat area, and the first pixel point to be interpolated is upsampled in both the horizontal and vertical directions using a bilinear interpolation algorithm to obtain a corresponding upsampled feature map.

[0047] Here, when the maximum edge gradient value of a pixel is less than or equal to the preset edge gradient threshold T, that is, When , it can be determined that this pixel is a pixel to be interpolated in the flat area, so the bilinear interpolation algorithm is directly used to upsample the first pixel to be interpolated in the horizontal and vertical directions to obtain the corresponding upsampled feature map.

[0048] The pixel point whose maximum edge gradient value is greater than the preset edge gradient threshold T is determined as the second pixel point to be interpolated in the edge area. Here, when the maximum edge gradient value of the pixel point is greater than the preset edge gradient threshold T, that is, , then it can be determined that this pixel is the second pixel to be interpolated in the edge area. Next, before upsampling, it is necessary to determine the direction of the second pixel to be interpolated. Here, the direction in which the maximum edge gradient value is obtained in the second pixel to be interpolated is taken as the gradient direction, which is expressed as follows: The direction perpendicular to the gradient direction is taken as the edge line direction, which is expressed as follows: For the second pixel point to be interpolated that is located in the non-gradient direction and the non-edge line direction, a bilinear interpolation algorithm is used for upsampling to obtain a corresponding upsampled feature map.

[0049] For the second pixel to be interpolated in the gradient direction, the gradient-weighted Lanczos window function is used for upsampling to obtain the corresponding upsampled feature map. The upsampled pixel is represented as w(d). The gradient-weighted Lanczos window function is expressed as the following formula (1): (1) In the above formula (1), d is the distance from the pixel to be interpolated to the edge pixel, r is the radius of the interpolation window, is the edge gradient value corresponding to the pixel point.

[0050] For the second pixel to be interpolated located in the direction of the edge line, a high-order Lanczos window function is used for upsampling to obtain the corresponding upsampled feature map. The upsampled pixel is represented as w(d), and the high-order Lanczos window function is expressed as the following formula (2): (2) In the above formula (2), d is the distance from the pixel to be interpolated to the edge pixel, and r is the radius of the interpolation window.

[0051] The embodiment of the present invention divides the pixels by calculating the edge gradient value of each pixel in the high-level feature map, and adopts different upsampling methods for processing, thereby solving the problem that the upsampling method in the prior art often adopts a bilinear interpolation algorithm, which can only uniformly amplify the whole, easily causing the loss of high-frequency components of the image and blurring the edges of the seedlings to a certain extent.

[0052] By performing edge-adaptive upsampling on the pixels in the high-level feature map, multiple upsampled feature maps are obtained, and then the upsampled feature maps are fused. The fused feature maps are input into the feature pyramid network for multiple iterative recursive calculations to obtain the first feature map corresponding to the patch.

[0053] Here, the output feature layer of the feature pyramid network is recorded as ,in, , S represents the number of levels of the backbone network of the first SwinTransformer model (that is, the number of layers of the first composite stage module, here 3 layers), thus It is expressed as the following formula (3) and formula (4): (3) (4) In the above formula (3), Represents the operation function of the top-down feature pyramid network of the backbone network (i.e. the first Swin Transformer model, which will not be described later), Represents the feature map of the i-th layer input in the backbone network, that is, the feature map obtained by fusion of multiple upsampled feature maps. Represents the top-down levels of the backbone network.

[0054] Therefore, the feature pyramid network outputs a set of features denoted as , and after adding feedback connection, the output feature layer of feature pyramid network is recorded as It can be expressed as the following formula (5) and formula (6): (5) (6) The explanation of the parameters in the above formula (5) can be referred to the above formula (3) and will not be repeated here. Represents the top-down level of the backbone network, represents the feature conversion function before connecting the converted feature layer back to the backbone network, so that the feature pyramid network can complete the recursive operation. The feature pyramid network is thus expanded into a sequential network. The formula for iterative calculation of the sequential network is expressed as follows: Formula (7) and Formula (8): (5) (6) Among them, in the above formula, , T represents the number of expansion iterations, which can be set to 2, and the superscript t represents the operation step of the feature pyramid network at the tth iteration. The structure of the feature pyramid network is implemented by including the atrous spatial convolutional pooling pyramid (ASPP) structure, such as Figure 2 ASPP in , so as to perform recursive calculation together.

[0055] After recursive calculation of the feature pyramid, the first feature map corresponding to the patch is finally obtained. Figure 2 As shown, in the first Swin Transformer model, each first composite stage module based on the RFP structure will output a corresponding first feature map, and the first feature map will be input to the second Swin Transformer model accordingly.

[0056] In the embodiment of the present invention, a feature pyramid network is added to the backbone network of the Swin Transformer model, and the feature representation capability of the network is enhanced by introducing a recursive feature pyramid structure. In the complex detection scenario where the stems of cabbage seedlings in plug trays are slender and the leaves of mature seedlings occlude each other, the effect of image feature extraction can also be guaranteed.

[0057] like Figure 2As shown, since the image to be processed (Image) will also be input into the second Swin Transformer model at the same time, the input image to be processed (size is H*W) will also be divided into non-overlapping patches (patch) through the Patch Partition module in the second Swin Transformer model, and the width, height and channel size are 4*4*3 respectively. Next, the patch is input into the 3-layer second composite stage module based on the RFP structure in the second Swin Transformer model. In the specific implementation, the first feature map output by the first Swin Transformer model is also input into the 3-layer second composite stage module based on the RFP structure in the second Swin Transformer model.

[0058] In this way, the second composite stage module extracts features from the patch and the first feature map to obtain a second feature map, which is used as the feature map of the binocular image. In this way, the multi-layer features of the network can be integrated to further enhance the feature extraction capability of the model.

[0059] In the second composite stage module based on the RFP structure, the first feature map and the corresponding patch are input. Since the model structures of the first Swin Transformer model and the second Swin Transformer model are the same, the specific networks of the second composite stage module and the first composite stage module are also similar. The method of extracting features can refer to the first composite stage module, and the specific feature extraction process will not be repeated here.

[0060] For the right image of the binocular image, similar to the left image, the above step 102 is also performed as the image to be processed, and the corresponding second feature map is finally extracted. In this way, the corresponding second feature maps are extracted from both the left image and the right image for subsequent target detection.

[0061] The embodiment of the present invention utilizes the Swin Transformer model based on the RFP structure to extract the features of the binocular image, which can significantly improve the feature extraction capability of the model. Even in a complex detection scenario where the stems of cabbage seedlings in plug trays are slender and the leaves of mature seedlings occlude each other, the effect of image feature processing can be guaranteed, the accuracy of feature extraction can be improved, and the foundation for accurately predicting the stem and leaf boundary boxes in the subsequent process can be laid.

[0062] Step 103: perform regression prediction on the feature map of the binocular image to obtain a stem bounding box and a leaf bounding box of the binocular image.

[0063] After extracting the features of the binocular image, the next step is to detect the target of the stem and leaves. The feature map of the binocular image is regressed and predicted to obtain the stem and leaf bounding boxes of the binocular image. Here, the feature map output by the Swin Transformer model based on the RFP structure is input into the detection head, which is used for target detection, that is, to locate the stem and leaves of the cabbage seedlings in the binocular image. The detection head specifically includes two modules: a convolutional neural network and a small regression network.

[0064] After the feature map of the binocular image is input into the detection head, the convolutional neural network is first called to predict the feature map of the binocular image to obtain the preliminary position and preliminary offset of the representative points of the stems and leaves of the cabbage seedlings in the plug tray. The role of the convolutional neural network is to calculate the representative points that can represent the stems and leaves of the cabbage seedlings in the plug tray according to the characteristic values ​​of the pixels in the feature map. The representative points of the stems and leaves may be the pixel points of the stems and leaves of the cabbage seedlings in the plug tray, and the preliminary offset is the position offset value that may be generated by the predicted preliminary position. Since the preliminary position of the representative points of the stems and leaves may be offset and need to be corrected, a small regression network is constructed here in the embodiment of the present invention. The preliminary position and preliminary offset of the representative points of the stems and leaves are corrected by the small regression network to obtain the corrected offset of the representative points of the stems and leaves.

[0065] The specific structure of the small regression network is as follows: the first layer is a 3x3 depthwise separable convolutional layer with 64 output channels, followed by a ReLU activation function. The second layer is a 1x1 depthwise separable convolutional layer with 128 output channels, followed by a ReLU activation function. The third layer is a 1x1 convolutional layer with n output channels, representing the number of seedling representative points, followed by a fully connected layer that outputs the corrected offset of the representative points.

[0066] The role of the small regression network is to further determine the position of the representative points of the stems and leaves. Therefore, the initial position of the representative points of the stems and leaves is corrected by correcting the offset to obtain the key points of the stems and leaves of the cabbage seedlings. These key points are the key points of the stems and leaves of the cabbage seedlings in the binocular image. Finally, the minimum bounding rectangle of the key points is determined as the stem bounding box and leaf bounding box of the binocular image. The area of ​​the minimum bounding rectangle can just cover all the key points of the stems and leaves, thereby forming the stem bounding box and leaf bounding box of the corresponding binocular image.

[0067] In the embodiment of the present invention, when performing regression prediction on the left image and the right image of the binocular image, a convolutional neural network and a small regression network are used to accurately locate the key points of the stems and leaves of the cabbage plug seedlings, forming a stem boundary box and a leaf boundary box, which is convenient for subsequent matching of the binocular images to determine the image coordinates of the key points, and then calculate the leaf uprightness and the maximum leaf span.

[0068] Step 104 , performing key point matching on the stem boundary box and the leaf boundary box of the binocular image, and determining the three-dimensional geometric position coordinates of the cabbage plug seedlings according to the image coordinates of the key points obtained by matching.

[0069] After target detection is performed in step 103, regression prediction is performed on the left image and the right image of the binocular image to obtain the corresponding stem boundary box and leaf boundary box. However, since the position coordinates of the key points in the binocular image are two-dimensional, the key points may have position differences in the binocular image, so it is necessary to perform stereo matching based on binocular information on the binocular image to screen out the real key points, and then obtain the three-dimensional geometric position information of the seedlings based on the real key points.

[0070] First, for the left image and the right image in the binocular image, candidate points in the stem bounding box and the leaf bounding box are extracted respectively. That is, the key points in the bounding box are determined from the binocular image as candidate points. Then, the candidate points of the left image and the right image are roughly matched based on the binocular vision information to obtain matching points. The rough matching process is realized by using binocular vision information, and finally the different position information of the candidate points in the two images is determined, such as the position coordinates in the horizontal and vertical directions. However, it is difficult for the detection information of the left and right images to be completely consistent. Only rough matching based on the corresponding position information cannot guarantee the matching effect.

[0071] Next, fine matching of matching points is performed. Fine matching is to calculate the Euclidean distance between two matching points in the left and right images. The smaller the Euclidean distance between two matching points, that is, the closer it is to 0, the better the matching effect is, and the more likely these two matching points are the key points of the stems and leaves of the cabbage seedlings. The process of calculating the Euclidean distance by fine matching is as follows (7): (7) In the above formula (7), Indicates the midpoint of the left image The pixel value at . Represented as the midpoint of the right image The pixel value at . The offset range is the size of the right image, that is, x, y.

[0072] After completing the Euclidean distance calculation for fine matching, it is necessary to filter based on the Euclidean distance and determine the matching points whose Euclidean distance is less than the distance threshold as the key points in the binocular image, because the smaller the Euclidean distance, that is, the closer it is to 0, the better the matching effect. Therefore, a distance threshold can be preset to measure the matching effect. When the calculated Euclidean distance is less than this distance threshold, it means that the two matching points corresponding to the left image and the right image have completed the match and can be determined as the key points of the stems and leaves of the cabbage seedlings. On the contrary, when the calculated Euclidean distance is greater than this distance threshold, it means that the two matching points corresponding to the left image and the right image have not been successfully matched, and the matching points in the two images are filtered out.

[0073] The embodiment of the present invention performs key point matching on the stem boundary box and the leaf boundary box of the binocular image, thereby calculating the key points in the binocular image, thereby achieving accurate prediction of the stem and leaf positions in the binocular image.

[0074] Finally, the three-dimensional geometric position coordinates of the cabbage seedlings are determined based on the image coordinates of the key points obtained by matching. Based on the key point matching, multiple key points can be determined, so that the image coordinates of each key point in the binocular image can also be determined one by one. The disparity information of the binocular image can be calculated based on the image coordinates. Next, the three-dimensional position coordinates of the key points in the real space coordinate system, that is, the three-dimensional geometric position coordinates, are calculated based on the disparity information of the binocular image.

[0075] Here, if Figure 4 As shown, the optical centers of the left and right lenses of the binocular camera are and , images of cabbage seedlings in plug trays are taken along the left and right optical axes, respectively, and the baseline B is and The focal length is f. On the left plane corresponding to the left image, the key point coordinates obtained by fine matching are , recorded as Correspondingly, on the right plane corresponding to the right image, the key point coordinates obtained by fine matching are , recorded as , d represents the Euclidean distance calculated according to formula (7).

[0076] and are the coordinates of the binocular camera imaging, that is, the coordinates of the key points on the binocular image. Therefore, it is also necessary to determine the three-dimensional position coordinates P of the key points in the real space coordinate system, recorded as ,in, express Figure 4 The coordinates of the midpoint P in the z-axis direction can be obtained from the geometric relationship of similar triangles as follows: (8) (9) (10) In the above formulas (8) to (10), B is the optical center of the left and right lenses of the binocular camera. and The distance between the two images is , and f is the focal length of the binocular camera.

[0077] According to the equation group composed of formula (8) and (10), the coordinates of point P are finally calculated as follows: (11) (12) (13) In the above formulas (11) to (13), D represents the parallax of the coordinates of point P, which is expressed as the following formula (14): (14) Substituting formula (14) into the equation group composed of formula (11) to formula (13) for calculation, the coordinates corresponding to point P can be obtained , which is the three-dimensional geometric position coordinates of the cabbage seedlings in the plug tray. The corresponding three-dimensional geometric position coordinates can be calculated for each key point P determined by the fine matching.

[0078] The embodiment of the present invention determines the three-dimensional geometric position coordinates of the cabbage seedlings in the plug tray according to the image coordinates of the key points obtained by matching, realizes the conversion of the two-dimensional image coordinates to the three-dimensional space coordinates, and facilitates the final calculation of the leaf uprightness and the maximum leaf span from the three-dimensional space stereo angle.

[0079] Step 105: Determine the leaf uprightness and maximum leaf span of the cabbage plug seedlings according to the three-dimensional geometric position coordinates.

[0080] After the three-dimensional geometric position coordinates of each key point of the cabbage seedling in the plug tray are determined in step 104, the leaf uprightness and the maximum leaf extension of the cabbage seedling in the plug tray are determined respectively according to the three-dimensional geometric position coordinates. By integrating the three-dimensional geometric position coordinates of these key points, these coordinates are all in the real world coordinate system. In the real world coordinate system, the polygonal area range that can encompass these three-dimensional geometric position coordinates can be determined in the image as the plant stem detection frame and the leaf detection frame of the cabbage seedling in the plug tray.

[0081] For example, Figure 5 As shown in the figure, after achieving the target detection of cabbage seedlings in plug trays, the plant stem detection frames are determined respectively ( Figure 5 The upper left corner coordinates of the blue rectangle And the lower right corner coordinates , leaf detection frame ( Figure 5 The upper left corner coordinates of the red rectangle And the lower right corner coordinates , and then calculate the leaf uprightness of cabbage seedlings in plug trays , expressed as the following formula (15): (15) Calculation of the maximum leaf extension of cabbage seedlings in plug trays , expressed as the following formula (16): (16) like Figure 5 As shown, in the mechanized transplanting agronomy scenario, when the seedling end effector clamps the substrate of the cabbage plug seedling, the mechanized cabbage plug seedling leaf uprightness and maximum leaf spread recognition method provided in the embodiment of the present application can be used to calculate the leaf uprightness and maximum leaf spread of the cabbage plug seedling in real time. Then, the transplanting process is adjusted according to the leaf uprightness and maximum leaf spread to avoid problems such as seedling cup blockage, seedling injury, and skewed seedling transplanting, thereby improving the efficiency of mechanized transplanting.

[0082] In the embodiment of the present invention, based on the binocular image of the cabbage seedlings in the plug tray, the SwinTransformer model based on the RFP structure is called to extract the feature map, and regression prediction is performed to obtain the stem boundary box and the leaf boundary box of the binocular image. Next, the boundary box is matched with key points, so as to determine the three-dimensional geometric position coordinates of the cabbage seedlings in the plug tray according to the image coordinates of the key points obtained by matching, and finally the leaf uprightness and maximum leaf span of the cabbage seedlings in the plug tray are calculated. Therefore, on the one hand, through the regression prediction of the model, the stem and leaf parts of the cabbage seedlings in the plug tray are detected, thereby improving the detection accuracy. On the other hand, by taking the image of the plug tray seedlings, the leaf uprightness and the maximum leaf span can be detected and determined in real time, forming a judgment standard for the mechanization of the plug tray seedlings, which can be applied to the actual production of mechanized transplanting agronomy.

[0083] In other embodiments, the calculation of the leaf uprightness and maximum leaf span of cabbage seedlings in plug trays mainly depends on the matching of key points in the binocular image, and the selection of key points can have a great impact on the calculation results, and there are high requirements on the accuracy of the key points, stem bounding boxes and leaf bounding boxes obtained by regression prediction of the detection head in the model.

[0084] Based on this, in order to ensure the calculation accuracy of leaf uprightness and maximum leaf spread, the embodiment of the present invention also trains a target detection model to more accurately predict the key points, stem bounding box and leaf bounding box of the cabbage plug seedlings. After obtaining the collected binocular image of the cabbage plug seedlings, the target detection model is called to detect the binocular image to obtain the stem bounding box and leaf bounding box of the binocular image, and finally the leaf uprightness and maximum leaf spread of the cabbage plug seedlings are determined based on the stem bounding box and leaf bounding box of the binocular image.

[0085] The determination process here can be to match the key points of the stem bounding box and the leaf bounding box of the binocular image, determine the three-dimensional geometric position coordinates of the cabbage plug seedlings according to the image coordinates of the matched key points, and finally determine the leaf uprightness and maximum leaf span of the cabbage plug seedlings according to the three-dimensional geometric position coordinates. The determination method is similar to the above steps 104 and 105, and can be referred to each other, and will not be repeated here.

[0086] Specifically, the target detection model includes a feature extraction network and a small multi-task network, wherein the feature extraction network can be the Swin Transformer model based on the RFP structure used in step 102, or other improved Swin Transformer models, and its main function is to extract the feature map of the binocular image. The small multi-task network includes a bounding box regression network and a multi-task regression network, and its main function is to predict the key points of the stem and leaves, the stem bounding box, and the leaf bounding box based on the extracted feature map. The training process of the target detection model is introduced below.

[0087] First, we can use a binocular camera to collect binocular image samples of cabbage seedlings in plug trays, and use the binocular image samples of cabbage seedlings in plug trays to train the target detection model. During training, the binocular image samples are divided into five parts, divided into training sets and test sets in a ratio of 4:1, and the model training is carried out using a five-fold cross-validation method. In each binocular image sample, the corresponding real key points and the real stem and leaf bounding boxes (that is, the real stem bounding box and the real leaf bounding box) are recorded as training labels, which are used to construct loss values ​​for parameter optimization.

[0088] At the beginning of training, the binocular image samples of cabbage seedlings in plug trays are input into the feature extraction network for feature extraction processing to obtain the output feature map F of the binocular image samples. This is also the process of forward propagation, and the process of feature extraction processing is similar to the above step 102, which will not be repeated here.

[0089] Next, the output feature map F obtained by the feature extraction network is input into the small multi-task network for regression prediction processing to obtain the predicted stem and leaf bounding box and the classification information of the stem and leaf of the cabbage seedlings in the tray. The small multi-task network is divided into a bounding box regression network and a multi-task network. The bounding box regression network first preliminarily predicts the corresponding preliminary stem and leaf bounding box based on the output feature map F obtained by the feature extraction network. The prediction process can refer to the process of performing regression prediction on the feature map of the binocular image in the above step 103, which will not be repeated here.

[0090] The multi-task network is used to further predict the output feature map F to correct the stem and leaf classification information of the preliminary stem and leaf bounding box. During the multi-task network processing, some initial representative points will be customized in the output feature map F, denoted as .

[0091] The multi-task network performs position correction tasks and classification tasks for these initial representative points. The first three layers of the network structure of the multi-task network are shared, and then divided into two network branches: the position correction branch and the classification branch. The position correction branch is used to perform regression tasks, process the representative point position correction tasks, and output the predicted stem and leaf bounding boxes (that is, the predicted stem bounding box and the predicted leaf bounding box). The classification branch is used to perform classification tasks and determine the classification information of the stem and leaf.

[0092] In the three-layer shared network layer, the first layer is a depth-separable convolutional layer that outputs the corresponding feature information , expressed as the following formula (17): (17) In the above formula (17), represents the processing function of the first depth-wise separable convolutional layer, represents the activation function of the first depth-wise separable convolutional layer, represents the output feature map F obtained by the feature extraction network, Represents the weight parameters of the first depthwise separable convolutional layer.

[0093] The second layer is also a depth-separable convolutional layer, which inputs the feature information output by the first depth-separable convolutional layer. , output corresponding feature information , expressed as the following formula (18): (18) In the above formula (18), represents the processing function of the second depth-wise separable convolutional layer, represents the activation function of the second depth-wise separable convolutional layer, Represents the feature information output by the first depth-separable convolutional layer, Represents the weight parameters of the second depthwise separable convolutional layer.

[0094] The third layer of the shared network layer is also a depth-wise separable convolutional layer, which inputs the feature information output by the second depth-wise separable convolutional layer. , output corresponding feature information , expressed as the following formula (19): (19) In the above formula (19), represents the processing function of the third depth-wise separable convolutional layer, represents the activation function of the third depth-wise separable convolutional layer, Represents the feature information output by the second depth-separable convolutional layer, Represents the weight parameters of the third depth-wise separable convolutional layer.

[0095] Next, it is divided into two branches, and the feature information output by the third layer of depth-separable convolutional layer is Input into two branches respectively, the position correction branch is a fully connected layer, expressed as , used to perform regression tasks, handle representative point position correction tasks, and output the correction values ​​of the initial key points , expressed as the following formula (20): (20) The classification branch is also a fully connected layer, expressed as , the feature information output by the third layer of depth-separable convolutional layer Mapping to obtain classification information of the stem and leaf parts of cabbage seedlings , expressed as the following formula (21): (twenty one) Next, use the correction value For the initial representative point Specifically, each initial representative point is corrected by the corresponding correction value Correction is performed, which is expressed as the following formula (22): (twenty two) After each initial representative point is corrected, the set of corrected representative points can be obtained. , as the predicted key points of the final output of the small regression network. Then the minimum bounding rectangle of these predicted key points is determined by , denoted as B, expressed as the following formula (23): (twenty three) Next, the standard bounding box regression technique can be used to further correct the minimum bounding rectangle B to obtain the predicted stem and leaf bounding box. , expressed as the following formula (24): (twenty four) In the above formula (24), It means that the bounding box regression network first preliminarily predicts the corresponding preliminary stem-leaf bounding box based on the output feature map F obtained by the feature extraction network.

[0096] Through the above steps, the output feature map F is input into the small multi-task network for regression prediction processing to obtain the predicted stem and leaf bounding box And classification information of the stems and leaves of cabbage seedlings in trays Next, we use the real stem and leaf bounding box and the predicted stem and leaf bounding box to calculate the , construct the bounding box regression loss, where the prediction results of the small multi-task network are compared with the true label to construct the corresponding bounding box regression loss .

[0097] The classification information of the stem and leaf parts of cabbage seedlings in plug trays here specifically represents the feature information of the stem key points and the leaf key points, that is, the classification result of judging whether each pixel point in the binocular image sample belongs to the stem key point or the leaf key point. By comparing this classification result with the real result of the stem and leaf key points in the binocular image sample, the classification loss of stem and leaf information is constructed. Specifically, the real key points in the real stem and leaf boundary box in the binocular image sample can be first extracted as the real results of the stem and leaf key points, and then based on the real key points and the classification information of the stem and leaf parts of the cabbage plug seedlings, the stem and leaf information classification loss is constructed, and the real key points in the real stem and leaf boundary box are compared with the classification results of the stem key points or the leaf key points to construct the stem and leaf information classification loss. The construction method can be to construct a cross entropy loss or a mean square error loss.

[0098] Finally, according to the bounding box regression loss And the stem and leaf information classification loss The final loss L of the target detection model is constructed as formula (25): (25) In the above formula (25), represents the bounding box regression loss And the stem and leaf information classification loss The balance weight is set according to actual needs, and is generally between 0 and 1.

[0099] In each round of training iteration, the final loss L is first calculated through the above steps, and then the final loss L is back-propagated in the feature extraction network and the small multi-task network to update the parameters of the target detection model. As mentioned above, the binocular image samples are divided into five parts, divided into training sets and test sets in a ratio of 4:1, and the model is trained using a five-fold cross-validation method. The sample batch size (Batch size) is set to 4. The number of iterative training (Epoch) is set to 1000 times, and the stochastic gradient descent optimization algorithm (SGD) is used to optimize the model. The initial learning rate is set to 0.01. When the loss function begins to converge or reaches 1000 iterations of training, stop training.

[0100] The trained object detection model can directly make predictions on the binocular images of cabbage seedlings in plug trays taken with binocular images, output the stem bounding box and the leaf bounding box, and then determine the leaf uprightness and maximum leaf span of the cabbage seedlings in plug trays.

[0101] The embodiment of the present invention trains a target detection model and uses the trained target detection model to perform stem and leaf bounding box detection and stem and leaf key point classification, so as to more accurately predict the key points, stem bounding box and leaf bounding box of the cabbage plug seedlings, and further ensure the prediction accuracy of the leaf uprightness and maximum leaf span of the cabbage plug seedlings.

[0102] The mechanized cabbage plug tray seedling leaf uprightness and maximum leaf extension recognition device provided by the present invention is described below. The mechanized cabbage plug tray seedling leaf uprightness and maximum leaf extension recognition device described below and the mechanized cabbage plug tray seedling leaf uprightness and maximum leaf extension recognition method described above can correspond to each other.

[0103] like Figure 6As shown, the mechanized cabbage plug seedling leaf uprightness and maximum leaf extension recognition device includes: an acquisition module 601, an extraction module 602, a prediction module 603, a matching module 604, and a determination module 605. Specifically, the acquisition module 601 is used to acquire the collected binocular image of the cabbage plug seedling; the extraction module 602 is used to call the SwinTransformer model based on the RFP structure to extract the features of the binocular image to obtain the feature map of the binocular image; the prediction module 603 is used to perform regression prediction on the feature map of the binocular image to obtain the stem bounding box and the leaf bounding box of the binocular image; the matching module 604 is used to match the key points of the stem bounding box and the leaf bounding box of the binocular image, and determine the three-dimensional geometric position coordinates of the cabbage plug seedling according to the image coordinates of the key points obtained by matching; the determination module 605 is used to determine the leaf uprightness and the maximum leaf extension of the cabbage plug seedling according to the three-dimensional geometric position coordinates.

[0104] It should be noted that the beneficial effects of the mechanized cabbage plug tray seedling leaf uprightness and maximum leaf spread identification device here and the mechanized cabbage plug tray seedling leaf uprightness and maximum leaf spread identification method mentioned above can correspond to each other, so the beneficial effects of the mechanized cabbage plug tray seedling leaf uprightness and maximum leaf spread identification device are not repeated here.

[0105] Figure 7 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 7 As shown, the electronic device may include: a processor (processor) 710 , a communication interface (Communications Interface) 720 , a memory (memory) 730 and a communication bus 740 , wherein the processor 710 , the communication interface 720 , and the memory 730 communicate with each other through the communication bus 740 . The processor 710 can call the logic instructions in the memory 730 to execute the mechanized cabbage plug seedling leaf uprightness and maximum leaf span recognition method, which includes: obtaining a collected binocular image of the cabbage plug seedling; calling the Swin Transformer model based on the RFP structure to extract features from the binocular image to obtain a feature map of the binocular image; performing regression prediction on the feature map of the binocular image to obtain a stem bounding box and a leaf bounding box of the binocular image; performing key point matching on the stem bounding box and the leaf bounding box of the binocular image, and determining the three-dimensional geometric position coordinates of the cabbage plug seedling according to the image coordinates of the matched key points; and determining the leaf uprightness and maximum leaf span of the cabbage plug seedling according to the three-dimensional geometric position coordinates.

[0106] In addition, the logic instructions in the above-mentioned memory 730 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0107] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the mechanized cabbage plug tray seedling leaf uprightness and maximum leaf span recognition method provided by the above methods, the method comprising: obtaining a collected binocular image of the cabbage plug tray seedling; calling a SwinTransformer model based on the RFP structure to perform feature extraction on the binocular image to obtain a feature map of the binocular image; performing regression prediction on the feature map of the binocular image to obtain a stem bounding box and a leaf bounding box of the binocular image; performing key point matching on the stem bounding box and the leaf bounding box of the binocular image, and determining the three-dimensional geometric position coordinates of the cabbage plug tray seedling according to the image coordinates of the key points obtained by matching; and determining the leaf uprightness and maximum leaf span of the cabbage plug tray seedling according to the three-dimensional geometric position coordinates.

[0108] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the mechanized cabbage plug tray seedling leaf uprightness and maximum leaf span recognition method provided by the above-mentioned methods, the method comprising: obtaining a collected binocular image of the cabbage plug tray seedling; calling a Swin Transformer model based on the RFP structure to perform feature extraction on the binocular image to obtain a feature map of the binocular image; performing regression prediction on the feature map of the binocular image to obtain a stem bounding box and a leaf bounding box of the binocular image; performing key point matching on the stem bounding box and the leaf bounding box of the binocular image, and determining the three-dimensional geometric position coordinates of the cabbage plug tray seedling according to the image coordinates of the matched key points; and determining the leaf uprightness and maximum leaf span of the cabbage plug tray seedling according to the three-dimensional geometric position coordinates.

[0109] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0110] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0111] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for identifying leaf uprightness and maximum leaf span of cabbage seedlings in mechanized trays, characterized in that: The method comprises: Obtain binocular images of collected cabbage plug seedlings; Calling a Swin Transformer model based on an RFP structure to perform feature extraction on the binocular image to obtain a feature map of the binocular image; Performing regression prediction on the feature map of the binocular image to obtain a stem bounding box and a leaf bounding box of the binocular image; Key point matching is performed on the stem boundary frame and the leaf boundary frame of the binocular image, and the three-dimensional geometric position coordinates of the cabbage plug seedlings are determined according to the image coordinates of the key points obtained by matching; The leaf uprightness and maximum leaf span of the cabbage plug seedlings are determined according to the three-dimensional geometric position coordinates.

2. The method for identifying leaf uprightness and maximum leaf span of mechanized cabbage plug seedlings according to claim 1, characterized in that: The calling of the Swin Transformer model based on the RFP structure to extract features from the binocular image to obtain a feature map of the binocular image includes: Divide the image to be processed into a plurality of patches, input the patches into a first composite stage module based on a three-layer RFP structure in a first Swin Transformer model, extract features of the patches through the first composite stage module, and obtain a first feature map corresponding to the patches, wherein the image to be processed is a left image or a right image in the binocular image; Input the patch into a second composite stage module based on a 3-layer RFP structure in a second Swin Transformer model; Inputting the first feature map into a three-layer second composite stage module based on the RFP structure in a second Swin Transformer model respectively, and extracting features from the patch and the first feature map through the second composite stage module to obtain a second feature map; The second feature map is used as the feature map of the binocular image.

3. The method for identifying leaf uprightness and maximum leaf span of mechanized cabbage plug seedlings according to claim 2, characterized in that: The step of extracting features from the patch by using the first composite stage module to obtain a first feature map corresponding to the patch includes: Embedding the patch to obtain a corresponding embedded patch sequence, and reshaping the embedded patch sequence into a two-dimensional feature map through the W-MSA module in the first composite stage module; The two-dimensional feature map is divided by non-overlapping detection windows, and the two-dimensional feature map divided in the detection window is grouped and self-attention is calculated by the W-MSA module and the SW-MSA module in the first composite stage module to obtain a window representation corresponding to the detection window; Perform global downsampling self-attention calculation on the window representation corresponding to each detection window, and input the calculated result into the Token fusion module for feature fusion to obtain the fused representation; Inputting the fused representation into an MLP module to obtain a high-level feature map of the patch; The high-level feature map is input into a feature pyramid network included in the first composite stage module for feature extraction to obtain a first feature map corresponding to the patch.

4. The method for identifying leaf uprightness and maximum leaf span of mechanized cabbage plug seedlings according to claim 3, characterized in that: The step of inputting the high-level feature map into a feature pyramid network included in the first composite stage module for feature extraction to obtain a first feature map corresponding to the patch includes: Performing edge-adaptive upsampling on the high-level feature map to obtain multiple upsampled feature maps, and fusing the upsampled feature maps; The fused feature map is input into a feature pyramid network for multiple iterative recursive calculations to obtain a first feature map corresponding to the patch.

5. The method for identifying leaf uprightness and maximum leaf span of mechanized cabbage plug seedlings according to claim 4, characterized in that: The edge-adaptively upsampling the high-level feature map to obtain a plurality of upsampled feature maps includes: Traversing each pixel point in the high-level feature map, and calling the Sobel operator to calculate the edge gradient value of each pixel point along the four directions of up, down, left and right, and determining the maximum edge gradient value of each pixel point; Determine a pixel point whose maximum edge gradient value is less than or equal to a preset edge gradient threshold as a first pixel point to be interpolated in the flat area, and upsample the first pixel point to be interpolated in both horizontal and vertical directions using a bilinear interpolation algorithm to obtain a corresponding upsampled feature map; Determine a pixel point whose maximum edge gradient value is greater than a preset edge gradient threshold as a second pixel point to be interpolated in the edge area; Taking the direction of the maximum edge gradient value in the second pixel to be interpolated as the gradient direction, and taking the direction perpendicular to the gradient direction as the edge line direction; For the second pixel point to be interpolated that is not in the gradient direction and not in the edge line direction, upsampling is performed using a bilinear interpolation algorithm to obtain a corresponding upsampled feature map; For the second pixel to be interpolated located in the gradient direction, up-sampling is performed using a gradient value-weighted Lanczos window function to obtain a corresponding up-sampled feature map; For the second pixel point to be interpolated located in the direction of the edge line, up-sampling is performed using a high-order Lanczos window function to obtain a corresponding up-sampled feature map.

6. The method for identifying leaf uprightness and maximum leaf span of mechanized cabbage plug seedlings according to claim 1, characterized in that: The step of performing regression prediction on the feature map of the binocular image to obtain a stem boundary box and a leaf boundary box of the binocular image includes: Calling a convolutional neural network to predict the feature map of the binocular image, obtaining the preliminary position and preliminary offset of the representative points of the stems and leaves of the cabbage seedlings in plug trays, and correcting the preliminary position and preliminary offset of the representative points of the stems and leaves through a small regression network to obtain the corrected offset of the representative points of the stems and leaves; Correcting the initial position of the representative point of the stem and leaf portion by the corrected offset to obtain the key points of the stem and leaf portion of the cabbage seedling in the plug tray; The minimum circumscribed rectangle of the key point is determined as the stem bounding box and the leaf bounding box of the binocular image.

7. The method for identifying leaf uprightness and maximum leaf span of mechanized cabbage plug seedlings according to claim 1, characterized in that: The step of performing key point matching on the stem bounding box and the leaf bounding box of the binocular image includes: For the left image and the right image in the binocular image, respectively extract candidate points in the stem bounding box and the leaf bounding box; Performing rough matching on the candidate points of the left image and the candidate points of the right image based on binocular vision information to obtain matching points; Determine the Euclidean distance between the matching points of the left image and the matching points of the right image; The matching points whose Euclidean distance is less than a distance threshold are determined as key points in the binocular image.

8. The method for identifying leaf uprightness and maximum leaf span of mechanized cabbage plug seedlings according to claim 1, characterized in that: The method of determining the three-dimensional geometric position coordinates of the cabbage plug seedlings according to the image coordinates of the key points obtained by matching includes: Determine the disparity value of the key point in the binocular image according to the image coordinates of the key point obtained by matching; The three-dimensional geometric position coordinates of the cabbage plug seedlings are determined according to the image coordinates and the parallax value.

9. The method for identifying leaf uprightness and maximum leaf span of mechanized cabbage plug seedlings according to claim 1, characterized in that: After acquiring the binocular image of the collected cabbage plug seedlings, the method further includes: Calling the target detection model to detect the binocular image to obtain a stem bounding box and a leaf bounding box of the binocular image; Determine the leaf uprightness and maximum leaf span of the cabbage plug seedlings based on the stem bounding box and the leaf bounding box of the binocular image; The target detection model includes a feature extraction network and a small multi-task network. The target detection model is trained by binocular image samples of cabbage seedlings in plug trays. The binocular image samples carry real stem and leaf boundary boxes. The training process of the target detection model includes: Inputting the binocular image sample of the cabbage seedlings in plug trays into the feature extraction network for feature extraction processing to obtain an output feature map of the binocular image sample; Inputting the output feature map into the small multi-task network for regression prediction processing to obtain predicted stem and leaf boundary boxes and classification information of the stem and leaf parts of cabbage seedlings in plug trays; Constructing a bounding box regression loss according to the true stem-and-leaf bounding box and the predicted stem-and-leaf bounding box; Extracting real key points within the real stem and leaf boundary box in the binocular image sample, and constructing stem and leaf information classification loss based on the real key points and the classification information of the cabbage plug seedling stem and leaf parts; The final loss of the target detection model is constructed according to the bounding box regression loss and the stem-leaf information classification loss, and the final loss is back-propagated in the feature extraction network and the small multi-task network to update the parameters of the target detection model.

10. A mechanized cabbage plug seedling leaf uprightness and maximum leaf extension identification device, characterized in that: The device comprises: An acquisition module, used for acquiring binocular images of the collected cabbage seedlings in plug trays; An extraction module, used for calling a Swin Transformer model based on an RFP structure to perform feature extraction on the binocular image to obtain a feature map of the binocular image; A prediction module, used for performing regression prediction on the feature map of the binocular image to obtain a stem bounding box and a leaf bounding box of the binocular image; A matching module, used for matching key points of the stem boundary box and the leaf boundary box of the binocular image, and determining the three-dimensional geometric position coordinates of the cabbage plug seedling according to the image coordinates of the key points obtained by matching; The determination module is used to determine the leaf uprightness and maximum leaf span of the cabbage plug seedlings according to the three-dimensional geometric position coordinates.

Citation Information

Patent Citations

  • Underwater weak and small target detection method based on Swin Transform

    CN117710801A

  • Plant multi-organ CT image phenotype analysis method based on label efficient learning

    CN118154555A

  • Using scene dependent object queries to generate bounding boxes

    WO2024081665A1