Field corn row detection method, device and equipment based on line Anchor and medium

Through the field corn row detection method based on line Anchor, using the characteristic pyramid network and the multi-head self-attention mechanism, the problem of corn row detection in the existing technology is not robust in complex environments, and higher detection accuracy and real-time performance are achieved.

CN120219961APending Publication Date: 2025-06-27SOUTH CHINA AGRICULTURAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510288300.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing corn row detection algorithm is susceptible to the influence of weeds and outliers in subsequent fitting straight lines, and it is difficult to obtain reliable results in complex field environments, which has the problem of poor robustness.

Method used

The field corn row detection method based on line Anchor is used to extract the corn field images through the feature pyramid network, generate the first line Anchor, and perform multi-head self-attention operations in the ROI alignment module and context enhancement module to generate an attention matrix, and finally regress through the preset loss function to determine the corn row detection result.

Benefits of technology

It significantly improves the accuracy and robustness of corn row detection, avoids detection errors caused by seedling shortage, reduces the impact of weeds and outliers on fitting, and improves the reliability and real-timeness of the detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219961A_ABST
    Figure CN120219961A_ABST
Patent Text Reader

Abstract

The invention relates to a field corn row detection method, device and equipment based on line Anchor and a medium, and the method comprises the steps: carrying out the multi-head self-attention operation of a horizontal compression matrix and a vertical compression matrix, so as to generate a first feature matrix and a second feature matrix, performing broadcasting operation on the first feature matrix and the second feature matrix, and adding the first feature matrix and the second feature matrix to generate a third feature matrix; the query matrix, the key matrix and the value matrix are spliced in a context enhancement module and then subjected to convolution operation to generate a fourth feature matrix, and the third feature matrix is subjected to convolution and then multiplied by the fourth feature matrix to determine an attention matrix; and adopting a preset loss function to regress seedling strip lines in the to-be-detected corn field image as a whole according to the first line Anchor and the attention matrix so as to determine a corn row detection result. According to the invention, the method can adaptively focus on the key area of the corn row, thereby improving the precise recognition of the corn row.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of agricultural production, and particularly to a method for detecting corn rows in the field based on line Anchor, a corresponding device, an electronic device, and a computer-readable storage medium. Background Art

[0002] Fine management needs to be achieved in agricultural production. Traditional manual farmland management has low efficiency, and using unmanned mechanical operations can save costs. Field management is an important link in corn production operations. Precise field management can reduce the losses caused by weed damage to corn production and improve the yield and quality of corn. During the process of unmanned mechanical operations, due to environmental factors such as errors in tractor navigation, irregular terrain, and complex crop distribution, it is of great significance for the actuator to accurately align with the rows to ensure the weeding rate and seedling injury rate. To achieve accurate alignment of the actuator with the rows, it is necessary to first detect the corn rows. Commonly used detection methods include: lidar detection and vision detection. Vision detection is widely used because of its advantages such as low cost, flexible operation, and rich information obtained.

[0003] Currently, the commonly used method at home and abroad is the machine vision navigation method, which usually includes two steps: identifying the corn rows and fitting the center line of the corn rows. Identifying the corn rows is mainly based on the crop feature method. The images of individual corn plants are segmented and clustered using the color and geometric shape features of the corn, and then the Hough transform method or the least squares method is used to fit the corn seedling line of the crop rows. The above-mentioned machine vision method has a simple algorithm and a fast detection speed, but it is easily affected by environmental factors, such as changes in field light and noise such as weeds, resulting in errors when fitting the center line of the corn. In recent years, scholars have used deep learning methods for the identification of corn row lines, effectively improving the detection accuracy. However, in these studies, the detection of corn rows is mostly still achieved based on the identification of individual corn plants. The missing seedling part will affect the clustering effect, and it is also easily affected by weeds and outlier corns in the subsequent step of fitting a straight line. Therefore, in a complex field environment, these methods are difficult to obtain reliable results, have the problem of poor robustness, and such algorithms have poor real-time performance and slow detection speed. In summary, designing a scheme that is compatible with the detection speed and detection accuracy is of great significance for improving the efficiency and quality of unmanned mechanical operations.

[0004] In summary, the corn row detection algorithms in the prior art are easily affected by weeds and outlier corns in the subsequent step of fitting a straight line, and it is difficult to obtain reliable results in a complex field environment, with problems such as poor robustness. The applicant has made corresponding explorations in consideration of solving this problem. Summary of the Invention

[0005] The purpose of this application is to solve the above problems and provide a method for detecting corn rows in the field based on line Anchors, a corresponding device, an electronic device, and a computer-readable storage medium.

[0006] To meet the various purposes of this application, the following technical solutions are adopted in this application:

[0007] A method for detecting corn rows in the field based on line Anchors proposed to meet one of the purposes of this application includes:

[0008] In response to a corn row detection instruction in the field, obtain an image of the corn field to be detected, call a trained corn row detection model in the field until it converges, and use a Feature Pyramid Network (FPN) to extract features from the corn field image to determine a first feature map and generate a first line Anchor;

[0009] Input the first feature map into the ROI Align module in the second detection head network to generate a query matrix, a key matrix, and a value matrix, and perform horizontal compression and vertical compression on the query matrix, the key matrix, and the value matrix respectively in the axial attention module to determine a horizontal compression matrix and a vertical compression matrix;

[0010] Perform multi-head self-attention operations on the horizontal compression matrix and the vertical compression matrix respectively to generate a first feature matrix and a second feature matrix, and perform a broadcast operation on the first feature matrix and the second feature matrix and then add them to generate a third feature matrix;

[0011] In the context enhancement module, splice the query matrix, the key matrix, and the value matrix and then perform a convolution operation to generate a fourth feature matrix. After convolving the third feature matrix and multiplying it with the fourth feature matrix, determine an attention matrix;

[0012] Use a preset loss function to perform regression on the seedling belt line in the corn field image to be detected as a whole based on the first line Anchor and the attention matrix to determine the corn row detection result, and complete the detection of corn rows in the field based on line Anchors.

[0013] Optionally, the corn row detection model in the field is constructed by a Feature Pyramid Network (FPN) and a second detection head network. Among them, the Feature Pyramid Network (FPN) includes three layers of ResNet convolutional neural networks, three convolutional layers, and three first detection head networks. The first detection head network and the second detection head network include an ROI Align module and a loss regression module. The ROI Align module includes an axial attention module and a context enhancement module.

[0014] Optionally, the step of using a Feature Pyramid Network (FPN) to extract features from the corn field image to determine a first feature map and generate a first line of anchors includes:

[0015] Use a Feature Pyramid Network (FPN) to extract features from the corn field image. Through three ResNet convolutional layers, successively extract feature maps containing multi-scale and multi-features, and pass the feature maps containing multi-scale and multi-features through three different convolutional layers in sequence to determine the first feature map. Input the first feature map into the three first detection head networks to determine the first line of anchors.

[0016] Optionally, before the step of calling a trained-to-converge field corn row detection model and using a Feature Pyramid Network (FPN) to extract features from the corn field image to determine a first feature map and generate a first line of anchors, it includes:

[0017] Use a polyline to annotate the corn field image and generate a segmentation map centered on the polyline;

[0018] Use the corn field image and the corresponding generated segmentation label map as a training dataset;

[0019] Input the dataset into the field corn row detection model for training to determine a trained-to-converge field corn row detection model.

[0020] Optionally, the step of obtaining a corn field image to be detected includes:

[0021] Connect a camera using a two-axis self-stabilizing pan-tilt head, mount it on an electric chassis, and cause the electric chassis to move forward slowly at a constant speed to capture a video. Intercept each frame from the captured video as the corn field image to be detected.

[0022] Optionally, the corn field image data includes one or any combination of corn field images at various growth stages, corn field images affected by light, corn field images affected by weeds, corn field images with missing seedlings, and corn field images affected by occlusion.

[0023] Optionally, the query matrix represents the features of the current region of interest; the key matrix represents the encoding of the input features and is used to calculate the correlation with the query region; the value matrix contains the true features and is used to obtain the final feature representation through weighted summation by the attention mechanism.

[0024] A field corn row detection device based on line anchors provided to meet another objective of the present application includes:

[0025] The first-line Anchor determination module is configured to respond to the field corn row detection instruction, obtain the corn field image to be detected, call the trained-to-converge state field corn row detection model, and use the Feature Pyramid Network to extract features from the corn field image to determine the first feature map and generate the first-line Anchor;

[0026] The feature matrix compression module is configured to input the first feature map into the ROI alignment module in the second detection head network to generate a query matrix, a key matrix, and a value matrix, and perform horizontal compression and vertical compression on the query matrix, the key matrix, and the value matrix respectively in the axial attention module to determine the horizontal compression matrix and the vertical compression matrix;

[0027] The feature matrix fusion module is configured to perform multi-head self-attention operations on the horizontal compression matrix and the vertical compression matrix respectively to generate a first feature matrix and a second feature matrix, and add them after performing a broadcast operation on the first feature matrix and the second feature matrix to generate a third feature matrix;

[0028] The attention matrix determination module is configured to splice the query matrix, the key matrix, and the value matrix in the context enhancement module and perform a convolution operation to generate a fourth feature matrix, and multiply the convolved third feature matrix by the fourth feature matrix to determine the attention matrix;

[0029] The corn row detection module is configured to use a preset loss function to perform regression on the seedling band line in the corn field image to be detected as a whole based on the first-line Anchor and the attention matrix to determine the corn row detection result and complete the field corn row detection based on the line Anchor.

[0030] An electronic device provided for another purpose of the present application includes a central processor and a memory. The central processor is configured to call and run a computer program stored in the memory to execute the steps of the field corn row detection method based on the line Anchor of the present application.

[0031] A computer-readable storage medium provided for another purpose of the present application stores a computer program implemented according to the field corn row detection method based on the line Anchor in the form of computer-readable instructions. When the computer program is called and run by the computer, it executes the steps included in the corresponding method.

[0032] Compared with the prior art, for the problems that the existing corn row detection algorithms in the prior art are easily affected by weeds and outlier corns in the subsequent straight line fitting and it is difficult to obtain reliable results in complex field environments, and there are problems such as weak robustness, the present application includes but is not limited to the following beneficial effects:

[0033] First, the present application can significantly improve the accuracy and robustness of maize row detection. Compared with the prior art method of identifying individual maize plants and then fitting a straight line, the present application detects based on an entire maize row, avoiding detection errors caused by missing seedlings. This means that areas with missing seedlings or sparse plants do not affect the detection of the entire maize row, avoiding the problem of error accumulation in individual maize plant detection. In addition, due to the use of a deep learning method based on line anchors, the influence of local noise (such as weeds, outlier maize plants, etc.) on maize row fitting is greatly reduced, ensuring the accuracy and robustness of the detection results.

[0034] Second, the present application uses line anchors for convolution, restricting the convolutional kernel to convolve only along the straight line of the maize row, thereby reducing the computational amount. Compared with the traditional global image convolution method, the line anchor method can more effectively utilize the convolutional operation to calculate only in the required area, which not only improves the processing speed but also greatly reduces the consumption of computing resources. This optimization enables the method to operate efficiently on mobile devices, with stronger real-time performance, and is suitable for rapid applications in complex field environments.

[0035] Third, in the feature extraction stage, the present application adopts a Feature Pyramid Network (FPN), which can effectively extract feature information of different scales from maize field images. The FPN ensures that maize row information at different scales can be extracted by generating the first feature map and line anchors, thereby optimizing the detection accuracy.

[0036] Fourth, the present application designs a context enhancement module. By splicing the query matrix, key matrix, and value matrix, and enhancing the context information of the features through convolutional operations, the model can effectively identify the spatial relationship of the maize row. This module can effectively eliminate the interference caused by complex environments (such as weeds, ground changes, etc.), ensuring that the maize row can be stably detected in field images. After combining with the attention matrix, the system can accurately focus on the seedling belt line of the maize row when performing the regression task, further improving the detection accuracy and robustness.

[0037] Furthermore, in the ROI alignment module and the axial attention module, for the horizontal and vertical compression operations of the query matrix, key matrix, and value matrix, combined with the multi-head self-attention mechanism, the correlation of features and context information are strengthened. Through this operation, the system can adaptively focus on the key areas of the corn rows, thereby improving the accurate recognition of the corn rows. The model of this application uses a relatively lightweight ResNet18 as the backbone network, which enables the entire system to be efficiently deployed on resource-constrained devices (such as mobile devices, drones, agricultural robots, etc.). Compared with traditional complex models, this lightweight design not only makes the deployment more flexible, but also enables real-time detection, reduces latency, and improves the efficiency of unmanned mechanical operations. Description of the Drawings

[0038] The above and / or additional aspects and advantages of this application will become apparent and be readily understood from the following description of the embodiments in conjunction with the drawings, where:

[0039] Figure 1 is a schematic flowchart of the method for detecting corn rows in the field based on line anchors in the embodiments of this application;

[0040] Figure 2 is a schematic diagram of using a polyline to annotate the collected corn pictures in the embodiments of this application;

[0041] Figure 3 is a schematic diagram of the semantic segmentation map of corn seedlings in the embodiments of this application;

[0042] Figure 4 is a schematic diagram of performing multi-scale feature extraction in the embodiments of this application;

[0043] Figure 5 is an exemplary network architecture of the ROI alignment module in the embodiments of this application;

[0044] Figure 6 is a schematic diagram of the detection result of corn rows in the embodiments of this application;

[0045] Figure 7 is a schematic diagram of the detection results of corn rows in different growth stages in the embodiments of this application;

[0046] Figure 8 is a schematic diagram of the detection results of corn rows in the case of weeds and missing seedlings in the embodiments of this application;

[0047] Figure 9 is a schematic block diagram of the device for detecting corn rows in the field based on line anchors in the embodiments of this application;

[0048] Figure 10 is a schematic diagram of the structure of the computer device in the embodiments of this application. Detailed Implementation Modes

[0049] The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary and are only used to explain the present application and should not be construed as limiting the present application.

[0050] Those skilled in the art of the present technology can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present application means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.

[0051] Those skilled in the art of the present technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present application belongs. It should also be understood that terms such as those defined in a general dictionary should be understood as having a meaning consistent with the meaning in the context of the prior art and will not be interpreted with an idealized or overly formal meaning unless specifically defined as herein.

[0052] Those skilled in the art can understand that the "client", "terminal", and "terminal device" used herein include both devices with a wireless signal receiver that only has the ability to receive and no ability to transmit, and devices with receiving and transmitting hardware that can perform two-way communication on a two-way communication link. Such devices can include: cellular or other communication devices such as personal computers, tablet computers, etc., which have a single-line display or a multi-line display or a cellular or other communication device without a multi-line display; PCS (Personal Communications Service), which can combine voice, data processing, fax, and / or data communication capabilities; PDA (Personal Digital Assistant), which can include a radio frequency receiver, pager, Internet / intranet access, web browser, notepad, calendar, and / or GPS (Global Positioning System) receiver; conventional laptop and / or palm-held computers or other devices, which are conventional laptop and / or palm-held computers or other devices with and / or including a radio frequency receiver. The "client", "terminal", and "terminal device" used herein can be portable, transportable, installed in a vehicle (air, sea, and / or land), or suitable for and / or configured to run locally, and / or run in a distributed form at any other location on the earth and / or in space. The "client", "terminal", and "terminal device" used herein can also be a communication terminal, an Internet access terminal, a music / video playback terminal, such as a PDA, MID (Mobile Internet Device), and / or a mobile phone with music / video playback function, or can also be a smart TV, a set-top box, and other devices.

[0053] The hardware referred to by names such as "server", "client", and "service node" in this application is essentially an electronic device with the equivalent capabilities of a personal computer, and is a hardware device with the necessary components disclosed by the von Neumann principle, including a central processing unit (including an arithmetic unit and a controller), a memory, an input device, and an output device. The computer program is stored in its memory, and the central processing unit loads the program stored in the external memory into the memory for execution, executes the instructions in the program, and interacts with the input / output devices to complete specific functions.

[0054] It should be noted that the concept of "server" in this application can similarly be extended to the case of server clusters. According to the network deployment principles understood by those skilled in the art, the servers should be logically divided. Physically, these servers can either be independent of each other but can be called through interfaces, or integrated into a single physical computer or a set of computer clusters. Those skilled in the art should understand this flexibility and should not be restricted by this when implementing the network deployment method of this application.

[0055] One or several technical features of this application, unless expressly specified, can either be deployed on the server for implementation and accessed by the client remotely invoking the online service interface provided by the server, or directly deployed and run on the client for implementation and access.

[0056] The neural network models cited or possibly cited in this application, unless expressly specified, can either be deployed on a remote server and remotely invoked on the client, or deployed on a client capable of handling the capabilities and directly invoked. In some embodiments, when it runs on the client, its corresponding intelligence can be obtained through transfer learning to reduce the requirements for the client's hardware operation resources and avoid excessive occupation of the client's hardware operation resources.

[0057] All kinds of data involved in this application, unless expressly specified, can either be remotely stored on the server or stored on the local terminal device, as long as it is suitable for being invoked by the technical solution of this application.

[0058] Those skilled in the art should be aware that although the various methods of this application are described based on the same concept and thus show commonality with each other, unless otherwise specified, these methods can all be executed independently. Similarly, for each of the embodiments disclosed in this application, they are all proposed based on the same inventive concept. Therefore, for concepts with the same expression, as well as concepts that are only appropriately transformed for convenience although the concept expressions are different, they should be equivalently understood.

[0059] For each of the embodiments to be disclosed in this application, unless expressly indicated that there is a mutually exclusive relationship between them, the relevant technical features involved in each embodiment can be cross-combined to flexibly construct new embodiments, as long as this combination does not deviate from the creative spirit of this application and can meet the requirements in the prior art or solve certain deficiencies in the prior art. Those skilled in the art should be aware of this flexibility.

[0060] Please refer to Figure 1 , in one embodiment of the method for detecting maize rows in the field based on line Anchors of this application, it includes:

[0061] Step S10: In response to the field corn row detection instruction, obtain the corn field image to be detected, call the field corn row detection model that has been trained to the convergence state, and use the Feature Pyramid Network to extract features from the corn field image to determine the first feature map and generate the first line of Anchors.

[0062] The field corn row detection system in the terminal device can respond to the field corn row detection instruction, obtain the corn field image to be detected, call the field corn row detection model that has been trained to the convergence state, and use the Feature Pyramid Network to extract features from the corn field image to determine the first feature map and generate the first line of Anchors. Among them, the field corn row detection model is constructed by the Feature Pyramid Network and the second detection head network. Among them, the Feature Pyramid Network includes three layers of ResNet convolutional neural networks, three convolutional layers, and three first detection head networks. The first detection head network and the second detection head network include an ROI alignment module and a loss regression module. The ROI alignment module includes an axial attention module and a context enhancement module. The corn field image data includes one or any combination of corn field images at various growth stages, corn field images affected by light, corn field images affected by weeds, corn field images with missing seedlings, and corn field images affected by occlusion.

[0063] In some embodiments, the step of obtaining the corn field image to be detected includes:

[0064] Connect a two-axis self-stabilizing pan-tilt to the camera and mount it on the electric chassis, so as to make the electric chassis move forward slowly at a constant speed to shoot a video, and intercept each frame from the shot video as the corn field image to be detected.

[0065] Specifically, use a two-axis self-stabilizing pan-tilt to connect the camera and mount it on the electric chassis, make the electric chassis move forward slowly at a constant speed to shoot a video, and intercept each frame from the shot video as the corn field image data. Among them, the collected corn field image data includes corn field images at different growth stages, affected by light and weeds, as well as corn field images with missing seedlings and occlusion. Before inputting the data set into the training of the field corn row detection model, it is necessary to annotate the collected corn field images. Please refer to Figure 2 and Figure 3 , use a broken line as shown in Figure 2 to pass through the center of the corn row, and annotate the four corn row lines in the middle of the corn field image. Use the annotated part as the position where the corn row is located, and the rest as the background to generate a binary semantic segmentation label map.

[0066] The polyline annotation method based on the entire corn row in this application is more efficient than other annotation methods for single corn plants, which can effectively reduce the manual effort and time spent in the annotation process. Using a polyline instead of a straight line can make the generated label map better fit the curvature of the corn row, thus better covering the corn row. The annotation method of this application starts from a global perspective and, based on the human understanding that corn rows form lines visually, enables the trained deep neural network model to effectively segment the corn row and the background, excluding the interference of weeds and missing seedlings.

[0067] In a further embodiment, the step of using a feature pyramid network to extract features from the corn field image to determine a first feature map and generate a first line of Anchors includes:

[0068] Using a feature pyramid network to extract features from the corn field image, successively extracting feature maps containing multi-scale and multi-features through three ResNet convolutional layers, and passing the feature maps containing multi-scale and multi-features through three different convolutional layers in sequence to determine the first feature map, and inputting the first feature map into the three first detection head networks to determine the first line of Anchors.

[0069] Specifically, in the feature pyramid network, the deep high-level features respond strongly to the entire object and have more semantic features, while the shallow low-level features have more location features. Extracting the deep features of the corn row can help the subsequent second detection head network develop more useful context information, for example, distinguishing weeds or outlier corns. At the same time, more detailed features contribute to detecting corn rows with high positioning accuracy.

[0070] Dividing the established corn row dataset into a training set, a validation set, and a test set according to a ratio of 7:2:1, and inputting it into a deep convolutional network model for training. In this embodiment, a combination of a feature pyramid network (FPN) and a ResNet lightweight neural network model is selected. Compared with other neural network models, this implementation scheme reduces the computational cost and memory consumption of the deep neural network while maintaining high accuracy, and can provide efficient real-time detection, facilitating the operation of the corn row detection method proposed in this application on mobile devices and low-power devices.

[0071] In some embodiments, before the step of using a feature pyramid network to extract features from the corn field image to determine a first feature map and generate a first line of Anchors by calling a trained corn field row detection model to convergence, it includes:

[0072] Step S101: Annotate the corn field image with a polyline and generate a segmentation map centered on the polyline;

[0073] Step S102: Use the corn field image and the corresponding generated segmentation label map as the training dataset.

[0074] Step S103: Input the dataset into the field corn row detection model for training to determine the field corn row detection model that has been trained to the convergence state.

[0075] Specifically, the parameter configuration of the training process is as follows: the pixel size of the input image is 1280*720; the model is implemented using the Pytorch API on the Linux-based Ubuntu20.04 operating system, and the computer hardware configuration for training includes a 12th Gen Intel(R) Core(TM) i5-12500 3.00GHz CPU and an NVIDIA GeForce RTX 3060 GPU; the field corn row detection model is evaluated through the Iou loss function, the training batch size is set to 40, and the field corn row detection model is trained for 70 iterations on the dataset, with the initial learning rate set to 0.0001.

[0076] After the field corn row detection model of this application is trained to the convergence state, it can be used to detect the corn rows in the corn field image.

[0077] Obtain the corn field image to be detected, call the field corn row detection model that has been trained to the convergence state, and use the Feature Pyramid Network to extract features from the corn field image to determine the first feature map and generate the first line of Anchors;

[0078] Step S20: Input the first feature map into the ROI alignment module in the second detection head network to generate a query matrix, a key matrix, and a value matrix, and perform horizontal compression and vertical compression on the query matrix, the key matrix, and the value matrix respectively in the axial attention module to determine the horizontal compression matrix and the vertical compression matrix;

[0079] Step S30: Perform multi-head self-attention operations on the horizontal compression matrix and the vertical compression matrix respectively to generate a first feature matrix and a second feature matrix, and add them after broadcasting operations on the first feature matrix and the second feature matrix to generate a third feature matrix;

[0080] Step S40: Concatenate the query matrix, the key matrix, and the value matrix in the context enhancement module and perform convolution operations to generate a fourth feature matrix, and multiply the convolved third feature matrix by the fourth feature matrix to determine the attention matrix;

[0081] Step S50: Using a preset loss function, perform regression on the seedling row line in the to-be-detected corn field image as a whole based on the first-line Anchor and the attention matrix to determine the corn row detection result, thus completing the field corn row detection based on line Anchor.

[0082] Specifically, please refer to Figure 4 . In this application, the to-be-detected corn field image is assigned to different levels through three ResNet convolutional layers, and corn row detection is performed through three different convolutional layers in sequence. Meanwhile, in order to collect more useful context information to better learn the corn row features, in this application, the deep feature P1 is sent into the detection head for ROI alignment and loss regression, and after adjusting the arrangement of the line Anchor, it is sent into the next-level feature extraction. After the feature map C2 is convolved through the adjusted line Anchor, it also enters the detection head for ROI alignment and loss regression, and the line Anchor that can better reflect the corn row is obtained and sent into the bottom-level feature extraction. After the feature map C3 is convolved, the feature map P3 is obtained and sent into the detection head to regress the corn row.

[0083] Please refer to Figure 5 , Figure 5It is a network schematic diagram of the ROI alignment module. The second detection head network is constructed by the ROI alignment module and the loss regression module. In order to achieve the best trade-off between detection accuracy and detection speed, this application abandons the Transformer model in the selection of the attention mechanism and innovatively adopts the squeeze enhancement method to select the region of interest. After the first feature map generated by the above-mentioned feature pyramid network enters the second detection head network, the ROI alignment module forms three feature matrices Q, feature matrix K, and feature matrix V through three different fully connected layers. Among them, the feature matrix Q represents the query matrix, the feature matrix K represents the key matrix, and the feature matrix V represents the value matrix; the ROI alignment module includes an axial attention module and a context enhancement module. In the axial attention module, in order to save computational cost and memory requirements, the feature matrix Q, the feature matrix K, and the feature matrix V are respectively horizontally compressed and vertically compressed to determine the horizontal compression matrix and the vertical compression matrix. After performing the multi-head self-attention operation (Multi-Head Attention) on the horizontally compressed matrix and the vertically compressed matrix after compression in both directions, the two new feature matrices, the first feature matrix and the second feature matrix, are broadcast and then added to restore the size before compression and form the feature matrix S. Among them, the feature matrix S represents the third feature matrix, and this feature matrix S collects the features of the pixels near the corn row. However, due to the squeeze enhancement, some context information may be lost. Therefore, in order to supplement this part of the context information, through the context enhancement module, the feature matrix Q, the feature matrix K, and the feature matrix V are concatenated into a matrix and then convolved twice to form the feature matrix T. Among them, the feature matrix T represents the fourth feature matrix. Finally, the convolved feature matrix S (the third feature matrix) is multiplied by the feature matrix T (the fourth feature matrix) and then normalized to form the attention matrix W. The attention matrix W can use more context information to learn better feature representations. The preset loss function is used to regress the seedling strip line in the to-be-detected corn field image as a whole according to the first line Anchor and the attention matrix W to determine the corn row detection result, and complete the detection of the field corn row based on the line Anchor.

[0084] In some embodiments, the query matrix represents the features of the current region of interest; the key matrix represents the encoding of the input features and is used to calculate the correlation with the query region; the value matrix contains the true features and is used to perform weighted summation through the attention mechanism to obtain the final feature representation.

[0085] In some embodiments, in the loss regression module, the present application first calculates the intersection over union (IOU) of the predicted corn rows, and uses a loss function to regress the seedling belt line in the corn field image to be detected as a whole based on the first-line Anchor and the attention matrix. The loss function is expressed as:

[0086] C LIOU = 1 - LIOU,

[0087] where -1 ≤ LIOU ≤ 1. When LIOU = 1, the predicted row completely overlaps with the ground truth. When LIOU converges to -1, the two lines are far apart. This loss function is simple and differentiable, easy to implement parallel computing, and predicting the corn row as a whole unit helps improve the overall performance.

[0088] In some embodiments, please refer to Figures 6 to 8 , where Figure 6 is a schematic diagram of the corn field image and its detection effect, Figure 7 is a schematic diagram of the detection results of corn rows in different growth stages, Figure 8 is a schematic diagram of the detection results of corn rows in the case of weeds and missing seedlings.

[0089] Please refer to Figure 9, a field corn row detection device provided to meet one of the purposes of the present application, includes a first line Anchor determination module 1100, a feature matrix compression module 1200, a feature matrix fusion module 1300, an attention matrix determination module 1400, and a corn row detection module 1500. Among them, the first line Anchor determination module 1100 is configured to respond to a field corn row detection instruction, obtain a corn field image to be detected, call a trained field corn row detection model to converge, and use a feature pyramid network to perform feature extraction on the corn field image to determine a first feature map and generate a first line Anchor; the feature matrix compression module 1200 is configured to input the first feature map into the ROI alignment module in the second detection head network to generate a query matrix, a key matrix, and a value matrix, and perform horizontal compression and vertical compression on the query matrix, the key matrix, and the value matrix respectively in the axial attention module to determine a horizontal compression matrix and a vertical compression matrix; the feature matrix fusion module 1300 is configured to perform multi-head self-attention operations on the horizontal compression matrix and the vertical compression matrix respectively to generate a first feature matrix and a second feature matrix, and add them after performing a broadcast operation on the first feature matrix and the second feature matrix to generate a third feature matrix; the attention matrix determination module 1400 is configured to splice the query matrix, the key matrix, and the value matrix in the context enhancement module and perform a convolution operation to generate a fourth feature matrix, and multiply the convolved third feature matrix by the fourth feature matrix to determine an attention matrix; the corn row detection module 1500 is configured to use a preset loss function to perform regression on the seedling belt line in the corn field image to be detected as a whole according to the first line Anchor and the attention matrix to determine a corn row detection result, and complete the field corn row detection based on the line Anchor.

[0090] Based on any embodiment of the present application, please refer to Figure 10 , another embodiment of the present application further provides an electronic device, which can be implemented by a computer device, such as Figure 10As shown, it is a schematic diagram of the internal structure of a computer device. The computer device includes a processor, a computer-readable storage medium, a memory, and a network interface connected through a system bus. Among them, the computer-readable storage medium of the computer device stores an operating system, a database, and computer-readable instructions. The database can store a control information sequence. When the computer-readable instructions are executed by the processor, the processor can implement a method for detecting field corn rows based on line Anchors. The processor of the computer device is used to provide computing and control capabilities to support the operation of the entire computer device. The memory of the computer device can store computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor can execute the method for detecting field corn rows based on line Anchors of the present application. The network interface of the computer device is used to connect and communicate with a terminal. Those skilled in the art can understand that Figure 10 The structure shown is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0091] In this embodiment, the processor is used to execute Figure 9 the specific functions of each module in. The memory stores the program code and various types of data required to execute the above modules or sub-modules. The network interface is used for data transmission between the user terminal and the server. The memory in this embodiment stores the program code and data required to execute all modules in the device for detecting field corn rows based on line ANCHOR of the present application. The server can call the program code and data of the server to execute the functions of all modules.

[0092] The present application also provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors are caused to execute the steps of the method for detecting field corn rows based on line ANCHOR according to any embodiment of the present application.

[0093] The present application also provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by one or more processors, the steps of the method for detecting field corn rows based on line Anchor according to any embodiment of the present application are implemented.

[0094] Those of ordinary skill in the art can understand that all or part of the processes in the above-described embodiments of the method of the present application can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-described methods. Among them, the aforementioned storage medium can be a computer-readable storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0095] The above are only some embodiments of the present application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A method for detecting corn rows in a field based on line anchors, characterized in that: include: In response to a field corn row detection instruction, a corn field image to be detected is obtained, a field corn row detection model that has been trained to a convergence state is called, and a feature pyramid network is used to extract features of the corn field image to determine a first feature map and generate a first line anchor; Inputting the first feature map into a ROI alignment module in a second detection head network to generate a query matrix, a key matrix, and a value matrix, and horizontally compressing and vertically compressing the query matrix, the key matrix, and the value matrix in an axial attention module to determine a horizontal compression matrix and a vertical compression matrix; Performing a multi-head self-attention operation on the horizontal compression matrix and the vertical compression matrix to generate a first feature matrix and a second feature matrix, and performing a broadcast operation on the first feature matrix and the second feature matrix and adding them to generate a third feature matrix; In the context enhancement module, the query matrix, the key matrix, and the value matrix are concatenated and then convolved to generate a fourth feature matrix, and the third feature matrix is ​​convolved and then multiplied with the fourth feature matrix to determine an attention matrix; The preset loss function is used to regress the seedling strip line in the corn field image to be detected as a whole according to the first line Anchor and the attention matrix to determine the corn row detection result, thereby completing the field corn row detection based on the line Anchor.

2. The method for detecting corn rows in a field based on line anchors according to claim 1, characterized in that: The field corn row detection model is constructed by a feature pyramid network and a second detection head network, wherein the feature pyramid network includes a three-layer ResNet convolutional neural network, three convolutional layers and three first detection head networks, the first detection head network and the second detection head network include a ROI alignment module and a loss regression module, and the ROI alignment module includes an axial attention module and a context enhancement module.

3. The method for detecting corn rows in a field based on line anchors according to claim 2, characterized in that: The step of extracting features from the corn field image using a feature pyramid network to determine a first feature map and generate a first line anchor includes: A feature pyramid network is used to extract features of the corn field image. Feature maps containing multiple scales and multiple features are extracted layer by layer through three layers of ResNet convolutional layers. The feature maps containing multiple scales and multiple features are passed through three different convolutional layers in turn to determine the first feature map. The first feature map is input into the three first detection head networks to determine the first line Anchor.

4. The method for detecting corn rows in a field based on line anchors according to claim 2, characterized in that: Before calling the field corn row detection model that has been trained to a converged state and extracting features from the corn field image using a feature pyramid network to determine a first feature map and generate a first line anchor, the method includes: Annotating the corn field image with a broken line, and generating a segmentation map with the broken line as the center; The corn field images and the corresponding generated segmentation label images are used as training data sets; The data set is input into a field corn row detection model for training to determine a field corn row detection model that has been trained to a convergent state.

5. The method for detecting corn rows in a field based on line anchors according to claim 1, characterized in that: The steps of obtaining the corn field image to be detected include: A two-axis self-stabilizing gimbal is used to connect the camera and is mounted on an electric chassis, so that the electric chassis moves forward slowly and uniformly to shoot a video, and each frame is captured from the shot video as a corn field image to be detected.

6. The method for detecting corn rows in a field based on line anchors according to claim 1, characterized in that: The corn field image data includes one or more of corn field images at various growth stages, corn field images affected by light, corn field images affected by weeds, corn field images with missing seedlings, and corn field images that are blocked.

7. The method for detecting corn rows in a field based on line anchors according to any one of claims 1 to 6, characterized in that: The query matrix represents the features of the current region of interest; the key matrix represents the encoding of the input features, which is used to calculate the relevance with the query region; the value matrix contains the real features, which are used for weighted summation through the attention mechanism to obtain the final feature representation.

8. A field corn row detection device based on line anchor, characterized in that: include: A first line Anchor determination module is configured to respond to a field corn row detection instruction, obtain a corn field image to be detected, call a field corn row detection model that has been trained to a convergence state, and use a feature pyramid network to extract features from the corn field image to determine a first feature map and generate a first line Anchor; a feature matrix compression module configured to input the first feature map into the ROI alignment module in the second detection head network to generate a query matrix, a key matrix, and a value matrix, and to horizontally compress and vertically compress the query matrix, the key matrix, and the value matrix in the axial attention module to determine a horizontal compression matrix and a vertical compression matrix; a feature matrix fusion module, configured to perform a multi-head self-attention operation on the horizontal compression matrix and the vertical compression matrix respectively to generate a first feature matrix and a second feature matrix, and perform a broadcast operation on the first feature matrix and the second feature matrix and add them together to generate a third feature matrix; an attention matrix determination module, configured to concatenate the query matrix, the key matrix, and the value matrix in the context enhancement module to generate a fourth feature matrix through a convolution operation, and to determine an attention matrix after convolving the third feature matrix and multiplying the third feature matrix with the fourth feature matrix; The corn row detection module is configured to use a preset loss function to regress the seedling strip line in the corn field image to be detected as a whole according to the first line Anchor and the attention matrix to determine the corn row detection result and complete the field corn row detection based on the line Anchor.

9. An electronic device, comprising a central processing unit and a memory, characterized in that: The central processing unit is used to call and run the computer program stored in the memory to execute the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: It stores a computer program implemented according to the method described in any one of claims 1 to 7 in the form of computer-readable instructions, and when the computer program is called and executed by a computer, the steps included in the corresponding method are executed.