Bean pod number detection method and device, computer equipment and computer program product

Through the recognition model of feature-focused diffuse pyramid network and aligned dynamic detection head, combined with Kalman filter and cost matrix function, the problem of low accuracy of machine vision algorithm in soybean pod number counting is solved, and accurate pod number counting is achieved.

CN120655684APending Publication Date: 2025-09-16SHANDONG LAB OF ADVANCED AGRI SCI AT WEIFANG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510778776.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing machine vision algorithms may cause recognition errors in soybean pod counting due to the complex field environment, resulting in low statistical accuracy.

Method used

The recognition model of the neck network with feature-focused diffuse pyramid network structure and aligned dynamic detection head is adopted, combined with Kalman filter and cost matrix function, to achieve accurate counting of the number of pods through image data recognition, tracking and matching.

Benefits of technology

This effectively avoids duplicate and missed counting, and improves the accuracy of pod number statistics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120655684A_ABST
    Figure CN120655684A_ABST
Patent Text Reader

Abstract

The invention discloses a pod number detection method and device, computer equipment and a computer program product. The method comprises the following steps: acquiring image data of a to-be-detected target area; identifying pods in the image data of the to-be-detected target area by adopting an identification model to obtain an identification result; tracking the pods in the recognition result to obtain the movement tracks of the pods in the recognition result; a counting result is obtained according to the recognition result and the matching result of the moving track, the recognition model comprises a neck network, a backbone network and a detection head, the neck network is of a feature focusing diffusion pyramid network structure, and the detection head is an alignment dynamic detection head. The technical problem that the accuracy of counting the number of pods is low in the prior art is at least solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and more specifically, to a method, apparatus, computer device, and computer program product for detecting the number of pods. Background Art

[0002] Soybeans are one of the most important sources of protein. With the growing population and increasing demand for agricultural economic development, improving soybean yield and quality is becoming increasingly urgent. Therefore, accelerating the selection and breeding of new high-yield, high-quality, and disease-resistant varieties is crucial. Pod number is a key factor in soybean yield estimation, providing crucial information for plant breeding.

[0003] Traditional manual counting methods are time-consuming and labor-intensive, and are prone to subjective errors, especially in large-scale planting scenarios where the workload becomes enormous and difficult to manage. In recent years, some automated technical methods have been applied to soybean pod counting, and soybean pod counting models have been constructed to count the soybean pods, enabling rapid soybean pod counting. Existing methods for counting soybean pods using machine vision algorithms rely solely on building network models to count the number of pods. However, due to the complex field environment, recognition errors may occur, resulting in low accuracy in counting the number of pods. Summary of the Invention

[0004] The embodiments of the present application provide a method, apparatus, computer device, and computer program product for detecting the number of pods, to at least address the technical problem of low accuracy in pod number statistics in related technologies.

[0005] According to one aspect of an embodiment of the present application, a method for detecting the number of pods is provided, comprising: acquiring image data of a target area to be detected; identifying pods in the image data of the target area to be detected using a recognition model to obtain a recognition result; tracking the pods in the recognition result to obtain a movement trajectory of the pods in the recognition result; and obtaining a counting result based on a matching result between the recognition result and the movement trajectory, wherein the recognition model comprises: a neck network, a backbone network and a detection head, the neck network being a feature-focused diffusion pyramid network structure, and the detection head being an aligned dynamic detection head.

[0006] Optionally, the recognition model is trained in the following manner, including: obtaining image data of pods in a field environment, and annotating the image data of the pods in the field environment to obtain annotated image data; performing image enhancement processing on the annotated image data to obtain processed image data; and constructing a data set based on the processed image data, wherein the image enhancement processing method includes at least one of the following: flipping, Gaussian blurring, color gamut change, and image size scaling; constructing an initial recognition model, and using the data set to train the initial recognition model to obtain the recognition model.

[0007] Optionally, constructing an initial recognition model includes: connecting the backbone network, the neck network and the detection head in sequence to obtain the initial recognition model, wherein the backbone network is used to extract features from the input data set to obtain feature maps of different scales, the neck network is used to use a feature focusing module to focus local features in the feature maps of different scales, and then fuse the focused feature maps of different scales to obtain a fused feature map, the detection head is used to identify the pod image from the fused feature map, the feature focusing module includes: a heuristic module and a downsampling module, wherein the heuristic module contains a group of parallel convolution kernels, each convolution kernel is used to process feature maps of different scales to obtain feature maps processed at different scales, the downsampling module is used to reduce the spatial dimensions of the feature maps processed at different scales, the detection head is used to use a shared convolution kernel to process multiple input feature maps, and use a scaling layer to scale the processed multiple input feature maps, multiple scaled feature maps, and obtain the joint features of the multiple scaled feature maps through a feature extractor.

[0008] Optionally, the method further includes: obtaining a task type performed by the recognition model, the task type including at least one of the following: a classification task and a positioning task; obtaining task features corresponding to the task type, and decomposing the task of the recognition model based on the task features.

[0009] Optionally, a counting result is obtained based on the recognition result and the matching result of the movement trajectory, including: using the recognition model to detect each frame of the image data of the target area to be detected in turn to obtain a pod detection frame in each frame of the image; using a Kalman filter to predict the pod trajectory of the next frame based on the pod trajectory of the current frame; using a cost matrix function to match the pod detection frame in the next frame of the image with the predicted value of the pod trajectory of the next frame to obtain a matching result; and determining the counting result based on the matching result.

[0010] Optionally, the counting result is determined based on the matching result, including: when the pod detection box in the next frame image matches the predicted value of the pod trajectory of the next frame, determining to update the parameters of the Kalman filter; when the pod detection box in the next frame image does not have a matching predicted value of the pod trajectory, determining that a new pod appears; when the predicted value of the pod trajectory of the next frame does not have a matching pod detection box, determining that a pod is not detected.

[0011] Optionally, the method further includes: determining a pod trajectory that has not been matched to the detection frame for more than a first time period as a first trajectory; inserting a first virtual trajectory at the position of the first trajectory, wherein the virtual trajectory is used to smooth the parameters of the Kalman filter; determining a pod trajectory that has not been matched to the detection frame for more than a second time period and is then re-matched to the detection frame as a second trajectory; and regenerating the second virtual trajectory based on the time point when tracking was lost and the time point when tracking was re-matched.

[0012] Optionally, regenerating a second virtual trajectory based on the time point at which tracking was lost and the time point at which tracking was rematched includes: obtaining a pod state observed at the time point at which tracking was lost and a pod state observed at the time point at which tracking was rematched; and determining the second virtual trajectory based on the pod state observed at the time point at which tracking was lost and the pod state observed at the time point at which tracking was rematched.

[0013] According to another aspect of an embodiment of the present application, a pod number detection device is also provided, including: a first acquisition module, used to acquire image data of a target area to be detected; a recognition module, used to use a recognition model to identify pods in the image data of the target area to be detected, and obtain a recognition result; a tracking module, used to use a pod tracking algorithm to track the pods in the recognition result, and obtain the movement trajectory of the pods in the recognition result; a matching module, used to obtain a counting result based on the matching result between the recognition result and the movement trajectory, wherein the recognition model includes: a neck network, a backbone network and a detection head, the neck network is a feature-focused diffusion pyramid network structure, and the detection head is an aligned dynamic detection head.

[0014] According to another aspect of an embodiment of the present application, a computer device is provided, comprising: a memory and a processor, wherein the memory is used to store program instructions; and the processor is connected to the memory and is used to execute the above-mentioned pod number detection method.

[0015] According to another aspect of the embodiments of the present application, a computer program product is provided, including computer instructions, which implement the above-mentioned pod number detection method when executed by a processor.

[0016] In an embodiment of the present application, image data of a target area to be detected is obtained; a recognition model is used to identify pods in the image data of the target area to be detected to obtain a recognition result; the pods in the recognition result are tracked to obtain a movement trajectory of the pods in the recognition result; and a counting result is obtained based on a matching result between the recognition result and the movement trajectory, wherein the recognition model includes: a neck network, a backbone network and a detection head, the neck network is a feature-focused diffusion pyramid network structure, and the detection head is an aligned dynamic detection head, thereby achieving the purpose of jointly determining the number of pods based on the recognition result of the recognition model and the trajectory of the pods in the image data, thereby avoiding repeated counting and missed counting, improving the technical effect of statistical accuracy, and thus solving the technical problem of low statistical accuracy of the number of pods in related technologies. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0018] Figure 1 This is a hardware structure block diagram of a computer terminal for implementing a method for detecting the number of pods according to an embodiment of the present application;

[0019] Figure 2 is a flow chart of a method for detecting the number of pods according to an embodiment of the present application;

[0020] Figure 3 is a schematic structural diagram of a neck network according to an embodiment of the present application;

[0021] Figure 4 is a structural diagram of a feature focusing module according to an embodiment of the present application;

[0022] Figure 5 is a structural diagram of an alignment dynamic detection head according to an embodiment of the present application;

[0023] Figure 6 This is a task decomposition process diagram according to an embodiment of the present application;

[0024] Figure 7 is a flowchart of a tracking algorithm according to an embodiment of the present application;

[0025] Figure 8 is a structural diagram of a pod detection device according to an embodiment of the present application. DETAILED DESCRIPTION

[0026] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0028] The information collected in the embodiments of the present application is information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with the relevant laws, regulations and standards of the relevant regions, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or reject the automated decision results; if the user chooses to reject, the expert decision-making process will be entered.

[0029] In order to solve the problems existing in the related art, the embodiment of the present application provides a method for detecting the number of pods, which can be run on Figure 1 In the computer terminal shown, the computer terminal is explained below.

[0030] The pod number detection method embodiment provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 The following is a hardware block diagram of a computer terminal for implementing a method for detecting the number of pods. Figure 1As shown, the computer terminal 10 may include one or more (illustrated by 102a, 102b, ..., 102n in the figure) processors (the processor may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission module 106 for communication functions connected via a wired and / or wireless network. In addition, it may also include: a display, a keyboard, a cursor control device, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, and a BUS bus. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.

[0031] It should be noted that the one or more processors and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry." The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuitry may be a single, independent processing module, or may be incorporated in whole or in part into any of the other components of the computer terminal 10. As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0032] Memory 104 can be used to store software programs and modules for application software, such as the program instructions / data storage device corresponding to the pod number detection method in the embodiments of the present application. The processor executes the software programs and modules stored in memory 104 to perform various functional applications and data processing, thereby implementing the pod number detection method described above. Memory 104 can include high-speed random access memory (RAM) and non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some embodiments, memory 104 may further include memory located remotely from the processor, which can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0033] The transmission module 106 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the computer terminal 10. In one embodiment, the transmission module 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission module 106 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.

[0034] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 .

[0035] It should be noted that, in some optional embodiments, the above Figure 1 The computer terminal shown may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of hardware elements and software elements. Figure 1 This is merely one example of a particular embodiment and is intended to illustrate the types of components that may be present in the computer terminal described above.

[0036] In the above operating environment, an embodiment of the present application provides an embodiment of a method for detecting the number of pods. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0037] Figure 2 This is a flow chart of a method for detecting the number of pods according to an embodiment of the present application. Figure 2 As shown, the method includes the following steps:

[0038] Step S202, obtaining image data of the target area to be detected;

[0039] In step S202 , image data of the target area to be detected can be collected by using a drone. For example, the drone is used to fly back and forth twice along the direction of the soybean planting rows in the target area to obtain image data of all soybeans in the target area.

[0040] Step S204: using the recognition model to identify the pods in the image data of the target area to be detected, and obtaining a recognition result;

[0041] Step S206: Track the pod in the recognition result to obtain the movement trajectory of the pod in the recognition result;

[0042] In step S208, a counting result is obtained based on the recognition result and the matching result of the movement trajectory, wherein the recognition model includes: a neck network, a backbone network and a detection head. The neck network is a feature-focused diffusion pyramid network structure, and the detection head is an aligned dynamic detection head.

[0043] Through the above steps S202 to S208, the image data of the target area to be detected is obtained; the recognition model is used to identify the pods in the image data of the target area to be detected to obtain a recognition result; the pods in the recognition result are tracked to obtain the movement trajectory of the pods in the recognition result; and the counting result is obtained based on the matching result of the recognition result and the movement trajectory, wherein the recognition model includes: a neck network, a backbone network and a detection head, the neck network is a feature-focused diffusion pyramid network structure, and the detection head is an aligned dynamic detection head, thereby achieving the purpose of jointly determining the number of pods based on the recognition result of the recognition model and the trajectory of the pods in the image data, thereby avoiding repeated counting and missed counting, improving the statistical accuracy, and thus solving the technical problem of low statistical accuracy of the number of pods in related technologies. The following is a detailed description.

[0044] In some embodiments of the present application, the recognition model is trained in the following manner, including: obtaining image data of pods in a field environment, and annotating the image data of the pods in the field environment to obtain annotated image data; performing image enhancement processing on the annotated image data to obtain processed image data; and constructing a data set based on the processed image data, wherein the image enhancement processing method includes at least one of the following: flipping, Gaussian blurring, color gamut change, and image size scaling; constructing an initial recognition model, and using the data set to train the initial recognition model to obtain a recognition model.

[0045] Specifically, the image data of the pods in the field environment is obtained in the same manner as the image data of the target area to be detected. The specific method of performing image enhancement processing on the annotated image data to obtain the processed image data is as follows: Horizontally flip the annotated image data: mirror-flip the image along the vertical axis, and add the flipped image to the dataset. Vertically flip the annotated image data: mirror-flip the image along the horizontal axis, and add the flipped image to the dataset. Or randomly flip the annotated image data: randomly select horizontal or vertical flipping with a preset probability to ensure that the dataset contains image samples of various flipping directions. Gaussian blur processing is performed on the annotated image data: a Gaussian kernel is selected, and the larger the standard deviation of the Gaussian kernel means the more obvious the blurring effect. The Gaussian kernel is applied to the image to simulate the blurring effect caused by different degrees of light or camera shake to obtain the processed image.

[0046] Perform color gamut changes on annotated image data, for example: brightness adjustment: randomly change the brightness of the image to simulate images taken under different lighting conditions; contrast adjustment: change the contrast of the image; color dithering: randomly adjust the RGB channels of the image; grayscale conversion: convert a color image into a grayscale image.

[0047] Resize the annotated image data: Random Scaling: Randomly select a scale factor to scale the image within a predefined scaling range (e.g., [0.5, 2]) to obtain the processed image. When scaling the image, adaptively adjust the size and position of the annotation box to ensure that the annotation information is consistent with the image content. Padding or Cropping: If the scaled image size does not match the preset size, padding or cropping can be performed to ensure that all augmented images meet the model input requirements. The augmented images are merged into the dataset along with the original images.

[0048] In some embodiments of the present application, an initial recognition model is constructed, including: connecting a backbone network, a neck network and a detection head in sequence to obtain an initial recognition model, wherein the backbone network is used to extract features from an input data set to obtain feature maps of different scales, the neck network is used to use a feature focusing module to focus local features in feature maps of different scales, and then fuse the focused feature maps of different scales to obtain a fused feature map, the detection head is used to identify pod images from the fused feature map, the feature focusing module includes: a heuristic module and a downsampling module, wherein the heuristic module contains a set of parallel convolution kernels, each convolution kernel is used to process feature maps of different scales to obtain feature maps processed at different scales, the downsampling module is used to reduce the spatial dimensions of feature maps processed at different scales, the detection head is used to process multiple input feature maps using a shared convolution kernel, and use a scaling layer to scale the processed multiple input feature maps, multiple scaled feature maps, and obtain joint features of multiple scaled feature maps through a feature extractor.

[0049] Figure 3 A schematic diagram of the structure of a neck network is shown. Figure 3 As shown in FIG, feature maps P3, P4, and P5 of different scales are input into the neck network. The feature focusing module in the neck network focuses the local features in the feature maps of different scales, and then the focused feature maps of different scales are fused to obtain the fused feature map.

[0050] Figure 4 A structural diagram of a feature focusing module is shown in FIG. Figure 4As shown in the figure, a set of parallel depth convolutions is used to capture rich information across multiple scale feature maps. After the convolution layer of the Adown (downsampling) module, the spatial dimension of the feature map is reduced by adjusting the stride, and the number of parameters in the convolution layer is optimized to reduce the complexity of the model. It contains an heuristic module (Inception-Style) that receives inputs of three scales and uses a set of parallel depth convolutions to capture rich information across multiple scales.

[0051] Figure 5 shows an aligned dynamic detection head, such as Figure 5 As shown, the use of shared convolution significantly reduces the number of parameters, making the model more lightweight. While using shared convolution, the Scale layer is used to scale features. The feature extractor learns task interaction features from multiple convolutional layers to obtain joint features. Its localization branch uses the offset and mask of DCNV2 (Deformable Convolutional Network), and the classification branch uses interactive features for dynamic feature selection. DCNV2 enhances the application of deformable convolution, using multi-layer offset learning to flexibly sample multiple layers of features. Its modulation mechanism adjusts the position and intensity of sampling points, allowing the network to dynamically adjust the sampling layout and influence weights, improving adaptability.

[0052] In some embodiments of the present application, the method also includes: obtaining the task type performed by the recognition model, the task type including at least one of the following: classification task and positioning task; obtaining task features corresponding to the task type, and decomposing the tasks of the recognition model based on the task features.

[0053] After decomposing the different tasks, specific neural network modules or architectures can be designed and integrated for each subtask (such as classification, localization, etc.). These modules may include, but are not limited to: convolutional layers, fully connected layers, attention mechanisms, deformable convolutions, etc. The selection and design of each architecture should consider how to best serve a specific subtask and define a corresponding loss function for each subtask. Independent training and joint training for each subtask: Each subtask module can be trained independently first, and its performance can be optimized before joint training to ensure that every part of the network can effectively contribute to the overall task.

[0054] Figure 6 A task decomposition flow chart is shown, such as Figure 6 As shown in the figure, by introducing attention weights, the task-spectific features (task features) of the classification task and the positioning task are calculated respectively. The specific calculation method is shown in the following formula:

[0055]

[0056] Where, ω kDenotes the attention weight ω∈R N The kth element of ω is determined based on the cross-layer task interaction characteristics:

[0057]

[0058] In the formula, σ represents the activation function, f c1 and f c2 There are two fully connected layers, x inter It is X inter Apply averagepooling (average pooling), and X inter It is X inter k Obtained by concatenate.

[0059] In some embodiments of the present application, a counting result is obtained based on the matching result of the recognition result and the movement trajectory, including: using a recognition model to detect each frame of the image data of the target area to be detected in turn to obtain a pod detection frame in each frame of the image; using a Kalman filter to predict the pod trajectory of the next frame based on the pod trajectory of the current frame; using a cost matrix function to match the pod detection frame in the next frame of the image with the predicted value of the pod trajectory of the next frame to obtain a matching result; and determining the counting result based on the matching result.

[0060] The counting result is determined according to the matching result, including: when the pod detection frame in the next frame image matches the predicted value of the pod trajectory in the next frame, determining to update the parameters of the Kalman filter; when the pod detection frame in the next frame image does not have a matching predicted value of the pod trajectory, determining that a new pod appears; when the predicted value of the pod trajectory in the next frame does not have a matching pod detection frame, determining that the pod is not detected.

[0061] like Figure 7 As shown, the target in each frame of the image data of the target area to be detected is detected to obtain the coordinates of the detection frame, confidence level and other information. The trajectory at time t (current frame) is predicted by Kalman filtering to obtain the trajectory prediction value at time t+1 (next frame). The detection value at time t+1 and the trajectory prediction value are then matched using the cost matrix function, resulting in three results: unmatched trajectory, unmatched detection frame, and matched. If matched, the parameters of the Kalman filter are updated. If the result after OCM is an unmatched detection frame, a new trajectory is divided for the detection frame. If the result after OCM is an unmatched trajectory, the trajectory is retained.

[0062] The matching process is as follows: Two empty lists are created: one for storing matched tracks and the other for storing unmatched bounding boxes. For each bounding box and each active track, the cost between the two is calculated. This cost can be the Euclidean distance, Intersection over Union (IoU), or other similarity or matching metrics. This forms a cost matrix, where rows represent tracks, columns represent bounding boxes, and each element in the matrix represents the cost between a track and a bounding box.

[0063] If the cost of a detection box and a trajectory is lower than the pre-set cost threshold, the detection box is considered to match the trajectory and the status of the trajectory is updated;

[0064] If the cost of the detection box and all trajectories is higher than the cost threshold, the detection box is considered to correspond to a newly appeared target, and a new trajectory ID is assigned to the detection box to confirm the appearance of a new pod.

[0065] If the cost of a track and all detection boxes is higher than the cost threshold, it is considered that the target corresponding to the track may have left the field of view or been blocked. The track is retained but its status is not updated.

[0066] For unmatched trajectories, we decide whether to keep them, delete them, or create virtual observations to maintain their status, based on their historical status.

[0067] For unmatched detection boxes, it may be due to the emergence of new targets or false positives of the detector, and judgments need to be made based on context and historical data.

[0068] For successfully matched trajectories and detection frames, the Kalman filter is used to update the trajectory status, including position, speed, etc. (Kalman filter parameters) to reflect the latest observation information.

[0069] For unmatched entities, the Kalman filter performs prediction updates and makes predictions based on their past motion states even without new observation information.

[0070] The above steps are repeated in consecutive image frames to achieve continuous target detection and tracking.

[0071] At time t+2, the Kalman filter is first applied to the unmatched, matched, and new trajectories at time t+1 to obtain an estimated value at time t+2. This value is then matched with the detected value at time t+2 using OCM (Optimal Cost Matching). If the predicted value of the unmatched trajectory at time t+1 matches the detected value, OCR (Object Consistency and Reassociation) is used to associate the trajectories based on the detection result at time t and the detection result at time t+2. The re-tracked trajectory triggers the Kalman filter parameters from time t to time t+2.

[0072] When a pod is not detected in a certain frame, but based on the Kalman filter prediction, the object should exist in the current field of view. In subsequent frames, if a new detection is found that is close to the predicted position of the previous frame, although the features of the new detection and the old track may be different due to lighting, occlusion, etc., the OCR mechanism will try to reassociate the two. By comparing the features of the pod (such as appearance, shape, motion pattern, etc.) and historical observations, it is decided whether to associate the new detection with the old track. If it is determined that the new detection belongs to the same object, the previous track will be updated to include the information of the new detection to maintain tracking consistency.

[0073] In some embodiments of the present application, the method further includes: determining a pod trajectory that has not been matched to a detection frame for more than a first duration as a first trajectory; inserting a first virtual trajectory at the position of the first trajectory, the virtual trajectory being used to smooth the parameters of the Kalman filter; determining a pod trajectory that has not been matched to a detection frame for more than a second duration and then is re-matched to the detection frame as a second trajectory; and regenerating the second virtual trajectory based on the time point at which tracking was lost and the time point at which tracking was re-matched.

[0074] In some embodiments of the present application, regenerating a second virtual trajectory based on the time point at which tracking was lost and the time point at which tracking was rematched includes: obtaining the pod state observed at the time point at which tracking was lost and the pod state observed at the time point at which tracking was rematched; and determining the second virtual trajectory based on the pod state observed at the time point at which tracking was lost and the pod state observed at the time point at which tracking was rematched.

[0075] In actual application scenarios, if the target is in an unmatched state for a long time, it will cause the error accumulation of the Kalman filter. At this stage, a virtual trajectory is established to smooth the parameters of the Kalman filter and reduce the error accumulation. Once a trajectory is associated with an observation again after a period of no tracking, OCSORT will backtrack to the period of its loss and re-update the parameters of the Kalman filter. The virtual trajectory is updated based on the "observation" of the virtual trajectory, and the virtual trajectory is generated with reference to the observation of the start and end time of the untracking period. By expressing the last observation seen before the untracking (the pod state observed at the time of the tracking loss) as O t1 The observation that triggers reassociation (the pod state observed at the rematched time point) is represented as O t2 , the mathematical model can be expressed as:

[0076] O t =Traj virtual (O t1 , O t2 , t), t1<t<t2

[0077] Where t1 represents the time point when tracking is lost, t2 represents the time point when tracking is re-matched, and t represents the current moment.

[0078] OCSPRT (a tracking algorithm) uses observations instead of estimates and adds momentum to the cost function to reduce the noise in the motion direction calculation:

[0079]

[0080] in, Represents the estimated matrix of the object, O is the observation state matrix at the new time, which contains the observation trajectories of all existing trajectories, λ represents the weight factor, represents the intersection-over-union ratio between the observed value and the predicted value, Indicates the calculation of the direction θ connecting two observations on the existing trajectory track and the direction θ formed by the historical observations of the trajectory and the observations at the new time intention , the algorithm takes radians as the direction of movement as follows:

[0081] Δθ=|θ track -θ intention |

[0082]

[0083] Among them, (μ1, ν1) and (μ2, ν2) are the observation values ​​of the target at two different times, μ1, ν1, μ2 and ν2 are the position information of the pod, Δθ represents radians, θ track ,θ intention All are calculated according to the calculation formula of θ.

[0084] Figure 8 A device for detecting the number of pods according to an embodiment of the present application includes:

[0085] An acquisition module 80 is used to acquire image data of a target area to be detected;

[0086] A recognition module 82 is configured to use a recognition model to recognize pods in the image data of the target area to be detected and obtain a recognition result;

[0087] A tracking module 84 is configured to track the pods in the recognition results using a pod tracking algorithm to obtain movement trajectories of the pods in the recognition results;

[0088] The matching module 86 is used to obtain the counting result based on the matching result of the recognition result and the movement trajectory, wherein the recognition model includes: a neck network, a backbone network and a detection head, the neck network is a feature-focused diffusion pyramid network structure, and the detection head is an aligned dynamic detection head.

[0089] Through the above-mentioned pod detection device, image data of the target area to be detected is obtained; the recognition model is used to identify the pods in the image data of the target area to be detected to obtain a recognition result; the pods in the recognition result are tracked to obtain the movement trajectory of the pods in the recognition result; and the counting result is obtained according to the matching result of the recognition result and the movement trajectory, wherein the recognition model includes: a neck network, a backbone network and a detection head, the neck network is a feature-focused diffusion pyramid network structure, and the detection head is an aligned dynamic detection head, thereby achieving the purpose of jointly determining the number of pods based on the recognition result of the recognition model and the trajectory of the pods in the image data, thereby avoiding repeated counting and missed counting, improving the statistical accuracy, and thus solving the technical problem of low statistical accuracy of the number of pods in the related art.

[0090] In some embodiments of the present application, the recognition model is trained in the following manner, including: obtaining image data of pods in a field environment, and annotating the image data of the pods in the field environment to obtain annotated image data; performing image enhancement processing on the annotated image data to obtain processed image data; and constructing a data set based on the processed image data, wherein the image enhancement processing method includes at least one of the following: flipping, Gaussian blurring, color gamut change, and image size scaling; constructing an initial recognition model, and using the data set to train the initial recognition model to obtain a recognition model.

[0091] In some embodiments of the present application, an initial recognition model is constructed, including: connecting a backbone network, a neck network and a detection head in sequence to obtain an initial recognition model, wherein the backbone network is used to extract features from an input data set to obtain feature maps of different scales, the neck network is used to use a feature focusing module to focus local features in feature maps of different scales, and then fuse the focused feature maps of different scales to obtain a fused feature map, the detection head is used to identify pod images from the fused feature map, the feature focusing module includes: a heuristic module and a downsampling module, wherein the heuristic module contains a set of parallel convolution kernels, each convolution kernel is used to process feature maps of different scales to obtain feature maps processed at different scales, the downsampling module is used to reduce the spatial dimensions of feature maps processed at different scales, the detection head is used to process multiple input feature maps using a shared convolution kernel, and use a scaling layer to scale the processed multiple input feature maps, multiple scaled feature maps, and obtain joint features of multiple scaled feature maps through a feature extractor.

[0092] The device also includes: a second acquisition module and a decomposition module, wherein the second acquisition module is used to obtain the task type performed by the recognition model, and the task type includes at least one of the following: classification task and positioning task; the decomposition module is used to obtain task features corresponding to the task type, and decompose the tasks of the recognition model based on the task features.

[0093] The matching module includes: a detection submodule, a prediction submodule, a matching submodule and a first determination submodule, wherein the detection submodule is used to use the recognition model to detect each frame of the image data of the target area to be detected in turn, and obtain the pod detection frame in each frame of the image; the prediction submodule is used to use the Kalman filter to predict the pod trajectory of the next frame based on the pod trajectory of the current frame; the matching submodule uses the cost matrix function to match the pod detection frame in the next frame of the image with the predicted value of the pod trajectory of the next frame to obtain the matching result; the first determination submodule determines the counting result based on the matching result.

[0094] The determination submodule includes: a first determination unit, a second determination unit and a third determination unit, wherein the first determination unit is used to determine the parameters of the updated Kalman filter when the pod detection box in the next frame image matches the predicted value of the pod trajectory in the next frame; the second determination unit is used to determine the presence of a new pod when the pod detection box in the next frame image does not have a matching predicted value of the pod trajectory; the third determination unit is used to determine that the pod is not detected when the predicted value of the pod trajectory in the next frame does not have a matching pod detection box.

[0095] The device also includes: a first determination module, an insertion module, a second determination module and a generation module, wherein the first determination module is used to determine the pod trajectory that has not been matched to the detection frame for more than a first time length as the first trajectory; the insertion module is used to insert a first virtual trajectory at the position of the first trajectory, and the virtual trajectory is used to smooth the parameters of the Kalman filter; the second determination module is used to determine the pod trajectory that has not been matched to the detection frame for more than a second time length and is then re-matched to the detection frame as the second trajectory; and the generation module is used to regenerate the second virtual trajectory based on the time point when tracking was lost and the time point when tracking was re-matched.

[0096] The generation module includes: an acquisition submodule and a second determination submodule, wherein the acquisition submodule is used to obtain the pod state observed at the time point when tracking is lost and the pod state observed at the time point when tracking is re-matched; the second determination submodule is used to determine the second virtual trajectory based on the pod state observed at the time point when tracking is lost and the pod state observed at the time point when tracking is re-matched.

[0097] It should be noted that Figure 8 The pod detection device shown is used to perform Figure 2 The pod number detection method shown in the figure is as follows, so the relevant explanations in the above pod number detection method are also applicable to the pod detection device and will not be repeated here.

[0098] An embodiment of the present application further provides a computer device comprising: a memory and a processor, wherein the memory is used to store program instructions; and the processor is connected to the memory and is used to execute the above-mentioned pod number detection method.

[0099] The pod number detection method executed by the above-mentioned computer device obtains image data of the target area to be detected; uses a recognition model to identify the pods in the image data of the target area to be detected to obtain a recognition result; tracks the pods in the recognition result to obtain the movement trajectory of the pods in the recognition result; obtains a counting result based on the matching result of the recognition result and the movement trajectory, wherein the recognition model includes: a neck network, a backbone network and a detection head, the neck network is a feature-focused diffusion pyramid network structure, and the detection head is an aligned dynamic detection head, thereby achieving the purpose of simultaneously determining the number of pods based on the recognition result of the recognition model and the trajectory of the pods in the image data, thereby avoiding repeated counting and missed counting, improving the statistical accuracy, and thus solving the technical problem of low pod number statistical accuracy in related technologies.

[0100] An embodiment of the present application further provides a non-volatile storage medium, which includes a stored computer program, wherein the device where the non-volatile storage medium is located executes the above-mentioned pod number detection method by running the computer program.

[0101] The above-mentioned pod number detection method stored in the non-volatile storage medium obtains image data of the target area to be detected; uses a recognition model to identify the pods in the image data of the target area to be detected to obtain a recognition result; tracks the pods in the recognition result to obtain the movement trajectory of the pods in the recognition result; obtains a counting result based on the matching result of the recognition result and the movement trajectory, wherein the recognition model includes: a neck network, a backbone network and a detection head, the neck network is a feature-focused diffusion pyramid network structure, and the detection head is an aligned dynamic detection head, thereby achieving the purpose of jointly determining the number of pods based on the recognition result of the recognition model and the trajectory of the pods in the image data, thereby avoiding repeated counting and missed counting, improving the statistical accuracy, and thus solving the technical problem of low statistical accuracy of the number of pods in the related art.

[0102] The embodiments of the present application also provide a computer program product, including computer instructions, which, when executed by a processor, implement the steps of the pod number detection method in the present application.

[0103] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0104] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0105] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0106] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0107] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0108] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0109] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A method for detecting the number of pods, characterized in that: include: Acquire image data of the target area to be detected; Using a recognition model to identify pods in the image data of the target area to be detected to obtain a recognition result; Tracking the pod in the recognition result to obtain a movement trajectory of the pod in the recognition result; A counting result is obtained based on the matching result of the recognition result and the movement trajectory, wherein the recognition model includes: a neck network, a backbone network and a detection head, the neck network is a feature-focused diffusion pyramid network structure, and the detection head is an aligned dynamic detection head.

2. The method according to claim 1, characterized in that The recognition model is trained by the following methods, including: Acquiring image data of bean pods in a field environment, and annotating the image data of the bean pods in the field environment to obtain annotated image data; Performing image enhancement processing on the annotated image data to obtain processed image data; and constructing a data set based on the processed image data, wherein the image enhancement processing method includes at least one of the following: flipping, Gaussian blurring, color gamut change, and image size scaling; An initial recognition model is constructed, and the data set is used to train the initial recognition model to obtain the recognition model.

3. The method according to claim 2, characterized in that Build an initial recognition model, including: The backbone network, the neck network and the detection head are connected in sequence to obtain the initial recognition model, wherein the backbone network is used to extract features from the input data set to obtain feature maps of different scales, the neck network is used to use a feature focusing module to focus local features in the feature maps of different scales, and then fuse the focused feature maps of different scales to obtain fused feature maps, the detection head is used to identify the pod image from the fused feature maps, the feature focusing module includes: a heuristic module and a downsampling module, wherein the heuristic module contains a group of parallel convolution kernels, each convolution kernel is used to process feature maps of different scales to obtain feature maps processed at different scales, the downsampling module is used to reduce the spatial dimensions of the feature maps processed at different scales, the detection head is used to use a shared convolution kernel to process multiple input feature maps, and use a scaling layer to scale the processed multiple input feature maps, multiple scaled feature maps, and obtain the joint features of the multiple scaled feature maps through a feature extractor.

4. The method according to claim 3, characterized in that The method further comprises: Acquire a task type performed by the recognition model, where the task type includes at least one of the following: a classification task and a positioning task; Task features corresponding to the task type are obtained, and tasks of the recognition model are decomposed based on the task features.

5. The method according to claim 1, wherein A counting result is obtained based on the recognition result and the matching result of the movement trajectory, including: Using the recognition model to detect each frame of the image data of the target area to be detected in sequence, and obtaining a pod detection frame in each frame of the image; The Kalman filter is used to predict the trajectory of the pod in the next frame based on the trajectory of the pod in the current frame. Using a cost matrix function, the bean pod detection frame in the next frame image is matched with the predicted value of the bean pod trajectory in the next frame to obtain a matching result; The counting result is determined according to the matching result.

6. The method according to claim 1, characterized in that Determining the counting result according to the matching result includes: When the pod detection frame in the next frame image matches the predicted value of the pod trajectory in the next frame image, determining to update the parameters of the Kalman filter; When the bean pod detection frame in the next frame image does not have a matching predicted value of the bean pod trajectory, determining that a new bean pod appears; In the case that there is no matching peapod detection frame for the predicted value of the peapod trajectory in the next frame image, it is determined that the peapod is not detected.

7. The method according to claim 6, characterized in that The method further comprises: The pod trajectory that is not matched to the detection frame for more than the first duration is determined as the first trajectory; Inserting a first virtual trajectory at the position of the first trajectory, wherein the virtual trajectory is used to smooth the parameters of the Kalman filter; Determine the pod trajectory that is not matched to the detection frame for more than the second time period and then re-matched to the detection frame as the second trajectory; A second virtual trajectory is regenerated based on the time point at which tracking was lost and the time point at which tracking was re-matched.

8. The method according to claim 7, characterized in that Regenerating a second virtual trajectory based on the lost time point and the re-matched time point, including: Obtain the pod status observed at the time point when the tracking was lost and the pod status observed at the time point when the tracking was re-matched; The second virtual trajectory is determined according to the pod state observed at the time point when the tracking is lost and the pod state observed at the time point when the tracking is re-matched.

9. A pod quantity detection device, characterized in that: include: A first acquisition module is used to acquire image data of a target area to be detected; a recognition module, configured to recognize pods in the image data of the target area to be detected using a recognition model to obtain a recognition result; a tracking module, configured to track the pods in the recognition result using a pod tracking algorithm to obtain movement trajectories of the pods in the recognition result; A matching module is used to obtain a counting result based on the recognition result and the matching result of the movement trajectory, wherein the recognition model includes: a neck network, a backbone network and a detection head, the neck network is a feature-focused diffusion pyramid network structure, and the detection head is an aligned dynamic detection head.

10. A computer device, characterized in that: include: A memory and a processor, wherein the memory is used to store program instructions; The processor is connected to the memory and is configured to execute the pod quantity detection method according to any one of claims 1 to 8.

11. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by a processor, the method for detecting the number of pods according to any one of claims 1 to 8 is implemented.