Mesh division and multi-scale feature fusion-based dense small target counting method and device

By dividing the grid according to the principle of "big near and small far" in the image and fusing multi-scale features, the problem of inappropriate grid size in traditional methods is solved, and the detection accuracy and calculation efficiency of dense small targets are improved.

CN119942449APending Publication Date: 2025-05-06BEIJING INFORMATION SCI & TECH UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510023206.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Traditional object detection methods and counting methods have problems with inappropriate grid size in dense small-object counting tasks, resulting in distant small-objectives being ignored or missed, affecting the counting accuracy.

Method used

The method of fusion of grid division and multi-scale features is adopted. The image is divided into near, middle and distant areas according to the principle of "near, large, far and small", and grids of different sizes are set up in each area. Multi-level features are extracted through the backbone network, and detection accuracy is improved through feature adaptive fusion.

Benefits of technology

The detection accuracy of long-distance small targets is improved, the computing efficiency is optimized, and the accuracy and robustness of dense small target counts are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942449A_ABST
    Figure CN119942449A_ABST
Patent Text Reader

Abstract

The invention discloses a dense small target counting method and device based on grid division and multi-scale feature fusion. The dense small target counting method comprises the steps of 1, processing a data set; step 2, dividing the input image into three areas according to the principle of large near and small far, setting grids in each area, and cutting according to the grids with different sizes; 3, performing feature extraction on the image slices to obtain a multi-level feature map of the image; step 4, each branch feature is aligned to other branch feature dimensions, feature fusion weights are dynamically generated, different branch features and corresponding weights are multiplied and then added, and low-layer branch fusion features, middle-layer branch fusion features and high-layer branch fusion features are output; and step 5, selecting the middle-layer branch fusion features and the high-layer branch fusion features for fusion, outputting position coordinates and confidence scores of prediction points, matching the position coordinates and the confidence scores with real points, and completing dense small target counting. According to the invention, the detection precision of the long-distance small target can be improved and the calculation efficiency can be optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision and deep learning technology, and in particular to a method and device for counting dense small targets by integrating grid division and multi-scale features. Background Art

[0002] In the task of counting dense small targets, the targets are usually densely distributed and have large size variations. Common applications include bird flock counting, traffic flow monitoring, pedestrian detection, etc. Traditional target detection and counting methods are mostly based on fixed grid image division, but this method has some significant disadvantages, especially when the target size varies greatly, it cannot adapt to the perspective law of "near big and far small", resulting in low detection accuracy of small targets in the distance and low computational efficiency.

[0003] The grid division method in the prior art usually adopts a uniform grid or a manually set grid of fixed size. However, the fixed grid cannot be adjusted adaptively according to the scale change of the target, which is an obvious limitation for targets with large differences in distance. In the task of counting dense small targets, especially in the detection of small targets in the distance, the traditional method often causes the small targets in the distance to be ignored or missed due to the inappropriate grid size, which affects the counting accuracy.

[0004] In addition, existing deep learning methods usually rely on feature extraction at a single scale when performing target detection, and it is difficult to effectively integrate features from different levels and scales, further affecting the detection accuracy of dense small targets. Therefore, how to improve the detection accuracy of targets of different scales while ensuring computational efficiency, especially the accuracy of small targets at long distances, has become a technical problem that needs to be solved urgently in the current target counting field. Summary of the invention

[0005] The purpose of the present invention is to provide a dense small target counting method and device by integrating grid division and multi-scale features, which can improve the detection accuracy of small targets at a long distance and optimize the calculation efficiency.

[0006] To achieve the above object, the present invention provides a dense small target counting method combining grid division and multi-scale feature fusion, which comprises:

[0007] Step 1: Process the data set: select a dense small object image data set and divide it into a training set and a validation set;

[0008] Step 2: Partition and grid the input image: According to the principle of "near is big and far is small", the input image is divided into three areas: near area, middle area and far area, and a grid of preset size is set in each area. Different areas of the image are cropped according to grids of different sizes to obtain image slices of different grid sizes;

[0009] Step 3: Extract features through the backbone network: The image slices in step 1 are subjected to feature extraction through the backbone network of the network model. Feature representations from low level to high level are gradually extracted through layer-by-layer convolution to obtain a multi-level feature map of the image: low-level branch features, middle-level branch features, and high-level branch features.

[0010] Step 4: Adaptive feature fusion: align each branch feature of step 3 to the dimension of other branch features, dynamically generate feature fusion weights, multiply different branch features and corresponding weights and add them together to obtain new branch fusion features, so that each branch fusion feature has other branch features, and output low-level branch fusion features, middle-level branch fusion features and high-level branch fusion features;

[0011] Step 5: Multi-scale feature aggregation and target prediction: Select the intermediate layer branch fusion features and the high-level branch fusion features in step 4 for fusion, output the predicted point location coordinates and confidence scores through the regression branch and classification branch of the network model, match them with the real points, and complete the dense small target counting.

[0012] Furthermore, the setting method of “near area, middle area and far area” in step 2 is: starting from the bottom of the image, the near area occupies 50% of the total height of the image, the middle area occupies 30% of the total height of the image, and the far area occupies 20% of the total height of the image;

[0013] The setting method of "Grid" in step 2 is: set a grid of 128*128 pixels in the near area, a grid of 64*64 pixels in the middle area, and a grid of 32*32 pixels in the far area.

[0014] Furthermore, the “branch fusion feature” in step 4 is expressed as Y, as shown in formula (1):

[0015]

[0016] Among them, X1 is one of the low-level branch features, middle-level branch features and high-level branch features as the alignment target, and X2 and X3 are other branch features. and are the branch features after X2 and X3 are aligned to X1, α, β and γ are the branch features after X1, and The weight parameters obtained after the preset convolution are in the range of [0,1]. α+β+γ=1.

[0017] Further, the low-level branch fusion features, the middle-level branch fusion features, and the high-level branch fusion features are represented as C3, C4, and C5, respectively. The method of "selecting the middle-level branch fusion features and the high-level branch fusion features of step 4 for fusion" in step 5 includes:

[0018] C4 and C5 are upsampled, and the resolutions of C4 and C5 are aligned.

[0019] Furthermore, the loss function L of the network model total It is expressed as formula (2):

[0020] L total =λ1L cls +λ2L loc (2)

[0021] Among them, L cls is the classification loss, λ1 is the weight factor of the negative prediction point, λ2 is the weight term for balancing the impact of regression loss, and L cls , L loc The calculation formulas are (3) and (4) respectively:

[0022]

[0023] The present invention also provides a dense small target counting device integrating grid division and multi-scale features, which comprises:

[0024] A data set processing unit, which is used to select a dense small object image data set and divide it into a training set and a validation set;

[0025] The image partitioning and grid division unit is used to divide the input image into three areas according to the principle of "near is big and far is small": near area, middle area and far area, and set a grid of preset size in each area, and crop different areas of the image according to grids of different sizes to obtain image slices of different grid sizes;

[0026] The feature extraction unit is used to extract features from the image slices in step 1 through the backbone network of the network model, and gradually extract feature representations from low-level to high-level through layer-by-layer convolution to obtain a multi-level feature map of the image: low-level branch features, middle-level branch features, and high-level branch features;

[0027] The feature adaptive fusion unit is used to align each branch feature of the feature extraction unit with the dimension of other branch features. By dynamically generating feature fusion weights, different branch features and corresponding weights are multiplied and then added to obtain new branch fusion features, so that each branch fusion feature has other branch features, and outputs low-level branch fusion features, middle-level branch fusion features and high-level branch fusion features;

[0028] The multi-scale feature aggregation and target prediction unit is used to select the intermediate layer branch fusion features and the high-level branch fusion features of the feature adaptive fusion unit for fusion, output the predicted point position coordinates and confidence scores through the regression branch and classification branch of the network model, match them with the real points, and complete the dense small target counting.

[0029] Furthermore, the setting method of “near area, middle area and far area” in the image partition and grid division unit is: starting from the bottom of the image, the near area occupies 50% of the total image height, the middle area occupies 30% of the total image height, and the far area occupies 20% of the total image height;

[0030] The setting method of "grid" in the image partition and grid division unit is: set a grid of 128*128 pixels in the near area, set a grid of 64*64 pixels in the middle area, and set a grid of 32*32 pixels in the far area.

[0031] Furthermore, the “branch fusion feature” in the feature adaptive fusion unit is expressed as Y, and the specific formula is as follows:

[0032]

[0033] Among them, X1 is one of the low-level branch features, middle-level branch features, and high-level branch features as the alignment target, and X2 and X3 are other branch features. and are the branch features after X2 and X3 are aligned to X1, α, β and γ are the branch features after X1, and The weight parameters obtained after the preset convolution are in the range of [0,1]. α+β+γ=1.

[0034] Further, the low-level branch fusion features, the middle-level branch fusion features, and the high-level branch fusion features are represented as C3, C4, and C5, respectively. The method of "selecting the middle-level branch fusion features and the high-level branch fusion features of step 4 for fusion" in step 5 includes:

[0035] C4 and C5 are upsampled, and the resolutions of C4 and C5 are aligned.

[0036] Furthermore, the loss function L of the network model total It is expressed as formula (2):

[0037] L total =λ1L cls +λ2L loc (2)

[0038] Among them, L clsis the classification loss, λ1 is the weight factor of the negative prediction point, λ2 is the weight term for balancing the impact of regression loss, and L cls , L loc The calculation formulas are (3) and (4) respectively:

[0039]

[0040] Among them, M is the number of candidate prediction points in each image, is the confidence score of the i-th predicted point, N is the number of real points in each image, and p i is the ith real point, is the best matching prediction point of the i-th true point.

[0041] The present invention has the following advantages due to the adoption of the above technical solution:

[0042] The present invention utilizes the principle that objects that are near appear larger and objects that are far appear smaller to partition and grid the image, integrates feature information of different scales, improves the detection accuracy of dense small targets, and improves the accuracy and robustness of dense small target counting. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 The present invention is a flowchart of a method for dense small target counting by integrating grid division and multi-scale features according to an embodiment of the present invention.

[0044] Figure 2 Schematic diagram of image partition and grid division diagram in an embodiment of the present invention.

[0045] Figure 3 4 is a structural diagram of the model network used in the embodiment of the present invention.

[0046] Figure 4 4 is a structural diagram of feature adaptive fusion in an embodiment of the present invention.

[0047] Figure 5 4 is a structural diagram of multi-scale feature aggregation in an embodiment of the present invention. DETAILED DESCRIPTION

[0048] In the drawings, the same or similar reference numerals are used to represent the same or similar elements or elements with the same or similar functions. The embodiments of the present invention are described in detail below in conjunction with the drawings.

[0049] In the description of the present invention, the terms "center", "longitudinal", "lateral", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate directions or positional relationships based on the directions or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific direction, be constructed and operated in a specific direction, and therefore should not be understood as limiting the scope of protection of the present invention.

[0050] like Figure 1 and Figure 3 As shown, the dense small target counting method of grid division and multi-scale feature fusion provided by the embodiment of the present invention includes:

[0051] Step 1: Process the dataset: select a dense small object image dataset, organize the annotation format, and divide it into a training set and a validation set. The small object image dataset can be obtained from a public crowd counting dataset.

[0052] Step 2: Partition and grid the input image: According to the principle of "near big and far small", divide the input image into three areas: near area, middle area and far area, and set a grid of preset size in each area. Crop different areas of the image according to grids of different sizes to obtain image slices of different grid sizes.

[0053] like Figure 2 As shown in the figure, the setting method of "Near Area, Middle Area and Far Area" is as follows:

[0054] Starting from the bottom of the image, the near area accounts for 50% of the total height of the image and usually contains more details and dense targets, such as large-scale targets. The middle area accounts for 30% of the total height of the image. The targets may be fewer and slightly sparse. Using a medium grid can balance detection accuracy and computational efficiency. The distant area accounts for 20% of the total height of the image. The targets are usually small and sparse. Using a larger grid reduces computational overhead and avoids redundant detection caused by overly dense grids. Of course, partitioning different areas can also reduce computational complexity. Using different grid sizes in different areas can effectively reduce the overall amount of calculation. For example, using a large grid for distant sparse target areas can avoid unnecessary high-resolution calculations and improve detection speed.

[0055] After the partitioning is completed, the "grid" of each area is further divided. The division method is: set a 128*128 pixel grid in the near area, a 64*64 pixel grid in the middle area, and a 32*32 pixel grid in the far area. Using a large grid in the near area can reduce the number of grids, thereby reducing the computational overhead. Small grids are used in the far area. Although the number of grids increases, the computational complexity remains controllable due to the small target. Dynamically allocate computing resources as a whole, significantly improving the detection speed.

[0056] Of course, the partition ratio and grid size of the input image can be adjusted appropriately according to actual conditions.

[0057] Step 3: Extract features through the backbone network: Use a network model, such as VGG16BN, as the backbone network to extract features from the image slices in step 1. Through layer-by-layer convolution, gradually extract feature representations from low-level to high-level, and obtain a multi-level feature map of the image: low-level branch features, middle-level branch features, and high-level branch features, providing multi-level and different dimensional feature support.

[0058] Step 4: Adaptive feature fusion: align each branch feature of step 3 to the dimension of other branch features, dynamically generate feature fusion weights, multiply different branch features and corresponding weights and add them together to obtain new branch fusion features, so that each branch fusion feature has other branch features, and output low-level branch fusion features, middle-level branch fusion features and high-level branch fusion features.

[0059] In one embodiment, the “branch fusion feature” is represented by Y, as shown in formula (1):

[0060]

[0061] Among them, X1 is one of the low-level branch features, middle-level branch features and high-level branch features as the alignment target, and X2 and X3 are other branch features. and are the branch features after X2 and X3 are aligned to X1, α, β and γ are the branch features after X1, and The weight parameters obtained after the preset convolution are in the range of [0,1]. α+β+γ=1.

[0062] like Figure 4 As shown in the figure, taking X1 as the alignment target as an example, the dynamic upsampling method based on CARAFE is used to upsample the middle-layer branch feature X2 and the low-layer branch feature X3, align them to the X1 branch, and generate the aligned branch feature and After alignment, the branch features of different scales are convolved with 1×1 to obtain weight parameters α, β, and γ, and the parameters α, β, and γ are passed through the softmax function so that their range is [0, 1] and their sum is 1. The other branches, namely X2 and X3, are used as alignment targets in the same way.

[0063] Step 5: Multi-scale feature aggregation and target prediction: Select the intermediate layer branch fusion features and the high-level branch fusion features of step 4 for fusion, output the predicted point location coordinates and confidence scores through the regression branch and classification branch of the network model, and match them with the real points, for example, using the existing Hungarian algorithm for matching, to complete the counting of dense small targets.

[0064] In one embodiment, Figure 5 As shown in the figure, the low-level branch fusion features, the middle-level branch fusion features and the high-level branch fusion features are represented as C3, C4 and C5 respectively, among which C3 is a low-level feature with the highest resolution and contains more spatial detail information; C4 and C5 are middle- and high-level features with strong semantic information, but the resolution gradually decreases. In order to effectively fuse these features, the lower-resolution C4 and C5 must be upsampled, and the CARAFE dynamic upsampling method is used to align the resolution of C4 and C5. CARAFE is a content-aware dynamic upsampling method that keeps C4, C5 and C3 at the same spatial resolution.

[0065] In one embodiment, the loss function L of the network model is total It is expressed as formula (2):

[0066] L total =λ1L cls +λ2L loc (2)

[0067] Among them, L cls L is the classification loss, which is used to measure the classification performance of the model. This embodiment uses the cross entropy loss, which helps the model optimize its classification performance by measuring the difference between the predicted probability distribution and the true label. loc is the regression loss, which is used to measure the regression error of the model. This embodiment adopts the Euclidean loss, which optimizes the performance of the model in the regression task by calculating the square difference between the predicted point and the true point; λ1 is the weight factor of the negative predicted point, such as 0.5, λ2 is the weight term for balancing the impact of the regression loss, such as 0.0002, L cls , L loc The calculation formulas are (3) and (4) respectively:

[0068]

[0069] Among them, M is the number of candidate prediction points in each image, is the confidence score of the i-th predicted point, N is the number of real points in each image, and p i is the ith real point, is the best matching prediction point of the i-th true point.

[0070] like Figure 3 As shown, the dense small target counting device for grid division and multi-scale feature fusion provided by the embodiment of the present invention includes a data set processing unit, an image partitioning and grid division unit, a feature extraction unit, a feature adaptive fusion unit and a multi-scale feature aggregation and target prediction unit, wherein:

[0071] The dataset processing unit is used to select a high-quality dense small object image dataset and divide it into a training set and a validation set.

[0072] The image partitioning and grid division unit is used to divide the input image into three areas according to the principle of "near big and far small": near area, middle area and far area, and set a grid of preset size in each area, and crop different areas of the image according to grids of different sizes to obtain image slices of different grid sizes.

[0073] The feature extraction unit is used to extract features from the image slices in step one through the backbone network of the network model, and gradually extract feature representations from low-level to high-level through layer-by-layer convolution to obtain a multi-level feature map of the image: low-level branch features, middle-level branch features, and high-level branch features.

[0074] The feature adaptive fusion unit is used to align each branch feature of the feature extraction unit to the dimension of other branch features. By dynamically generating feature fusion weights, different branch features and corresponding weights are multiplied and added to obtain new branch fusion features, so that each branch fusion feature has other branch features, and outputs low-level branch fusion features, middle-level branch fusion features and high-level branch fusion features.

[0075] The multi-scale feature aggregation and target prediction unit is used to select the intermediate layer branch fusion features and the high-level branch fusion features of the feature adaptive fusion unit for fusion, and output the predicted point position coordinates and confidence scores through the regression branch and classification branch of the network model, which are matched with the real points to complete the dense small target counting.

[0076] In one embodiment, the setting method of “near area, middle area and far area” in the image partition and grid division unit is: starting from the bottom of the image, the near area occupies 50% of the total height of the image, the middle area occupies 30% of the total height of the image, and the far area occupies 20% of the total height of the image;

[0077] The setting method of "grid" in the image partition and grid division unit is: set a grid of 128*128 pixels in the near area, set a grid of 64*64 pixels in the middle area, and set a grid of 32*32 pixels in the far area.

[0078] In one embodiment, the “branch fusion feature” in the feature adaptive fusion unit is represented by Y, and the specific formula is shown in formula (1).

[0079] Finally, it should be pointed out that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them. Those skilled in the art should understand that the technical solutions described in the above embodiments may be modified, or some of the technical features thereof may be replaced by equivalents; these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A dense small target counting method based on grid division and multi-scale feature fusion, characterized in that: include: Step 1: Process the data set: select a dense small object image data set and divide it into a training set and a validation set; Step 2: Partition and grid the input image: According to the principle of "near is big and far is small", the input image is divided into three areas: near area, middle area and far area, and a grid of preset size is set in each area. Different areas of the image are cropped according to grids of different sizes to obtain image slices of different grid sizes; Step 3: Extract features through the backbone network: The image slices in step 1 are subjected to feature extraction through the backbone network of the network model. Feature representations from low level to high level are gradually extracted through layer-by-layer convolution to obtain a multi-level feature map of the image: low-level branch features, middle-level branch features, and high-level branch features. Step 4: Adaptive feature fusion: align each branch feature of step 3 to the dimension of other branch features, dynamically generate feature fusion weights, multiply different branch features and corresponding weights and add them together to obtain new branch fusion features, so that each branch fusion feature has other branch features, and output low-level branch fusion features, middle-level branch fusion features and high-level branch fusion features; Step 5: Multi-scale feature aggregation and target prediction: Select the intermediate layer branch fusion features and the high-level branch fusion features in step 4 for fusion, output the predicted point location coordinates and confidence scores through the regression branch and classification branch of the network model, match them with the real points, and complete the dense small target counting.

2. The dense small target counting method of grid division and multi-scale feature fusion according to claim 1 is characterized in that: The setting method of "near area, middle area and far area" in step 2 is: starting from the bottom of the image, the near area occupies 50% of the total image height, the middle area occupies 30% of the total image height, and the far area occupies 20% of the total image height; The setting method of "Grid" in step 2 is: set a grid of 128*128 pixels in the near area, a grid of 64*64 pixels in the middle area, and a grid of 32*32 pixels in the far area.

3. The dense small target counting method of grid division and multi-scale feature fusion as described in claim 1 or 2, characterized in that: In step 4, the "branch fusion feature" is expressed as Y, as shown in formula (1): Among them, X1 is one of the low-level branch features, middle-level branch features and high-level branch features as the alignment target, and X2 and X3 are other branch features. and are the branch features after X2 and X3 are aligned to X1, α, β and γ are the branch features after X1, and The weight parameters obtained after the preset convolution are in the range of [0,1]. a+b+c=1.

4. The dense small target counting method of grid division and multi-scale feature fusion as claimed in claim 3 is characterized in that: The low-level branch fusion features, the middle-level branch fusion features, and the high-level branch fusion features are represented as C3, C4, and C5, respectively. The method of "selecting the middle-level branch fusion features and the high-level branch fusion features of step 4 for fusion" in step 5 includes: C4 and C5 are upsampled, and the resolutions of C4 and C5 are aligned.

5. The dense small target counting method of grid division and multi-scale feature fusion as claimed in claim 3 is characterized in that: The loss function L of the network model total It is expressed as formula (2): L total =λ1L cls +λ2L loc (2) Among them, L cls is the classification loss, λ1 is the weight factor of the negative prediction point, λ2 is the weight term for balancing the impact of regression loss, and L cls , L loc The calculation formulas are (3) and (4) respectively:

6. A dense small target counting device integrating grid division and multi-scale features, characterized in that: include: A data set processing unit, which is used to select a dense small object image data set and divide it into a training set and a validation set; The image partitioning and grid division unit is used to divide the input image into three areas according to the principle of "near is big and far is small": near area, middle area and far area, and set a grid of preset size in each area, and crop different areas of the image according to grids of different sizes to obtain image slices of different grid sizes; The feature extraction unit is used to extract features from the image slices in step 1 through the backbone network of the network model, and gradually extract feature representations from low-level to high-level through layer-by-layer convolution to obtain a multi-level feature map of the image: low-level branch features, middle-level branch features, and high-level branch features; The feature adaptive fusion unit is used to align each branch feature of the feature extraction unit with the dimension of other branch features. By dynamically generating feature fusion weights, different branch features and corresponding weights are multiplied and then added to obtain new branch fusion features, so that each branch fusion feature has other branch features, and outputs low-level branch fusion features, middle-level branch fusion features and high-level branch fusion features; The multi-scale feature aggregation and target prediction unit is used to select the intermediate layer branch fusion features and the high-level branch fusion features of the feature adaptive fusion unit for fusion, output the predicted point position coordinates and confidence scores through the regression branch and classification branch of the network model, match them with the real points, and complete the dense small target counting.

7. The dense small target counting device integrating grid division and multi-scale features as claimed in claim 6, characterized in that: The setting method of "near area, middle area and far area" in the image partition and grid division unit is: starting from the bottom of the image, the near area occupies 50% of the total image height, the middle area occupies 30% of the total image height, and the far area occupies 20% of the total image height; The setting method of "grid" in the image partition and grid division unit is: set a grid of 128*128 pixels in the near area, a grid of 64*64 pixels in the middle area, and a grid of 32*32 pixels in the far area.

8. The dense small target counting device integrating grid division and multi-scale features as claimed in claim 6 or 7, characterized in that: The "branch fusion feature" in the feature adaptive fusion unit is represented by Y, and the specific formula is as follows: Among them, X1 is one of the low-level branch features, middle-level branch features, and high-level branch features as the alignment target, and X2 and X3 are other branch features respectively. and are the branch features after X2 and X3 are aligned to X1, α, β and γ are the branch features after X1, and The weight parameters obtained after the preset convolution are in the range of [0,1].

9. The dense small target counting device integrating grid division and multi-scale features as claimed in claim 8, characterized in that: The low-level branch fusion features, the middle-level branch fusion features, and the high-level branch fusion features are represented as C3, C4, and C5, respectively. The method of "selecting the middle-level branch fusion features and the high-level branch fusion features of step 4 for fusion" in step 5 includes: C4 and C5 are upsampled, and the resolutions of C4 and C5 are aligned.

10. The dense small target counting method of grid division and multi-scale feature fusion according to claim 9, characterized in that: The loss function L of the network model total It is expressed as formula (2): L total =λ1L cls +λ2L loc (2) Among them, L cls is the classification loss, λ1 is the weight factor of the negative prediction point, λ2 is the weight term for balancing the impact of regression loss, and L cls , L loc The calculation formulas are (3) and (4) respectively: Among them, M is the number of candidate prediction points in each image, is the confidence score of the i-th predicted point, N is the number of real points in each image, and p i is the ith real point, is the best matching prediction point of the i-th true point.

Citation Information

Cited By

  • Crop phenotype in-situ analysis method and device, electronic equipment and storage medium

    CN121214383A