Feature fusion network construction method, sample selection method, target detection method and device

By adopting feature fusion network and sample selection method based on fusion information in a single-stage object detector, the problems of complex model structure, large memory burden and redundant detection results in the prior art are solved, and the detection speed and performance are improved.

CN113989601BActive Publication Date: 2025-06-06北京轩宇空间科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111224557.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-21
Publication Date
2025-06-06
Estimated Expiration
2041-10-21

AI Technical Summary

Technical Problem

The existing single-stage object detector uses feature pyramid networks to lead to complex model structure and high memory burden, and many-to-one sample selection strategies that lead to redundant detection results and long running time of non-maximum suppression algorithms.

Method used

A feature fusion network is used to replace the feature pyramid network, feature fusion is performed through projection layer, residual module, downsampling and upsampling convolutional layer, and a single-layer feature map is output. At the same time, a sample selection method based on fusion information is provided, and the optimal positive sample is selected for each target using the Hungarian algorithm.

Benefits of technology

The model structure is simplified, the memory burden and computational complexity are reduced, the detection speed and performance are improved, the detection results are sparse, and the non-maximum suppression algorithm is removed, which significantly speeds up the model's computing speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113989601B_ABST
    Figure CN113989601B_ABST
Patent Text Reader

Abstract

The present invention provides a feature fusion network, a sample selection method, a target detection method and a device, which are applied to a single-stage target detector. The feature fusion network is arranged between a backbone network and a head detection network, and includes three projection layers, two residual modules, a downsampling convolution layer, an upsampling convolution layer and two mergings, which replaces a feature pyramid network, fuses the features extracted by the backbone network, and outputs a single-layer feature map, thereby reducing model parameters and the complexity of the network structure. The sample selection method fully considers the influence of the network's prediction information and prior information on sample selection, and uses the Hungarian algorithm to select the optimal sample for each target to participate in model training, so that the detection model output is more sparse, which is a key step in removing the non-maximum suppression algorithm. The target detection method and device can improve the performance of target detection and accelerate the operating efficiency of the detection algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and in particular to target detection, and is specifically related to a feature fusion network, a sample selection method, a target detection method and a device. Background Art

[0002] Object detection is a basic research in the field of computer vision and is widely used in fields such as video surveillance and autonomous driving. In recent years, with the development of deep learning technology, object detection algorithms based on deep neural networks have made great progress. At present, object detection networks based on deep learning are mainly divided into two-stage detectors and single-stage detectors.

[0003] At present, the single-stage detector mainly includes four core modules: data preprocessing, network model construction, sample selection and model training. Single-stage detectors mostly use the model structure of backbone network, feature pyramid network and head detection network, such as Figure 1 As shown in the figure, the backbone network is mainly used to extract the feature information of the image. The feature pyramid network fuses the features of each layer extracted by the backbone network and outputs feature maps of different resolutions respectively. The head detection network uses the feature maps output by the feature pyramid for target classification and position regression. This combination can assign targets of different scales to different feature layers to improve the performance of target detection. In addition, single-stage detectors often adopt a many-to-one sample selection strategy, selecting multiple positive samples for each target to participate in training, which can provide rich feature information for model training.

[0004] The single-stage detector has achieved good detection results in practical applications based on the existing network model and sample selection strategy. However, the single-stage detector currently uses a feature pyramid network to fuse the feature information extracted by the backbone network, and then outputs feature maps of different resolutions. Each feature map needs to be input into a head detection network with independent parameters for detection, such as Figure 1 As shown in the figure, the structure of the single-stage detector is relatively complex, which brings a large memory burden and reduces the speed of the detector. On the other hand, the single-stage detector uses a many-to-one sample selection strategy to select multiple positive samples for each target to participate in model training, so that the detection model will produce redundant detection results in the inference stage, and the non-maximum suppression algorithm needs to be used to screen the results, which greatly increases the algorithm running time. Summary of the invention

[0005] In view of the above situation, the present invention provides a feature fusion network to replace the feature pyramid network, fuses the features extracted by the backbone network, and outputs a single-layer feature map, thereby reducing model parameters and reducing the complexity of the network structure; at the same time, a sample selection method and device based on fusion information are provided, which fully considers the influence of the network's prediction information and prior information on sample selection, uses the Hungarian algorithm to select the optimal sample for each target to participate in model training, and makes the detection model output more sparse, which is a key step in removing the non-maximum suppression algorithm; at the same time, a target detection method and device are provided to improve the performance of target detection and speed up the operation efficiency of the detection algorithm.

[0006] In order to achieve the purpose of the present invention, the following technical solutions are adopted:

[0007] A feature fusion network, applied to a single-stage object detector, is placed between the backbone network and the head detection network, including:

[0008] Projection layer, used to extract original feature maps of different resolutions from the backbone network x 3 , feature map x 4 , feature map x 5 After processing respectively, the feature maps are output respectively f 3 , feature map f 4 , feature map f 5 , so that the feature map provides more feature information required for target detection; among them, the feature map f 3 The resolution is larger than the feature map f 4 Resolution, feature map f 4 The resolution is larger than the feature map f 5 Resolution

[0009] Two residual modules are used to respectively f 4 Processing is performed to increase the receptive field. The dilated convolutions in the two residual modules use different dilated rates. One residual module outputs a feature map with an enlarged receptive field. f 6 , another residual module outputs the feature map after the receptive field is enlarged f 7 ;

[0010] Downsampling convolution layer, used to perform feature map f 3 Perform downsampling convolution to obtain the samef 4 Feature maps with the same resolution f 3 ';

[0011] Upsampling convolution layer, used to upsample feature maps f 5 Perform downsampling convolution to obtain the same f 4 Feature maps with the same resolution;

[0012] The first merging unit is used to process the upsampling convolution layer to obtain the feature map f 4 Feature maps and feature maps with the same resolution f 6 Merge to generate a new feature map f 5 ';

[0013] The second merging unit is used to merge the feature map f 3 ', Feature map f 4 , feature map f 5 ', Feature map f 7 Merge into a new feature map to input into the head detection network.

[0014] A sample selection method is applied to a single-stage target detector. The model structure of the single-stage target detector includes a backbone network, the feature fusion network mentioned above, and a head detection network. The sample selection method includes the steps of:

[0015] S1. Calculate the IoU value between the regression box predicted by the network and the target box based on the regression information output by the head detection network. q ;

[0016] S2. Calculate the combined cost of the classification information of the grid prediction and the IoU value calculated in step S1 based on the classification information output by the head detection network C Lc ;

[0017] S3. Calculate the distance cost between the grid point in the network and the center point of the target box based on the prior information C L1 , and C Lc Perform weighted combination to form a cost matrix C ;

[0018] S4. Based on prior information, from the cost matrix CFilter out valid candidate samples for each target to obtain the cost loss C o ;

[0019] S5. Based on cost loss C o , the Hungarian algorithm is used to assign the best positive sample to each target for model training.

[0020] This application proposes an optimal sample selection method / strategy based on fusion information, where each target corresponds to only one positive sample and the rest are negative samples.

[0021] A sample selection device is applied to a single-stage target detector. The model structure of the single-stage target detector includes a backbone network, the feature fusion network mentioned above, and a head detection network. The sample selection device includes:

[0022] The first calculation unit is used to calculate the IoU value between the regression box predicted by the network and the target box according to the regression information output by the head detection network q ;

[0023] The second calculation unit is used to calculate the combined cost of the classification information predicted by the grid and the IoU value calculated in step S1 according to the classification information output by the head detection network C Lc ;

[0024] The cost loss unit is used to calculate the distance cost between the grid point in the network and the center point of the target box based on the prior information. C L1 , and C Lc Perform weighted combination to form a cost matrix C ;

[0025] The screening unit is used to select the cost matrix based on the prior information. C Filter out valid candidate samples for each target to obtain the cost loss C o ;

[0026] Allocation unit for loss according to consideration C o , the Hungarian algorithm is used to assign the best positive sample to each target for model training.

[0027] An end-to-end object detection method for detecting an object of interest in an image and outputting detection information comprises the steps of:

[0028] Preprocess the data;

[0029] Constructing a single-stage detector network model, the network model comprising a backbone network, a head detection network, and a feature fusion network disposed between the backbone network and the head detection network;

[0030] The sample selection method described above was used for sample selection;

[0031] Perform model training based on the selected samples;

[0032] Use the trained model for object detection.

[0033] An end-to-end object detection device, for detecting an object of interest in an image and outputting detection information, comprises:

[0034] A preprocessing module is used to preprocess the data;

[0035] A construction module, used to construct a single-stage detector network model, wherein the network model includes a backbone network, a head detection network, and a feature fusion network disposed between the backbone network and the head detection network;

[0036] A sample selection module, used to select samples using the sample selection method described above;

[0037] The training module is used to train the model based on the selected samples;

[0038] The detection module is used to perform target detection using the trained model.

[0039] An electronic device comprises: at least one processor and a memory; wherein the memory stores computer-executable instructions; the computer-executable instructions stored in the memory are executed on the at least one processor, so that the at least one processor executes a sample selection method, or executes an end-to-end target detection method.

[0040] A computer-readable storage medium stores a computer program, which controls a device where the storage medium is located to execute a sample selection method or an end-to-end target detection method when the computer program is executed by a processor.

[0041] The beneficial effects of the present invention are:

[0042] 1. The new feature fusion network uses the idea of ​​divide and conquer to fuse the features of different layers and compress them into one layer of feature output. By using dilated convolution, the receptive field of the model is increased, and the feature information of the upper and lower layers is fused, so that the output layer features can cover a wider range of target scales, thereby achieving the effect of better feature extraction for targets of different scales;

[0043] 2. When using the optimal sample selection strategy based on fusion information to select a suitable positive sample for each target, the influence of network prediction information and prior information on sample selection is considered at the same time. The traditional sample selection method only considers regression loss, but target detection includes two major tasks: classification and regression. Only considering a single loss to select positive samples cannot select the optimal sample for training. The sample selection strategy based on fusion information in this application fully considers the classification and regression information predicted by the network, and considers the correlation between them. Prior information is provided according to the true value of the target, and auxiliary prediction information selects more stable samples as positive samples. Selecting an optimal positive sample for each target to participate in model training through the Hungarian algorithm makes the output of the detection network more sparse, which is the key to removing the non-maximum suppression algorithm and speeds up the operation of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] The drawings described herein are for illustrative purposes only of selected embodiments and not all possible implementations, and are not intended to limit the scope of the present invention.

[0045] Figure 1 It is the network model of the single-stage detector in the prior art.

[0046] Figure 2 This is a schematic diagram of the feature fusion network structure of this application.

[0047] Figure 3 Schematic diagram of the sample selection method based on fusion information of the present application.

[0048] Figure 4 This is a schematic diagram of the structure of a sample selection device based on fusion information of the present application.

[0049] Figure 5 This is a flow chart of the end-to-end target detection method of this application.

[0050] Figure 6 This is a structural block diagram of the end-to-end target detection device of the present application. DETAILED DESCRIPTION

[0051] In order to make the objectives, technical solutions and advantages of the present invention more clear, the present invention will be further described in detail below with reference to the accompanying drawings.

[0052] At present, the single-stage detector is mainly composed of three parts: the backbone network, the feature pyramid network and the head detection network. The feature pyramid has two main core benefits: on the one hand, the feature pyramid can perform multi-scale feature fusion, which combines feature maps of multiple scales to obtain a better representation; on the other hand, it is a divide-and-conquer strategy, detecting targets on feature maps of different levels according to the different scales of the target. Therefore, a series of artificial designs or NAS searches for more complex fusion feature pyramid networks have been triggered. However, while the feature pyramid network brings better performance to the single-stage detector, it also brings a relatively large memory burden, making the detector structure more complex, such as Figure 1 As shown, the detector speed is reduced.

[0053] In order to make full use of the advantages of the feature pyramid and reduce the complexity of the model, one aspect of the embodiment of the present application provides a new feature fusion network structure, which is applied to a single-stage target detector and is arranged between a backbone network and a head detection network. Figure 2 As shown in Figure 2, it is used to fuse multiple layers of features and compress them into a layer of feature map output.

[0054] Specifically, it includes: three projection layers, two residual modules, a downsampling convolution layer, an upsampling convolution layer and two merging (corresponding to two merging units: the first merging unit and the second merging unit).

[0055] The projection layer includes a layer of convolution kernels with 1x1 convolution and a layer of convolution kernels with 3x3 convolution, which extracts feature maps of different resolutions from the backbone network. x 3 , feature map x 4 , feature map x 5 After processing respectively, the feature maps are output respectively f 3 , feature map f 4 , feature map f 5 ; Among them, the feature map f 3 The resolution is larger than the feature map f 4 Resolution, feature map f 4 The resolution is larger than the feature map f 5 Resolution

[0056] The two residual modules respectively perform f 4The residual module consists of three layers of convolution. The first layer has a 1x1 convolution kernel, which is used to reduce the channel dimension, generally reduced to 1 / 4 of the original channel. The second layer has a 3x3 dilated convolution kernel, which is used to increase the receptive field. The last layer has a 1x1 convolution kernel, which is used to restore the channel dimension.

[0057] The dilated convolutions in the two residual modules use different dilated rates. The dilated rate of one residual module is r 1 Set to 2 to output the feature map after increasing the receptive field f 6 , the void rate of another residual module r 2 Set to 4 to output the feature map after increasing the receptive field f 7 .

[0058] Downsampling convolutional layer to the feature map with the largest resolution f 3 Perform downsampling convolution to obtain the same f 4 Feature maps with the same resolution f 3 '. Here the downsampling convolution layer uses a convolution with a kernel size of 3x3 and a convolution stride of 2.

[0059] Upsampling convolutional layer for the smallest resolution feature map f 5 Perform downsampling convolution to obtain the same f 4 The first merging unit processes the upsampling convolution layer to obtain a feature map with the same resolution as the feature map f 4 Feature maps and feature maps with the same resolution f 6 Merge to generate a new feature map f 5 '.

[0060] The second merging unit combines the feature map f 3 ', Feature map f 4 , feature map f 5 ', Feature map f 7 Merge into a new feature map to input into the head detection network.

[0061] In this implementation, we first select appropriate feature maps and pass them through a residual module with dilated convolution to obtain feature maps with different receptive fields. Then, these feature maps are fused with the multi-layer features extracted by the backbone network and compressed into a layer of feature maps for output. This fully meets the features required for target detection of different scales while reducing the number of model parameters, simplifying the model structure, and speeding up the algorithm.

[0062] The feature fusion network implemented in this paper replaces the feature pyramid network, fuses the features extracted by the backbone network, and outputs a single-layer feature map, reducing model parameters and reducing the complexity of the network structure.

[0063] At present, the single-stage detector uses a many-to-one sample selection strategy to select multiple positive samples for each target. The many-to-one sample selection strategy provides rich feature information for model training, but in the inference stage, it will produce redundant detection results, and the detection results need to be screened using a non-maximum suppression algorithm, which greatly increases the running time of the algorithm. In order to remove the non-maximum suppression algorithm and improve the running efficiency of the detection algorithm, another aspect of the embodiment of the present application provides a sample selection method, specifically an optimal sample selection strategy based on fusion information.

[0064] Suppose the classification information output by the detection network is p , the regression information is s First, calculate the IoU value between the regression box and the target box based on the regression information q , and then according to the classification information output by the network p , calculate the combined cost of classification information and IoU value C Lc , such as the formula:

[0065] ;

[0066] ;

[0067] in, q i Represents the network regression box and the i The IoU matrix of the target box, , p i Indicates i The prediction information of the category to which the target belongs, , a It is a regulating factor used to adjust the importance of classification and regression information. Indicates i The combined cost matrix corresponding to each target. N represents the total number of targets contained in the input image.

[0068] Calculate the coordinates of each grid point on the feature map and the center point of the target box L 1 distance, get the center point distance cost C L1 , and get the final cost C:

[0069] .

[0070] in, λ Lc Indicates the cost C Lc The weight coefficient of λ L1 Represents the center point distance cost C L1 The weight coefficient of .

[0071] Since the network prediction information is very unstable in the initialization stage, it is proposed to further integrate the prior information, screen the candidate samples, select the samples in the target box as valid samples, and the rest as invalid samples, as shown in the formula:

[0072] ;

[0073] Among them, Ω represents the prior information, and the grid coordinate point of the candidate sample is within the target box, then Ω i ∈Ω is 1, otherwise it is infinite. The cost of the candidate sample in the target box is a finite value, and the sample is a valid sample, while the cost of the candidate sample outside the target box is infinite, so the sample is an invalid sample.

[0074] Finally, according to the cost C o , use the Hungarian algorithm to assign a suitable positive sample to each target to participate in model training.

[0075] Another aspect of the embodiments of the present application provides a sample selection device, which is applied to a single-stage target detector. The model structure of the single-stage target detector includes a backbone network, a feature fusion network, and a head detection network. The feature fusion network is arranged between the backbone network and the head detection network. Specifically, the feature fusion network can adopt the feature fusion network described in the previous embodiment.

[0076] Sample selection device such as Figure 4 As shown, it includes a first calculation unit, a second calculation unit, a cost loss unit, a screening unit, and an allocation unit.

[0077] The first calculation unit calculates the IoU value between the regression box predicted by the network and the target box according to the regression information output by the head detection network. q The second calculation unit calculates the combined cost of the classification information of the grid prediction and the IoU value calculated in step S1 according to the classification information output by the head detection network. CLc The cost loss unit calculates the distance cost between the grid point in the network and the center point of the target box based on the prior information. C L1 , and C Lc Perform weighted combination to form a cost matrix C The screening unit selects the cost matrix according to the prior information. C Filter out valid candidate samples for each target to obtain the cost loss C o . Allocate units according to cost loss C o , the Hungarian algorithm is used to assign the best positive sample to each target for model training.

[0078] The sample selection method and device based on fusion information provided in the embodiment of the present application fully considers the influence of the classification and regression information predicted by the network on the sample selection, and also utilizes the prior information of the target to enhance the stability of the sample selection. The cost function is constructed by considering the prediction information and the prior information at the same time, and the Hungarian algorithm is used to select the optimal sample for each target, and the model is trained to make the output of the detection algorithm more sparse. During reasoning, the non-maximum suppression algorithm can be directly removed to achieve end-to-end target detection, which greatly improves the running speed of the algorithm.

[0079] In another aspect of the embodiments of the present application, an end-to-end object detection method is provided for detecting an object of interest in an image and outputting detection information, such as Figure 5 As shown, the following steps are included:

[0080] (1) Preprocess the data;

[0081] (2) constructing a single-stage detector network model, wherein the network model includes a backbone network, a head detection network, and a feature fusion network as described in the above embodiment disposed between the backbone network and the head detection network;

[0082] (3) Selecting samples using the sample selection method described in the previous embodiment;

[0083] (4) Perform model training based on the selected samples;

[0084] (5) Use the trained model to perform target detection.

[0085] This example can improve the performance of target detection and speed up the running efficiency of the detection algorithm.

[0086] In another aspect of the embodiments of the present application, an end-to-end object detection device is provided for detecting an object of interest in an image and outputting detection information, such as Figure 6As shown, it includes: a preprocessing module, a construction module, a sample selection module, a training module, and a detection module.

[0087] The preprocessing module preprocesses the data. The construction module constructs a single-stage detector network model, which includes a backbone network, a head detection network, and a feature fusion network as described in the previous embodiment arranged between the backbone network and the head detection network. The sample selection module selects samples using the sample selection method as described in the previous embodiment. The training module performs model training based on the selected samples. The detection module performs target detection using the trained model. This example can improve the performance of target detection and speed up the running efficiency of the detection algorithm.

[0088] According to another aspect of the embodiments of the present application, an electronic device is provided, comprising: at least one processor and a memory; wherein the memory stores computer-executable instructions; the computer-executable instructions stored in the memory are executed by the at least one processor, so that the at least one processor executes the sample selection method described in the foregoing embodiments, or executes the end-to-end target detection method described in the foregoing embodiments.

[0089] In another aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the device where the storage medium is located controls the sample selection method described in the previous embodiments to execute, or the end-to-end target detection method described in the previous embodiments to execute.

[0090] The above are only preferred embodiments of the present invention, and are not intended to be the only or limiting embodiments of the present invention. Those skilled in the art should understand that various changes or equivalent substitutions made to the present invention without departing from the scope of the present invention are within the scope of protection of the present invention.

Claims

1. A feature fusion network construction method is applied to a single-stage target detector to detect the target of interest in the image. It is set between the backbone network and the head detection network. The backbone network is used to extract the feature information of the image, and the head detection network is used to perform target classification and position regression. It is characterized in that The feature fusion network construction method includes constructing: Projection layer, used to extract original feature maps of different resolutions from the backbone network x 3 , feature map x 4 , feature map x 5 After processing respectively, the feature maps are output respectively f 3 , feature map f 4 , feature map f 5 , so that the feature map can provide more feature information required for target detection; among them, the feature map f 3 The resolution is larger than the feature map f 4 Resolution, feature map f 4 The resolution is larger than the feature map f 5 Resolution; Two residual modules are used to respectively f 4 Processing is performed to increase the receptive field. The dilated convolutions in the two residual modules use different dilated rates. One residual module outputs a feature map with an enlarged receptive field. f 6 , another residual module outputs the feature map after the receptive field is enlarged f 7 ; Downsampling convolution layer, used to perform feature map f 3 Perform downsampling convolution to obtain the same f 4 Feature maps with the same resolution f 3 '; Upsampling convolution layer, used to upsample feature maps f 5 Perform downsampling convolution to obtain the same f 4 Feature maps with the same resolution; The first merging unit is used to process the upsampling convolution layer to obtain the feature map f 4 Feature maps and feature maps with the same resolution f 6 Merge to generate a new feature map f 5 '; The second merging unit is used to merge the feature map f 3 ', Feature map f 4 , feature map f 5 ', Feature map f 7 Merge into a new feature map to input into the head detection network.

2. The feature fusion network construction method according to claim 1, It is characterized in that The projection layer includes a layer of convolution kernel with 1x1 convolution and a layer of convolution kernel with 3x3 convolution.

3. The feature fusion network construction method according to claim 1, It is characterized in that The residual module consists of three layers of convolution. The first layer has a 1x1 convolution kernel, which is used to reduce the channel dimension. The second layer has a 3x3 dilated convolution kernel, which is used to increase the receptive field. The last layer has a 1x1 convolution kernel, which is used to restore the channel dimension.

4. The feature fusion network construction method according to claim 1, It is characterized in that The downsampling convolution layer uses a convolution with a kernel size of 3x3 and a convolution stride of 2.

5. A sample selection method for detecting objects of interest in an image using a single-stage object detector. It is characterized in that The model structure of the single-stage target detector includes a backbone network, a feature fusion network constructed by the feature fusion network construction method according to any one of claims 1 to 4, and a head detection network. The sample selection method includes the steps of: S1. Calculate the IoU value between the regression box predicted by the network and the target box based on the regression information output by the head detection network. q ; S2. Calculate the combined cost of the classification information of the grid prediction and the IoU value calculated in step S1 based on the classification information output by the head detection network C Lc ; S3. Calculate the distance cost between the grid point in the network and the center point of the target box based on the prior information C L1 , and with C Lc Perform weighted combination to form a cost matrix C ; S4. Based on prior information, from the cost matrix C Filter out valid candidate samples for each target to obtain the cost loss C o ; S5. Based on cost loss C o , the Hungarian algorithm is used to assign the best positive sample to each target for model training.

6. A sample selection device, applied to a single-stage object detector for detecting an object of interest in an image, It is characterized in that The model structure of the single-stage target detector includes a backbone network, a feature fusion network constructed by the feature fusion network construction method according to any one of claims 1 to 4, and a head detection network, and the sample selection device includes: The first calculation unit is used to calculate the IoU value between the regression box predicted by the network and the target box according to the regression information output by the head detection network q ; The second calculation unit is used to calculate the combined cost of the classification information predicted by the grid and the IoU value calculated in step S1 according to the classification information output by the head detection network C Lc ; The cost loss unit is used to calculate the distance cost between the grid point in the network and the center point of the target box based on the prior information. C L1 , and with C Lc Perform weighted combination to form a cost matrix C ; The screening unit is used to select the cost matrix based on the prior information. C Filter out valid candidate samples for each target to obtain the cost loss C o ; Allocation unit for loss according to consideration C o , the Hungarian algorithm is used to assign the best positive sample to each target for model training.

7. An end-to-end object detection method for detecting objects of interest in an image and outputting detection information, It is characterized in that Includes steps: Preprocess the data; Constructing a single-stage detector network model, the network model comprising a backbone network, a head detection network, and a feature fusion network constructed by the feature fusion network construction method according to any one of claims 1 to 4 and arranged between the backbone network and the head detection network; Using the sample selection method as claimed in claim 5 to select samples; Perform model training based on the selected samples; Use the trained model for object detection.

8. An end-to-end object detection device for detecting an object of interest in an image and outputting detection information, It is characterized in that include: A preprocessing module is used to preprocess the data; A construction module, used to construct a single-stage detector network model, wherein the network model includes a backbone network, a head detection network, and a feature fusion network constructed by the feature fusion network construction method according to any one of claims 1 to 4 and arranged between the backbone network and the head detection network; A sample selection module, configured to select samples using the sample selection method according to claim 5; The training module is used to train the model based on the selected samples; The detection module is used to perform target detection using the trained model.

9. An electronic device, include: At least one processor and a memory; wherein the memory stores computer-executable instructions; characterized in that the computer-executable instructions stored in the memory are executed by the at least one processor, so that the at least one processor executes the sample selection method as described in claim 5, or executes the end-to-end target detection method as described in claim 7.

10. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by the processor, the device where the storage medium is located is controlled to execute the sample selection method as claimed in claim 5, or to execute the end-to-end target detection method as claimed in claim 7.

Citation Information

Patent Citations

  • Small target detection method based on deep learning

    CN112488220A

  • Target detection network and method based on mixed cavity convolution pyramid

    CN113392960A