Small target detection method and device, computer equipment and medium
By setting up multiple levels of upsampling and downsampling processing in the object detection method and using sparse feature maps for target prediction, the problem of small object detection accuracy and low efficiency is solved, and efficient and accurate small object detection is achieved.
Patent Information
- Application Number
- CN202311728621.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-14
- Publication Date
- 2025-06-17
AI Technical Summary
The existing object detection methods have low detection accuracy and low detection efficiency in the targets with a small proportion in the detection image.
By setting up the upsampling and downsampling processing of multiple levels, each level is set up with a target detection network and a target positioning network, and the target prediction is used to use sparse feature maps to reduce the calculation amount and improve detection efficiency.
Improves the accuracy and detection efficiency of small object detection, reduces computational costs, increases frames per second (FPS), and ensures the retention of high-resolution features.
Smart Images

Figure CN120164007A_ABST
Abstract
Description
Technical Field
[0001] The present invention is applicable to the field of object detection, and particularly relates to a small object detection method, device, computer device and medium. Background Art
[0002] In the prior art, the FPN (Feature Pyramid Network) algorithm is usually used for object detection. Its disadvantage is that for some small object detection scenarios, such as in construction sites, especially in industrial parks, docks and other scenarios, the monitoring perspectives in these scenarios are generally large and the coverage area is wide, resulting in workers looking very small in the image, and the targets corresponding to the reflective vests and helmets on them are even smaller. When using the FPN algorithm to detect small objects, due to the too small proportion of the object in the image, the accuracy of object detection is low.
[0003] To solve the above problems, small object detection is improved by scaling the size of the input image or reducing the downsampling rate of the image to maintain high-resolution features. The effect is only to increase the resolution of the feature map, but it will generate a relatively large computational cost, reduce the FPS (frames per second) of detection, and the improvement effect is poor.
[0004] Therefore, how to improve the accuracy and detection efficiency of small object detection has become an urgent problem to be solved at present. Summary of the Invention
[0005] In view of this, embodiments of the present invention provide a small object detection method, device, computer device and medium to solve the problems of low detection accuracy and low detection efficiency of the existing object detection method for detecting objects with a relatively small proportion in the image.
[0006] In a first aspect, a small object detection method is provided, including the following steps:
[0007] Obtain a target image, perform downsampling on the target image according to a preset level to obtain a downsampled feature map of each level;
[0008] Input the downsampled feature map of each level into a target detection network with the preset level for upsampling processing to obtain an upsampled feature map of each level;
[0009] Extract N key positions of the target to be predicted in the upsampled feature map of each level output by the target detection network, where N is an integer greater than zero;
[0010] Extract sparse features in the upsampled feature map of the next level according to the key positions to obtain a sparse feature map of the next level;
[0011] Input the sparse feature map of the next level into the target localization network preset for the corresponding level to obtain the target detection result of the next level.
[0012] In a second aspect, a small target detection device is provided, including:
[0013] A downsampling module, configured to obtain a target image, perform downsampling on the target image according to a preset level to obtain a downsampled feature map for each level;
[0014] A target detection network module, configured to input the downsampled feature map of each level into a target detection network with the preset level for upsampling processing to obtain an upsampled feature map for each level;
[0015] A target position localization module, configured to extract N key positions of a target to be predicted in the upsampled feature map of each level output by the target detection network, where N is an integer greater than zero;
[0016] A feature extraction module, configured to extract sparse features in the upsampled feature map of the next level according to the key positions to obtain a sparse feature map of the next level;
[0017] A target localization network module, configured to input the sparse feature map of the next level into the target localization network preset for the corresponding level to obtain the target detection result of the next level.
[0018] In a third aspect, an embodiment of the present invention provides a computer device, where the computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the small target detection method described in the first aspect is implemented.
[0019] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the small target detection method described in the first aspect is implemented.
[0020] The beneficial effects of the present invention compared with the prior art are:
[0021] The small target detection method, device, computer device and medium of the present invention set up multiple levels of upsampling processing, downsampling processing, and a target detection network and a target localization network are correspondingly set for each level. Among them, the target detection network of each level is used to input the downsampled feature map obtained by the downsampling processing of the corresponding level, and then perform upsampling processing to obtain the upsampled feature map of each level. Then, several key positions of small targets in the upsampled feature map of each level are predicted to extract features from the upsampled feature map of the next level, and a sparse feature map containing the key positions of small targets is obtained as the input of the target localization network of the next level. The target localization network is used to generate the target detection result of the corresponding level according to the sparse feature map.
[0022] Compared with the prior art, since the small target detection method of the present invention performs target prediction on the sparse feature map containing the key positions of small targets instead of directly using the upsampled feature map for target prediction, the calculation amount of prediction is small, the prediction speed is fast, and the detection efficiency is increased. Moreover, since the sparse feature map of each level contains more reliable key position information of small targets from higher levels, the accuracy of prediction can be guaranteed to be relatively high. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings in the following descriptions are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0024] Figure 1 It is a schematic diagram of an application environment of a small target detection method provided by an embodiment of the present invention;
[0025] Figure 2 It is a schematic flowchart of a small target detection method provided by an embodiment of the present invention;
[0026] Figure 3 It is a schematic flowchart of the training process of the target localization head network of each level provided by an embodiment of the present invention;
[0027] Figure 4 It is provided by an embodiment of the present invention in Figure 3 Based on this, it is a schematic flowchart of the training process of the target localization network of each level;
[0028] Figure 5 It is provided by an embodiment of the present invention in Figure 4 Based on this, it is a schematic flowchart of calculating the total target detection loss value according to the target detection loss value of each level;
[0029] Figure 6 It is a schematic structural diagram of a small target detection device provided by an embodiment of the present invention;
[0030] Figure 7 It is a schematic structural diagram of a computer device provided by an embodiment of the present invention. Specific Embodiments
[0031] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present invention. However, those skilled in the art should clearly understand that the present invention can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present invention.
[0032] It should be understood that when used in the specification and appended claims of the present invention, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0033] It should also be understood that the term "and / or" as used in the specification and appended claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0034] As used in the specification and appended claims of the present invention, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" according to the context.
[0035] In addition, in the description of the specification and appended claims of the present invention, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0036] References to "one embodiment" or "some embodiments" in the description of the present invention mean that a specific feature, structure, or characteristic described in connection with the embodiment is included in one or more embodiments of the present invention. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having" and their variants mean "including but not limited to", unless otherwise specifically emphasized.
[0037] It should be understood that the magnitudes of the sequence numbers of the steps in the following embodiments do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0038] To illustrate the technical solution of the present invention, specific embodiments are used for illustration below.
[0039] A small target detection method provided in the first embodiment of the present invention can be applied in an application environment such as Figure 1 where the client communicates with the server. The client includes, but is not limited to, terminal devices such as a personal digital assistant (PDA), a desktop computer, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a cloud terminal device, and a personal digital assistant (PDA). The server can be implemented by an independent server or a server cluster composed of multiple servers.
[0040] See Figure 2 , which is a flowchart of a small target detection method provided in an embodiment of the present invention. The above abnormal display detection method can be applied to the client in Figure 1 . The terminal device corresponding to the client connects to the target database through a preset application programming interface (API). When the target data is driven to run to execute corresponding tasks, corresponding task logs will be generated, and the above task logs can be collected through the API. As shown in Figure 2 , the small target detection method may include the following steps:
[0041] Step S201, obtain a target image, perform downsampling on the target image according to a preset level to obtain a downsampled feature map for each level;
[0042] Among them, downsampling the target image according to a preset level, and the obtained downsampled feature maps of each level include:
[0043] Preset M levels, where M is an integer greater than 1. Input the obtained target image into the convolutional neural network of the first level for feature extraction to obtain a feature map with the first resolution; use the convolutional neural network of the second level to downsample the feature map with the first resolution to obtain a feature map with the second resolution; and so on, until using the convolutional neural network of the Mth level to downsample the feature map with the (M - 1)th resolution to obtain a feature map with the Mth resolution.
[0044] In this step, the resolutions of the feature map with the first resolution, the feature map with the second resolution,..., and the feature map with the Mth resolution decrease in turn, that is, as the downsampling levels increase from low to high, the resolution of the input feature map decreases from high to low.
[0045] In this step, the specific level M can be 4, or other integers greater than 1, such as 2, 3, or 5, etc.
[0046] Step S202: Input the downsampled feature maps of each level into the target detection network with the preset level for upsampling processing to obtain the upsampled feature maps of each level;
[0047] Among them, the target detection network can be FPN, that is, the Feature Pyramid Network. The upsampling processing of the downsampled feature maps of each level by this target detection network to obtain the upsampled feature maps of each level includes:
[0048] Obtain the feature map with the Mth resolution from the convolutional neural network of the Mth level as the upsampled feature map of the Mth level in this target detection network. Since the Mth level is the top level, it is actually not upsampled, but only called the upsampled feature map.
[0049] The depth convolutional neural network of the (M - 1)th level in this target detection network obtains the upsampled feature map of the Mth level, and performs upsampling processing on the feature map of the Mth level to obtain the feature map of the (M - 1)th level. Superimpose the feature map of the (M - 1)th level with the feature map of the (M - 1)th resolution of the same level determined in step S201 to obtain the upsampled feature map of the (M - 1)th level.
[0050] And so on, the depth convolutional neural network of the next level performs upsampling processing on the upsampled feature map of the previous level to obtain the feature map of the next level. Superimpose the feature map of the next level with the feature map of the same level determined in step S201 to obtain the upsampled feature map of the next level; and so on, until obtaining the upsampled feature map of the first level.
[0051] In one embodiment, when M is 4, the feature maps obtained respectively from the fourth layer to the first layer of the above-mentioned object detection network are successively the upsampled feature map of the fourth layer, the upsampled feature map of the third layer, the upsampled feature map of the second layer, and the upsampled feature map of the first layer. The sizes of the upsampled feature maps output by these four layers are 1 / 32, 1 / 16, 1 / 8, and 1 / 4 of the size of the original object image in sequence.
[0052] Step S203: Extract N key positions of the object to be predicted in the upsampled feature map of each layer output by the object detection network, where N is an integer greater than zero;
[0053] Among them, several key positions of the object to be predicted in the upsampled feature maps of each layer can be extracted in sequence according to the order from the highest layer to the lowest layer, for example, N. And, according to the order from the highest layer to the lowest layer, the number of key positions of the object to be predicted in the upsampled feature maps of each layer increases successively. The reason is that the sample quantity of the upsampled feature maps from the first layer to the fourth layer gradually decreases, so the number of key positions detected in the upsampled feature maps from the first layer to the fourth layer also decreases successively.
[0054] For example, when M is 4, the pre-trained object position prediction network of the fourth layer can predict N1 key positions of N1 objects from the upsampled feature map of the fourth layer, where N1 is an integer greater than zero; the pre-trained object position prediction network of the third layer can predict N2 key positions of N2 objects from the upsampled feature map of the third layer, where N2 is an integer and N2 > N1; the pre-trained object position prediction network of the second layer can predict N3 key positions of N3 objects from the upsampled feature map of the second layer, where N3 is an integer and N3 > N2.
[0055] Step S204: Extract sparse features in the upsampled feature map of the next layer according to the key positions, and obtain the sparse feature map of the next layer;
[0056] Among them, using the N key positions of the object to be predicted output by the object position prediction network of the previous layer Ls, the coordinates of these N key positions can be mapped to 4N nearest neighbor coordinates in the upsampled feature map of the next layer Lx as the coordinates of the 4N object key positions of the upsampled feature map of the next layer Lx. Feature extraction is performed on the upsampled feature map of the next layer Lx according to the coordinates of these 4N object key positions to obtain sparse features, and the preset sparse convolution operation model is used to calculate the sparse features to obtain the sparse feature map of the next layer Lx.
[0057] In one embodiment, when the preset level M is 4, the key positions of N1 targets are predicted in the upsampled feature map of the fourth level obtained according to the process of the previous step, the key positions of N2 targets are predicted in the upsampled feature map of the third level, and the key positions of N3 targets are predicted in the upsampled feature map of the second level. On this basis, according to the process of this step, the sparse feature maps of the third level, the second level, and the first level can be correspondingly obtained.
[0058] It should be noted that in this step, the upsampled feature map of the first level is the upsampled feature map of the lowest level, and it is not necessary to predict the key positions of the targets in the upsampled feature map of the first level and transfer them to the next level. Therefore, in the previous step S203, it is not necessary to determine the key positions of the targets to be predicted in the upsampled feature map of the lowest level.
[0059] Step S205: Input the sparse feature map of the next level into the target localization network preset for the corresponding level to obtain the target detection result of the next level.
[0060] Among them, the corresponding number of target localization networks is set according to the preset number of levels. For example, M-level target localization networks can be set, and one target localization network is set for each level.
[0061] In one embodiment, the target localization network can adopt a region generation network, that is, an RPN (Region Proposal Network) Head network. The region generation network of each level is used to input the sparse feature map of the current level and output the target detection result of the current level.
[0062] The small target detection method of the present invention, by setting upsampling processing and downsampling processing of multiple levels, and setting a target detection network and a target localization network for each level. Among them, the target detection network of each level is used to input the downsampled feature map obtained by the downsampling processing of the corresponding level, and then perform upsampling processing to obtain the upsampled feature map of each level. Then, several key positions of small targets in the upsampled feature map of each level are predicted to extract features from the upsampled feature map of the next level, and a sparse feature map containing the key positions of small targets is obtained as the input of the target localization network of the next level. The target localization network is used to generate the target detection result of the corresponding level according to this sparse feature map.
[0063] Compared with the prior art, since the small target detection method of the present invention performs target prediction on the sparse feature map containing the key positions of small targets, rather than directly using the feature map for target prediction, the computational amount of prediction is smaller and the prediction speed is faster, thereby increasing the detection efficiency. Moreover, since each level of the sparse feature map contains more reliable key position information of small targets from higher levels, the accuracy of prediction can be guaranteed to be relatively high.
[0064] The small target detection method of the above embodiment is particularly applicable to the target detection scenarios of safety helmets and reflective vests of workers under a large monitoring coverage area, which can improve the detection speed and accuracy of safety helmets and reflective vests, and effectively assist in the safety supervision of construction sites.
[0065] In one embodiment, in Figure 2 step S203 of, extracting N key positions of the target to be predicted in the upsampled feature map of each level output by the target detection network includes:
[0066] Input the upsampled feature map of each level output by the target detection network into the corresponding pre-trained target localization head network respectively to obtain N key positions of the target to be predicted in the upsampled feature map of each level.
[0067] Among them, the target localization head network of each level can also adopt a region proposal network, which is used to input the upsampled feature map of the corresponding level and output several key positions of the target to be predicted at the corresponding level. The specific processing process of the target localization head network of each level includes:
[0068] Obtain the upsampled feature map of the current level, and obtain an intermediate feature map through a 3x3 convolutional layer; input the intermediate feature map into two 1x1 convolutional layers to obtain the prediction results of classification and regression respectively, that is, several key positions of the target to be predicted and the corresponding position prediction values.
[0069] Among them, for each position on the upsampled feature map, generate k anchor boxes with different sizes and aspect ratios. For each anchor box, calculate its IoU value (the ratio between the intersection area and the union area of the anchor box and the bounding box) with all real bounding boxes, and use it as the label of positive and negative samples.
[0070] Then, input the feature vectors of all anchor boxes into the classification and regression networks to obtain the final classification and regression prediction results; calculate the classification loss according to the classification prediction results and labels; calculate the regression loss according to the regression prediction results and labels; finally, add the classification loss and the regression loss to obtain the final loss; then, update the parameters of the region proposal network through backpropagation.
[0071] The small target detection method of this embodiment can accurately locate several key positions of small targets in the upsampled feature map of the current level relative to the upsampled feature map of the size of the current level by setting corresponding target localization head networks at each level. These key positions are used to map and obtain more accurate key positions of small targets in the next level, and then the key positions of small targets in the next level are used to extract features from the upsampled feature map of the next level to obtain a sparse feature map, which is used as the input of the target localization network, reducing the computational amount of localization and increasing the detection efficiency of localization.
[0072] In one embodiment, in Figure 2 step S204 of, extracting the sparse features in the upsampled feature map of the next level according to the key positions to obtain the sparse feature map of the next level, includes:
[0073] Obtain the N key position coordinates (X, Y) of the target to be predicted in the upsampled feature map of the current level, and determine the 4N key position coordinates (2X + i, 2Y + j) in the upsampled feature map of the next level according to the N key position coordinates (X, Y), where i, j ∈ {0, 1};
[0074] Extract the sparse features in the upsampled feature map of the next level according to the 4N key position coordinates (2X + i, 2Y + j) to obtain the sparse feature map of the next level.
[0075] Among them, use the target localization head network of each level to output a position prediction probability map with a size of H*W, where H*W is the size of the upsampled feature map; use the position prediction probability map of each level to determine the N key position coordinates of each current level.
[0076] For example, obtain the position prediction probability map output by the target localization head network of the fourth level, and set an appropriate threshold T, for example, it can be set to 0.4; query the positions in the position prediction probability map with probability score values greater than the threshold T as the key positions, and output all the determined N key position coordinates (X, Y).
[0077] Then, determine the N key position coordinates (2X, 2Y), the N key position coordinates (2X, 2Y + 1), the N key position coordinates (2X + 1, 2Y), and the N key position coordinates (2X + 1, 2Y + 1) in the upsampled feature map of the third level according to the N key position coordinates (X, Y). Finally, extract the sparse features in the upsampled feature map of the third level according to the above determined 4N key position coordinates (2X + i, 2Y + j) to obtain the sparse feature map of the third level. By analogy, the sparse feature maps of the second level and the first level can also be obtained.
[0078] In the small target detection method of this embodiment, considering that the proportion of targets in the upsampled feature maps at higher levels is relatively large, while the proportion of targets in the upsampled feature maps at lower levels is relatively small, the key positions extracted from the upsampled feature maps at higher levels are relatively accurate and are used to guide the small target localization in the next level, improving the small target localization accuracy in the next level and thus improving the overall small target localization accuracy.
[0079] In one embodiment, as Figure 3 shown, the training process of the target localization head network at each level includes:
[0080] Step S301: Input the upsampled feature maps at each level output by the target detection network into the target localization head network with corresponding initialized parameters respectively, to obtain the position prediction values of each target in the upsampled feature maps at each level.
[0081] Step S302: Calculate the target localization loss value corresponding to each level according to the position prediction values of each target in the upsampled feature maps at each level.
[0082] Among them, the sizes of the small targets detected by the target localization head network at each level are different, that is, in the order from high to low levels, the sizes of the small targets detected by the target localization head network at each level increase from small to large. When the target localization head network at each level detects N key positions of small targets in the upsampled feature map of this level, multiple detection frames (anchors) with different sizes are set for the upsampled feature maps at different levels. The size of the smallest detection frame in this level is used as the size threshold for detecting small targets. If the size of the detected target is smaller than this size threshold, it is considered that an effective small target is detected, and the position prediction value of this small target is marked as a key position.
[0083] The small target detection method of this embodiment can effectively detect the position prediction values of each target in the upsampled feature maps at each level by using the size of the smallest detection frame in each level as the size threshold to train the target localization head network at each level.
[0084] In one embodiment, as Figure 4 shown, on the basis of Figure 3 , the training process of the target localization network at each level includes:
[0085] Step S401: Input the sparse feature map of the previous level into the target localization network with the initialized parameters of the current level to obtain the target classification prediction value and the target regression prediction value of the current level.
[0086] Among them, the specific processing process of the target localization network at each level includes:
[0087] Obtain the sparse feature map of the current level, and obtain an intermediate feature map through a 3x3 convolutional layer; input the intermediate feature map into two 1x1 convolutional layers to obtain the target classification prediction value and the target regression prediction value of the current level respectively.
[0088] Step S402, calculate the target classification loss value of the current level according to the target classification prediction value of the current level, and calculate the target regression loss value of the current level according to the target regression prediction value of the current level;
[0089] Among them, the loss functions of the target classification loss value and the target regression loss value of the current level can be set according to specific situations. For example, the cross-entropy loss function can be selected.
[0090] Step S403, calculate the target detection loss value of the current level according to the target localization loss value, the target classification loss value and the target regression loss value of the current level;
[0091] Among them, use the target localization loss value, the target classification loss value and the target regression loss value of the current level to jointly affect the target detection loss value of this level. Exemplarily, the target localization loss value, the target classification loss value and the target regression loss value of the current level can be directly added to obtain the target detection loss value of the current level.
[0092] Step S404, calculate the total target detection loss value according to the target detection loss value of each level;
[0093] Among them, use the target detection loss values of each level to jointly affect the total target detection loss value of the overall network model.
[0094] Step S405, if the total target detection loss value does not meet the preset conditions, update the parameters of the target localization head network and the target localization network, and repeat the above training process until the total target detection loss value meets the preset conditions, and obtain the trained target localization head network and target localization network.
[0095] Among them, the target localization loss value, the target classification loss value and the target regression loss value can all be implemented by using a dynamically scaled cross-entropy loss function. The calculation formula of the dynamically scaled cross-entropy loss function is as follows:
[0096] FL(Pt)=-(1-Pt) γ log(Pt)
[0097]
[0098] Among them, FL(Pt) represents the value of the dynamically scaled cross-entropy loss function, that is, the Focal Loss value, Pt represents the probability value, γ represents the hyperparameter of the network, log represents the logarithmic function, P represents the probability that the sample is predicted to be 1 for the y category, and y represents the sample category.
[0099] Taking the implementation of the dynamically scaled cross-entropy loss function with the above formula for the target localization loss value as an example to illustrate:
[0100] After the target localization head network at a certain level inputs the upsampled feature map of that level, it can locate all the small targets on the upsampled feature map of that level. Then, calculate the minimum Euclidean distance between the position of each small target on the upsampled feature map and the center of the positions of all small targets. Then, compare the obtained minimum Euclidean distance with the size threshold of the detection box with the smallest size at the current level. If the minimum Euclidean distance is less than the size threshold, set the GT value of the small target at this feature map position to 1, otherwise set the GT value of the small target at this feature map position to 0, where the GT value is the category y in the above formula.
[0101] When the target localization loss value, the target classification loss value, and the target regression loss value all adopt the above cross-entropy loss function, the calculation formula for the target detection loss value at each level is as follows:
[0102] L(Cl, Rl, Vl) = L FL (Cl, Cl * ) + L r (Rl, Rl * ) + L FL (Vl, Vl * )
[0103] Among them, L(Cl, Rl, Vl) represents the target detection loss value at each level, L FL (Cl, Cl * ) represents the target classification loss value at each level, L r (Rl, Rl * ) represents the target regression loss value at each level, L FL (Vl, Vl * ) represents the target localization loss value at each level. Among them, Cl represents the target classification prediction value at each level, Cl * represents the target classification label value at each level, Rl represents the target regression prediction value at each level, Rl * represents the target regression label value at each level, Vl represents the position prediction value of each target in the upsampled feature map at each level, Vl * represents the position label value of each target in the upsampled feature map at each level.
[0104] In the small target detection method of this embodiment, through unified joint training of target localization head networks and target localization networks at multiple levels, during the training process, the target localization head network at the upper level and the target localization network at the lower level cooperate with each other. That is, the target localization head network at the upper level outputs several key position information with relatively high credibility and transmits it to the lower level, and uses the key position information at the upper level to determine the key position information in the upsampled feature map at the lower level.
[0105] Then, using this key position information to generate a sparse feature map for reducing the computational amount and serving as the input for the target localization network at the lower level. Finally, the target localization head networks and target localization networks at each level obtained through training can more quickly and accurately locate small targets in the upsampled feature map at this level, thereby achieving efficient small target detection.
[0106] In one embodiment, as Figure 5 shown, on the basis of Figure 4 , calculating the total target detection loss value according to the target detection loss value at each level includes:
[0107] Step S501, the preset level is four levels, and preset the weights of the target detection loss values of the first level, the second level, the third level, and the fourth level;
[0108] Step S502, using the weights of the target detection loss values of the first level, the second level, the third level, and the fourth level, perform superposition summation on the target detection loss values of the first level, the second level, the third level, and the fourth level to obtain the total target detection loss value.
[0109] Among them, after obtaining the target detection loss value at each level, the calculation formula of the total target detection loss value is as follows:
[0110] LOSS = β1L p1 +β2L p2 +β3L p3 +β4L p4
[0111] Among them, LOSS represents the target detection loss value at each level, β1 represents the weight of the target detection loss value of the first level, L p1 represents the target detection loss value of the first level, β2 represents the weight of the target detection loss value of the second level, L p2represents the target detection loss value of the second level, β3 represents the weight of the target detection loss value of the third level, L p3 represents the target detection loss value of the third level, β4 represents the weight of the target detection loss value of the fourth level, L p4 represents the target detection loss value of the fourth level.
[0112] In the small target detection method of this embodiment, by jointly training the target localization head networks and target localization networks of multiple levels in a unified manner, and comprehensively considering the influence of the target detection losses of each level on the total target detection loss during the training process, the target localization head networks and target localization networks of each level are trained to obtain a small target detection result with higher prediction accuracy for the original target image.
[0113] In one embodiment, on the basis of Figure 5 the weight of the target detection loss value of the first level and the weight of the target detection loss value of the second level are less than the weight of the target detection loss value of the third level and the weight of the target detection loss value of the fourth level.
[0114] In one implementation manner, the weights β1 of the target detection loss value of the first level, β2 of the target detection loss value of the second level, β3 of the target detection loss value of the third level, and β4 of the target detection loss value of the fourth level can be set to increase in sequence, that is, β1 < β2 < β3 < β4, and the sum of β1, β2, β3, and β4 is 1. Exemplarily, β1 can be 0.2, β2 can be 0.24, β3 can be 0.26, and β4 can be 0.3; in other examples, β1 can be 0.15, β2 can be 0.25, β3 can be 0.29, and β4 can be 0.31.
[0115] In one implementation manner, the weights β1 of the target detection loss value of the first level and β2 of the target detection loss value of the second level can also be set to be equal, and the weights β3 of the target detection loss value of the third level and β4 of the target detection loss value of the fourth level can be set to be equal, that is, β1 = β2 < β3 = β4, and the sum of β1, β2, β3, and β4 is 1. Exemplarily, β1 can be 0.2, β2 can be 0.2, β3 can be 0.3, and β4 can be 0.3.
[0116] In this embodiment, considering that the weights of the target detection loss values of the high levels (such as the fourth level or the third level) are set to be small, when adding the high-resolution upsampled feature maps of the low levels (such as the first level or the second level) and the low-resolution upsampled features of the high levels Figure 1 as training samples at the same time, the total number of training samples of the high-resolution upsampled feature maps of the low levels is much larger than the total number of training samples of the low-resolution upsampled feature maps of the high levels.
[0117] In the above case, if the weights of the target detection loss values are set to be small in the target localization head network and the target localization network, the overall joint training process will be dominated by small targets, resulting in a decrease in the detection of normal and large targets. Although the purpose of the method of the present invention is to improve the detection of small targets, the preferred model training method is to make the trained target localization head network and target localization network models not overly focus on the prediction of small targets, but reduce or even ignore the prediction of normal and large targets.
[0118] In the small target detection method of this embodiment, by setting the weight of the target detection loss value at the lower level to be less than the weight of the target detection loss value at the higher level, it is possible to prevent the trained target localization head network and target localization network from overly focusing on the detection of small targets, and still ensure the detection accuracy for normal and large targets.
[0119] Corresponding to the method in the above embodiment, Figure 6 The structural block diagram of a small target detection device provided by an embodiment of the present invention is shown. The above small target detection device is applied to a computer device, and the computer device is connected to a target database through a preset application program interface. When the target database is driven to run to execute corresponding tasks, corresponding task logs will be generated, and the above task logs can be collected through the API. For the sake of simplicity, only the parts related to the embodiments of the present invention are shown.
[0120] See Figure 6 , the small target detection device includes:
[0121] A downsampling module 61, configured to obtain a target image, perform downsampling on the target image according to a preset level, and obtain a downsampled feature map for each level;
[0122] A target detection network module 62, configured to respectively input the downsampled feature map of each level into a target detection network with the preset level for upsampling processing, and obtain an upsampled feature map for each level;
[0123] A target position localization module 63, configured to extract N key positions of a target to be predicted in the upsampled feature map of each level output by the target detection network, where N is an integer greater than zero;
[0124] A feature extraction module 64, configured to extract sparse features in the upsampled feature map of the next level according to the key positions, and obtain a sparse feature map of the next level;
[0125] A target localization network module 65, configured to input the sparse feature map of the next level into a preset target localization network of the corresponding level, and obtain a target detection result of the next level.
[0126] Optionally, the above-mentioned target position positioning module 63 includes:
[0127] A target positioning head network unit, configured to input the upsampled feature maps of each level output by the target detection network into the corresponding pre-trained target positioning head network respectively, and obtain N key positions of the targets to be predicted in the upsampled feature maps of each level.
[0128] Optionally, the above-mentioned feature extraction module 64 includes:
[0129] A coordinate calculation unit, configured to obtain the coordinates (X, Y) of N key positions of the targets to be predicted in the upsampled feature map of the current level, and determine the coordinates (2X + i, 2Y + j) of 4N key positions in the upsampled feature map of the next level according to the N key position coordinates (X, Y), where i, j ∈ {0, 1};
[0130] A feature extraction unit, configured to extract the sparse features in the upsampled feature map of the next level according to the 4N key position coordinates (2X + i, 2Y + j), and obtain the sparse feature map of the next level.
[0131] Optionally, in the above-mentioned target position positioning module 63, the training process of the target positioning head network of each level includes:
[0132] Input the upsampled feature maps of each level output by the target detection network into the target positioning head network with corresponding initialized parameters respectively, and obtain the position prediction values of each target in the upsampled feature maps of each level;
[0133] Calculate the target positioning loss value of the corresponding level according to the position prediction values of each target in the upsampled feature map of each level.
[0134] Optionally, in the above-mentioned target positioning network module 65, the training process of the target positioning network of each level includes:
[0135] Input the sparse feature map of the previous level into the target positioning network with initialized parameters of the current level, and obtain the target classification prediction value and the target regression prediction value of the current level;
[0136] Calculate the target classification loss value of the current level according to the target classification prediction value of the current level, and calculate the target regression loss value of the current level according to the target regression prediction value of the current level;
[0137] Calculate the target detection loss value of the current level according to the target positioning loss value, the target classification loss value and the target regression loss value of the current level;
[0138] Calculate the total object detection loss value according to the object detection loss value of each level;
[0139] If the total object detection loss value does not meet the preset condition, update the parameters of the object localization head network and the object localization network, and repeat the above training process until the total object detection loss value meets the preset condition, obtaining a trained object localization head network and object localization network.
[0140] Optionally, in the above object localization network module 65, during the training process of the object localization network of each level, the calculating the total object detection loss value according to the object detection loss value of each level includes:
[0141] The preset level is four levels, and the weights of the object detection loss values of the preset first level, the second level, the third level, and the fourth level are preset;
[0142] Use the weights of the object detection loss values of the first level, the second level, the third level, and the fourth level to perform superposition summation on the object detection loss values of the first level, the second level, the third level, and the fourth level to obtain the total object detection loss value.
[0143] Optionally, the weights of the object detection loss values of the first level and the second level are less than the weights of the object detection loss values of the third level and the fourth level.
[0144] It should be noted that the information interaction, execution process, etc. between the above modules, due to being based on the same concept as the method embodiment of the present invention, for their specific functions and the technical effects brought, please refer to the method embodiment part specifically, and will not be elaborated here.
[0145] Figure 7 This is a schematic structural diagram of a computer device provided by an embodiment of the present invention. As Figure 7 shown, the computer device of this embodiment includes: at least one processor( Figure 7 only one is shown in the figure), a memory, and a computer program stored in the memory and executable on at least one processor. When the processor executes the computer program, it implements the steps in any of the above-mentioned abnormal display detection method embodiments.
[0146] The computer device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand, Figure 7The examples of computer devices are merely illustrative and do not constitute a limitation on computer devices. A computer device may include more or fewer components than shown in the figure, or combine certain components, or have different components. For example, it may also include a network interface, a display screen, an input device, etc.
[0147] The so-called processor may be a CPU, and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0148] The memory includes a readable storage medium, an internal memory, etc. Among them, the internal memory may be the memory of the computer device, and the internal memory provides an environment for the operation of the operating system and computer-readable instructions in the readable storage medium. The readable storage medium may be the hard disk of the computer device, and in some other embodiments, it may also be an external storage device of the computer device. For example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Further, the memory may also include both the internal storage unit of the computer device and the external storage device. The memory is used to store the operating system, application programs, boot loaders, data, and other programs, such as the program code of computer programs. The memory may also be used to temporarily store the data that has been output or will be output.
[0149] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present invention. The specific working processes of the units and modules in the above device can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0150] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above method embodiments of the present invention, a computer program can be used to instruct the relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps in the above method embodiments can be implemented.
[0151] Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, USB flash drive, mobile hard disk, magnetic disk or optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.
[0152] The present invention can also implement all or part of the processes in the above method embodiments through a computer program product. When the computer program product runs on a computer device, the computer device can be made to execute the steps in the above method embodiments.
[0153] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0154] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.
[0155] In the embodiments provided by the present invention, it should be understood that the disclosed device / computer device and method can be implemented in other ways. For example, the device / computer device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the device or unit can be electrical, mechanical or other forms.
[0156] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0157] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A small target detection method, characterized in that, It includes the following steps: Obtain a target image, and perform downsampling on the target image according to a preset level to obtain downsampled feature maps of each level; Input the downsampled feature maps of each level into a target detection network with the preset level for upsampling processing to obtain upsampled feature maps of each level; Extract N key positions of the target to be predicted in the upsampled feature maps of each level output by the target detection network, where N is an integer greater than zero; Extract sparse features in the upsampled feature map of the next level according to the key positions to obtain the sparse feature map of the next level; Input the sparse feature map of the next level into a preset target localization network of the corresponding level to obtain the target detection result of the next level.
2. The small target detection method according to claim 1, characterized in that, Extracting N key positions of the target to be predicted in the upsampled feature maps of each level output by the target detection network includes: Input the upsampled feature maps of each level output by the target detection network into a corresponding pre-trained target localization head network respectively to obtain N key positions of the target to be predicted in the upsampled feature maps of each level.
3. The small target detection method according to claim 1, characterized in that, Extracting sparse features in the upsampled feature map of the next level according to the key positions to obtain the sparse feature map of the next level includes: Obtain the coordinates (X, Y) of N key positions of the target to be predicted in the upsampled feature map of the current level, and determine the coordinates (2X + i, 2Y + j) of 4N key positions in the upsampled feature map of the next level according to the N key position coordinates (X, Y), where i, j ∈ {0, 1}; Extract sparse features in the upsampled feature map of the next level according to the 4N key position coordinates (2X + i, 2Y + j) to obtain the sparse feature map of the next level.
4. The small target detection method according to claim 2, characterized in that, The training process of the target localization head network of each level includes: Input the upsampled feature maps of each level output by the target detection network into a target localization head network with corresponding initialized parameters respectively to obtain the position prediction values of each target in the upsampled feature maps of each level; Calculate the target localization loss value of the corresponding level according to the position prediction values of each target in the upsampled feature maps of each level.
5. The small target detection method according to claim 4, characterized in that, The training process of the target localization network of each level includes: Input the sparse feature map of the previous level into a target localization network with the initialized parameters of the current level to obtain the target classification prediction value and the target regression prediction value of the current level; Calculate the target classification loss value of the current level according to the target classification prediction value of the current level, and calculate the target regression loss value of the current level according to the target regression prediction value of the current level; Calculate the target detection loss value of the current level according to the target localization loss value, the target classification loss value and the target regression loss value of the current level; Calculate the total target detection loss value according to the target detection loss values of each level; If the total target detection loss value does not meet the preset condition, update the parameters of the target localization head network and the target localization network, and repeat the above training process until the total target detection loss value meets the preset condition, so as to obtain the trained target localization head network and target localization network.
6. The small target detection method according to claim 5, characterized in that, Calculating the total object detection loss value according to the object detection loss values of each level, including: The preset levels are four levels, and the weights of the object detection loss values of the preset first level, the weights of the object detection loss values of the second level, the weights of the object detection loss values of the third level, and the weights of the object detection loss values of the fourth level are preset; Using the weights of the object detection loss values of the first level, the weights of the object detection loss values of the second level, the weights of the object detection loss values of the third level, and the weights of the object detection loss values of the fourth level, perform superposition summation on the object detection loss values of the first level, the object detection loss values of the second level, the object detection loss values of the third level, and the object detection loss values of the fourth level to obtain the total object detection loss value.
7. The small target detection method according to claim 6, characterized in that, The weights of the object detection loss values of the first level and the weights of the object detection loss values of the second level are less than the weights of the object detection loss values of the third level and the weights of the object detection loss values of the fourth level.
8. A small target detection device, characterized in that, Including: A downsampling module for obtaining a target image and downsampling the target image according to a preset level to obtain a downsampled feature map of each level; An object detection network module for respectively inputting the downsampled feature map of each level into an object detection network with the preset level for upsampling processing to obtain an upsampled feature map of each level; An object position localization module for extracting N key positions of the target to be predicted in the upsampled feature map of each level output by the object detection network, where N is an integer greater than zero; A feature extraction module for extracting sparse features in the upsampled feature map of the next level according to the key positions to obtain a sparse feature map of the next level; An object localization network module for inputting the sparse feature map of the next level into a preset object localization network of the corresponding level to obtain the object detection result of the next level.
9. A computer device, characterized in that, The computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the small object detection method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the small object detection method according to any one of claims 1 to 7.