Target detection method, device, equipment, readable storage medium and program product

By preprocessing, feature extraction, feature fusion and two-way branch feature processing on the input image, the problem that traditional object detection methods cannot achieve end-to-end detection is solved, and efficient object detection is achieved.

CN119206377BActive Publication Date: 2025-06-06CHINA TELECOM CLOUD TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411689387.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-25
Publication Date
2025-06-06
Estimated Expiration
2044-11-25

AI Technical Summary

Technical Problem

Traditional object detection methods require the object detection process to be divided into two stages (prediction and deduplication), and end-to-end detection cannot be achieved, resulting in slower object detection speed.

Method used

End-to-end object detection is achieved by preprocessing, feature extraction, feature fusion, two-way branch feature processing and matching relationship determination of the input image.

Benefits of technology

End-to-end object detection is realized, and the target detection speed is improved, with fast detection speed and high accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119206377B_ABST
    Figure CN119206377B_ABST
Patent Text Reader

Abstract

The present application relates to a target detection method, device, equipment, readable storage medium and program product. The method comprises: preprocessing an input image to obtain a preprocessed image; extracting features from the preprocessed image to obtain a first set of feature pyramids; fusing features of the first set of feature pyramids through different transmission channels to obtain a fused feature image; performing feature processing on the fused feature image through two branches to obtain a predicted target category and predicted target coordinates; determining the matching relationship between the predicted target category and the predicted target coordinates to obtain a target detection result. This can reduce redundant inference pixel content, improve the inference speed in video surveillance scenarios, and directly complete end-to-end target detection with fast detection speed and high accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a target detection method, apparatus, computer equipment, computer-readable storage medium, and computer program product. Background Art

[0002] With the development of artificial intelligence technology, visual tasks for various scenarios have emerged to achieve rapid recognition and tracking of targets, etc. Among them, object detection is a common task in visual tasks, such as accurately predicting the location of targets in surveillance videos.

[0003] In traditional technology, the prediction network generally generates a certain number of candidate boxes (one-to-many) on the image, and then performs prediction target deduplication processing (for example, non-maximum suppression processing) on ​​the generated multiple candidate boxes, and finally selects a unique target candidate box for each object (many-to-one) to achieve one-to-one target detection.

[0004] However, the above target detection method needs to divide the target detection process into two stages (prediction and deduplication), which cannot achieve end-to-end detection, resulting in a slower target detection speed. Summary of the invention

[0005] Based on this, it is necessary to provide a target detection method, device, computer equipment, computer-readable storage medium and computer program product that can achieve end-to-end detection and improve the target detection speed in response to the above technical problems.

[0006] In a first aspect, the present application provides a target detection method, comprising:

[0007] Preprocessing the input image to obtain a preprocessed image;

[0008] Performing feature extraction on the preprocessed image to obtain a first set of feature pyramids;

[0009] Performing feature fusion on the first set of feature pyramids through different transmission channels to obtain a fused feature image;

[0010] The fused feature image is processed through two branches respectively to obtain a predicted target category and a predicted target coordinate;

[0011] A matching relationship between the predicted target category and the predicted target coordinates is determined to obtain a target detection result.

[0012] In one embodiment, preprocessing the input image to obtain a preprocessed image includes:

[0013] According to a preset size, the width and height of the input image are adjusted to obtain a preprocessed image.

[0014] In one embodiment, extracting features from the preprocessed image to obtain a first set of feature pyramids includes:

[0015] Build a deep residual network;

[0016] The preprocessed image is subjected to feature extraction through the deep residual network to obtain a first set of feature pyramids.

[0017] In one embodiment, the step of fusing the first set of feature pyramids through different transmission channels to obtain a fused feature image includes:

[0018] Inputting the first set of feature pyramids into a first transmission channel for slicing and aggregation processing to obtain a feature image of a first size;

[0019] Inputting the first set of feature pyramids into a second transmission channel for slicing and aggregation processing to obtain a feature image of a second size;

[0020] Inputting the first set of feature pyramids into a third transmission channel for slicing and aggregation processing to obtain a feature image of a third size;

[0021] After respectively fusing the feature image of the first size, the feature image of the second size, and the feature image of the third size with the feature images of corresponding sizes in the first group of feature pyramids, a second group of feature pyramids is obtained;

[0022] Down-sampling is performed on the second group of feature pyramids to obtain a fused feature image.

[0023] In one embodiment, the step of performing feature processing on the fused feature image through two branches to obtain a predicted target category and a predicted target coordinate includes:

[0024] Inputting the fused feature image into two parallel convolution layers respectively to obtain a first image feature and a second image feature;

[0025] Performing repeated prediction suppression processing on the first image feature and the second image feature through 3D maximum pooling to obtain a predicted target category;

[0026] Based on the second image feature, the target coordinates are predicted.

[0027] In one embodiment, determining the matching relationship between the predicted target category and the predicted target coordinates to obtain the target detection result includes:

[0028] Determine the matching quality between each predicted target category and the predicted target coordinates, and use the matching quality as a predicted label;

[0029] A matching cost matrix is ​​determined according to the predicted label, and a one-to-one matching relationship between the predicted target category and the predicted target coordinates is determined according to the matching cost matrix to obtain a target detection result.

[0030] In a second aspect, the present application also provides a target detection device, comprising:

[0031] A preprocessing module, used for preprocessing the input image to obtain a preprocessed image;

[0032] A feature extraction module, used to extract features from the preprocessed image to obtain a first set of feature pyramids;

[0033] A feature fusion module, used for performing feature fusion on the first set of feature pyramids through different transmission channels to obtain a fused feature image;

[0034] A prediction module, used for performing feature processing on the fused feature image through two branches respectively to obtain a predicted target category and predicted target coordinates;

[0035] The matching module is used to determine the matching relationship between the predicted target category and the predicted target coordinates to obtain a target detection result.

[0036] In a third aspect, the present application further provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0037] Preprocessing the input image to obtain a preprocessed image;

[0038] Performing feature extraction on the preprocessed image to obtain a first set of feature pyramids;

[0039] Performing feature fusion on the first set of feature pyramids through different transmission channels to obtain a fused feature image;

[0040] The fused feature image is processed through two branches respectively to obtain a predicted target category and a predicted target coordinate;

[0041] A matching relationship between the predicted target category and the predicted target coordinates is determined to obtain a target detection result.

[0042] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the following steps are implemented:

[0043] Preprocessing the input image to obtain a preprocessed image;

[0044] Performing feature extraction on the preprocessed image to obtain a first set of feature pyramids;

[0045] Performing feature fusion on the first set of feature pyramids through different transmission channels to obtain a fused feature image;

[0046] The fused feature image is processed through two branches respectively to obtain a predicted target category and a predicted target coordinate;

[0047] A matching relationship between the predicted target category and the predicted target coordinates is determined to obtain a target detection result.

[0048] In a fifth aspect, the present application further provides a computer program product, including a computer program, which implements the following steps when executed by a processor:

[0049] Preprocessing the input image to obtain a preprocessed image;

[0050] Performing feature extraction on the preprocessed image to obtain a first set of feature pyramids;

[0051] Performing feature fusion on the first set of feature pyramids through different transmission channels to obtain a fused feature image;

[0052] The fused feature image is processed through two branches respectively to obtain a predicted target category and a predicted target coordinate;

[0053] A matching relationship between the predicted target category and the predicted target coordinates is determined to obtain a target detection result.

[0054] The above-mentioned target detection method, device, computer equipment, computer-readable storage medium and computer program product obtain a preprocessed image by preprocessing the input image; thereby, the aspect ratio of the size of the video stream input can be matched, the redundant inference pixel content can be reduced, and the inference speed in the video surveillance scene can be improved. Feature extraction is performed on the preprocessed image to obtain a first set of feature pyramids; thereby, a set of images with different resolutions can be obtained, which is convenient for obtaining feature images of different sizes later. The first set of feature pyramids are feature fused through different transmission channels to obtain a fused feature image; thereby, the information contained in the fused feature image can be made more comprehensive and have stronger representation capabilities. The fused feature image is feature processed through two branches respectively to obtain a predicted target category and a predicted target coordinate; thereby, the category and the center coordinates of the annotation box can be directly obtained, and the prediction speed is faster. The matching relationship between the predicted target category and the predicted target coordinates is determined to obtain a target detection result. Thus, end-to-end target detection can be directly completed, with fast detection speed and high accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the drawings required for use in the embodiments of the present application or related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0056] Figure 1 A diagram showing an application environment of a target detection method in an embodiment;

[0057] Figure 2 Schematic diagram of a target detection method in one embodiment;

[0058] Figure 3 It is a schematic diagram of the principle of performing feature fusion processing on a feature pyramid in one embodiment;

[0059] Figure 4 Schematic diagram of target detection principle in a video surveillance scenario in one embodiment;

[0060] Figure 5 is a flow chart of a target detection method in another embodiment;

[0061] Figure 6 is a structural block diagram of a target detection device in one embodiment;

[0062] Figure 7 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0063] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0064] The target detection method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or it can be placed on the cloud or other network servers. The terminal 102 can obtain the input image and transmit the input image to the server 104 for processing, or it can be processed by the processor of the terminal 102 itself, and the processing result is displayed on the display interface of the terminal. Exemplarily, the terminal 102 or the server 104 pre-processes the input image to obtain a pre-processed image; the pre-processed image is subjected to feature extraction to obtain a first set of feature pyramids; the first set of feature pyramids are subjected to feature fusion through different transmission channels to obtain a fused feature image; the fused feature image is subjected to feature processing through two branches respectively to obtain a predicted target category and a predicted target coordinate; the matching relationship between the predicted target category and the predicted target coordinate is determined to obtain a target detection result. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, smart car devices, projection devices, etc. The portable wearable device may be a smart watch, a smart bracelet, a head-mounted device, etc. The head-mounted device may be a virtual reality (VR) device, an augmented reality (AR) device, smart glasses, etc. The server 104 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0065] In an exemplary embodiment, Figure 2 As shown, a target detection method is provided, which is applied to Figure 1 The terminal in is taken as an example to illustrate, including the following steps 201 to 206. Among them:

[0066] Step 201, preprocessing the input image to obtain a preprocessed image.

[0067] In this embodiment, since the sizes of video images collected by different terminals are inconsistent, the input image may be pre-processed to facilitate the subsequent processed image to fit the aspect ratio of the mainstream video surveillance image.

[0068] Exemplarily, the width and height of the input image can be adjusted according to a preset size to obtain a preprocessed image. For example, when a picture is input, the width w and height h of the picture can be scaled to obtain an image of 640x384 size, thereby fitting the aspect ratio of mainstream video surveillance images.

[0069] In this embodiment, by preprocessing the input image, the aspect ratio of the video stream input size can be matched as much as possible, redundant inference pixel content can be reduced, and the inference speed in the video surveillance scenario can be improved.

[0070] Step 202: extract features from the preprocessed image to obtain a first set of feature pyramids.

[0071] In this embodiment, the pyramid of an image is a set of images arranged in a pyramid shape with gradually decreasing resolutions and originating from the same original image, which are obtained by stepwise down sampling until a certain termination condition is reached.

[0072] In this embodiment, a deep residual network can be constructed; and features of the preprocessed image are extracted through the deep residual network to obtain a first set of feature pyramids.

[0073] Exemplarily, Resnet34 (a deep residual network) is used as the basic architecture network to extract features from images of uniform size, thereby obtaining a set of images of different resolutions, namely, the first set of feature pyramids.

[0074] Step 203 , performing feature fusion on the first set of feature pyramids through different transmission channels to obtain a fused feature image.

[0075] In this embodiment, different transmission channels refer to different sizes of images output after the transmission channels perform feature segmentation and aggregation.

[0076] Exemplarily, assuming that there are three transmission channels, the first group of feature pyramids is input into the first transmission channel for segmentation and aggregation processing to obtain a feature image of a first size; the first group of feature pyramids is input into the second transmission channel for segmentation and aggregation processing to obtain a feature image of a second size; the first group of feature pyramids is input into the third transmission channel for segmentation and aggregation processing to obtain a feature image of a third size; the feature image of the first size, the feature image of the second size, and the feature image of the third size are respectively fused with the feature images of corresponding sizes in the first group of feature pyramids to obtain a second group of feature pyramids; the second group of feature pyramids is downsampled to obtain a fused feature image.

[0077] In this embodiment, the transmission modules included in the three transmission channels are used to perform feature fusion processing on the first set of input feature pyramids. Figure 3 As shown, assuming that the first set of feature pyramids includes images of three different sizes, 1 / 8, 1 / 16, and 1 / 32, the three images of different sizes are respectively input into three transmission channels for cross-space and cross-scale feature fusion. Among them, the first transmission channel includes: ST module and GT module; the second transmission channel includes: RT module, ST module and GT module; the third transmission channel includes: RT module and ST module.

[0078] For example, the ST module (self-transformer) aims to capture the object features that appear simultaneously on the same feature image. Its essence is a spatial feature interaction. The ST module divides the input feature image into N patches and maps them into three parts q, k, and v through a transformation function (q represents q module, k represents k module, and v represents v module). express right The similarity value after attention calculation, represents the i-th q module, represents the jth k module. The calculated similarity value is then normalized using hybrid softmax (MoS). The specific calculation formula is as follows:

[0079]

[0080] in, represents the i-th aggregation weight, Indicates the number of modules, represents the jth noticed module, Represents the attention weight of the i-th module to the j-th module.

[0081] The attention weight After weighted calculation with v, the feature image output by the ST module is the same as the input scale of the ST module.

[0082] For example, the GT module (Grounding Transformer) is a top-down feature fusion that integrates the information in the high-level feature image into the low-level feature image. The output feature image has the same size as the input low-level feature image. The GT module also divides the input feature image into N pieces and maps them into three parts q, k, and v through a transformation function. Due to the cross-scale problem, in the GT module, right When doing attention calculation, Euclidean distance is used as the similarity value. The specific calculation formula is as follows:

[0083]

[0084] in, represents the i-th Euclidean distance.

[0085] The calculated similarity value is then normalized using hybrid softmax (MoS) as the normalization function to obtain , After weighted calculation with v, the feature image output by the GT module is the same size as the low-level feature image input to the GT module.

[0086] For example, the RT module (Rendering Transformer) is a bottom-up feature fusion, and the output scale is the same as the size of the high-level feature image. The RT module divides the input feature image into N pieces and maps them into three parts q, k, and v through a transformation function. express right After the attention calculation, the similarity value is normalized using the hybrid softmax (MoS) function. After weighted calculation with v, the feature image output by the RT module is the same size as the feature image input by the high-level RT module.

[0087] In this embodiment, the RT module, the ST module and the GT module are used to calculate the feature images of the corresponding cross-space and cross-scale information, and the obtained feature images are reordered according to the size and fused with the feature images of the corresponding size of the first group of feature pyramids. Finally, the dimension is reduced by 3x3 conv to finally obtain the fused feature image.

[0088] Exemplarily, the second set of feature pyramids obtained after network inference of the first set of feature pyramids is downsampled (stride=8), and the size of the final output fused feature image is H / 8, W / 8.

[0089] In this embodiment, the target rectangular frame is predicted by predicting the target center point and width and height, which can reduce the problem of poor generalization of the artificially designed anchor frame in the target detection process. Through three different modules, cross-space and cross-scale feature fusion is performed, so that the fused feature image output by the model can fuse multiple layers of features, making the fused feature image more capable of representation.

[0090] Step 204 , the fused feature image is subjected to feature processing through two branches respectively to obtain a predicted target category and predicted target coordinates.

[0091] Exemplarily, the fused feature image is input into two parallel convolutional layers to obtain a first image feature and a second image feature; the first image feature and the second image feature are subjected to repeated prediction suppression processing through 3D maximum pooling to obtain a predicted target category; and based on the second image feature, the target coordinates are predicted.

[0092] In this embodiment, it can be combined with Figure 4 As shown in FIG. 1 , after the 1 / 8 fused feature image obtained in step 203 passes through two parallel three 3x3 convolutional layers, it is divided into two branches: a regression branch and a classification branch. Among them, the target coordinates are predicted based on the feature output of the regression branch. In the classification branch, the features extracted by the convolutional layer are input to the 3D maximum pooling module, and the 3D maximum pooling is used to perform 3D pooling on the same channel input to obtain the output feature. , output features It is fused with the output features of the classification branch, and then the dimension is reduced through 3x3 convolution to finally obtain the predicted target category.

[0093] In this embodiment, embedding a 3D maximum pooling module can suppress the problem of repeated prediction in the most reliable predicted spatial area near the real labeled target. The 3D in the 3D maximum pooling module can suppress repeated prediction suppression between different scale layers of FPN.

[0094] Step 205, determining the matching relationship between the predicted target category and the predicted target coordinates to obtain a target detection result.

[0095] In this embodiment, when there is only one predicted target category and predicted target coordinates in the image, a one-to-one matching relationship between the predicted target category and the predicted target coordinates can be directly obtained. When there are two or more predicted target categories and / or predicted target coordinates, a one-to-one matching relationship between each predicted target category and predicted target coordinates must be clarified.

[0096] Exemplarily, the matching quality between each predicted target category and the predicted target coordinates can be determined, and the matching quality can be used as a prediction label; a matching cost matrix is ​​determined based on the prediction label, and a one-to-one matching relationship between the predicted target category and the predicted target coordinates is determined based on the matching cost matrix to obtain a target detection result.

[0097] In this embodiment, the calculation formula for the matching quality between each predicted target category and predicted target coordinates is as follows:

[0098]

[0099] in, Indicates the relationship between the i-th predicted target category and the The matching quality of the predicted target coordinates, represents the candidate prediction set of the i-th prediction target category, i.e., the spatial prior; represents the i-th spatial prior probability, represents the i-th prediction confidence, represents the i-th predicted detection box, Represents the i-th real detection box. At the same time, the classification probability and regression IOU output by the network are used A weighted geometric average was performed. Then the Hungarian algorithm was used to solve the matching cost matrix to complete the one-to-one matching relationship between the predicted target category and the predicted target coordinates, so that each predicted target coordinate corresponds to a predicted target category.

[0100] In the above target detection method, the input image is preprocessed to obtain a preprocessed image; thereby, the aspect ratio of the video stream input can be matched, the redundant inference pixel content can be reduced, and the inference speed in the video surveillance scenario can be improved. Feature extraction is performed on the preprocessed image to obtain a first set of feature pyramids; thereby, a set of images with different resolutions can be obtained, which is convenient for obtaining feature images of different sizes later. The first set of feature pyramids are feature fused through different transmission channels to obtain a fused feature image; thereby, the information contained in the fused feature image can be made more comprehensive and have stronger representation capabilities. The fused feature image is feature processed through two branches respectively to obtain a predicted target category and a predicted target coordinate; thereby, the category and the center coordinates of the annotation box can be directly obtained, and the prediction speed is faster. The matching relationship between the predicted target category and the predicted target coordinates is determined to obtain a target detection result. Thus, end-to-end target detection can be directly completed, with fast detection speed and high accuracy.

[0101] In another exemplary embodiment, Figure 5 As shown, a target detection method is provided, which is applied to Figure 1 The terminal in is used as an example to illustrate, including: first preprocessing the input image, extracting features from the preprocessed image, obtaining a set of feature images of different resolutions, fusing these feature images across space and scale, and generating attention fusion feature images. The fused feature images are then transmitted to two parallel convolutional layers (for example, three 3x3 convolutional layers) for feature extraction. One of the feature extraction results uses 3D maximum pooling to suppress repeated predictions, and the output features obtained are fused with the features output by the other convolutional layer to obtain the predicted target category. The output result of the other convolutional layer is the predicted target coordinates. Finally, the matching cost matrix is ​​solved according to the predicted target category and the predicted target coordinates to complete the one-to-one matching. Finally, end-to-end target detection is achieved.

[0102] In this embodiment, by preprocessing the input image, the model input size can be redesigned for the mainstream aspect ratio of existing video stream images to fit the video stream input size aspect ratio as much as possible, reduce redundant inference pixel content, and improve the inference speed in video surveillance scenarios. By predicting the target center point and width and height, the target rectangular box (bbox) is predicted, thereby solving the problem of poor generalization of the artificially designed anchor box (anchor) in the target detection process.

[0103] In this embodiment, by adding a 3D maximum pooling module, the problem of repeated prediction introduced by the FPN (multi-scale fusion) feature sharing structure is suppressed. The maximum pooling in the 3D maximum pooling module can suppress the repeated prediction problem in the spatial field near the most reliable prediction of the real annotated target, and the 3D in the 3D maximum pooling module can suppress the repeated prediction suppression between different scale layers of FPN. In this embodiment, by performing cross-space and cross-scale feature fusion on image features of different resolutions, multi-layer features can be fused, making the representation ability of the fused feature image stronger. In addition, the attention similarity calculation uses hybrid softmax (MoS) as the normalization function, which makes the attention information more comprehensive, which helps to improve the end-to-end target detection effect.

[0104] Compared with the existing method, the method in this embodiment does not require NMS processing, and directly completes the target detection (one-to-one) task in one network at the same time, realizing end-to-end target detection. Therefore, the model training converges faster, the model detection effect is better, and the reasoning speed is greatly improved.

[0105] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0106] Based on the same inventive concept, the embodiment of the present application also provides a target detection device for implementing the target detection method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more target detection device embodiments provided below can refer to the limitations on the target detection method above, and will not be repeated here.

[0107] In an exemplary embodiment, Figure 6 As shown, a target detection device is provided, including: a preprocessing module 601, a feature extraction module 602, a feature fusion module 603, a prediction module 604 and a matching module 605, wherein:

[0108] A preprocessing module 601 is used to preprocess the input image to obtain a preprocessed image;

[0109] A feature extraction module 602 is used to extract features from the preprocessed image to obtain a first set of feature pyramids;

[0110] A feature fusion module 603 is used to perform feature fusion on the first set of feature pyramids through different transmission channels to obtain a fused feature image;

[0111] A prediction module 604 is used to perform feature processing on the fused feature image through two branches to obtain a predicted target category and a predicted target coordinate;

[0112] The matching module 605 is used to determine the matching relationship between the predicted target category and the predicted target coordinates to obtain a target detection result.

[0113] Exemplarily, the preprocessing module 601 is specifically configured to adjust the width and height of the input image according to a preset size to obtain a preprocessed image.

[0114] Exemplarily, the feature extraction module 602 is specifically used to construct a deep residual network; and extract features from the preprocessed image through the deep residual network to obtain a first set of feature pyramids.

[0115] Exemplarily, the feature fusion module 603 is specifically used to input the first group of feature pyramids into the first transmission channel for segmentation and aggregation processing to obtain a feature image of a first size; input the first group of feature pyramids into the second transmission channel for segmentation and aggregation processing to obtain a feature image of a second size; input the first group of feature pyramids into the third transmission channel for segmentation and aggregation processing to obtain a feature image of a third size; fuse the feature image of the first size, the feature image of the second size, and the feature image of the third size with the feature images of corresponding sizes in the first group of feature pyramids respectively to obtain a second group of feature pyramids; and downsample the second group of feature pyramids to obtain a fused feature image.

[0116] Exemplarily, the prediction module 604 is specifically used to input the fused feature image into two parallel convolutional layers respectively to obtain a first image feature and a second image feature; suppress repeated prediction processing on the first image feature and the second image feature through 3D maximum pooling to obtain a predicted target category; and predict the target coordinates based on the second image feature.

[0117] Exemplarily, the matching module 605 is specifically used to determine the matching quality between each predicted target category and the predicted target coordinates, and use the matching quality as a prediction label; determine a matching cost matrix based on the prediction label, and determine a one-to-one matching relationship between the predicted target category and the predicted target coordinates based on the matching cost matrix to obtain a target detection result.

[0118] Each module in the above target detection device can be implemented in whole or in part by software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute operations corresponding to each module.

[0119] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in FIG. Figure 7 As shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface, the display unit and the input device are connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and the external device. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (Near Field Communication, NFC) or other technologies. When the computer program is executed by the processor, a target detection method is implemented. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device shell, or an external keyboard, touchpad or mouse.

[0120] Those skilled in the art will understand that Figure 7 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0121] In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.

[0122] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0123] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0124] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0125] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but are not limited to this.

[0126] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0127] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A target detection method, characterized in that: The method comprises: Preprocessing the input image to obtain a preprocessed image; Performing feature extraction on the preprocessed image to obtain a first set of feature pyramids; Performing feature fusion on the first set of feature pyramids through different transmission channels to obtain a fused feature image; The fused feature image is subjected to feature processing through two branches respectively to obtain a predicted target category and a predicted target coordinate; wherein, the fused feature image is subjected to feature processing through two branches respectively to obtain a predicted target category and a predicted target coordinate, including: inputting the fused feature image into two parallel convolution layers respectively to obtain a first image feature and a second image feature; performing repeated prediction suppression processing on the first image feature and the second image feature through 3D maximum pooling to obtain a predicted target category; and predicting the target coordinate based on the second image feature; Determine the matching relationship between the predicted target category and the predicted target coordinates to obtain a target detection result; wherein the matching relationship between the predicted target category and the predicted target coordinates is determined by the matching quality between each predicted target category and the predicted target coordinates; the matching quality is calculated as follows: in, Indicates i The predicted target category and The matching quality of the predicted target coordinates, Indicates i The candidate prediction set of predicted target categories, Indicates i A spatial prior probability, Indicates i prediction confidence, Indicates i Predicted detection boxes, Indicates i A real detection box; IOU represents the regression loss function, represents the geometric mean weighted.

2. The method according to claim 1, characterized in that The step of preprocessing the input image to obtain a preprocessed image includes: According to a preset size, the width and height of the input image are adjusted to obtain a preprocessed image.

3. The method according to claim 1, characterized in that The step of extracting features from the preprocessed image to obtain a first set of feature pyramids includes: Build a deep residual network; The preprocessed image is subjected to feature extraction through the deep residual network to obtain a first set of feature pyramids.

4. The method according to claim 1, characterized in that: The step of fusing the first set of feature pyramids through different transmission channels to obtain a fused feature image includes: Inputting the first set of feature pyramids into a first transmission channel for slicing and aggregation processing to obtain a feature image of a first size; Inputting the first set of feature pyramids into a second transmission channel for slicing and aggregation processing to obtain a feature image of a second size; Inputting the first set of feature pyramids into a third transmission channel for slicing and aggregation processing to obtain a feature image of a third size; After respectively fusing the feature image of the first size, the feature image of the second size, and the feature image of the third size with the feature images of corresponding sizes in the first group of feature pyramids, a second group of feature pyramids is obtained; Down-sampling is performed on the second group of feature pyramids to obtain a fused feature image.

5. The method according to any one of claims 1 to 4, characterized in that: The determining the matching relationship between the predicted target category and the predicted target coordinates to obtain the target detection result includes: Determine the matching quality between each predicted target category and the predicted target coordinates, and use the matching quality as a predicted label; A matching cost matrix is ​​determined according to the predicted label, and a one-to-one matching relationship between the predicted target category and the predicted target coordinates is determined according to the matching cost matrix to obtain a target detection result.

6. A target detection device, characterized in that: The device comprises: A preprocessing module, used for preprocessing the input image to obtain a preprocessed image; A feature extraction module, used to extract features from the preprocessed image to obtain a first set of feature pyramids; A feature fusion module, used for performing feature fusion on the first set of feature pyramids through different transmission channels to obtain a fused feature image; A prediction module is used to perform feature processing on the fused feature image through two branches respectively to obtain a predicted target category and a predicted target coordinate; wherein the prediction module is specifically used to: input the fused feature image into two parallel convolution layers respectively to obtain a first image feature and a second image feature; perform repeated prediction suppression processing on the first image feature and the second image feature through 3D maximum pooling to obtain a predicted target category; and predict the target coordinates based on the second image feature; The matching module is used to determine the matching relationship between the predicted target category and the predicted target coordinates to obtain the target detection result; wherein the matching relationship between the predicted target category and the predicted target coordinates is determined by the matching quality between each predicted target category and the predicted target coordinates; the matching quality is calculated as follows: in, Indicates i The predicted target category and The matching quality of the predicted target coordinates, Indicates i The candidate prediction set of predicted target categories, Indicates i A spatial prior probability, Indicates i prediction confidence, Indicates i Predicted detection boxes, Indicates i A real detection box; IOU represents the regression loss function, represents the geometric mean weighted.

7. The device according to claim 6, characterized in that The matching module is specifically used to: determine the matching quality between each predicted target category and the predicted target coordinates, and use the matching quality as a prediction label; A matching cost matrix is ​​determined according to the predicted label, and a one-to-one matching relationship between the predicted target category and the predicted target coordinates is determined according to the matching cost matrix to obtain a target detection result.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Satellite image target intelligent identification system and method based on improved SSD algorithm

    CN113963274A

  • Target tracking method and device, equipment and storage medium

    CN116109674A

  • Remote sensing image target detection method and component

    CN116824388A

  • Pest detection and identification method and system, electronic equipment and storage medium

    CN117373055A

  • Image matching method and device, equipment and storage medium

    CN118115765A