Target detection method and device, computing equipment and storage medium
By combining the skeleton network and the prediction head, and using self-supervised training methods and Gaussian heat maps to construct the loss function, the scale and rotation sensitivity problems of template matching detection are solved, efficient and real-time target detection is achieved, and the interpretability and adaptability of the model are improved.
Patent Information
- Application Number
- CN202511101361.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-08-07
AI Technical Summary
Template matching detection in existing technologies is sensitive to scale and rotation, has high computational cost, and deep learning target detection algorithms are highly data-dependent, have poor interpretability, and perform poorly with small samples.
The first skeleton network and the second skeleton network are used to extract the features of the template and the image to be detected. The pre-trained prediction head is used to predict the target confidence, bounding box and rotation angle. Zero-sample generalization capability is achieved through a small number of templates. The loss function is constructed by combining self-supervised training methods and Gaussian heat maps to improve the interpretability and real-time performance of the model.
It achieves real-time and robust target detection in different application scenarios, can realize efficient target detection with a small number of templates, improves the interpretability and practicality of the model, and does not require changing model parameters when adapting to new target types.
Smart Images

Figure CN120599237A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a target detection method, apparatus, computing device, and storage medium. Background Art
[0002] Template matching is a common algorithm in machine vision. It uses the correlation between a template and an object to determine information such as the template's position and angle on the object. Its core concept is to perform pixel-by-pixel comparisons using a sliding window, calculating the similarity between the template and a local area in the image to find the optimal matching position. This method is simple to implement and requires no training data. This technology is widely used in industrial quality inspection, medical imaging, and traffic monitoring. However, its drawback is that it is sensitive to scale. Scaling the object by a certain ratio can cause the matching to fail, necessitating a multi-scale search by constructing a scale pyramid. Furthermore, rotating the object by a certain angle can also cause the matching to fail, necessitating a one-to-one matching between the template and the object from multiple angles. Using template matching to detect objects with certain scale and rotation increases the computational cost.
[0003] Currently commonly used deep learning object detection algorithm architectures, such as Facer-RCNN and YOLO, essentially enable the model to learn the mapping relationship from pixels to semantics. Its core process includes feature extraction, region proposal, classification, and regression. Compared with template matching, it has advantages such as strong generalization, strong deformation robustness, multi-object processing, high computational efficiency, and strong adaptability to the target environment. However, its disadvantages include high data dependence and poor interpretability, which means that it is difficult to adjust parameters for optimization in real time and poor performance with small samples. Summary of the Invention
[0004] In view of this, the embodiments of the present application provide a target detection method to address the technical deficiencies in the prior art. The embodiments of the present application also provide a target detection apparatus, a computing device, and a computer-readable storage medium.
[0005] According to a first aspect of an embodiment of the present application, a target detection method is provided, comprising: Receive the template image and the image to be detected; Performing feature extraction on the template image based on the first skeleton network to obtain template features, and performing feature extraction on the image to be detected based on the second skeleton network to obtain features to be detected; Using the template feature as a convolution kernel and the feature to be detected as a convolution input, extracting the response output of the image to be detected with respect to the template image; Based on the pre-trained prediction head, the target confidence, bounding box, and rotation angle of the response output are predicted to obtain the target detection result of the template image in the image to be detected.
[0006] Optionally, the first skeleton network and the second skeleton network have the same network structure or different network structures; when the first network skeleton and the second network skeleton have the same network structure, the first network skeleton and the second network skeleton share weights or independently train weight parameters.
[0007] Optionally, the prediction head uses upsampling to align the prediction results of the target confidence, the bounding box, and the rotation angle with the resolution of the image to be detected.
[0008] Optionally, the prediction of the rotation angle is divided into positive and negative sign classification and rotation value regression.
[0009] Optionally, the training process of the prediction head includes: Based on the sample picture set, build a training picture set; Using the training picture set as the input of the prediction head to be trained, and outputting the prediction result; Constructing a loss function based on the training image set and the prediction result; The parameters of the prediction head to be trained are adjusted according to the loss function until the training conditions are met, thereby obtaining the pre-trained prediction head.
[0010] Optionally, constructing a training picture set based on the sample picture set includes: Randomly select a template center point in a sample template image in the sample image; Cutting the template according to the center point of the template, and randomly scaling and rotating the template to obtain a template to be pasted; Paste the template to be pasted to the sample picture to be detected in the sample picture set to obtain a training sample picture set, wherein the pasting process adopts Poisson fusion.
[0011] Optionally, constructing a loss function based on the training image set and the prediction result includes: Based on the training image set, constructing a training truth value through an adaptive Gaussian heat map; Constructing a first loss function corresponding to the center position of the template according to the prediction result and the training true value, wherein the first loss function adopts a cross entropy loss function; According to the training results and the training picture set, determine the second loss function corresponding to sub-pixel offset regression, the third loss function for scaling size regression, and the fourth loss function for rotation angle regression, wherein the second loss function and the third loss function adopt Smooth L1, and the loss function includes the first loss function, the second loss function, the third loss function and the fourth loss function.
[0012] According to a second aspect of an embodiment of the present application, there is provided an object detection device, comprising: A receiving module is configured to receive a template image and a picture to be detected; an extraction module configured to perform feature extraction on the template image based on the first skeleton network to obtain template features, and to perform feature extraction on the image to be detected based on the second skeleton network to obtain features to be detected; A convolution module is configured to use the template feature as a convolution kernel and the feature to be detected as a convolution input, and extract a response output of the image to be detected with respect to the template image; The prediction module is configured to perform target confidence, bounding box, and rotation angle prediction on the response output based on the pre-trained prediction head to obtain the target detection result of the template image in the image to be detected.
[0013] According to a third aspect of an embodiment of the present application, a computing device is provided, including: memory and processor; The memory is used to store computer-executable instructions, and the processor implements the steps of the target detection method when executing the computer-executable instructions.
[0014] According to a fourth aspect of an embodiment of the present application, a computer-readable storage medium is provided, which stores computer-executable instructions, and when the instructions are executed by a processor, the steps of the target detection method are implemented.
[0015] According to a fifth aspect of an embodiment of the present application, a chip is provided, which stores a computer program, and when the computer program is executed by the chip, the steps of the target detection method are implemented.
[0016] The target detection method provided by the present application receives a template image and a target image to be detected; extracts features of the template image based on a first skeleton network to obtain template features, and extracts features of the target image based on a second skeleton network to obtain features to be detected; uses the template features as convolution kernels and the features to be detected as convolution inputs to extract the response output of the target image in the target image; and predicts the target confidence, bounding box, and rotation angle of the response output based on a pre-trained prediction head to obtain the target detection result of the target image in the target image. A method is provided to achieve a detection effect close to that of supervised learning through a small number of templates. Specifically, by extracting the deep features of the template and the image, and then paying attention to each other, the position, size, and angle of the template target in the image are predicted, thereby improving the interpretability of the model. In addition, the practicality of target detection in practical applications can be greatly improved. That is, when a new target type to be detected is introduced, only the new target type needs to be provided as the input of the target detection model, and no model parameters need to be changed. At the same time, the real-time and robustness of target detection are guaranteed. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0018] Figure 1 This is a flow chart of a target detection method provided by an embodiment of the present application; Figure 2 This is a network structure diagram of a target detection method provided by an embodiment of the present application; Figure 3 This is a schematic structural diagram of a target detection device provided in one embodiment of the present application; Figure 4 This is a structural block diagram of a computing device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0019] The following description sets forth many specific details to facilitate a thorough understanding of the present application. However, the present application can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of the present application. Therefore, the present application is not limited to the specific implementations disclosed below.
[0020] The terms used in one or more embodiments of the present application are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of the present application. The singular forms "a", "the" and "the" used in one or more embodiments of the present application and the appended claims are also intended to include plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of the present application refers to and includes any or all possible combinations of one or more associated listed items.
[0021] It should be understood that although the terms "first," "second," and the like may be used to describe various information in one or more embodiments of the present application, such information should not be limited to these terms. These terms are merely used to distinguish information of the same type from one another. For example, "first" may also be referred to as "second," and similarly, "second" may also be referred to as "first," without departing from the scope of one or more embodiments of the present application.
[0022] This application provides a target detection method, a target detection apparatus, a computing device, and a computer-readable storage medium, which are described in detail in the following embodiments.
[0023] Figure 1 A flow chart of a target detection method according to an embodiment of the present application is shown, which specifically includes the following steps: Step S102: receiving a template image and a picture to be detected; Step S104: extracting features from the template image based on the first skeleton network to obtain template features, and extracting features from the image to be detected based on the second skeleton network to obtain features to be detected; Step S106: using the template feature as a convolution kernel and the feature to be detected as a convolution input, extracting a response output of the image to be detected with respect to the template image; Step S108: Based on the pre-trained prediction head, the target confidence, bounding box, and rotation angle prediction are performed on the response output to obtain the target detection result of the template image in the image to be detected.
[0024] Among them, the picture to be detected is the picture that needs to be recognized, and the template picture is the picture of the target that needs to be recognized; the first skeleton network is BackboneT and the second skeleton network is BackboneS, BackboneT extracts the deep features of the template picture, that is, the template features, and BackboneS extracts the deep features of the picture to be detected, that is, the features to be detected.
[0025] Based on this, the template features are used as convolution kernels, the features to be detected are used as convolution inputs, and the template features and the features to be detected are fused to achieve target detection. This can train a model with zero-shot generalization capability, and convert the template features into convolution kernels to participate in reasoning, so that the model only learns "how to match" rather than "remembering the template", giving the model zero-shot generalization capabilities across templates and scenarios.
[0026] The prediction head is then used to predict the target confidence, bounding box, and rotation angle. For target detection applications in different application scenarios, only a small number of target templates need to be provided to complete target detection in the image to be detected, that is, small sample target detection.
[0027] Furthermore, in step S104, the first skeleton network and the second skeleton network have the same network structure or different network structures; when the first network skeleton and the second network skeleton have the same network structure, the first network skeleton and the second network skeleton share weights or independently train weight parameters.
[0028] Furthermore, in step S108 , the prediction head uses upsampling to align the prediction results of the target confidence, the bounding box, and the rotation angle with the resolution of the image to be detected.
[0029] Furthermore, in step S108, the prediction of the rotation angle is divided into positive and negative sign classification and rotation value regression.
[0030] Furthermore, in step S108, the training process of the prediction head is specifically implemented as follows in this embodiment: Based on the sample picture set, a training picture set is constructed; the training picture set is used as the input of the prediction head to be trained, and the prediction result is output; a loss function is constructed based on the training picture set and the prediction result; the parameters of the prediction head to be trained are adjusted according to the loss function until the training conditions are met, thereby obtaining the pre-trained prediction head.
[0031] Furthermore, the process of constructing a training picture set based on the sample picture set is specifically implemented as follows in this embodiment: In the sample template picture in the sample picture, a template center point is randomly selected; the template is cropped according to the template center point, and the template is randomly scaled and randomly rotated to obtain a template to be pasted; the template to be pasted is pasted to the sample picture to be detected in the sample picture set to obtain a training sample picture set, wherein the pasting process adopts Poisson fusion.
[0032] Furthermore, the process of constructing the loss function based on the training image set and the prediction results is implemented as follows in this embodiment: Based on the training picture set, a training true value is constructed through an adaptive Gaussian heat map; according to the prediction result and the training true value, a first loss function corresponding to the template center position is constructed, wherein the first loss function adopts a cross entropy loss function; according to the training result and the training picture set, a second loss function corresponding to sub-pixel offset regression, a third loss function for scaling size regression, and a fourth loss function for rotation angle regression are determined, wherein the second loss function and the third loss function adopt Smooth L1, and the loss function includes the first loss function, the second loss function, the third loss function and the fourth loss function.
[0033] Among them, Figure 2 As shown in the network structure diagram of a target detection method, the template and the image to be detected are respectively extracted through BackboneT and BackboneS. Subsequently, DW_Conv (depthwise separable convolution) is used. The template features are used as the convolution kernel and the features to be detected are used as the convolution input to obtain the response of the template on the image to be detected. This achieves mutual attention between the template features and the features to be detected, fully utilizes the computational advantages of channel-by-channel convolution to extract spatial features, and achieves efficient matching between features. Head (the prediction head of the network structure) is then used to predict the target confidence, bounding box, and rotation angle of the DW_Conv response.
[0034] It should be noted that the above steps will cause the feature map to have a lower resolution than the original image. The Head not only needs to learn information such as the target position, size, angle, etc., but also needs to upsample the feature map to ensure that the result of the prediction head is aligned with the resolution of the image to be detected. Upsampling can be performed through PixelShuffle, deconvolution, interpolation, etc. The specific upsampling method adopted is determined by the actual usage scenario, and this embodiment does not limit it. Finally, based on the above results, the classification head and regression head output the prediction results, that is, the prediction of the target confidence, border, and rotation angle. In addition, in this embodiment, the rotation angle prediction is divided into two parts: positive and negative sign classification and rotation value regression, which effectively solves the problems of angle periodicity (-180° to +180°) and the influence of the value range scale.
[0035] Furthermore, a self-supervised training method is employed during training, using the target box regions annotated in open-source datasets such as CoCo and CUB-200 as templates. Pixel weight reorganization is then used to randomly paste the templates onto dataset images. Specifically, the templates are scaled and rotated to a certain extent before pasting, and Poisson fusion is introduced to avoid unnatural boundary artifacts in the pasted images. Specifically, each batch undergoes random data generation and processing before training. The basic principle is to randomly select the center point of the template on the first image, crop the template, and then paste it onto the images in the batch. During the pasting process, the templates need to be randomly scaled and rotated. The upsampling process uses pixel reorganization to perform feature upsampling, effectively avoiding the checkerboard artifacts seen in traditional upsampling methods, retaining the original position information, and upsampling to the original image size.
[0036] And the training truth value uses an adaptive Gaussian heat map as the GT, and the confidence of the template center is classified for each pixel on the output feature. At the same time, only the area where the GT confidence is greater than the confidence threshold is calculated to reduce the amount of calculation. The first loss function uses the cross entropy loss function. The result of this branch obtains the position of the template center. The first loss function is constructed by estimating the position of the center point based on the template target detection, avoiding problems such as uneven positive and negative samples and difficulty in setting anchor box hyperparameters. It should be noted that the confidence threshold value range is [0.3, 0.7], and preferably, the execution threshold value is 0.5. The application of sub-pixel positioning technology, taking advantage of the floating-point precision characteristics of the Gaussian heat map, and calculating the weighted center of all pixel positions around the predicted hotspot that exceed the set threshold, achieves more accurate target positioning than a single integer coordinate.
[0037] The prediction of the rotation angle is divided into two parts: sign classification and cosine value regression of the rotation angle. Specifically, the loss function of the rotation angle, that is, the fourth loss function, is divided into sign loss and angle value loss. Since the template target may be rotated at an angle of (-180, 180) in the image to be detected, the angle value predicted by the model is defined as the cosine value of the actual rotation angle. The model can fully express the rotation angle of the template target in the image to be detected by adding the sign prediction of the angle value, which effectively solves the problem of the mismatch between the angle periodicity (-180°, 180°) and the value range scale.
[0038] In addition, in order to avoid the problem of imbalance between positive and negative samples, that is, the serious imbalance between positive and negative samples in the prediction of the template target center point, logistic regression is performed through focal loss with penalty to effectively balance the training samples.
[0039] Corresponding to the above method embodiment, the present application also provides an embodiment of a target detection device, Figure 3FIG. 1 shows a schematic diagram of the structure of a target detection device provided by an embodiment of the present application. Figure 3 As shown, the device includes: The receiving module 302 is configured to receive a template image and a picture to be detected; The extraction module 304 is configured to perform feature extraction on the template image based on the first skeleton network to obtain template features, and perform feature extraction on the image to be detected based on the second skeleton network to obtain features to be detected; The convolution module 306 is configured to use the template feature as a convolution kernel and the feature to be detected as a convolution input, and extract a response output of the image to be detected with respect to the template image; The prediction module 308 is configured to perform target confidence, bounding box, and rotation angle prediction on the response output based on the pre-trained prediction head, and obtain the target detection result of the template image in the image to be detected.
[0040] In an optional embodiment, the receiving module 302 is further configured such that the first skeleton network and the second skeleton network have the same network structure or different network structures; when the first network skeleton and the second network skeleton have the same network structure, the first network skeleton and the second network skeleton share weights or independently train weight parameters.
[0041] In an optional embodiment, the prediction module 308 is further configured such that the prediction head uses upsampling to align the prediction results of the target confidence, the bounding box, and the rotation angle with the resolution of the image to be detected.
[0042] In an optional embodiment, the prediction module 308 is further configured to divide the prediction of the rotation angle into positive and negative sign classification and rotation value regression.
[0043] In an optional embodiment, the target detection device further includes: The training module is configured to construct a training picture set based on a sample picture set; use the training picture set as the input of the prediction head to be trained and output a prediction result; construct a loss function based on the training picture set and the prediction result; adjust the parameters of the prediction head to be trained according to the loss function until the training conditions are met, thereby obtaining the pre-trained prediction head.
[0044] In an optional embodiment, the training module is further configured to: In the sample template picture in the sample picture, a template center point is randomly selected; the template is cropped according to the template center point, and the template is randomly scaled and randomly rotated to obtain a template to be pasted; the template to be pasted is pasted to the sample picture to be detected in the sample picture set to obtain a training sample picture set, wherein the pasting process adopts Poisson fusion.
[0045] In an optional embodiment, the training module is further configured to: Based on the training picture set, a training true value is constructed through an adaptive Gaussian heat map; according to the prediction result and the training true value, a first loss function corresponding to the template center position is constructed, wherein the first loss function adopts a cross entropy loss function; according to the training result and the training picture set, a second loss function corresponding to sub-pixel offset regression, a third loss function for scaling size regression, and a fourth loss function for rotation angle regression are determined, wherein the second loss function and the third loss function adopt Smooth L1, and the loss function includes the first loss function, the second loss function, the third loss function and the fourth loss function.
[0046] The target detection device provided by the present application receives a template image and a target image to be detected; extracts features of the template image based on a first skeleton network to obtain template features, and extracts features of the target image based on a second skeleton network to obtain features to be detected; uses the template features as convolution kernels and the features to be detected as convolution inputs to extract the response output of the target image in the target image; and predicts the target confidence, bounding box, and rotation angle of the response output based on a pre-trained prediction head to obtain the target detection result of the target image in the target image. A method is provided to achieve a detection effect close to that of supervised learning using a small number of templates. Specifically, by extracting the deep features of the template and the image, and then performing mutual attention between the two, the position, size, and angle of the template target in the image are predicted, thereby improving the interpretability of the model. In addition, the practicality of target detection in practical applications can be greatly improved. That is, when a new target type to be detected is introduced, only the new target type needs to be provided as the input of the target detection model, and no model parameters need to be changed. At the same time, the real-time and robustness of target detection are guaranteed.
[0047] The above is a schematic scheme of a target detection device of this embodiment. It should be noted that the technical solution of the target detection device and the technical solution of the target detection method mentioned above belong to the same concept. For details that are not described in detail in the technical solution of the target detection device, please refer to the description of the technical solution of the target detection method mentioned above. In addition, the various components in the device embodiment should be understood as functional modules that must be established to implement each step of the program flow or each step of the method, and each functional module is not an actual functional division or separation definition. The device claim defined by such a group of functional modules should be understood as a functional module architecture that mainly implements the solution through the computer program recorded in the specification, and should not be understood as a physical device that mainly implements the solution through hardware.
[0048] Figure 4 4 shows a block diagram of a computing device 400 according to an embodiment of the present application. Components of the computing device 400 include, but are not limited to, a memory 410 and a processor 420. The processor 420 is connected to the memory 410 via a bus 430, and a database 450 is used to store data.
[0049] Computing device 400 also includes an access device 440 that enables computing device 400 to communicate via one or more networks 460. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. Access device 440 may include one or more of any type of network interface (e.g., a network interface card (NIC)), whether wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, a near field communication (NFC) interface, and the like.
[0050] In one embodiment of the present application, the above components of the computing device 400 and Figure 4 Other components not shown in the figure may also be connected to each other, for example, via a bus. Figure 4 The computing device structure block diagram shown is for illustrative purposes only and is not intended to limit the scope of the present application. Those skilled in the art may add or replace other components as needed.
[0051] Computing device 400 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or PC. Computing device 400 can also be a mobile or stationary server.
[0052] The processor 420 is configured to execute computer executable instructions for each step of the target detection method.
[0053] The above is a schematic diagram of a computing device according to this embodiment. It should be noted that the technical solution of the computing device and the technical solution of the target detection method described above are of the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the target detection method described above.
[0054] An embodiment of the present application further provides a computer-readable storage medium storing computer instructions, which are used to implement the steps of the target detection method when executed by a processor.
[0055] The above is a schematic diagram of a computer-readable storage medium of this embodiment. It should be noted that the technical solution of the storage medium and the technical solution of the target detection method described above are based on the same concept. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the target detection method described above.
[0056] An embodiment of the present application further provides a chip storing a computer program, which implements the steps of the target detection method when executed by the chip.
[0057] The foregoing description describes specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0058] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content of the computer-readable medium may be appropriately increased or decreased based on the requirements of legislation and patent practice within a jurisdiction. For example, in some jurisdictions, based on legislation and patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.
[0059] It should be noted that for the aforementioned method embodiments, for ease of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0060] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0061] The preferred embodiments of the present application disclosed above are intended only to help illustrate the present application. The optional embodiments do not describe all details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made based on the content of this application. This application selects and describes these embodiments in detail in order to better explain the principles and practical applications of this application, so that those skilled in the art can better understand and utilize this application. This application is limited only by the claims and their full scope and equivalents.
Claims
1. A target detection method, characterized in that: include: Receive the template image and the image to be detected; Performing feature extraction on the template image based on the first skeleton network to obtain template features, and performing feature extraction on the image to be detected based on the second skeleton network to obtain features to be detected; Using the template feature as a convolution kernel and the feature to be detected as a convolution input, extracting the response output of the image to be detected with respect to the template image; Based on the pre-trained prediction head, the target confidence, bounding box, and rotation angle of the response output are predicted to obtain the target detection result of the template image in the image to be detected.
2. The target detection method according to claim 1, characterized in that The first skeleton network and the second skeleton network have the same network structure or different network structures; when the first network skeleton and the second network skeleton have the same network structure, the first network skeleton and the second network skeleton share weights or independently train weight parameters.
3. The target detection method according to claim 1, characterized in that The prediction head uses upsampling to align the prediction results of the target confidence, the bounding box, and the rotation angle with the resolution of the image to be detected.
4. The target detection method according to claim 1, characterized in that The prediction of the rotation angle is divided into positive and negative sign classification and rotation value regression.
5. The target detection method according to claim 1, characterized in that: The training process of the prediction head includes: Based on the sample picture set, build a training picture set; Using the training picture set as the input of the prediction head to be trained, and outputting the prediction result; Constructing a loss function based on the training image set and the prediction result; The parameters of the prediction head to be trained are adjusted according to the loss function until the training conditions are met, thereby obtaining the pre-trained prediction head.
6. The target detection method according to claim 5, characterized in that: The step of constructing a training picture set based on the sample picture set includes: Randomly select a template center point in a sample template image in the sample image; Cutting the template according to the center point of the template, and randomly scaling and rotating the template to obtain a template to be pasted; Paste the template to be pasted to the sample picture to be detected in the sample picture set to obtain a training sample picture set, wherein the pasting process adopts Poisson fusion.
7. The target detection method according to claim 5, characterized in that: The constructing a loss function according to the training picture set and the prediction result includes: Based on the training image set, constructing a training truth value through an adaptive Gaussian heat map; Constructing a first loss function corresponding to the center position of the template according to the prediction result and the training true value, wherein the first loss function adopts a cross entropy loss function; According to the training results and the training picture set, determine the second loss function corresponding to sub-pixel offset regression, the third loss function for scaling size regression, and the fourth loss function for rotation angle regression, wherein the second loss function and the third loss function adopt Smooth L1, and the loss function includes the first loss function, the second loss function, the third loss function and the fourth loss function.
8. A target detection device, characterized in that: include: A receiving module is configured to receive a template image and a picture to be detected; an extraction module configured to perform feature extraction on the template image based on the first skeleton network to obtain template features, and to perform feature extraction on the image to be detected based on the second skeleton network to obtain features to be detected; A convolution module is configured to use the template feature as a convolution kernel and the feature to be detected as a convolution input, and extract a response output of the image to be detected with respect to the template image; The prediction module is configured to perform target confidence, bounding box, and rotation angle prediction on the response output based on the pre-trained prediction head to obtain the target detection result of the template image in the image to be detected.
9. A computing device, characterized in that include: memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the steps of the target detection method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing computer instructions, characterized in that: When the instruction is executed by the processor, the steps of the target detection method described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
End-to-end image template matching method based on twin network
CN113705731A
Target template graph matching and positioning method based on twin network and central position estimation
CN115330876A
Fast food rotating target detection method based on LR-Center Net
CN115713760A
Target following model training method and device, equipment and storage medium
CN117274296A
Object detection method and device, intelligent equipment and storage medium
CN117994551A