Transform-based target detection pre-training method
Through a self-supervised pre-training framework based on multi-view contrastive learning, the problem of dependence on labeled data in target detection is solved, and efficient seamless migration from self-supervised pre-training to supervised learning is achieved, which improves detection accuracy and training speed and is suitable for a variety of visual tasks.
Patent Information
- Application Number
- CN202510927570.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-10-10
AI Technical Summary
The existing technology's reliance on large-scale labeled data in target detection tasks leads to high labor costs and time expenditure, and self-supervised pre-training methods have room for improvement in the micro-structural features and semantic interpretation capabilities of natural images.
A self-supervised pre-training framework based on a multi-view contrastive learning strategy is adopted. Images are processed through standardization and data augmentation to construct multi-view self-supervised contrastive learning samples. Combined with a four-stage pyramid structure backbone network, end-to-end training is performed to optimize the model and achieve seamless migration from self-supervised pre-training to supervised learning.
It effectively reduces the dependence on large-scale labeled data, improves the detection accuracy and training speed of the model, has good task scalability, and is suitable for visual tasks such as target detection, image classification, and semantic segmentation.
Smart Images

Figure CN120766030A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of self-supervised learning in deep learning, specifically a method based on Target detection pre-training method. Background Art
[0002] Technological advancements are advancing at a rapid pace, with deep learning driving innovation in fields such as computer vision, natural language processing, and speech recognition. The significant success of these models generally relies on large quantities of labeled samples. Taking image classification as an example, with the widespread application of deep learning in visual tasks, high-precision models are becoming increasingly reliant on large amounts of labeled data. This is especially true in object detection tasks, which require not only labeling image categories but also precise location of objects. Faced with the challenges of data growth and diverse scenarios, traditional supervised learning methods face significant labor costs and time overhead. Therefore, obtaining effective image features in the absence of labeled data has become a critical issue that needs to be addressed. Currently, research on self-supervised pre-training methods for object detection without labeled data has attracted considerable attention from scholars.
[0003] Self-supervised learning uses the structure of the data itself to mine supervision clues, without the need for manual labeling in advance, and uses a masked autoencoder It performs well in unsupervised representation learning. It uses a method of randomly blocking most areas of the image and then reconstructing the blocked parts. Even without labeling, it can obtain good feature expression. This feature expression can effectively improve the accuracy of downstream tasks such as image classification, target detection and semantic segmentation. The results show significant effects, but there is still room for further improvement in the ability to capture the microstructural features of natural images and interpret semantics, especially when multi-view contrast learning and When performing integration, how to adjust the comparison strategy and view construction to improve the quality of feature expression still needs further optimization. Summary of the Invention
[0004] The purpose of the present invention is to provide a The target detection pre-training method is based on An efficient self-supervised pre-training framework based on the architecture and multi-view contrastive learning strategy is designed to solve the dependence on large-scale labeled data and the high labor cost and time overhead in target detection based on traditional supervised learning, and to enable the framework to have good task scalability and can be applied to visual tasks such as target detection, image classification, semantic segmentation, and instance segmentation.
[0005] The purpose of the present invention is achieved through the following technical solutions:
[0006] A kind of A target detection pre-training method, the method comprising:
[0007] Step 1 At the beginning of the encoding phase, each input image is normalized, the input image is adjusted to a pixel matrix of fixed size, and the pixel values are normalized to a specific range; each input standardized image is data augmented to construct its enhanced view, which is consistent with the original view. Figure One The original view and the enhanced view are then divided into multiple image blocks of fixed size, and the blocks are randomly masked with the same mask position and then input into the two branches respectively. encoder and ;
[0008] Step 2 The output of is processed by the projector and predictor, and The output is used for similarity calculation to maximize the feature similarity between different views of the same image and minimize the feature similarity between different views of different images, where The total loss function comprehensively considers the reconstruction task and the contrastive learning task; the original image branch is updated by back propagation, and the enhanced branch is updated by Dynamic update, in each iteration, The parameters will be frozen. pass right The parameters of are updated to gradually optimize the model;
[0009] Step 3: Use a four-stage pyramid structure As the backbone network to build the downstream detection framework, Pre-trained Direct initialization of encoder weights The corresponding layer weights are then fine-tuned to adapt to the downstream target detection task; Each stage produces feature maps of different resolutions, effectively alleviating The computational complexity is too high when processing high-resolution images, and The pre-trained encoder is transferred to the object detection task;
[0010] Step 4: Adopt The structure further The output multi-scale features are fused to meet the detection requirements of targets of different sizes; the detection head uses shared classification and regression branches, the classification branch is responsible for target category prediction, and uses As the loss function, the regression branch is responsible for bounding box position prediction, using Optimize the accuracy of bounding box regression; through this end-to-end training method, the pre-trained general features can adapt to the target detection task, realizing seamless migration from self-supervised pre-training to supervised learning.
[0011] The beneficial effects of the present scheme are: to solve the dependence on large-scale labeled data in target detection based on traditional supervised learning and the high labor cost and time overhead it brings, and to make the framework have good task expansion to be applied in various scenes.
[0012] Further, in step 1, the beginning of the CL-MAE encoding stage, first, the input image is subjected to standardization preprocessing operation, the image size is adjusted to pixels, and the pixel value is normalized to range; four data enhancement means of random cropping, random rotation, color jittering and Gaussian blur are applied to each input image for data enhancement, to construct the original view V pure and the enhanced view V change .Subsequently, the input image is divided into fixed-size image blocks (Patches) according to The size of each is pixels, so The image will be divided into image blocks. Next, 75% of the high mask ratio is used to randomly mask the input image blocks, and 75% of the image blocks are marked as masked state, only 25% of the visible original image blocks and enhanced image blocks are input into the two branches Encoder and In the double-branch architecture, in order to ensure the fairness and effectiveness of contrastive learning, the same mask position is used in both branches.
[0013] Further, in step 2, after encoding, two feature vectors in the latent space are obtained, denoted as and , then is projected into another feature space, and the projected feature vector is denoted as , then a projection head is used to further process through a series of fully connected layers and activation functions, first projected into an intermediate dimension, then projected into the target dimension, and nonlinearly transformed using the activation function. The output of is processed by the projector and the predictor, and the similarity between the two is calculated, which maximizes the feature similarity between different views of the same image, improving the model's understanding of image semantic information. The decoder part adopts a lightweight design and is mainly responsible for the image reconstruction task in the pre-training stage. At the same time, it also needs to cooperate with the contrastive learning module to ensure that the features extracted by the encoder can support both the reconstruction task and the contrastive learning requirements. The total loss function takes into account the reconstruction task and the contrastive learning task: .in is the contrastive learning loss, , , ; To rebuild the losses, , and They refer to the masked area in the original image and the corresponding reconstruction result, represents the set of masked pixel positions, Represents the total number of mask pixels; , is the balance coefficient. The original branch is updated by back propagation, and the enhanced branch is updated by Technology is updated dynamically, in each iteration, The parameters will be frozen. pass right The parameters of are updated to gradually optimize the model. ,in, is the smoothing coefficient (also called attenuation factor), and its value range is , is the previous time step The exponential moving average of .
[0014] Furthermore, in step 3, As the backbone network to build the downstream detection framework, The backbone network adopts a four-stage pyramid structure, and each stage produces feature maps of different resolutions. Specifically, the input image size after preprocessing is , after four stages of processing, the feature map sizes are 、 、 and , the corresponding feature dimensions are 64, 128, 320 and 512 respectively. In the process of transfer learning, Pre-trained Direct initialization of encoder weights The corresponding layer of is fine-tuned to adapt to the feature extraction requirements of the target detection task.
[0015] Furthermore, in step 4, Structural pair The multi-scale features of the output are fused; The output feature maps of the four stages are obtained by The convolution is unified to 256 channels, and then passed Perform upsampling and feature fusion operations; the detection head uses shared classification and regression branches, the classification branch is responsible for target category prediction, and uses As the loss function to deal with the problem of class imbalance; the regression branch is responsible for bounding box position prediction, using Optimize bounding box regression accuracy; the loss function of the entire detection network is designed to be ,in, is the classification loss, is the regression loss, is the balancing coefficient; through this end-to-end training method, the general features of the pre-trained encoder are effectively adapted to the specific needs of the target detection task, achieving seamless migration from self-supervised pre-training to supervised learning. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0017] Figure 1 In the specific embodiment of the present invention, Flowchart of target detection pre-training method;
[0018] Figure 2 In the specific embodiment of the present invention, Target detection pre-training method Pre-training self-supervised framework diagram;
[0019] Figure 3 In the specific embodiment of the present invention, Model training flowchart of the target detection pre-training method;
[0020] Figure 4 In the specific embodiment of the present invention Pre-trained and non-pre-trained models on Change trend chart. DETAILED DESCRIPTION
[0021] The following is a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments, and do not constitute a limitation of the present invention. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0022] like Figure 1 、 Figure 2 、 Figure 3 As shown, this example provides The target detection pre-training method includes the following steps:
[0023] Step 1 At the beginning of the encoding phase, each input image is normalized, the input image is adjusted to a pixel matrix of fixed size, and the pixel values are normalized to a specific range; each input standardized image is data augmented to construct its enhanced view, which is consistent with the original view. Figure One The original view and the enhanced view are then divided into multiple image blocks of fixed size, and the blocks are randomly masked with the same mask position and then input into the two branches respectively. encoder and ;
[0024] In this step, the input image is preprocessed to adjust the image size to pixels and normalize the pixel values to Range; Apply random cropping, random rotation, color jittering, and Gaussian blur to enhance the data of each input image and construct the original view V pure and Enhanced View V change Then follow The processing method is to divide the input image into image blocks of fixed size ( ), each The size is pixels, so The image will be split into Next, a high mask ratio of 75% is used to randomly mask the input image blocks, marking 75% of the image blocks as masked, and only 25% of the visible original image blocks and enhanced image blocks are retained and input into the two branches respectively. encoder and In the dual-branch architecture, both branches use the same mask position.
[0025] Step 2 and the output of the projector and the predictor are used to calculate the similarity between the two feature vectors The similarity between the two feature vectors is calculated, and the feature similarity between different views of the same image is maximized and the feature similarity between different views of different images is minimized The total loss function of the model takes into account both the reconstruction task and the contrastive learning task; the original image branch is updated through backpropagation, and the enhanced branch is updated through dynamic updates, in each iteration, the parameters of the model are frozen, the parameters of the model are updated through the parameters of the model are updated through the parameters of the model are updated through
[0026] In this step, the feature vectors in the two latent spaces obtained after encoding are denoted as and First, project into another feature space, and the projected feature vector is denoted as . Then use a predictor to further process through a series of fully connected layers and activation functions, first project into an intermediate dimension, then project into the target dimension, and use the activation function for nonlinear transformation. and the output of the projector and the predictor are used to calculate the similarity between the two feature vectors The similarity between the two feature vectors is calculated, and the feature similarity between different views of the same image is maximized to improve the model's understanding of image semantic information. The decoder part of the model adopts a lightweight design, taking on the image reconstruction task in the pre-training stage and cooperating with the contrastive learning module. The total loss function of the model takes into account both the reconstruction task and the contrastive learning task: . Among them is the contrastive learning loss, , , ; is the reconstruction loss, , and respectively represent the masked region in the original image and the corresponding reconstruction result, denotes the set of masked pixel positions, denotes the total number of masked pixels; , is the balance coefficient. The original image branch is updated through backpropagation, and the enhanced branch is updated through technique dynamic updates, in each iteration, the parameters of the model are frozen, the parameters of the model are updated through the parameters of the model are updated through the parameters of the model are updated through ,in, is the smoothing coefficient (also called attenuation factor), and its value range is , is the previous time step The exponential moving average of .
[0027] Step 3: Use a four-stage pyramid structure As the backbone network to build the downstream detection framework, Pre-trained Direct initialization of encoder weights The corresponding layer weights are then fine-tuned to adapt to the downstream target detection task; Each stage produces feature maps of different resolutions, effectively alleviating The computational complexity is too high when processing high-resolution images, and The pre-trained encoder is transferred to the object detection task;
[0028] In this step, the As the backbone network to build the downstream detection framework, The backbone network adopts a four-stage pyramid structure, and each stage produces feature maps of different resolutions. The input image size after preprocessing is , after four stages of processing, the feature map sizes are 、 、 and , the corresponding feature dimensions are 64, 128, 320 and 512 respectively. In the process of transfer learning, Pre-trained Direct initialization of encoder weights The corresponding layer of is fine-tuned to adapt to the feature extraction requirements of the target detection task.
[0029] Step 4: Adopt The structure further The output multi-scale features are fused to meet the detection requirements of targets of different sizes; the detection head uses shared classification and regression branches, the classification branch is responsible for target category prediction, and uses As the loss function, the regression branch is responsible for bounding box position prediction, using Optimize bounding box regression accuracy; through this end-to-end training method, the general features obtained by pre-training can be adapted to the target detection task, achieving seamless migration from self-supervised pre-training to supervised learning.
[0030] In this step, the Structural pair The multi-scale features of the output are fused; The output feature maps of the four stages are obtained by The convolution is unified to 256 channels, and then passed Perform upsampling and feature fusion operations; the detection head uses shared classification and regression branches, the classification branch is responsible for target category prediction, and uses As the loss function to deal with the problem of class imbalance; the regression branch is responsible for bounding box position prediction, using Optimize bounding box regression accuracy; the loss function of the entire detection network is designed to be ,in, is the classification loss, is the regression loss, is the balance coefficient.
[0031] Compared with the untrained model, the self-supervised pre-training of the embodiment has higher detection accuracy and faster convergence speed.
[0032] Detection accuracy and convergence speed analysis: From Figure 4 It can be analyzed that the self-supervised pre-training model and the non-pre-training model have the same training time within 100 training cycles. The trend of changes in indicators. It can be observed that the pre-trained model Inside The improvement is significant, and the convergence speed is significantly better than that of the non-pre-trained model; the latter grows slowly in the early stage of training, and Then it stabilized and eventually About 79%. In contrast, the pre-trained model It is close to the optimal level and eventually reaches more than 80% , showing higher detection accuracy and faster convergence performance. This result shows that The pre-training strategy can not only significantly improve the performance of downstream object detection models, but also effectively accelerate the model training process, especially for application scenarios with limited resources or limited training time.
[0033] The above is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes based on the technical solutions and concepts of the present invention within the scope disclosed by the present invention, which fall within the scope of protection of the present invention.
Claims
1. A method based on The target detection pre-training method is characterized by: The method comprises: Step 1: Perform data augmentation on each normalized image to construct its enhanced view, which is used together with the original view as a training sample for multi-view self-supervised contrastive learning. The original view and the enhanced view are then split into multiple image blocks of fixed size, which are randomly masked with the same mask position and then input into two branches respectively. encoder and ; Step 2 The output of is processed by the projector and predictor, and The output is used for similarity calculation to maximize the feature similarity between different views of the same image and minimize the feature similarity between different views of different images, where The total loss function comprehensively considers the reconstruction task and the contrastive learning task; the original image branch is updated by back propagation, and the enhanced branch is updated by Dynamic update, in each iteration, The parameters will be frozen. pass right The parameters of are updated to gradually optimize the model; Step 3: Use a four-stage pyramid structure As the backbone network to build the downstream detection framework, Pre-trained Direct initialization of encoder weights The corresponding layer weights are then fine-tuned to adapt to the downstream target detection task; Each stage produces feature maps of different resolutions, effectively alleviating The computational complexity is too high when processing high-resolution images, and The pre-trained encoder is transferred to the object detection task; Step 4: Adopt The structure further The multi-scale features of the output are fused to meet the detection requirements of targets of different sizes; the detection head uses shared classification and regression branches, the classification branch is responsible for target category prediction, and uses As the loss function, the regression branch is responsible for bounding box position prediction, using Optimize bounding box regression accuracy; through this end-to-end training method, the general features obtained by pre-training can be adapted to the target detection task, achieving seamless migration from self-supervised pre-training to supervised learning.
2. According to claim 1 The target detection pre-training method is characterized by: The process of step 1 is specifically as follows: Apply random cropping, random rotation, color jittering, and Gaussian blur to enhance the data of each input standardized image and construct the original view V pure and Enhanced View V change ; Then follow The input image is divided into image blocks of fixed size. Then, a high mask ratio is used to randomly mask the input image blocks, and the high-ratio image blocks are marked as masked. Only the low-ratio visible original image blocks and enhanced image blocks are retained and input into two branches respectively. encoder and In the dual-branch architecture, to ensure the fairness and effectiveness of contrastive learning, the two branches use the same mask position.
3. The method according to claim 1 The target detection pre-training method is characterized by: The process of step 2 is specifically as follows: After encoding, we get two feature vectors in the latent space, denoted as and , then first Projected into another feature space, the feature vector after projection is , then uses a projection head to pass through a series of fully connected layers and activation functions Further processing is performed by projecting it into an intermediate dimension, then projecting it into the target dimension, and performing nonlinear transformation using activation functions; The output of is processed by the projector and predictor, and The output is used for similarity calculation, which improves the model's ability to understand image semantic information by maximizing the feature similarity between different views of the same image. The decoder part adopts a lightweight design and is mainly responsible for the image reconstruction task in the pre-training stage. It also needs to cooperate with the contrastive learning module to ensure that the features extracted by the encoder can support both the reconstruction task and the contrastive learning requirements. The total loss function takes into account the reconstruction task and the contrastive learning task: ;in is the contrastive learning loss, , , ; To rebuild the losses, , and They refer to the masked area in the original image and the corresponding reconstruction result, represents the set of masked pixel positions, represents the total number of mask pixels, , is the balance coefficient; The original branch is updated by back propagation, and the enhanced branch is updated by Technology is updated dynamically, in each iteration, The parameters will be frozen. pass right The parameters of are updated to gradually optimize the model. ,in, is the smoothing coefficient, and its value range is , is the previous time step The exponential moving average of .
4. The method according to claim 1 The target detection pre-training method is characterized by: The process of step 3 is specifically as follows: use As the backbone network to build the downstream detection framework, The backbone network adopts a four-stage pyramid structure, and each stage produces feature maps of different resolutions; specifically, the input image size after preprocessing is , after four stages of processing, the feature map sizes are 、 、 and , the corresponding feature dimensions are 64, 128, 320 and 512 respectively; in the process of transfer learning, Pre-trained Direct initialization of encoder weights The corresponding layer of is fine-tuned to adapt to the feature extraction requirements of the target detection task.
5. The method according to claim 1 The target detection pre-training method is characterized by: The process of step 4 is specifically as follows: use Structural pair The output multi-scale features are fused to meet the detection requirements of targets of different sizes. The output feature maps of the four stages are obtained by The convolution is unified to 256 channels, and then passed Perform upsampling and feature fusion operations; the detection head uses shared classification and regression branches, the classification branch is responsible for target category prediction, and uses As the loss function to deal with the problem of class imbalance; the regression branch is responsible for bounding box position prediction, using Optimize bounding box regression accuracy; the loss function of the entire detection network is designed to be ,in, is the classification loss, is the regression loss, is the balancing coefficient; through this end-to-end training method, the general features of the pre-trained encoder are effectively adapted to the specific needs of the target detection task, achieving seamless migration from self-supervised pre-training to supervised learning.
Citation Information
Cited By
Tunnel surface detection method and device and electronic equipment
CN121353258A
Pattern recognition method based on adaptive target enhancement and contrast learning framework
CN121686178A
Double-stage pre-training system for reading of industrial inspection instrument
CN121960652A
Weak supervision image enhancement method based on land codebook prior and contrast constraint
CN122243839A
Weakly supervised image enhancement method based on terrestrial codebook prior and contrast constraint
CN122243839B