Data processing method, device, equipment and medium based on multi-level contrastive learning
Through multi-level contrastive learning and dynamic weight adjustment methods, the problems of occlusion, scale change and inter-domain differences in target detection are solved, the detection accuracy and generalization ability of the model are improved, and the dependence on labeled data is reduced.
Patent Information
- Application Number
- CN202511038423.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-07-28
AI Technical Summary
Existing target detection methods have insufficient ability to detect occluded targets, strong dependence on labeled data, poor adaptability to multi-scale targets, and inconsistency between pre-training source domain and target domain, resulting in reduced detection accuracy.
By obtaining the unlabeled first data set, the initial feature extractor is adjusted through multi-level contrastive learning, an intermediate feature extractor is constructed, and the initial target recognition model is adjusted using the labeled second data set. The final target recognition model is constructed by combining multi-level feature extraction and dynamic weight adjustment.
It improves the accuracy of target detection, enhances the ability to understand complex scenes, reduces dependence on large-scale labeled data, and improves the model's generalization ability and detection accuracy.
Smart Images

Figure CN120544204B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and in particular to a data processing method, device, equipment and medium based on multi-level contrastive learning. Background Art
[0002] Object detection aims to locate and identify objects of interest in images or videos. With the advancement of deep learning technology, object detection methods based on convolutional neural networks and Transformers have made significant progress. However, existing object detection methods still face the following challenges: Inadequate detection of occluded objects: In complex scenes, objects may be partially occluded, making it difficult for the model to accurately identify and localize them. Traditional object detection methods typically rely on global features and have limited sensitivity to local features, making them difficult to handle occlusion. Strong reliance on labeled data: Existing object detection methods typically require a large amount of accurately labeled training data, which is expensive and time-consuming to obtain, limiting their application in real-world scenarios. Poor adaptability to multi-scale objects: When the scale of an object varies significantly within an image, the model struggles to simultaneously capture both local details and global semantic information, resulting in reduced detection accuracy. Inconsistency between the pre-trained source and target domains: In real-world applications, pre-trained models are often trained on the source domain (e.g., public datasets), while the data distribution of the target domain (e.g., in real-world scenarios) may differ significantly from that of the source domain. This inter-domain difference will lead to a decrease in the performance of the model in the target domain, making it difficult to directly transfer applications. Summary of the Invention
[0003] In view of the above technical problems, the technical solution adopted by the present invention is:
[0004] According to a first aspect of the present invention, a data processing method based on multi-level contrastive learning is provided, the method comprising the following steps:
[0005] S100, obtaining a dataset to be processed, wherein the dataset to be processed includes a first dataset and a second dataset, wherein the first dataset is an unlabeled image set including original images and preprocessed images obtained by preprocessing the original images, and the second dataset is a labeled image set.
[0006] S200, using the first data set to adjust the adjustable variables of the initial feature extractor, to obtain the initial feature extractor with the adjusted adjustable variables as an intermediate feature extractor; wherein, in the process of adjusting the adjustable variables of the initial feature extractor, the adjustable variables of the initial feature extractor are adjusted based on the target adjustment value, and the target adjustment value includes a first adjustment value obtained based on contrastive learning and multi-level features of the original image, a second adjustment value obtained based on contrastive learning and the feature distribution of the original image, and a third adjustment value obtained based on the features of the original image and the preprocessed image.
[0007] S300: Building an initial target recognition model for identifying targets of interest based on the intermediate feature extractor.
[0008] S400: Using the second data set, adjust the adjustable variables of the initial target recognition model to obtain a final target recognition model.
[0009] According to a second aspect of the present invention, there is provided a data processing device based on multi-level contrastive learning, the device comprising:
[0010] The data acquisition module is used to acquire a data set to be processed, wherein the data set to be processed includes a first data set and a second data set, wherein the first data set is an unlabeled image set, including original images and preprocessed images obtained by preprocessing the original images, and the second data set is an labeled image set.
[0011] The first adjustment module is used to adjust the adjustable variables of the initial feature extractor using the first data set to obtain the initial feature extractor with the adjusted adjustable variables as an intermediate feature extractor; wherein, in the process of adjusting the adjustable variables of the initial feature extractor, the adjustable variables of the initial feature extractor are adjusted based on the target adjustment value, and the target adjustment value includes a first adjustment value obtained based on contrastive learning and multi-level features of the original image, a second adjustment value obtained based on contrastive learning and the feature distribution of the original image, and a third adjustment value obtained based on the features of the original image and the preprocessed image.
[0012] A model building module is used to build an initial target recognition model for identifying targets of interest based on the intermediate feature extractor.
[0013] The second adjustment module adjusts the adjustable variables of the initial target recognition model using the second data set to obtain a final target recognition model.
[0014] According to a third aspect of the present invention, there is provided an electronic device comprising a processor and a memory; the processor is configured to execute the steps of the method according to the first aspect by calling a program or instruction stored in the memory.
[0015] According to a fourth aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium is used to store a program or instruction, wherein the program or instruction enables a computer to execute the steps of the method described in the first aspect.
[0016] The present invention has at least the following beneficial effects:
[0017] The data processing method based on multi-level contrastive learning provided by an embodiment of the present invention includes: adjusting the adjustable variables of an initial feature extractor using an unlabeled first data set to obtain an initial feature extractor with adjusted adjustable variables as an intermediate feature extractor; constructing an initial target recognition model for identifying targets of interest based on the intermediate feature extractor; and adjusting the adjustable variables of the initial target recognition model using a labeled second data set to obtain a final target recognition model, wherein, in the process of adjusting the adjustable variables of the initial feature extractor, the adjustable variables of the initial feature extractor are adjusted based on target adjustment values, the target adjustment values including a first adjustment value obtained based on contrastive learning and multi-level features of the original image, a second adjustment value obtained based on contrastive learning and feature distribution of the original image, and a third adjustment value obtained based on features of the original image and the preprocessed image. The present invention can improve the accuracy of the target recognition model in target detection tasks.
[0018] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other aspects of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0020] Figure 1 This is a flowchart of a data processing method based on multi-level contrastive learning provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0021] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.
[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this invention pertains. The terms used herein in the specification of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0023] It should be noted that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the steps as sequential processes, many of the steps can be performed in parallel, concurrently, or simultaneously. In addition, the order of the steps can be rearranged. A process can be terminated when its operation is completed, but can also have additional steps not included in the accompanying drawings. A process can correspond to a method, function, procedure, subroutine, subprogram, etc.
[0024] The embodiment of the present invention provides a data processing method based on multi-level contrast learning, such as Figure 1 As shown, the method may include the following steps:
[0025] S100: Obtain a data set to be processed.
[0026] In an embodiment of the present invention, the data set to be processed may include a first data set and a second data set, wherein the first data set is an unlabeled image set, including original images and preprocessed images obtained by preprocessing the original images, and the second data set is an labeled image set.
[0027] In the embodiment of the present invention, the original image can be an image from a dataset such as AGVS-R, COCO, ImageNet, or other self-built datasets in business scenarios. In the embodiment of the present invention, the size of the original image is uniformly adjusted to an image with a resolution of 288×288.
[0028] In an embodiment of the present invention, the original image may be preprocessed using one or more of the following preprocessing methods:
[0029] Random cropping: Randomly select different areas in the image for cropping to capture target features from different perspectives.
[0030] Rotation and flipping: Rotate the image at random angles and enhance the data by horizontal / vertical flipping.
[0031] Scaling and resizing: Adjusting the image size to improve the model’s generalization ability to objects of different scales.
[0032] Color transformation: Randomly change brightness, contrast, and saturation to enhance the model’s robustness to lighting changes.
[0033] In the embodiment of the present invention, one or more preprocessing methods may be randomly selected to preprocess each original image to obtain a corresponding preprocessed image. If multiple preprocessing methods are selected, the resulting preprocessed image is a fusion image of the preprocessed images obtained by the multiple preprocessing methods, that is, each original image has only one preprocessed image.
[0034] In the embodiment of the present invention, by preprocessing the image using multiple preprocessing methods, the diversity of samples can be improved, and diverse samples can be provided for subsequent comparative learning, which is conducive to alleviating the model's dependence on specific image distribution.
[0035] In an embodiment of the present invention, the annotated image may be an image in a dataset such as AGVS-R, COCO, or ImageNet, or may be an annotated image in a business scenario.
[0036] S200 , adjusting the adjustable variables of the initial feature extractor using the first data set to obtain the initial feature extractor with the adjusted adjustable variables as an intermediate feature extractor.
[0037] Furthermore, S200 specifically includes:
[0038] S201, input the current batch of input image sets into the current feature extractor to obtain multi-level features of each image.
[0039] The input image set of the current batch may include the original image and the corresponding pre-processed image. The number of input image sets of the current batch can be set based on actual needs. The initial value of the current feature extractor is the initial feature extractor.
[0040] In an embodiment of the present invention, the initial feature extractor is a model for extracting semantic information from an image, for example, a model using a network architecture of a deep convolutional neural network (CNN) or a Transformer structure, or an architecture combining the two, specifically a ResNet50 model. The ResNet50 model uses residual connections or connection technology to transmit information between different layers, thereby achieving efficient fusion of low-level edge features and high-level semantic features.
[0041] In an embodiment of the present invention, features at different levels are used to present local features and global features of an image. In an illustrative embodiment, the first-level features, second-level features, and third-level features of an image can be obtained, that is, the multi-level features include first-level features, second-level features, and third-level features. The levels of semantic information represented by the first-level features to the third-level features increase in sequence. Among them, the first-level features are used to capture the edge, texture, and detail information in the image, and provide fine-grained feature support for local target relationship modeling. The second-level features are used to extract the semantic features of mesoscale targets, taking into account the information expression of local components of the target and global targets. The third-level features are used to extract the overall semantic information of the image, including the category, position, and spatial context of the target.
[0042] It is known to those skilled in the art that, for a given network structure, they clearly know which network layer outputs which features and at which level. In the case where the initial feature extractor is a ResNet50 model, the first-level features may be features output by stage 2 of the ResNet50 model, the second-level features may be features output by stage 3 of the ResNet50 model, and the third-level features may be features output by stage 4 of the ResNet50 model.
[0043] Those skilled in the art will recognize that any method for extracting features at different levels from image features obtained by a feature extractor falls within the scope of the present invention. Taking the ResNet50 model as an example, a predefined extraction dictionary, return_nodes, can be created, with key-value pairs representing the names of the layers in the ResNet50 model and the corresponding names of the extracted layers. For example, the key-value pair (5, 1) represents the fifth layer in the ResNet50 model, and 1 represents the first-level features. Next, the create_feature_extractor function is used to extract the corresponding hierarchical features from the features obtained from the ResNet50 model.
[0044] S202 : Based on the multi-level features of each image, obtain a target adjustment value for adjusting an adjustable variable of a feature extractor.
[0045] In the embodiment of the present invention, the adjustable variables of the feature extractor are model parameters of the feature extractor.
[0046] In an embodiment of the present invention, the target adjustment value may include a first adjustment value obtained based on contrastive learning and multi-level features of the original image, a second adjustment value obtained based on contrastive learning and the feature distribution of the original image, and a third adjustment value obtained based on the features of the original image and the preprocessed image.
[0047] In an embodiment of the present invention, the target adjustment value can be expressed as: Loss = f1 × L1 + f2 × L2 + f3 × L3. Where Loss is the target adjustment value, L1 is the first adjustment value, L2 is the second adjustment value, L3 is the third adjustment value, f1 to f3 are the first to third hyperparameters, respectively, and f1 + f2 + f3 = 1. In one exemplary embodiment, f1 = 0.8, and f2 = f3 = 0.1.
[0048] Furthermore, the first adjustment value corresponding to each original image is a weighted sum of the adjustment values corresponding to all hierarchical features corresponding to the image, wherein the adjustment value corresponding to each hierarchical feature is obtained based on contrast learning and the corresponding hierarchical feature.
[0049] In an embodiment of the present invention, the first adjustment value corresponding to the input image set of the current batch is equal to the sum of the first adjustment values corresponding to all original images in the input image set of the current batch.
[0050] Those skilled in the art know that each level of features in each image can be regarded as a feature map.
[0051] In this embodiment of the present invention, the first adjustment value of each original image in the current batch of input images can be expressed as:
[0052] .
[0053] Among them, L 1-single is the first adjustment value corresponding to each original image, L ci is the adjustment value corresponding to the i-th level feature of the image, α i For L ci The corresponding weight, i ranges from 1 to k, where k is the number of hierarchical features.
[0054] Furthermore, the adjustment value corresponding to each level feature of each original image can be obtained based on the following steps:
[0055] S10, dividing each level feature of each original image into multiple grid areas.
[0056] Specifically, S10 includes:
[0057] S1001 , performing dimensionality reduction processing on each hierarchical feature of each original image to obtain the corresponding hierarchical feature after dimensionality reduction processing as the dimensionality reduction feature of the hierarchical feature.
[0058] In an embodiment of the present invention, the unflod function can be used to perform dimensionality reduction processing on each level feature. In an illustrative embodiment, the initial sizes of the first to third level features may be [N, 36, 36, 512], [N, 18, 18, 1024], [N, 9, 9, 2048]. N is the number of original images in the input image set of the current batch. The size of the feature after dimensionality reduction may be [N, nc × b × b, np]. nc is the number of input channels, b is the size of the image patch in the hierarchical feature, and np is the number of image patches in the hierarchical feature. b can be determined based on actual conditions, as long as the divided image patches can cover the key areas of targets of different scales.
[0059] S1002: Divide the dimension reduction features of the hierarchical features into multiple grid areas.
[0060] In this embodiment of the present invention, each grid area may be an image area in the corresponding feature map and may be formed by c×c grids. One grid corresponds to one image patch, that is, each grid area is composed of the most adjacent c×c image patches in the corresponding feature map, and no two grid areas have the same image patch.
[0061] In an embodiment of the present invention, grid regions are obtained based on spatial local a priori rules, so that the subsequently obtained first-type grids and second-type grids can effectively represent local and global information. In one exemplary embodiment, c=3, that is, each grid region can be a rectangular area formed by 9 image patches.
[0062] S11 , based on the reference grid, the first type of grid, and the second type of grid in each grid area, using contrastive learning to obtain an adjustment value corresponding to the grid area.
[0063] Among them, the reference grid of each grid area is the central grid of the grid area, that is, the central image patch; the first type of grids of each grid area are the other grids corresponding to the first H similarities in the similarity set corresponding to the grid area; the second type of grids of each grid area are the other grids except the first type of grids among all other grids corresponding to the grid area; the similarity set corresponding to each grid area is the similarity set obtained by sorting the similarities between all other grids corresponding to the grid area and the reference grid in descending order; the other grids of each grid area are the grids except the reference grid in the grid area.
[0064] In the embodiment of the present invention, H may be 4. The similarity between each other grid and the reference grid may be the cosine similarity between the corresponding feature vectors.
[0065] Furthermore, the adjustment value L corresponding to the j-th grid area of the i-th level feature of each image isij The following conditions are met:
[0066] .
[0067] Among them, q j is the characteristic vector of the reference grid in the jth grid area, p r1 is the eigenvector corresponding to the r1th first-class grid in the jth grid area, r1 ranges from 1 to h1, and h1 is the number of first-class grids. r2 is the feature vector corresponding to the r2th second-class grid in the jth grid area. r1 ranges from 1 to h2, where h2 is the number of second-class grids. τ is the temperature coefficient used to adjust the distribution of similarity scores. exp() is an exponential function. j ranges from 1 to f(i), where f(i) is the number of grid areas corresponding to the i-th level feature.
[0068] S12, add up the adjustment values of all grid areas corresponding to each level feature of each original image to obtain the adjustment value corresponding to the level feature, i.e., L ci =L i1 +……+L ij +……+L if(i) .
[0069] In an embodiment of the present invention, since the adjustment value corresponding to each hierarchical feature takes into account the similarity between the feature vectors of the reference grid and the first and second types of grids, the model can learn the relationship between different categories of targets and local objects of the same type in the latent space, thereby enhancing the target recognition ability of the model. Furthermore, in an embodiment of the present invention, in the process of adjusting the adjustable variables of the initial feature extractor, the weight corresponding to the adjustment value corresponding to each hierarchical feature is adjusted according to a set adjustment cycle. The set update cycle can be an empirical value. In an exemplary embodiment, the set update cycle can be equal to 20 rounds, that is, the adjustable variables of each pair of feature extractors are adjusted for 20 rounds and updated once.
[0070] In the embodiment of the present invention, the current weight corresponding to the adjustment value corresponding to each level feature is determined based on the previous weight and the ratio of the change rate of the current adjustment value. Specifically, the adjustment corresponding to the i-th level feature to the corresponding current weight is adjusted as follows:
[0071] .
[0072] in, For L ci The corresponding current weight, For L ci The corresponding last weight, λ is a smoothing coefficient, which can be an empirical value. In an illustrative embodiment, λ=0.9. is the current adjustment value change rate corresponding to the i-th level feature, , is the rate of change of the current adjustment value corresponding to the i-th level feature, , is the current adjustment value corresponding to the i-th level feature, is the previous adjustment value corresponding to the i-th level feature, is the rate of change of the current adjustment value corresponding to the s-th level feature, , is the current adjustment value corresponding to the s-th level feature, is the previous adjustment value corresponding to the s-th level feature, and the value of s ranges from 1 to k.
[0073] In an embodiment of the present invention, in the process of adjusting the adjustable variables of the feature extractor, the adjustment result shows that the layer that converges quickly, that is, the layer whose current loss decreases quickly, is assigned a higher weight, while the layer that converges stably has its weight reduced to allow other layers to learn, thereby ensuring the optimal balance between the initial global feature learning and the later local feature refinement, thereby further improving the feature extraction capability of the feature extractor.
[0074] In the embodiment of the present invention, the second adjustment value may include an interaction adjustment value and a covariance adjustment value.
[0075] The interactive adjustment values corresponding to the input image set of the current batch can be obtained through the following steps:
[0076] S20, designating an image in the current batch of input image sets as a reference scene and other images as candidate scenes.
[0077] In an embodiment of the present invention, the reference scene may be the first original image in the input image set of the current batch.
[0078] S21 , obtaining an interaction value between each candidate scene and the reference scene, and sorting all obtained interaction values in ascending order to obtain sorted interaction values.
[0079] In this embodiment of the present invention, the interaction value between the reference scene and the candidate scene may satisfy the following conditions:
[0080] .
[0081] Where D is the interaction value between the benchmark scene and the candidate scene, q dw is the value of the wth element in the final feature of the benchmark scene, w ranges from 2 to R, R is the feature dimension of the final feature, p xw is the value of the wth element in the final feature of the candidate scene.
[0082] In the embodiment of the present invention, the final features of the image are the features finally output by the feature extractor, that is, the features output by the last layer of the feature extractor.
[0083] S22: The candidate scenarios corresponding to the first A interaction values among the sorted interaction values are taken as the first scenario, and the candidate scenarios corresponding to the interaction values other than the first A interaction values among the sorted interaction values are taken as the second scenario.
[0084] In an embodiment of the present invention, A may be equal to N / 8.
[0085] S23 , obtaining interaction adjustment values corresponding to the input image set of the current batch based on the final features of the reference scene, the first scene, and the second scene.
[0086] In this embodiment of the present invention, the interactive adjustment values corresponding to the current batch of input image sets meet the following conditions:
[0087] .
[0088] Among them, L KL is the interactive adjustment value corresponding to the input image set of the current batch, q d is the final feature of the benchmark scene, p u1 is the final feature of the first scene u1, the value of u1 ranges from 1 to g1, g1 is the number of the first scene, p u2 is the final feature of the u2th second scene, the value of u2 is g2, and g2 is the number of second scenes.
[0089] In an embodiment of the present invention, the interaction values of different scenes represented by different images in the same batch are included in the target adjustment value, so that the feature extractor can bring similar scenes closer to each other in the feature space, thereby clustering different scenes more accurately.
[0090] In this embodiment of the present invention, the covariance adjustment value is used to prevent feature collapse in the model. Feature collapse refers to the phenomenon in which the feature vectors of all samples converge to the same small region (all feature representations tend to 0) during contrastive learning, resulting in a loss of diversity in the feature space, an inability of the model to distinguish between samples of different categories, and a sharp decline in downstream task performance.
[0091] In this embodiment of the present invention, the covariance adjustment value corresponding to the current batch of input image sets satisfies the following conditions:
[0092] .
[0093] Among them, L cov is the covariance adjustment value corresponding to the input image set of the current batch, C vis the covariance matrix corresponding to the x-th original image in the input image set of the current batch, where x ranges from 1 to N and I is the identity matrix.
[0094] In an embodiment of the present invention, ,q e is the final feature of the e-th original image in the input image set of the current batch, the value of e ranges from 1 to N, Avg is the mean of the final features of N images, and T is the transposed sign.
[0095] In this embodiment of the present invention, the covariance adjustment value is used to force the off-diagonal elements of the covariance matrix to approach zero. This means that the vectors (different samples) in a batch do not contain redundant information. Using the identity matrix can force the feature variance of each dimension to be 1, thereby avoiding model collapse and achieving an efficient representation of the feature space.
[0096] In this embodiment of the present invention, the second adjustment value L2 corresponding to the input image set of the current batch is L KL +L cov .
[0097] In an embodiment of the present invention, the third adjustment value is adopted to introduce a potential representation consistency loss so that the feature representations of the original image and the data preprocessed image after passing through the backbone network model remain close, thereby accelerating the convergence of model training.
[0098] In this embodiment of the present invention, the third adjustment value corresponding to the current batch of input image sets satisfies the following conditions:
[0099] .
[0100] Among them, L3 is the third adjustment value corresponding to the input image set of the current batch, u a is the feature vector of the ath original image corresponding to the current input image set, where a ranges from 1 to N, and v a is the feature vector of the pre-processed image corresponding to the a-th original image, for u a The norm of v a The norm of , • represents the dot product.
[0101] In an embodiment of the present invention, the third adjustment value can promote more stable convergence of the feature extractor during the training process, improve the consistency of feature representation, improve the robustness of the model to different visual changes, and enable the model to have stronger generalization ability.
[0102] S203: If the target adjustment value is in the first state, the adjustable variable of the current feature extractor is adjusted using the target adjustment value, and S201 is executed. If the target adjustment value is in the second state, a feature extractor with adjusted adjustable variables is obtained as an intermediate feature extractor, and S300 is executed.
[0103] In this embodiment of the present invention, the first state is a non-converged state, and the second state is a converged state. If the target adjustment value converges, the adjustable variable of the feature extractor is no longer adjusted. Those skilled in the art will appreciate that any method of adjusting the adjustable variable of the feature extractor using the target adjustment value falls within the scope of protection of the present invention. If the target adjustment value does not converge, S201 is continued, performing iterative adjustments until the target adjustment value converges.
[0104] In an embodiment of the invention, the adjustable variables of the intermediate feature extractor can be used as pre-trained weights.
[0105] S300: Building an initial target recognition model for identifying targets of interest based on the intermediate feature extractor.
[0106] S300 may specifically include: replacing the backbone network model in the specified target recognition model with an intermediate feature extractor to obtain an initial target recognition model.
[0107] In an embodiment of the present invention, the designated target recognition model may be an RT-DETR model. It is known to those skilled in the art that, since the backbone network model in the designated target detection model is replaced with an intermediate feature extractor, the initial target recognition model has pre-trained weights.
[0108] S400: Using the second data set, adjust the adjustable variables of the initial target recognition model to obtain a final target recognition model.
[0109] In this embodiment of the present invention, labeled sample images are images that have been annotated with the category and bounding box of the object of interest within the image. The number of labeled sample images can be significantly smaller than the number of sample images used to train the initial feature extractor. Specifically, the initial object recognition model is fine-tuned end-to-end using a small amount of labeled data. The model parameters are optimized using supervised learning signals, enabling the output of object category and bounding box information after inputting labeled image data, thereby achieving high-precision detection and localization of objects in complex scenes.
[0110] In an embodiment of the present invention, in the process of adjusting the adjustable variables of the initial target recognition model, the model parameters can be jointly optimized through category label prediction and bounding box regression to improve the accuracy of the model in the target detection task.
[0111] Furthermore, in an embodiment of the present invention, during the adjustment of the adjustable variables of the initial target recognition model, some underlying network layers of the intermediate feature extractor (such as shallow convolutional layers or low-order Transformer modules) can be frozen. That is, during the update of the parameters of the initial target recognition model, only some parameters of the intermediate feature extractor are updated to maintain the stability of feature extraction. At the same time, only high-level task-related parameters are optimized to avoid overfitting, thereby ensuring that the model performs more stably and accurately in target detection tasks. The specific number of frozen underlying network layers can be obtained through experiments. In an exemplary embodiment, the first 40 layers of the ResNet50 model can be frozen. At this time, good results are achieved in both model convergence speed and accuracy.
[0112] In an embodiment of the present invention, by loading an intermediate feature extractor with pre-adjusted parameters onto the RT-DETR model, the features extracted by the RT-DETR model loaded with pre-trained weights will be more effectively represented in the image than those with ordinary initialized weights, enabling accurate modeling of target details and effectively improving the model's detection accuracy in complex scenarios. In addition, by freezing some of the underlying weights of the feature extractor and only fine-tuning the subsequent high-level parts and the detection head to adapt to specific tasks, it is possible to retain the captured image details while obtaining the features of different targets in the space.
[0113] In summary, the data processing method based on multi-level contrastive learning provided by the present invention can effectively solve problems such as occlusion, scale change, and inconsistency between the source and target domains in object detection. It can significantly improve the detection accuracy and generalization ability of the model while reducing the dependence on large-scale annotated data. It has broad application prospects, which are specifically reflected in:
[0114] (1) At the target object level, through multi-level feature extraction and hierarchical contrast learning, the model can simultaneously learn the local details and global semantic information of the target, improving the ability to understand complex targets.
[0115] (2) At the learning target relationship level, through the dynamic weight adjustment mechanism, the model can automatically adjust the weights of local and global features at different training stages, so as to achieve effective clustering of local and global features in the feature space.
[0116] (3) In terms of improving the stability of the pre-trained model during training, through feature distribution constraints and consistency learning, the model can avoid feature collapse and improve its robustness to different data augmentation strategies, thereby having better understanding capabilities in complex scenarios.
[0117] (4) In terms of improving target detection performance, the pre-trained model can obtain more detailed representations of the image in the feature space, so that during fine-tuning, the model can focus on some non-salient target areas in the image, enhance the perception of local details, and thus improve the detection accuracy in complex scenes such as when the target is partially occluded.
[0118] Based on the same inventive concept, an embodiment of the present invention further provides a data processing device based on multi-level contrastive learning, the device comprising:
[0119] The data acquisition module is used to acquire a data set to be processed, wherein the data set to be processed includes a first data set and a second data set, wherein the first data set is an unlabeled image set, including original images and preprocessed images obtained by preprocessing the original images, and the second data set is an labeled image set.
[0120] The first adjustment module is used to adjust the adjustable variables of the initial feature extractor using the first data set to obtain the initial feature extractor with the adjusted adjustable variables as an intermediate feature extractor; wherein, in the process of adjusting the adjustable variables of the initial feature extractor, the adjustable variables of the initial feature extractor are adjusted based on the target adjustment value, and the target adjustment value includes a first adjustment value obtained based on contrastive learning and multi-level features of the original image, a second adjustment value obtained based on contrastive learning and the feature distribution of the original image, and a third adjustment value obtained based on the features of the original image and the preprocessed image.
[0121] A model building module is used to build an initial target recognition model for identifying targets of interest based on the intermediate feature extractor.
[0122] The second adjustment module adjusts the adjustable variables of the initial target recognition model using the second data set to obtain a final target recognition model.
[0123] The device can be used to perform Figure 1 The method shown in the embodiment shown, therefore, for the functions that can be realized by each functional module of the device, please refer to Figure 1 The description of the illustrated embodiment is omitted for brevity.
[0124] An embodiment of the present invention also provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are configured to execute the method described in the embodiment of the present invention.
[0125] An embodiment of the present invention further provides a computer-readable storage medium storing computer-executable instructions, wherein the computer instructions are used to execute the method described in the embodiment of the present invention.
[0126] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present invention can be achieved. This is not limited herein.
[0127] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A data processing method based on multi-level contrastive learning, characterized in that: The method comprises the following steps: S100, obtaining a dataset to be processed, the dataset to be processed comprising a first dataset and a second dataset, wherein the first dataset is a set of unlabeled images, including original images and preprocessed images obtained by preprocessing the original images, and the second dataset is a set of labeled images; S200, adjusting the adjustable variables of the initial feature extractor using the first data set to obtain the initial feature extractor with the adjusted adjustable variables as the intermediate feature extractor; wherein, in the process of adjusting the adjustable variables of the initial feature extractor, the adjustable variables of the initial feature extractor are adjusted based on target adjustment values, the target adjustment values including a first adjustment value obtained based on contrastive learning and multi-level features of the original image, a second adjustment value obtained based on contrastive learning and feature distribution of the original image, and a third adjustment value obtained based on features of the original image and the preprocessed image; S300, constructing an initial target recognition model for identifying targets of interest based on the intermediate feature extractor; S400: Using the second data set, adjust the adjustable variables of the initial target recognition model to obtain a final target recognition model.
2. The method according to claim 1, characterized in that in, The first adjustment value corresponding to each original image is a weighted sum of the adjustment values corresponding to all hierarchical features corresponding to the image, wherein the adjustment value corresponding to each hierarchical feature is obtained based on contrast learning and the corresponding hierarchical feature.
3. The method according to claim 2, characterized in that In the process of adjusting the adjustable variables of the initial feature extractor, the weight corresponding to the adjustment value corresponding to each level feature is adjusted according to the set adjustment period.
4. The method according to claim 2, characterized in that The adjustment value corresponding to each level feature of each original image is obtained based on the following steps: Each level feature of each original image is divided into multiple grid regions; each grid region consists of c×c grids; Based on the reference grid, first-category grids, and second-category grids in each grid area, comparative learning is used to obtain the adjustment value corresponding to the grid area; wherein the reference grid of each grid area is the central grid of the grid area, the first-category grids of each grid area are the first H other grids in the similarity set corresponding to the grid area, the second-category grids of each grid area are the grids other than the first-category grids among all other grids corresponding to the grid area, the similarity set corresponding to each grid area is the similarity set obtained by sorting the similarities between all other grids corresponding to the grid area and the reference grid in descending order, and the other grids of each grid area are the grids other than the reference grid in the grid area; The adjustment values of all grid areas corresponding to each hierarchical feature of each original image are added together to obtain the adjustment value corresponding to the hierarchical feature.
5. The method according to claim 3, characterized in that The current weight corresponding to the adjustment value corresponding to each level feature is determined based on the ratio of the previous weight and the change rate of the current adjustment value.
6. A data processing device based on multi-level contrastive learning, characterized in that: The device comprises: a data acquisition module, configured to acquire a data set to be processed, wherein the data set to be processed includes a first data set and a second data set, wherein the first data set is a set of unlabeled images, including original images and preprocessed images obtained by preprocessing the original images, and the second data set is a set of labeled images; a first adjustment module, configured to adjust adjustable variables of an initial feature extractor using a first data set to obtain an initial feature extractor having adjusted adjustable variables, as an intermediate feature extractor; wherein, in the process of adjusting the adjustable variables of the initial feature extractor, the adjustable variables of the initial feature extractor are adjusted based on target adjustment values, the target adjustment values including a first adjustment value obtained based on contrastive learning and multi-level features of the original image, a second adjustment value obtained based on contrastive learning and a feature distribution of the original image, and a third adjustment value obtained based on features of the original image and the preprocessed image; A model building module, configured to build an initial target recognition model for identifying targets of interest based on the intermediate feature extractor; The second adjustment module adjusts the adjustable variables of the initial target recognition model using the second data set to obtain a final target recognition model.
7. An electronic device, characterized in that: including processor and memory; The processor is configured to execute the steps of the method according to any one of claims 1 to 5 by calling the program or instructions stored in the memory.
8. A computer-readable storage medium, characterized in that The computer-readable storage medium is used to store a program or instruction, and the program or instruction enables a computer to execute the steps of the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Man-machine asynchronous recognition model construction method and system based on active comparative learning
CN117828317A
State prediction method and apparatus, computer device, and storage medium
WO2025112186A1