Training method for pre-training image feature extraction model and related device
By introducing SAM and OCR models to generate multidimensional local labels, and combining ViT structure and masking mechanism, the problem of traditional image pre-trained models being unable to handle the joint representation of visual objects and OCR text in complex visual scenes is solved, and the model's ability to capture fine-grained information and its robustness are improved.
Patent Information
- Application Number
- CN202511021913.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-10-31
AI Technical Summary
Traditional image pre-training models struggle to achieve joint representation of visual objects and OCR text when handling dense prediction tasks, especially in scenarios with complex visual elements and textual information.
By introducing the Segmentation All Model (SAM) and Optical Character Recognition (OCR) model, multidimensional local labels are automatically generated. Combined with the Vision Transformer (ViT) structure and masking mechanism, image feature extraction and parameter adjustment are performed to achieve joint annotation and supervision of regional visual and textual information in the image.
It enhances the model's ability to capture fine-grained information, better handles dense prediction tasks, achieves joint representation of visual objects and OCR text, and improves the model's accuracy and robustness in complex visual scenes.
Smart Images

Figure CN120877302A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision, specifically to a training method and related apparatus for a pre-trained image feature extraction model. Background Technology
[0002] Image pre-trained models are one of the core technologies in computer vision, and their function is to extract image features. Transferring image pre-trained models, which are pre-trained on large-scale datasets, to specific tasks can significantly improve the accuracy and robustness of downstream tasks, while reducing dependence on labeled data and accelerating model development and deployment.
[0003] However, traditional image pre-training models mostly focus on learning global features and lack sufficient mining of fine-grained information in images. As a result, they have limited performance when dealing with dense prediction tasks that require fine localization and recognition. In particular, they are difficult to achieve joint representation of visual objects and optical character recognition (OCR) text, which restricts their widespread application in scenarios containing complex visual elements and text information. Summary of the Invention
[0004] To address the aforementioned issues, this application provides a training method and related apparatus for a pre-trained image feature extraction model. This method can reduce the reliance of the pre-trained image model on fine label annotation during the pre-training stage, thereby reducing training costs and shortening the training cycle, while maintaining or improving the model's performance and robustness.
[0005] The embodiments of this application disclose the following technical solutions:
[0006] A training method for a pre-trained image feature extraction model, the method comprising:
[0007] A pre-training dataset is obtained; the pre-training dataset includes multiple first sample images; the multiple first sample images include a first target image and a second target image; the first target image is a sample image labeled with region recognition boxes and corresponding first local labels; the second target image is a sample image labeled with region recognition boxes and corresponding second local labels; the region recognition boxes labeled in the first target image are obtained by the Segmentation All Model (SAM); the first local label is a category identifier in numerical or symbolic form, obtained by clustering the region recognition boxes identified by the SAM model; the second local label is a text description consistent with the visual content of the image within the region recognition box, obtained by the annotation based on the Optical Character Recognition (OCR) model.
[0008] Using an initial model, multiple local features are extracted from each first sample image in the pre-training dataset. The initial model employs a Vision Transformer (ViT) structure and includes a masking mechanism. This masking mechanism converts the global features extracted by the initial model into one or more local features. The formula for converting global features into local features is as follows: R = σ( )·v; σ(·) represents the softmax normalization function; Q is the query matrix; K is the key matrix; V is the global feature; M is the mask value corresponding to the region recognition box; d is the scaling factor;
[0009] By combining the loss function, the parameters of the initial model are adjusted based on all local features and the local labels corresponding to each local feature, to obtain a pre-trained image feature extraction model.
[0010] In one possible implementation, the loss function is L = λ1·L object +λ2·L ocr λ1 and λ2 are weight hyperparameters; L object For object detection loss; L ocr This is due to OCR loss;
[0011] N is the number of region recognition boxes in the first target image. Let be the feature vector of the first local label of the i-th region recognition box. The feature vector of the first local label excluding the first local label corresponding to the i-th region recognition box. Let be the feature vector of the image within the i-th region recognition box in the first target image, and sim(·,·) be the feature similarity function. It is a sampled subset of the set of non-first local labels of the i-th region;
[0012] M represents the number of region recognition boxes in the second target image. Let the feature vector be the second local label of the j-th region bounding box. Let be the feature vector of similar labels to the second local label corresponding to the j-th region recognition box. Let be the feature vector of the image within the bounding box of the j-th region in the second target image. Let j be the set of all similar labels corresponding to the second local label of the region recognition boxes. It is a sampled subset of the set of non-second local labels of the j-th region.
[0013] In one possible implementation, the process of constructing the pre-trained dataset includes:
[0014] Obtain an unlabeled training set; the unlabeled training set includes multiple second sample images;
[0015] For each second sample image in the unlabeled training set, region recognition is performed on the second sample image based on the SAM model to obtain the first target image; for each second sample image in the unlabeled training set, region recognition is performed on the second sample image based on the OCR model to obtain the second target image.
[0016] For each first target image, the identified regions are labeled with a first local label to obtain the first sample image; for each second target image, the identified regions are labeled with a second local label to obtain the first sample image.
[0017] By integrating all the first sample images, the pre-training dataset is obtained.
[0018] In one possible implementation, the step of performing region recognition on each second sample image in the unlabeled training set based on the SAM model to obtain the first target image includes:
[0019] For each second sample image in the unlabeled training set, the SAM model is used to generate a mask for the second sample image to obtain multiple first candidate images; the multiple first candidate images include second sample images with region recognition boxes and / or second sample images without region recognition boxes.
[0020] All first candidate images are filtered to obtain multiple filtered images; the filtered images include the second sample images marked with region recognition boxes;
[0021] For each selected image, the region identification box with the smallest side length less than the target pixel in the selected image is removed to obtain the first target image.
[0022] In one possible implementation, the step of performing first local labeling on the identified regions of each first target image to obtain the first sample image includes:
[0023] Cluster the images within all region recognition boxes in all first target images to obtain multiple clusters;
[0024] Each cluster is assigned a corresponding similarity label; the similarity label is used to indicate the category number to which the cluster belongs;
[0025] For each region recognition box in each first target image, a similar label corresponding to that region recognition box is marked as a first local label within that region recognition box.
[0026] In one possible implementation, the step of performing region recognition on each second sample image in the unlabeled training set based on the OCR model to obtain the second target image includes:
[0027] For each second sample image in the unlabeled training set, the OCR model is used to perform region recognition on the second sample image to obtain multiple second candidate images; the multiple second candidate images include second sample images with region recognition boxes and / or second sample images without region recognition boxes.
[0028] For each region recognition box in each second candidate image, the OCR model is used to extract text information from the region recognition box to obtain the text recognition result;
[0029] For each second candidate image, the text recognition results corresponding to each region recognition box are filtered based on a confidence threshold. Region recognition boxes with a confidence level higher than the confidence threshold and their corresponding text recognition results are retained to obtain the second target image.
[0030] In one possible implementation, the step of performing second local labeling on the identified regions of each second target image to obtain the first sample image includes:
[0031] For each region recognition box in each second target image, the text recognition result corresponding to that region recognition box is used as a second local label and labeled within that region recognition box.
[0032] In one possible implementation, the masking mechanism in the initial model includes setting the mask value inside the region identification box to 0 and setting the mask value outside the region identification box to negative infinity.
[0033] A training apparatus for a pre-trained image feature extraction model, the apparatus comprising:
[0034] The first acquisition unit is used to acquire a pre-training dataset; the pre-training dataset includes multiple first sample images; the multiple first sample images include a first target image and a second target image; the first target image is a sample image labeled with region recognition boxes and corresponding first local labels; the second target image is a sample image labeled with region recognition boxes and corresponding second local labels; the region recognition boxes labeled in the first target image are obtained by the Segmentation All Model (SAM); the first local label is a category identifier in numerical or symbolic form, and the first local label is obtained by clustering the region recognition boxes identified by the SAM model; the second local label is a text description consistent with the visual content of the image within the region recognition box, and the second local label is obtained by labeling based on the Optical Character Recognition (OCR) model;
[0035] The feature extraction unit is used to extract multiple local features from each first sample image in the pre-training dataset using the initial model; the initial model adopts a ViT structure; the initial model includes a masking mechanism; the masking mechanism is used to convert the global features extracted by the initial model into one or more local features, and the formula for converting global features into local features is as follows: R = σ( )·v; σ(·) represents the softmax normalization function; Q is the query matrix; V is the global feature; M is the mask value corresponding to the region recognition box; d is the scaling factor;
[0036] The parameter adjustment unit is used to adjust the parameters of the initial model based on all local features and the local labels corresponding to each local feature, in conjunction with the loss function, to obtain a pre-trained image feature extraction model.
[0037] A computer-readable storage medium storing instructions that, when executed on a terminal device, cause the terminal device to perform the training method for the pre-trained image feature extraction model as described above.
[0038] Compared with the prior art, this application has the following beneficial effects:
[0039] This application provides a training method and related apparatus for a pre-trained image feature extraction model. Specifically, when executing the training method for the pre-trained image feature extraction model provided in this application embodiment, firstly, a pre-training dataset is obtained: this dataset contains multiple sample images, divided into first sample images and second sample images. The first sample images are obtained by annotating region recognition boxes and their corresponding first local labels. The region recognition boxes annotated in the first target image are obtained by identifying them using a Segment Anything Model (SAM), and the local labels corresponding to each region recognition box in the first target image are numerical or symbolic category identifiers generated by clustering the region recognition boxes identified by the SAM model. The second sample images are obtained by annotating region recognition boxes and their corresponding second local labels. These local labels are text descriptions generated based on an OCR model and are consistent with the visual content within the region recognition boxes. Then, feature extraction is performed: the initial model is used to extract features from each first sample image to obtain multiple local features. The initial model adopts a Vision Transformer (ViT) structure and has a built-in masking mechanism. This mechanism can convert global features into local features. Finally, parameter tuning is performed: Based on all local features and their corresponding local labels, the parameters of the initial model are adjusted using the loss function to obtain the pre-trained image feature extraction model. This application achieves joint annotation and supervision of region-level visual and textual information in images by introducing multi-dimensional local labels automatically generated by the SAM and OCR models. Simultaneously, leveraging the masking mechanism in the ViT structure, global features are effectively transformed into corresponding local features, improving the model's ability to capture fine-grained information in dense prediction tasks. Furthermore, parameter optimization combining multi-label local features and the loss function enables the model to maintain overall semantic understanding while also considering the visual and textual feature representation of different regions. Attached Figure Description
[0040] To more clearly illustrate the technical solutions in this embodiment or the prior art, the drawings used in the description of the embodiment or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0041] Figure 1 A flowchart illustrating a training method for a pre-trained image feature extraction model provided in this application embodiment;
[0042] Figure 2 A flowchart illustrating a method for constructing a pre-trained dataset, as provided in an embodiment of this application;
[0043] Figure 3 A flowchart illustrating a target image acquisition method provided in this application embodiment;
[0044] Figure 4 A flowchart illustrating a local labeling method provided in this application embodiment;
[0045] Figure 5 A flowchart illustrating another target image acquisition method provided in this application embodiment;
[0046] Figure 6 This is a schematic diagram of the structure of a training device for a pre-trained image feature extraction model provided in an embodiment of this application. Detailed Implementation
[0047] To facilitate understanding of the technical solutions provided in the embodiments of this application, the background technology involved in the embodiments of this application will be described below.
[0048] Dense prediction tasks refer to computer vision tasks that require detailed prediction and analysis of each pixel or local region in an image, with the output being a high-resolution result similar in size to the input image, rather than a single global label.
[0049] The joint representation of visual objects and OCR text refers to the unified modeling and expression of visual targets (such as people, vehicles, objects, etc.) and text information (text content obtained through OCR recognition) in an image within the same model or feature space.
[0050] Image pre-trained models are one of the core technologies in computer vision, primarily used for extracting image features. By pre-training these models on large-scale datasets, they can be transferred to specific tasks, significantly improving the accuracy and robustness of downstream tasks while reducing reliance on labeled data and accelerating model development and deployment. However, traditional image pre-trained models mainly focus on global image representation and cannot effectively handle dense prediction tasks, meaning they struggle to uniformly handle the joint representation of visual objects and OCR text.
[0051] To address this issue, this application provides a training method and related apparatus for a pre-trained image feature extraction model. First, a pre-training dataset comprising multiple first sample images is acquired. The first sample images include a first target image and a second target image. The first target image is a sample image annotated with region recognition boxes and corresponding first local labels, and the second target image is a sample image annotated with region recognition boxes and corresponding second local labels. The region recognition boxes annotated in the first target image are identified using a SAM model. The first local labels are category identifiers in numerical or symbolic form, obtained by clustering the region recognition boxes identified by the SAM model. The second local labels are textual descriptions consistent with the visual content of the image within the region recognition boxes, obtained based on an OCR model, and supplemented with rules to obtain filtered data. Then, using an initial model with a ViT structure, multiple local features are extracted for each first sample image in the pre-training dataset. Next, combined with a loss function, the parameters of the initial model are adjusted based on all local features and their corresponding local labels to obtain the pre-trained image feature extraction model. This application not only improves the model's ability to capture fine-grained information, but also better handles dense prediction tasks, achieving joint representation of visual objects and OCR text, thus demonstrating higher accuracy and robustness in complex visual scenes.
[0052] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0053] See Figure 1 This figure is a flowchart of a training method for a pre-trained image feature extraction model provided in an embodiment of this application. Figure 1 As shown, the training method for this pre-trained image feature extraction model may include steps S101-S103:
[0054] S101: Obtain the pre-trained dataset.
[0055] To achieve comprehensive capture of fine-grained visual regions and their corresponding textual information in images, a structured and diverse pre-training dataset is first required. This dataset contains multiple first sample images, which are further divided into two categories: first target images and second target images.
[0056] The first target image is generated by labeling region bounding boxes and their corresponding first local labels. The labeled region bounding boxes in the first target image are obtained through a Segmentation All Model (SAM). The first local labels appear as category identifiers in numerical or symbolic form, and these labels are obtained by clustering the region bounding boxes identified by the SAM model. For example, assuming an image containing multiple objects, the SAM model can identify the location of each object and assign it a corresponding category identifier, such as "001" indicating it might be a cat, "002" indicating it might be a dog, etc.
[0057] The second target image is generated by labeling the region recognition box and its corresponding second local label. The second local label is a label in text form describing the visual content of the image within the region recognition box. These labels are obtained based on the Optical Character Recognition (OCR) model. For example, if there is an image containing numbers and letters, the OCR model (such as the Baidu Paddle Optical Character Recognition (Paddle OCR) model) can recognize and label the specific text content, such as "1234" or "ABC".
[0058] For example, suppose we have a set of images, one of which contains a cat and a sign. Using the SAM model, a region bounding box can be annotated for the cat, and a category label "001" can be assigned to it. Meanwhile, using the OCR model, the text content on the sign, such as "pet shop," can be recognized and annotated.
[0059] In this way, the pre-training dataset not only contains images of different types and their corresponding fine-grained annotations, but also combines visual recognition and text recognition capabilities, thus providing rich data support for subsequent feature extraction and model training. This comprehensive approach significantly improves the model's ability to understand and accurately interpret complex visual scenes.
[0060] S102: Using the initial model, perform feature extraction on each first sample image in the pre-training dataset to obtain multiple local features.
[0061] To achieve efficient image feature extraction, the initial model is used to extract features from each first sample image in the pre-training dataset. This initial model employs the ViT architecture, a Transformer-based architecture particularly well-suited for processing image data. The ViT model captures both global and local image features by segmenting the input image into multiple small patches and transforming them into a series of vectors.
[0062] Within the initial model, a masking mechanism is implemented to transform the global features extracted by the model into one or more local features. Specifically, the masking mechanism utilizes the interaction between the query matrix Q in the attention mechanism to extract multiple local features for each first sample image in the pre-training dataset and the global feature V to extract features for each first sample image in the pre-training dataset. By combining the mask value M corresponding to the region recognition box and the scaling factor d, this mechanism can effectively distinguish and extract features from different regions.
[0063] In one possible implementation, the masking mechanism converts global features into local features using the following formula:
[0064] R=σ( )·v;
[0065] σ(·) represents the softmax normalization function, used to normalize the attention weights;
[0066] Q is the query matrix, which is obtained by linear transformation of the input feature map and is used to represent the query information of each local region.
[0067] K extracts multiple local features for each first sample image in the pre-training dataset to form a key matrix, which is usually related to Q and used to calculate attention weights.
[0068] V represents the global feature, indicating the features of the entire image;
[0069] M is the mask value corresponding to the region identification box, used to mask irrelevant regions;
[0070] d is a scaling factor, equal to the length of a single vector in the global features (i.e., the feature dimension), used to adjust the size of the attention weights.
[0071] S103: Combine the loss function and adjust the parameters of the initial model based on all local features and the local labels corresponding to each local feature to obtain a pre-trained image feature extraction model.
[0072] After extracting local features from each sample image in the pre-training dataset, a suitable loss function is designed to compare and constrain all extracted local features with their corresponding local labels, thereby effectively adjusting the initial model parameters. Specifically, the loss function not only measures the matching degree between local features and numerical or symbolic category labels (such as the first local label obtained by clustering region recognition boxes based on the SAM model), but also evaluates the conformity between local features and textual description labels consistent with visual content (the second local label generated by the OCR model). This multi-dimensional, multi-task supervision signal prompts the model to consider both the classification accuracy of visual objects and the semantic relevance of textual information during the optimization process, thereby improving the model's ability to express fine-grained local information. Through the backpropagation algorithm, the error in the loss function is gradually fed back to the parameters of each layer of the model, guiding its iterative update so that the model parameters continuously approach the optimal state. After sufficient training, this process ultimately yields a pre-trained image feature extraction model that can accurately capture and jointly represent visual regions and their corresponding textual information. While maintaining overall image understanding, this model places greater emphasis on the accurate recognition of local details, providing a solid feature foundation and robust performance for subsequent downstream intensive prediction tasks.
[0073] In one possible implementation, the loss function is L = λ1·L object +λ2·L ocr λ1 and λ2 are weight hyperparameters used to balance the contributions of different loss terms; L object The object detection loss measures the difference between the local features extracted by the model and the first local label; L ocr The OCR loss measures the difference between the local features extracted by the model and the second local label.
[0074] .
[0075] Where N is the number of region recognition boxes in the first target image, that is, the total number of independent region recognition boxes for object detection identified by the SAM model in a first target image, which determines L. object The number of regions traversed during loss calculation.
[0076] The feature vector of the first local label of the i-th region recognition box is obtained by extracting features from the first local label corresponding to the i-th region recognition box, and is used to compare with the true feature vector. Calculate the similarity to determine the accuracy of the prediction.
[0077] It is the feature vector of the first local label other than the first local label corresponding to the i-th region identification box. It represents the first local label features associated with other region identification boxes. In the loss calculation, it is used to construct "negative sample" comparison to distinguish the label difference between the current region and other regions.
[0078] Let be the feature vector of the image within the i-th region bounding box in the first target image. This vector represents the feature representation of the real image content corresponding to the bounding box, serving as the "real baseline" for judging the accuracy of predictions in object detection tasks. and Calculate the similarity score separately and measure the prediction error.
[0079] sim(·,·) is a feature similarity function used to calculate the similarity between two feature vectors (e.g., ...). and , and The similarity between features (e.g., features) can be measured by various functions such as cosine similarity and Euclidean distance (which can be converted into a similarity form). The larger the value, the more similar the features are, which helps to determine whether the model's prediction matches the real situation.
[0080] This is a sampled subset of the set of non-first local labels for the i-th region. This set is sampled from all sets that do not belong to the first local labels of the i-th region and is used for filtering. These "other label feature vectors" reduce computational cost while ensuring training effectiveness (avoiding traversing all irrelevant labels) and support the loss calculation logic for "negative sample" comparison.
[0081] .
[0082] Where M is the number of region recognition boxes in the second target image, that is, the total number of region recognition boxes related to OCR recognition in a single second target image, which determines L. ocr The number of regions traversed during loss calculation.
[0083] This is the feature vector of the second local label of the j-th region recognition box. It is obtained after feature extraction of the second local label (a text description consistent with the visual content of the image within the region) corresponding to the j-th region recognition box, and is used to compare with the ground truth feature vector. Calculate the similarity to determine whether the OCR recognition is accurate.
[0084] This parameter represents the feature vector of similar labels to the second local label corresponding to the j-th region bounding box. It indicates the feature vectors of other labels that are semantically or visually similar to the second local label (text description) of the j-th region bounding box. For example, if the second local label of the j-th region is "red sedan", then... It may include feature vectors with semantically similar labels such as "red car" and "crimson sedan".
[0085] Let be the feature vector of the image within the j-th region recognition box in the second target image. This vector represents the feature representation of the real image content corresponding to the region recognition box, serving as the "real benchmark" for judging the accuracy of predictions in the OCR recognition task. , Instead of calculating similarity, measure the OCR recognition error.
[0086] Let J be the set of all similar labels corresponding to the second local labels of the j-th region recognition boxes. This set includes the second local label corresponding to the j-th region recognition box in the current second target image, and all similar labels corresponding to that region recognition box. This is for obtaining... These "other label feature vectors" provide the source and support the comparative logic for calculating OCR loss.
[0087] This is a sampled subset of the set of non-second local labels for the j-th region. This set is sampled from all sets of second local labels that do not belong to the j-th region. It is used to construct "negative sample" comparisons (samples that differ greatly from the true label features) so that the model can learn to distinguish between the target label and irrelevant labels and avoid incorrect associations.
[0088] Based on the content of S101-S103, we first acquire a pre-training dataset containing multiple first sample images. The first sample images consist of two types of target images: the first target image and the second target image. Region recognition boxes labeled in the first target image are obtained using a Segmentation All Model (SAM). First local labels appear as category identifiers in numerical or symbolic form; these labels are obtained by clustering the region recognition boxes identified by the SAM model. The second target image also contains region recognition boxes, but its corresponding second local label is a text description corresponding to the visual content within the region; this text label is obtained through an OCR model. Next, using an initial model employing the ViT architecture, features are extracted from each sample image in the pre-training dataset, obtaining multiple local features. This initial model incorporates a masking mechanism to convert the extracted global features into one or more local features representing different regions. Finally, combined with the designed loss function, the parameters of the initial model are optimized and adjusted based on all local features and their corresponding local labels, thereby training a pre-trained image feature extraction model with good feature extraction capabilities. This application not only enhances the model's ability to capture fine-grained information but also significantly improves its performance in dense prediction tasks. By achieving a joint representation of visual objects and OCR text, this application demonstrates higher accuracy and robustness in complex visual scenes.
[0089] In one possible implementation, this application also provides a method for constructing a pre-trained dataset, see [link to relevant documentation]. Figure 2 , Figure 2 A flowchart of a pre-training dataset construction method provided in this application embodiment is shown, which can be implemented through steps S201-S204:
[0090] S201: Obtain the unlabeled training set.
[0091] The first step in building a pre-training dataset is to obtain an unlabeled training set. This unlabeled training set refers to a dataset containing a large number of images, each without pre-labeled information. These images are called second-sample images; they are the raw image data without any manual annotation.
[0092] The purpose of the unlabeled training set is to provide a large amount of raw image data so that data that can be generated for training models can be obtained through a series of automated processing and annotation methods. Specifically, each second sample image in the unlabeled training set is unlabeled, meaning that these images do not contain any specific descriptions or classification information about the image content before being processed.
[0093] By using unlabeled training sets, the expensive and time-consuming manual annotation process can be avoided. Instead, advanced machine learning and computer vision techniques are used to automatically generate the label information needed for training. For example, in subsequent steps, SAM and OCR models will be used to process these unlabeled images to identify different regions and text, and add the necessary local labels. This not only efficiently generates a large amount of training data but also ensures the diversity and breadth of the data, thereby improving the generalization ability of the trained model.
[0094] In summary, obtaining an unlabeled training set and extracting second sample images from it is a fundamental step in building a high-quality pre-trained dataset, providing a rich source of raw data for subsequent region recognition and local labeling.
[0095] S202: For each second sample image in the unlabeled training set, perform region recognition on the second sample image based on the SAM model to obtain the first target image; for each second sample image in the unlabeled training set, perform region recognition on the second sample image based on the OCR model to obtain the second target image.
[0096] To fully exploit the potential information in each second sample image in the unlabeled training set and further improve the data annotation quality and model training effect, two advanced automatic recognition technologies were employed to process the images. First, region recognition was performed on each second sample image based on the SAM model. SAM can efficiently and accurately segment semantically meaningful visual object regions from complex images and generate corresponding bounding boxes for these regions, thus obtaining first target images with visual object segmentation information. These images carry rich, fine-grained visual structural information, laying the foundation for subsequent visual category labeling. Simultaneously, for the same batch of second sample images, region detection and text recognition were performed based on the OCR model, automatically locating and extracting text regions contained in the images, converting them into parsable text content, thereby forming second target images with text descriptions. Through the parallel processing of these two paths, accurate segmentation of visual objects and effective extraction of text information were ensured, achieving comprehensive capture of multimodal information from unlabeled images. Ultimately, these first target images, processed by the SAM and OCR models, together with the second target images, form an important foundation for subsequent local labeling and pre-training dataset construction, providing the model with rich and accurate supervision signals.
[0097] S203: For each first target image, the identified regions are labeled with a first local label to obtain the first sample image; for each second target image, the identified regions are labeled with a second local label to obtain the first sample image.
[0098] After completing the region identification of the second sample images in the unlabeled training set, the next step is to perform detailed local labeling on the obtained first and second target images to construct high-quality training samples. Specifically, for each first target image, the visual regions automatically identified by the SAM model are assigned first local labels using feature clustering methods. These labels are usually represented in numerical or symbolic form to clearly indicate the category or semantic attribute of each region. This process not only endows each segmented region in the image with specific identity information but also provides the model with accurate visual supervision signals, helping to improve the model's classification and segmentation capabilities. Simultaneously, for each second target image, second local labels are applied to the text regions detected by the OCR model. These labels exist in the form of textual descriptions highly consistent with the visual content within the region, accurately reflecting the semantic information expressed by the text. Through this dual labeling strategy, both clear classification of visual objects and semantic expression of textual information are ensured, achieving an organic combination of multimodal information. Ultimately, these two types of target images with corresponding local labels are collectively referred to as "first sample images," ensuring that each training sample possesses rich and accurate label information in both visual and textual dimensions, providing a solid data foundation for the multi-task learning of subsequent pre-trained models.
[0099] S204: Integrate all the first sample images to obtain the pre-training dataset.
[0100] After completing the local labeling of each first and second target image, the next step is to systematically integrate all these first sample images with rich annotation information to construct a complete and structured pre-training dataset. This integration process not only involves unifying visual category labels and textual description labels into the same data framework, but also ensures the consistency of data format and the accurate correspondence of annotation information, thereby achieving effective fusion of multimodal labels. By centrally aggregating a large number of high-quality first sample images, the pre-training dataset covers a variety of visual objects and their fine-grained classifications, while also including detailed descriptions of textual content within the regions, providing the model with joint supervision signals that take into account both visual features and linguistic information. This comprehensive pre-training dataset not only meets the multidimensional information requirements of complex visual tasks, but also greatly improves the model's generalization ability to different scenes and diverse content. Ultimately, this dataset serves as a core training resource, providing solid data support for subsequent image feature extraction models and promoting superior performance in various downstream applications.
[0101] In one possible implementation, this application also provides a method for acquiring a target image, see [link to relevant documentation]. Figure 3 , Figure 3This is a flowchart of a target image acquisition method provided in an embodiment of this application. Accordingly, in step S202, for each second sample image in the unlabeled training set, region recognition is performed on the second sample image based on the SAM model to obtain the first target image. Specifically, this can be implemented through steps S301-S303:
[0102] S301: For each second sample image in the unlabeled training set, the SAM model is used to generate a mask for the second sample image to obtain multiple first candidate images.
[0103] When processing each second sample image in the unlabeled training set, SAM is used to generate a mask to obtain multiple first candidate images. This process not only helps to identify various regions in the image, but also provides a foundation for subsequent local labeling.
[0104] Specifically, the SAM model, through its powerful segmentation capabilities, can perform detailed region detection and segmentation on the input second sample image. In this process, the SAM model generates multiple masks on the image, each corresponding to a specific region. These masks can completely cover the pixel range of a region, thus forming an accurate region recognition bounding box. The generated first candidate images include two types:
[0105] Second sample images with region bounding boxes: In this type of image, the SAM model successfully identified and labeled multiple regions in the image. Each region has a corresponding bounding box, which specifies the exact location and size of the region. For example, if the input second sample image is a scene containing vehicles and pedestrians, the SAM model might generate several masks corresponding to the regions of vehicles and pedestrians, thus generating a first candidate image with region bounding boxes.
[0106] Second sample image without labeled region bounding boxes: In some cases, the SAM model may fail to identify salient regions, or the generated mask may be insufficient to form clear bounding boxes. In such cases, the generated first candidate image will not contain any region bounding boxes, preserving the state of the original second sample image. This situation typically occurs when the image lacks clear segmentation boundaries or has a complex background.
[0107] For example, suppose we have a second sample image containing a car, a person, and some buildings. When generating a mask using the SAM model, the model might identify the areas of the vehicle and pedestrian, generate masks for these areas, and label the corresponding region bounding boxes. Thus, the resulting first candidate image will include an image with region bounding boxes, where the vehicle and pedestrian areas are clearly marked. On the other hand, building areas, due to complex backgrounds or indistinct boundaries, may not form valid bounding boxes; therefore, this part of the image may be retained as a second sample image without labeled region bounding boxes.
[0108] This process generates multiple first candidate images from each second sample image. These images include both valid region images with labeled region recognition boxes and the original images without labeled region recognition boxes. This approach not only helps extract and label salient features from the images but also provides rich material for subsequent dataset construction and model training.
[0109] S302: Filter all first candidate images to obtain multiple filtered images.
[0110] When processing the first candidate images, further filtering is required to ensure that the final images have practical significance and value. Specifically, all first candidate images are filtered to remove those that do not provide valuable regional information, resulting in multiple filtered images. These filtered images mainly contain second sample images labeled with region recognition boxes.
[0111] The core objective of the screening process is to retain images that provide clear regional information and remove those without effective bounding boxes or with unclear masks. This ensures that subsequent data processing and model training are based on more representative and reliable image data.
[0112] The specific operating steps are as follows:
[0113] Initial assessment: First, check if each first candidate image contains an labeled region bounding box. If so, these images will be retained; otherwise, they may be excluded or subject to further examination.
[0114] Quality Assessment: For images with labeled region bounding boxes, further evaluate the quality of these boxes. For example, do the boxes accurately cover the target area, and can they clearly distinguish different objects? This step helps to remove images that have bounding boxes but are of poor quality.
[0115] Filtering invalid images: For images without labeled bounding boxes, further analysis is conducted to determine whether they should be retained. Typically, if there are no valid bounding boxes, these images may be considered invalid or suboptimal and thus excluded.
[0116] For example, suppose that after processing a second sample image using the SAM model, multiple first candidate images are obtained, including:
[0117] An image marked with a vehicle identification frame;
[0118] An image marked with a pedestrian crossing identification box;
[0119] A raw image without any bounding boxes.
[0120] During the screening process, each image is first checked for the presence of region identification boxes. For images with labeled vehicle and pedestrian identification boxes, the quality of these boxes is further evaluated. If these boxes are accurate and clear, the images will be retained for screening. Raw images without any labeled boxes may be excluded because they do not provide additional regional information.
[0121] Through this screening process, the final selected images all include second sample images labeled with region recognition boxes. These images have greater practical application value and can be used for subsequent tasks such as dataset construction and model training, thereby improving the overall system performance and effectiveness.
[0122] S303: For each filtered image, remove the region recognition box whose minimum side length is less than the target pixel in the filtered image to obtain the first target image.
[0123] When processing the filtered images, further quality control can be performed on each region bounding box to ensure that the final generated first target image has high usability and accuracy. Specifically, we need to remove region bounding boxes in the filtered images whose minimum side length is smaller than the target pixel (e.g., 28 pixels) to obtain a higher quality first target image.
[0124] The main purpose of this process is to eliminate overly small and potentially meaningless region bounding boxes, as these small regions typically offer little in the way of valuable segmentation information or features. By removing these small regions, it is ensured that the final generated first target image contains larger and more representative regions, which helps improve the effectiveness of subsequent dataset construction and model training.
[0125] The specific steps are as follows:
[0126] Region recognition box detection: First, detect all region recognition boxes in each selected image and record the size of each recognition box (including width and height).
[0127] Size check: For each region bounding box, check if its minimum side length is less than the target pixel value (e.g., 28 pixels). If the minimum side length of a bounding box is less than 28 pixels, the bounding box is considered too small and needs to be removed.
[0128] Remove small bounding boxes: Remove bounding boxes that meet the above conditions from the filtered images, retaining only those that are the correct size. This ensures that the final image contains only larger and more meaningful bounding boxes.
[0129] For example, suppose there is a screening image containing multiple region recognition boxes, including:
[0130] A human figure recognition frame, measuring 100x50 pixels;
[0131] A bounding box for recognizing a small object, measuring 20x30 pixels;
[0132] The recognition frame for a car measures 150x80 pixels.
[0133] According to the regulations, recognition boxes with a minimum side length of less than 28 pixels will be removed. Therefore, the minimum side length of each recognition box is checked sequentially:
[0134] The minimum side length of the person recognition box is 50 pixels; if it is greater than 28 pixels, it will be retained.
[0135] The minimum side length of the small object recognition box is 20 pixels, which is less than 28 pixels, so it is removed.
[0136] The minimum side length of the vehicle recognition frame is 80 pixels. If it is greater than 28 pixels, it will be retained.
[0137] Through this process, only two bounding boxes remain in the final target image: the person and the vehicle. These bounding boxes are not only larger in size but also more practically meaningful, and can be used for subsequent dataset construction and model training tasks.
[0138] This meticulous screening and removal process ensures that the final generated first target image has high accuracy and practicality, providing a better foundation for subsequent deep learning and other vision tasks.
[0139] In one possible implementation, this application also provides a method for annotating local labels, see [link to relevant documentation]. Figure 4 , Figure 4This is a flowchart of a local labeling method provided in an embodiment of this application. Accordingly, in step S203, for each first target image, the identified region is labeled with a first local label to obtain the first sample image. This can be specifically implemented through steps S401-S403:
[0140] S401: Cluster the images within the recognition boxes of all regions in all first target images to obtain multiple clusters.
[0141] To effectively classify and organize the visual content within bounding boxes in a large number of target images, it is first necessary to extract the feature information of these regions and then perform cluster analysis on all regions based on these features. Specifically, the image fragment corresponding to each bounding box is input into the feature extraction module, and its high-dimensional visual feature representation is obtained through methods such as deep neural networks. These features can accurately reflect the color, texture, shape, and semantic information of the region. Subsequently, unsupervised clustering processing is performed on the feature vectors of all regions using algorithms such as K-means, hierarchical clustering, or density clustering. These regions are divided into multiple non-overlapping clusters according to the principle of visual similarity or semantic similarity. The regions within each cluster have high consistency, representing visual targets of the same or similar categories, thus achieving effective summarization and structuring of a large number of regions. For example, suppose the dataset contains a large number of local regions of street view images. During the clustering process, vehicle regions may be grouped into one cluster, pedestrian regions into another, and buildings or road signs into different clusters. Such clustering not only simplifies the complexity of subsequent label assignment, but also helps the model better understand and distinguish visual objects of different categories, achieving more accurate local labeling.
[0142] S402: Assign a corresponding similar label to each cluster; the similar label is used to indicate the category number to which the cluster belongs.
[0143] After clustering the images within the region recognition boxes in all the first target images, the next step is to assign corresponding similarity labels to each cluster. These labels indicate the category number to which the cluster belongs. Specifically, each cluster represents a group of image regions with similar visual features or semantic correlation. Therefore, by assigning uniform similarity labels, all regions within the same cluster can be effectively classified into the same category, thereby achieving structured management and classification of massive amounts of unlabeled data. These similarity labels usually exist in the form of numerical numbers, symbols, or codes, facilitating automatic identification and processing by the system and serving as supervision signals during model training. For example, in a clustering result containing street scene image regions, vehicle regions may be divided into one cluster, which is assigned the similarity label "Category 1"; pedestrian regions may form another cluster, corresponding to "Category 2"; and building regions may form another cluster, labeled "Category 3". This labeling not only improves the organization efficiency of the dataset but also provides a clear category basis for subsequent local labeling, enabling each image region to accurately reflect its visual category and promoting the performance improvement and generalization ability of the pre-trained model in multi-task learning.
[0144] S403: For each region recognition box in each first target image, label the region recognition box with a similar label as the first local label.
[0145] After assigning corresponding similar labels to each cluster, the next crucial step is to apply these labels to the bounding boxes of each region in the first target image, achieving precise local labeling. Specifically, for each bounding box in each first target image, the cluster to which the region belongs is first determined, and then the similar label corresponding to that cluster is directly labeled inside the bounding box as the first local label. This not only ensures the consistency and accuracy of the labels but also ensures that each visual region carries clear and structured category information, providing rich and fine-grained training signals for the model's supervised learning. For example, assuming a street scene first target image contains three bounding boxes, corresponding to three different clusters: vehicles, pedestrians, and buildings, the system will, based on the previously assigned similar labels, label "Category 1" in the vehicle region's bounding box, "Category 2" in the pedestrian region's bounding box, and "Category 3" in the building region's bounding box. In this way, each region obtains a clear semantic classification, which greatly improves the expressive power and supervision effect of the training samples, and helps the pre-trained model achieve higher accuracy and generalization performance when dealing with complex visual tasks.
[0146] In one possible implementation, this application also provides a method for acquiring the target image, see [link to relevant documentation]. Figure 5 , Figure 5 This is a flowchart of another target image acquisition method provided in this application embodiment. Accordingly, in step S202, for each second sample image in the unlabeled training set, region recognition is performed on the second sample image based on the OCR model to obtain the second target image. Specifically, this can be implemented through steps S501-S503:
[0147] S501: For each second sample image in the unlabeled training set, the OCR model is used to perform region recognition on the second sample image to obtain multiple second candidate images.
[0148] For each second sample image in the unlabeled training set, an OCR model is first used for region recognition. This process aims to automatically detect text regions that may be present in the image. Specifically, the OCR model analyzes features such as texture, edges, and character shapes to accurately locate potential text regions and generate corresponding region recognition boxes on the image. After this step, each second sample image is divided into multiple second candidate images. These candidate images can be divided into two categories: one category consists of images with clearly labeled region recognition boxes, meaning the model has successfully recognized and bounded the text regions, thus effectively capturing text information; the other category consists of images without labeled region recognition boxes, indicating that the OCR model has not detected obvious text in these regions, which may be background or low-confidence areas. For example, assuming the input second sample image is a street scene photo containing shop signs and billboards, the OCR model may detect text regions on the signs and generate recognition boxes around them. However, it will not generate any boxes for textless areas such as streets or the sky. The final second candidate images include both areas with text recognition boxes and blank or irrelevant areas that are not bounded. In this way, the OCR model not only helps to automatically extract potential text information regions, but also provides basic data support for subsequent text recognition and filtering steps.
[0149] S502: For each region recognition box in each second candidate image, the OCR model is used to extract text information from the region recognition box to obtain the text recognition result.
[0150] After obtaining the second candidate images and their corresponding region recognition boxes, the next step is to use an OCR model to extract text from the visual information within each region recognition box in each second candidate image. Specifically, the OCR model converts the text in the image into a readable text string by detecting and recognizing characters within the region recognition box. This process involves not only character shape recognition but also understanding character order, font style, and text layout, thereby improving the accuracy and completeness of text recognition. Through this step, accurate text recognition results can be obtained from each located text region, providing crucial information for subsequent labeling, confidence evaluation, and filtering operations. For example, suppose a second candidate image contains a billboard area marked with a region recognition box. Within this region box, the text "SUMMER extracts multiple local features for each first sample image in the pre-training dataset; SALE extracts multiple local features for each first sample image in the pre-training dataset; 50% extracts multiple local features for each first sample image in the pre-training dataset; OFF" is displayed. The OCR model will automatically recognize this text and convert it into the digitized string "SUMMER extracts multiple local features for each first sample image in the pre-training dataset; SALE extracts multiple local features for each first sample image in the pre-training dataset; 50% extracts multiple local features for each first sample image in the pre-training dataset; OFF," and may also provide the corresponding recognition confidence score. By extracting text information from all region recognition boxes, not only can the text content in the image be decoded, but the validity and category of the text can also be determined, which helps improve the accuracy and practical value of the entire target image acquisition method.
[0151] S503: For each second candidate image, the text recognition results corresponding to each region recognition box are filtered in combination with the confidence threshold, and the region recognition boxes with a confidence level higher than the confidence threshold and their corresponding text recognition results are retained to obtain the second target image.
[0152] After extracting text information from the bounding boxes of each region in each second candidate image, the text recognition results need to be rigorously screened using a predefined confidence threshold to ensure high accuracy and reliability of the final retained regions and text. Specifically, the OCR model not only outputs the text content but also assigns a confidence score to each recognized text segment, reflecting the model's confidence level in the text recognition result. By using a predefined confidence threshold, text regions with low confidence levels, potentially containing misidentifications or noise, can be filtered out, thus preventing erroneous information from being introduced into subsequent processing stages. During the screening process, only bounding boxes with confidence levels higher than the threshold and their corresponding text recognition results are retained and form the final second target image. This ensures data quality while improving the downstream task's understanding of the text content. For example, suppose a second candidate image contains three region recognition boxes with text recognition confidence scores of 0.95, 0.60, and 0.30, respectively. If the confidence threshold is set to 0.80, only the region box with a confidence score of 0.95 and its text content, such as "OPEN extracts multiple local features for each first sample image in the pre-training dataset, and HOURS extracts multiple local features for each first sample image in the pre-training dataset," while the two regions with lower confidence scores are discarded. The final second target image contains rigorously selected, high-confidence text regions, providing a solid and accurate data foundation for subsequent text analysis, semantic understanding, and even vision-language joint training.
[0153] In one possible implementation, step S203 involves performing second local labeling on the identified regions of each second target image to obtain the first sample image, including:
[0154] For each region recognition box in each second target image, the text recognition result corresponding to that region recognition box is used as a second local label and labeled within that region recognition box.
[0155] Specifically, for each region recognition box in each second target image, the text recognition result extracted by the OCR model is directly used as the second local label for that region and annotated inside the corresponding region recognition box. This method not only ensures a high degree of consistency between the local label and the image region content, but also endows each text region in the image with semantically rich and precise label information. For example, assuming a second target image contains a region recognition box and its text recognition result is "WELCOME", then the text "WELCOME" is used as the second local label for that region and annotated inside the recognition box, forming a first sample image with a clear text label. This local labeling method based on text recognition results helps subsequent vision-language joint modeling, improving the model's ability to understand and utilize text information in images.
[0156] In one possible implementation, the masking mechanism in the initial model is achieved by differentially processing the region recognition boxes in the image. Specifically, for each pixel location within a region recognition box, its mask value is set to 0, indicating that these regions can be normally focused on and utilized during model computation; while for pixels outside the recognition box, their mask value is set to negative infinity, meaning that these regions are completely masked in the model's attention mechanism or subsequent computation and cannot have any impact. This masking strategy effectively guides the model to focus on the target region, avoids interfering information, and improves the targeting and accuracy of feature extraction. For example, when processing an image containing multiple text regions, the model is only allowed to focus on pixels within each text region, while other background pixels are masked with a negative infinity mask value, thereby enhancing the model's ability to recognize text content and its robustness.
[0157] See Figure 6 , Figure 6 This is a schematic diagram of the structure of a training device for a pre-trained image feature extraction model provided in an embodiment of this application. Figure 6 As shown, the training device for this pre-trained image feature extraction model includes:
[0158] The first acquisition unit 601 is used to acquire a pre-training dataset; the pre-training dataset includes multiple first sample images; the multiple first sample images include a first target image and a second target image; the first target image is a sample image labeled with region recognition boxes and corresponding first local labels; the second target image is a sample image labeled with region recognition boxes and corresponding second local labels; the region recognition boxes labeled in the first target image are obtained by the Segmentation All Model (SAM); the first local label is a category identifier in numerical or symbolic form, and the first local label is obtained by clustering the region recognition boxes identified by the SAM model; the second local label is a text description consistent with the visual content of the image within the region recognition box, and the second local label is obtained by labeling based on the Optical Character Recognition (OCR) model;
[0159] Feature extraction unit 602 is used to extract multiple local features from each first sample image in the pre-training dataset using an initial model; the initial model adopts a ViT structure; a masking mechanism is set within the initial model; the masking mechanism is used to convert the global features extracted by the initial model into one or more local features, and the formula for converting global features into local features is as follows: R=σ( )·v; σ(·) represents the softmax normalization function; Q is the query matrix; V is the global feature; M is the mask value corresponding to the region recognition box; d is the scaling factor;
[0160] The parameter adjustment unit 603 is used to combine the loss function and adjust the parameters of the initial model based on all local features and the local labels corresponding to each local feature to obtain a pre-trained image feature extraction model.
[0161] In one possible implementation, the loss function is L = λ1·L object +λ2·L ocr λ1 and λ2 are weight hyperparameters; L object For object detection loss; L ocr This is due to OCR loss;
[0162] N is the number of region recognition boxes in the first target image. Let be the feature vector of the first local label of the i-th region recognition box. The feature vector of the first local label excluding the first local label corresponding to the i-th region recognition box. Let be the feature vector of the image within the i-th region recognition box in the first target image, and sim(·,·) be the feature similarity function. It is a sampled subset of the set of non-first local labels of the i-th region;
[0163] M represents the number of region recognition boxes in the second target image. Let the feature vector be the second local label of the j-th region bounding box. Let be the feature vector of similar labels to the second local label corresponding to the j-th region recognition box. Let be the feature vector of the image within the bounding box of the j-th region in the second target image. Let j be the set of all similar labels corresponding to the second local label of the region recognition boxes. It is a sampled subset of the set of non-second local labels of the j-th region.
[0164] In one possible implementation, the device further includes:
[0165] The second acquisition unit is used to acquire an unlabeled training set; the unlabeled training set includes multiple second sample images;
[0166] The first region recognition unit is used to perform region recognition on each second sample image in the unlabeled training set based on the SAM model to obtain the first target image.
[0167] The second region recognition unit is used to perform region recognition on each second sample image in the unlabeled training set based on the OCR model to obtain the second target image.
[0168] The first annotation unit is used to perform first local label annotation on the identified regions of each first target image to obtain the first sample image;
[0169] The second annotation unit is used to perform second local label annotation on the identified regions of each second target image to obtain the first sample image;
[0170] An integration unit is used to integrate all the first sample images to obtain the pre-trained dataset.
[0171] In one possible implementation, the first region identification unit specifically includes:
[0172] The mask generation unit is used to generate a mask for each second sample image in the unlabeled training set using the SAM model to obtain multiple first candidate images; the multiple first candidate images include second sample images with region recognition boxes and / or second sample images without region recognition boxes.
[0173] The first filtering unit is used to filter all first candidate images to obtain multiple filtered images; the filtered images include the second sample images marked with region recognition boxes;
[0174] The second filtering unit is used to remove the region recognition box with the smallest side length less than the target pixel in each filtered image to obtain the first target image.
[0175] In one possible implementation, the first annotation unit specifically includes:
[0176] The clustering unit is used to cluster the images within all region recognition boxes in all first target images to obtain multiple clusters;
[0177] A similar labeling unit is used to assign a corresponding similar label to each cluster; the similar label is used to indicate the category number to which the cluster belongs;
[0178] The first setting unit is used to mark a similar label corresponding to the region recognition box as a first local label within the region recognition box for each region recognition box in each first target image.
[0179] In one possible implementation, the second region identification unit specifically includes:
[0180] The third region recognition unit is used to perform region recognition on each second sample image in the unlabeled training set using the OCR model to obtain multiple second candidate images; the multiple second candidate images include second sample images with region recognition boxes and / or second sample images without region recognition boxes.
[0181] The text information extraction unit is used to extract text information from each region recognition box in each second candidate image using the OCR model to obtain the text recognition result.
[0182] The second setting unit is used to filter the text recognition results corresponding to each region recognition box for each second candidate image in combination with a confidence threshold, and retain the region recognition boxes with a confidence level higher than the confidence threshold and their corresponding text recognition results to obtain the second target image.
[0183] In one possible implementation, the second annotation unit is specifically used for:
[0184] For each region recognition box in each second target image, the text recognition result corresponding to that region recognition box is used as a second local label and labeled within that region recognition box.
[0185] In one possible implementation, the masking mechanism in the initial model includes setting the mask value inside the region identification box to 0 and setting the mask value outside the region identification box to negative infinity.
[0186] In addition, this application embodiment also provides a computer-readable storage medium storing instructions that, when executed on a terminal device, cause the terminal device to perform the training method of the pre-trained image feature extraction model as described above.
[0187] This application introduces multi-dimensional local labels automatically generated by the SAM and OCR models, enabling joint annotation and supervision of visual features and textual information in different regions of an image. Simultaneously, utilizing the masking mechanism in the ViT structure, global features can be effectively converted into precise local features, significantly enhancing the model's ability to capture fine-grained information in dense prediction tasks. Furthermore, by combining multi-label local features with the loss function for end-to-end parameter optimization, the model not only maintains a deep understanding of the overall semantics but also considers the visual and textual feature representation of each region, improving the multimodal information fusion effect and task performance.
[0188] The training method and related apparatus for a pre-trained image feature extraction model provided in this application have been described in detail above. The various embodiments in the specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and relevant parts can be referred to in the method section. It should be noted that those skilled in the art can make several improvements and modifications to this application without departing from the principles of this application, and these improvements and modifications also fall within the protection scope of the claims of this application.
[0189] It should also be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
Claims
1. A training method for a pre-trained image feature extraction model, characterized in that, The method includes: A pre-training dataset is obtained; the pre-training dataset includes multiple first sample images; the multiple first sample images include a first target image and a second target image; the first target image is a sample image labeled with region recognition boxes and corresponding first local labels; the second target image is a sample image labeled with region recognition boxes and corresponding second local labels; the region recognition boxes labeled in the first target image are obtained by the Segmentation All Model (SAM); the first local label is a category identifier in numerical or symbolic form, obtained by clustering the region recognition boxes identified by the SAM model; the second local label is a text description consistent with the visual content of the image within the region recognition box, obtained by the annotation based on the Optical Character Recognition (OCR) model. Using an initial model, multiple local features are extracted from each first sample image in the pre-training dataset. The initial model employs a Vision Transformer (ViT) structure and includes a masking mechanism. This masking mechanism converts the global features extracted by the initial model into one or more local features. The formula for converting global features into local features is as follows: R = σ( )·v; σ(·) represents the softmax normalization function; Q is the query matrix; K is the key matrix; V is the global feature; M is the mask value corresponding to the region recognition box; d is the scaling factor; By combining the loss function, the parameters of the initial model are adjusted based on all local features and the local labels corresponding to each local feature, to obtain a pre-trained image feature extraction model.
2. The method according to claim 1, characterized in that, Loss function L=λ1·L object +λ2·L ocr λ1 and λ2 are weight hyperparameters; L object For object detection loss; L ocr This is due to OCR loss; N is the number of region recognition boxes in the first target image. Let be the feature vector of the first local label of the i-th region recognition box. The feature vector of the first local label excluding the first local label corresponding to the i-th region recognition box. Let be the feature vector of the image within the i-th region recognition box in the first target image, and sim(·,·) be the feature similarity function. It is a sampled subset of the set of non-first local labels of the i-th region; ; M is the number of region recognition boxes in the second target image. Let the feature vector be the second local label of the j-th region bounding box. Let be the feature vector of similar labels to the second local label corresponding to the j-th region recognition box. Let be the feature vector of the image within the bounding box of the j-th region in the second target image. Let j be the set of all similar labels corresponding to the second local label of the region recognition boxes. It is a sampled subset of the set of non-second local labels of the j-th region.
3. The method according to claim 1, characterized in that, The process of constructing the pre-trained dataset includes: Obtain an unlabeled training set; the unlabeled training set includes multiple second sample images; For each second sample image in the unlabeled training set, region recognition is performed on the second sample image based on the SAM model to obtain the first target image; for each second sample image in the unlabeled training set, region recognition is performed on the second sample image based on the OCR model to obtain the second target image. For each first target image, the identified regions are labeled with a first local label to obtain the first sample image; for each second target image, the identified regions are labeled with a second local label to obtain the first sample image. By integrating all the first sample images, the pre-training dataset is obtained.
4. The method according to claim 3, characterized in that, For each second sample image in the unlabeled training set, region recognition is performed on the second sample image based on the SAM model to obtain the first target image, including: For each second sample image in the unlabeled training set, the SAM model is used to generate a mask for the second sample image to obtain multiple first candidate images; the multiple first candidate images include second sample images with region recognition boxes and / or second sample images without region recognition boxes. All first candidate images are filtered to obtain multiple filtered images; the filtered images include the second sample images marked with region recognition boxes; For each selected image, the region identification box with the smallest side length less than the target pixel in the selected image is removed to obtain the first target image.
5. The method according to claim 3, characterized in that, For each first target image, the identified regions are labeled with a first local label to obtain the first sample image, including: Cluster the images within all region recognition boxes in all first target images to obtain multiple clusters; Each cluster is assigned a corresponding similarity label; the similarity label is used to indicate the category number to which the cluster belongs; For each region recognition box in each first target image, a similar label corresponding to that region recognition box is marked as a first local label within that region recognition box.
6. The method according to claim 3, characterized in that, For each second sample image in the unlabeled training set, the second target image is obtained by performing region recognition on the second sample image based on the OCR model, including: For each second sample image in the unlabeled training set, the OCR model is used to perform region recognition on the second sample image to obtain multiple second candidate images; the multiple second candidate images include second sample images with region recognition boxes and / or second sample images without region recognition boxes. For each region recognition box in each second candidate image, the OCR model is used to extract text information from the region recognition box to obtain the text recognition result; For each second candidate image, the text recognition results corresponding to each region recognition box are filtered based on a confidence threshold. Region recognition boxes with a confidence level higher than the confidence threshold and their corresponding text recognition results are retained to obtain the second target image.
7. The method according to claim 6, characterized in that, For each second target image, the identified regions are annotated with second local labels to obtain the first sample image, including: For each region recognition box in each second target image, the text recognition result corresponding to that region recognition box is used as a second local label and labeled within that region recognition box.
8. The method according to claim 1, characterized in that, The masking mechanism in the initial model includes setting the mask value inside the region identification box to 0 and setting the mask value outside the region identification box to negative infinity.
9. A training device for a pre-trained image feature extraction model, characterized in that, The device includes: The first acquisition unit is used to acquire a pre-training dataset; the pre-training dataset includes multiple first sample images; the multiple first sample images include a first target image and a second target image; the first target image is a sample image labeled with region recognition boxes and corresponding first local labels; the second target image is a sample image labeled with region recognition boxes and corresponding second local labels; the region recognition boxes labeled in the first target image are obtained by the Segmentation All Model (SAM); the first local label is a category identifier in numerical or symbolic form, and the first local label is obtained by clustering the region recognition boxes identified by the SAM model; the second local label is a text description consistent with the visual content of the image within the region recognition box, and the second local label is obtained by labeling based on the Optical Character Recognition (OCR) model; The feature extraction unit is used to extract multiple local features from each first sample image in the pre-training dataset using the initial model; the initial model adopts a ViT structure; the initial model includes a masking mechanism; the masking mechanism is used to convert the global features extracted by the initial model into one or more local features, and the formula for converting global features into local features is as follows: R = σ( )·v; σ(·) represents the softmax normalization function; Q is the query matrix; V is the global feature; M is the mask value corresponding to the region recognition box; d is the scaling factor; The parameter adjustment unit is used to adjust the parameters of the initial model based on all local features and the local labels corresponding to each local feature, in conjunction with the loss function, to obtain a pre-trained image feature extraction model.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed on a terminal device, cause the terminal device to perform the training method for the pre-trained image feature extraction model as described in any one of claims 1-8.