A Transformer-based visual large model training system

By adopting the Transformer structure in the visual big model training system, dynamically distinguishing fuzzy and clear areas, automatically generating pseudo-labels, dynamically adjusting feature weights, and achieving multi-scale feature fusion through cross-layer information transmission, the problems of poor detailed feature recognition effect and low learning efficiency of label-free samples in the existing technology are solved, and the model's recognition and analysis capabilities are significantly improved.

CN119169414BActive Publication Date: 2025-05-16PINGYI TECH (HANGZHOU) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411666890.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-21
Publication Date
2025-05-16
Estimated Expiration
2044-11-21

AI Technical Summary

Technical Problem

The prior art is difficult to accurately distinguish feature weights in the processing of image blur areas and clear areas, resulting in poor recognition of detailed features by the model and insufficient classification guidance in feature learning of label-free samples, which affects the model's learning efficiency and accurate understanding of local details.

Method used

The visual big model training system based on Transformer is adopted, and the area weighted graph is obtained through the fuzzy area selection module, which dynamically distinguishes fuzzy and clear areas and gives clear areas higher weights; the similarity pseudo-label assignment module assigns pseudo-labels to label-free samples through the similarity matrix; the collaborative feature association module dynamically adjusts the weights of local and global features through the context correlation matrix; the hierarchical multi-scale modeling module embeds small-scale features into large-scale areas through cross-layer information transmission.

Benefits of technology

The model's discrimination ability in detailed rich scenarios is improved, the feature learning accuracy and efficiency of label-free samples is improved, the precise correlation matching between local details and the global structure is achieved, and the model's ability to coordinate the processing of target features at different levels is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119169414B_ABST
    Figure CN119169414B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of model training technology, specifically a Transformer-based visual large model training system, the system comprising: a fuzzy area selection module obtains distribution information of fuzzy areas and clear areas based on input image data, performs weight distribution according to the fuzzy areas and clear areas of the image, obtains a regional weighted map, and uses the regional weighted map in the Transformer self-attention layer to generate an attention distribution map after weight adjustment. In the present invention, by processing feature information such as image brightness, color change, contrast, edge clarity and texture density, fuzzy and clear areas are dynamically distinguished, and clear areas are given higher weights, so that the model focuses more on areas with high information content, and improves the resolution ability in scenes with rich details. Based on similarity pseudo-labels, the accuracy and efficiency of unlabeled samples in the feature learning process are improved by the feature similarity relationship between labeled samples and unlabeled samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of model training technology, and in particular to a Transformer-based large-scale visual model training system. Background Art

[0002] The field of model training technology refers to the process of training with large amounts of data to gradually learn tasks such as recognition, classification, and prediction. Model training covers steps from data preparation, feature extraction, model selection, training process, parameter optimization to verification, evaluation, and deployment. With the development of artificial intelligence technology, the types of algorithms involved in the field of model training are increasing, such as neural networks, support vector machines, decision trees, clustering algorithms, etc. Various algorithms show their own advantages in different scenarios. Through model training, different types of data (such as text, images, videos, audio, etc.) can be converted into structured data with semantic information.

[0003] Among them, the visual large model training system refers to a training model specifically for visual data (such as images and videos). It is usually used to train large visual models to enable them to have visual perception capabilities such as image recognition, object detection, and scene understanding. The use of the visual large model training system is very wide, including application scenarios such as autonomous driving, security monitoring, face recognition, image retrieval, and medical image analysis. Through deep learning of visual data, the model can gradually achieve high-precision recognition and classification capabilities.

[0004] The existing technology has difficulty in accurately distinguishing the feature weights of blurred and clear areas of an image, resulting in poor recognition of detail features by the model under the interference of blurred areas. For example, in visual inspection tasks, the presence of blurred areas will weaken the ability to resolve object edges, resulting in deviations in object contour recognition. In terms of feature learning of unlabeled samples, the existing technology does not provide sufficient guidance for the classification of unlabeled samples during large-scale data processing, which affects the learning efficiency of the model for unlabeled data. In terms of the association between local and global features, it is difficult to achieve dynamic association between details and overall structures, which affects the model's accurate understanding of local details in tasks such as image segmentation. In addition, the existing technology has difficulty in achieving unified scale fusion in the processing of multi-scale features, which makes the model less adaptable to targets of different scales and affects the recognition effect of the model in complex scenes. Summary of the invention

[0005] The purpose of the present invention is to solve the shortcomings of the prior art and propose a Transformer-based visual large model training system.

[0006] In order to achieve the above object, the present invention adopts the following technical solution: A Transformer-based visual large model training system includes:

[0007] The fuzzy area selection module obtains the distribution information of fuzzy areas and clear areas based on the input image data, assigns weights according to the fuzzy areas and clear areas of the image, obtains the area weighted map, and uses the area weighted map in the Transformer self-attention layer to generate an attention distribution map after weight adjustment;

[0008] The similarity pseudo-label assignment module refers to the weight of the image area in the attention distribution map after weight adjustment, constructs a similarity matrix between samples, assigns pseudo-labels to unlabeled samples, and assigns unlabeled samples to the category to which the most similar labeled samples belong, to obtain pseudo-label assignment results;

[0009] The collaborative feature association module constructs a context association matrix based on the local detail areas of the image in the high-weighted areas in the region weighted map and the information of the similarity matrix between samples in the pseudo-label assignment result, dynamically adjusts the weights of local features and global features, and obtains a cross-level feature association map;

[0010] The hierarchical multi-scale modeling module scales the input image features based on the cross-level feature association graph, embeds the small-scale region features into the large-scale region to achieve feature fusion through cross-layer information transmission, and inputs the fused features into the subsequent Transformer layer to generate multi-scale feature training results.

[0011] As a further solution of the present invention, the step of obtaining the regional weighted map is specifically as follows:

[0012] Based on the input image data, the image is divided into several sized blocks using the formula:

[0013]

[0014] Calculate the blur score , get the fuzziness evaluation value of the area;

[0015] in, , is the combined weight coefficient used to adjust the relative influence of brightness, color change, contrast and edge clarity, and spatial frequency. is the brightness value of the area, is the upper limit of brightness normalization, that is, the maximum value of brightness, is the color change value, is the normalized upper limit of color variation, is the contrast difference value, is the normalized upper limit of contrast, is the edge definition, is the normalized upper limit of edge sharpness, is the spatial frequency, is the normalized upper limit of spatial frequency, , , , , is the feature weight coefficient;

[0016] According to the blur evaluation value of the area, the blur degree of the area block is analyzed and the corresponding weight is set. The weight data of each area block is filled into a weight matrix with the same size as the area, and the weight matrix is ​​spliced ​​and combined according to the position of the area in the image to obtain a regional weighted map.

[0017] As a further solution of the present invention, the step of obtaining the attention distribution map after weight adjustment is specifically as follows:

[0018] Based on the region weighted graph, the weighted value of each region is mapped to the input matrix of the Transformer self-attention layer, and the pixel position is adjusted according to the weighted value of the region where the pixel is located to obtain an input matrix with regional weight distribution;

[0019] According to the input matrix with regional weight distribution, the attention coefficient of the attention head to the clear and blurred areas is adjusted, the attention to the clear area is increased according to the weighted value, and the attention to the blurred area is reduced. The attention matrix of each attention head is integrated to obtain the attention distribution map after weight adjustment.

[0020] As a further solution of the present invention, the step of obtaining the similarity matrix between samples is specifically as follows:

[0021] According to the weight-adjusted attention distribution map, the weight information of the image area is mapped to the image feature space using the formula:

[0022]

[0023] Calculate the similarity value of the feature vector of the labeled sample and the unlabeled sample , get the similarity information of feature vectors;

[0024] in, represents each component in the eigenvector, is the dimension of the feature vector, is the feature vector of the labeled sample, is the first labeled sample eigenvalues, is the feature vector of the unlabeled sample, is the first unlabeled sample eigenvalues, is the weight adjustment factor;

[0025] According to the similarity information of the feature vectors, the feature vectors of the unlabeled samples and the labeled samples are compared item by item, and the similarity values ​​of each pair of samples are filled into the corresponding matrix positions to construct a similarity matrix between the samples.

[0026] As a further solution of the present invention, the step of obtaining the pseudo label allocation result is specifically:

[0027] Based on the data of the similarity matrix between the samples, extract the similarity information of each unlabeled sample, compare the similarity values ​​of the unlabeled sample and the labeled sample, select the labeled sample with the highest similarity, assign the unlabeled sample to the corresponding category, and generate an initial classification result of the label;

[0028] Based on the initial division results of the labels, unlabeled samples and labeled samples with pseudo labels are input into the Transformer model, the feature vectors of the unlabeled samples are gradually optimized through self-supervised learning, and the pseudo labels are verified and fine-tuned in training iterations to obtain pseudo label assignment results.

[0029] As a further solution of the present invention, the step of obtaining the context association matrix is ​​specifically as follows:

[0030] According to the similarity matrix information in the pseudo-label assignment result, a high-weight region is identified from the region weighted map, color, texture and edge features in the region are extracted, and the overall image features are processed through a global pooling operation to obtain a set of local features and a global feature;

[0031] According to the local feature and global feature set, the formula is adopted:

[0032]

[0033] Calculate the normalized correlation , used for the matching degree between the local feature set and the global feature set, and constructing the context association matrix;

[0034] in, Used to indicate the components of local and global features, represents the number of components contained in the local and global eigenvectors, is the local eigenvector Quantity, is the first feature in the global feature set Quantity, is the adjustment factor, Represents each component of the local feature set, Represents the components of the global feature set, is the first in the local feature set Quantity, is the first in the global feature set A quantity.

[0035] As a further solution of the present invention, the step of obtaining the cross-level feature association graph is specifically as follows:

[0036] According to the information of the context association matrix, referring to the association degree between each local feature and the global feature, dynamically adjusting the weight of the feature according to the level of association, and generating an adjusted feature set;

[0037] According to the adjusted feature set, a context association graph is constructed layer by layer in the multi-layer self-attention layer structure of the Transformer model, and the association information of each layer is accumulated and transmitted layer by layer to obtain a cross-layer feature association graph.

[0038] As a further solution of the present invention, the step of obtaining the multi-scale feature training result is specifically:

[0039] Based on the cross-layer feature association map, the input image features are scaled, large-scale and small-scale feature regions are identified and separated, detail features in the small-scale region are extracted, and gradually embedded into the large-scale region through cross-layer information transfer to obtain a fused feature map;

[0040] Based on the fused feature map, it is input into the subsequent Transformer layer, and the image features of each layer are dynamically updated through inter-layer information transmission, and the details and global structural features are accumulated layer by layer to generate multi-scale feature training results.

[0041] Compared with the prior art, the advantages and positive effects of the present invention are:

[0042] In the present invention, by processing the feature information of the image such as brightness, color change, contrast, edge clarity and texture density, the fuzzy and clear areas are dynamically distinguished, and the clear areas are given higher weights, so that the model focuses more on the areas with high information content, and improves the resolution ability in the scenes with rich details. Based on the similarity pseudo-label, the guidance label is automatically generated by the feature similarity relationship between the labeled sample and the unlabeled sample, which improves the accuracy and efficiency of the unlabeled sample in the feature learning process. In the feature association processing, the local details of the image can be effectively combined with the global structure, and the accurate association matching of the local and overall information is achieved through the dynamic context association matrix, so that the model can fully understand the semantic relationship across levels. Through the transmission and fusion of cross-level information, small-scale features can be embedded in large-scale areas, further enhancing the model's ability to coordinate the processing of detail features and overall structures, which helps to more accurately capture target features of different levels in image segmentation and detection tasks, and enhances the recognition and analysis capabilities of the model in complex visual tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 is a system flow chart of the present invention;

[0044] Figure 2 A flow chart for obtaining a regional weighted graph for the present invention;

[0045] Figure 3 A flow chart of obtaining a weight-adjusted attention distribution map for the present invention;

[0046] Figure 4 A flow chart for constructing a similarity matrix between samples for the present invention;

[0047] Figure 5 A flowchart of the pseudo-label assignment result obtained by the present invention;

[0048] Figure 6 A flowchart for constructing a context association matrix for the present invention;

[0049] Figure 7 A flow chart for obtaining a cross-level feature association graph for the present invention;

[0050] Figure 8 A flow chart for generating multi-scale feature training results for the present invention. DETAILED DESCRIPTION

[0051] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0052] In the description of the present invention, it should be understood that the terms "length", "width", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate positions or positional relationships based on the positions or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, in the description of the present invention, "multiple" means two or more, unless otherwise clearly and specifically defined.

[0053] See also Figure 1 , a Transformer-based visual large model training system includes:

[0054] Based on the input image data, the fuzzy area selection module obtains the brightness, color change, contrast difference, edge clarity and texture density information of the image area, analyzes the degree of blur in the area, obtains the distribution information of the fuzzy area and the clear area, calculates the weight according to the fuzzy area and the clear area of ​​the image, sets the fuzzy area to a low weight and the clear area to a high weight, summarizes the regional weight distribution of the entire image, obtains the regional weighted map, uses the regional weighted map in the Transformer self-attention layer, adjusts the attention of each attention head to the area, and generates the attention distribution map after weight adjustment;

[0055] The similarity pseudo-label assignment module refers to the weight of the image region in the attention distribution map after weight adjustment, maps the weight information of the region to the image feature space, extracts the feature vectors of the labeled samples and the unlabeled samples, calculates the cosine similarity of the feature vectors between the labeled samples and the unlabeled samples, clusters the unlabeled samples according to the similarity values, constructs the similarity matrix between the samples, assigns pseudo-labels to the unlabeled samples according to the similarity matrix between the samples, assigns the unlabeled samples to the category to which the most similar labeled samples belong, and obtains the pseudo-label assignment results, which are used to guide the unlabeled sample feature learning process in the Transformer model;

[0056] The collaborative feature association module extracts the local image detail areas of the high-weighted areas in the region weighted map, including color, texture and edge features, to obtain a local feature set. Referring to the overall information of the region weighted map, the average feature value of the entire image is obtained through a global pooling operation to reflect the overall structure and semantic background of the image, and a global feature set is obtained. Based on the information of the similarity matrix between samples in the pseudo-label assignment result, the local features of each unlabeled sample are associated and matched with the global features. By calculating the correlation between the local feature set and the global feature set, a context association matrix is ​​constructed, and the weights of the local features and the global features are dynamically adjusted. A multi-layer context association structure is constructed for each feature map to obtain a cross-layer feature association map, which is used as the input of each layer of Transformer. The image feature weights of each layer are dynamically updated, and context association information is accumulated layer by layer between each layer.

[0057] The hierarchical multi-scale modeling module divides the input image features into scales based on the cross-level feature association graph, extracts large-scale feature regions and small-scale feature regions, embeds the small-scale region features into the large-scale region through cross-layer information transmission to achieve feature fusion, and inputs the fused features into the subsequent Transformer layer to generate multi-scale feature training results for Transformer visual large model training and optimization;

[0058] The attention distribution map after weight adjustment includes the low weight distribution of blurred areas, the high weight distribution of clear areas and the regional weight structure of the overall image. The pseudo-label assignment results include the category assignment information of unlabeled samples, the matching degree with labeled samples and the feature vector mapping information of unlabeled samples. The cross-level feature association map includes local feature weight adjustment, global feature weight assignment and the dynamic mapping relationship between local and global features. The multi-scale feature training results include small-scale features embedded in large-scale areas, the fused multi-scale feature mapping and the information transfer structure between multi-scale features.

[0059] See also Figure 2 , the specific steps for obtaining the regional weighted graph are:

[0060] Based on the input image data, the image is divided into several sized blocks using the formula:

[0061]

[0062] Calculate the blur score , get the fuzziness evaluation value of the area;

[0063] in, Used to determine whether the area is blurred. , is the combined weight coefficient used to adjust the relative influence of brightness, color change, contrast and edge clarity, and spatial frequency to meet , adjusted through experiments and The test model is tested in different and The effect of distinguishing blurry areas and clear areas under the combination, for example: first set and , observe the model's accuracy in distinguishing different clarity areas, record the classification accuracy, and then Adjusted to 0.6, Adjust to 0.4, re-evaluate the classification effect of the model, record the accuracy, and further adjust , , Repeat the above process, after multiple rounds of experiments and result analysis, observe the performance of the model under different combinations, and finally determine the optimal parameter combination so that the model can achieve the best classification effect in the identification of fuzzy areas and clear areas. It is the brightness value of the area, reflecting the overall brightness level of the area. The brightness value is obtained through image processing tools, such as accumulating the grayscale values ​​of pixels in the area and taking the average to obtain the brightness of the current area. is the upper limit of brightness normalization, that is, the maximum value of brightness, used for normalization processing. It is the color change value, which reflects the degree of color change within the area. The color change is analyzed by image processing tools to analyze the standard deviation of the RGB channels in the area and quantify the degree of color difference. is the normalized upper limit of color change, used for standardization, It is the contrast difference value, which is used to measure the change of pixel brightness in the area. The standard deviation of pixel brightness reflects the contrast difference. The larger the value, the more significant the contrast change in the area. is the normalized upper limit of contrast, used for normalization of contrast difference values, It is the edge clarity, which indicates the density of edges in the area. It is obtained through edge detection tools (such as Sobel operator) and uses the distribution density of edge pixels as a measurement standard. is the normalized upper limit of edge sharpness, used to normalize edge density, It is the spatial frequency, which indicates the density of regional texture. It is obtained through frequency domain analysis tools (such as Fourier transform) and reflects the texture density with characteristic frequency information. is the normalized upper limit of spatial frequency, used for standardization of spatial frequency, , , , , It is the feature weight coefficient, which is used to adjust the influence of each feature on the fuzzy degree calculation. The weight value is determined by model experiments. In the experiment, the weight of each feature is adjusted to observe its effect on the regional fuzzy degree classification, and finally the weight is evenly distributed.

[0064] If the brightness of the area , color change , contrast difference , edge clarity , spatial frequency ; The normalized upper limit is: , , , , ; Combination weight , , and the feature weights are all 0.2.

[0065] Substitute into the formula:

[0066]

[0067]

[0068]

[0069]

[0070] The result shows a blur score of 0.1856.

[0071] According to the fuzziness evaluation value of the region, the fuzziness of the region block is analyzed and the corresponding weight is set. The weight data of each region block is filled into a weight matrix with the same size as the region, and the weight matrix is ​​spliced ​​and combined according to the position of the region in the image to obtain a regional weighted map;

[0072] After obtaining the blur score, in order to assign weights to the image regions, the blur regions with higher scores are set to low weights, while the clear regions with lower scores are set to high weights. First, the blur score of each region block is obtained. Specifically, the blur score of each region block is based on multiple features (such as brightness, color change, contrast, etc.), so as to obtain the clarity evaluation of each region. For example, the brightness of a region block is 150, the color change is 50, the contrast is 30, etc., and the blur score of the region block is 0.1856. Then, the region blocks are weighted according to the blur score. Assuming that the clear region threshold is set to 0.1 and the blur region threshold is set to 0.2, the region blocks with scores less than 0.1 are divided into clear regions, the region blocks with scores greater than 0.2 are divided into blur regions, and the region blocks with scores between 0.1 and 0.2 are medium blur regions. According to the division results, a weight of 1 is assigned to the clear region block, a weight of 0.5 is assigned to the medium blur region block, and a weight of 0.1 is assigned to the blur region block. Then, for each region block, a weight matrix with the same size as the region is generated based on the weight values ​​after division. For example, if a region block is classified as a clear region, the weights of all pixels in the region are set to 1; if the region block belongs to a blurred region, the weights of all pixels are set to 0.1. Then, these weight matrices are spliced ​​together according to their positions in the image, and superimposed block by block to form a complete region weighted map. The region weighted map reflects the different weight information of the clear and blurred regions in the image.

[0073] See also Figure 3 , the steps to obtain the attention distribution map after weight adjustment are as follows:

[0074] Based on the region weighted graph, the weighted value of each region is mapped to the input matrix of the Transformer self-attention layer, and the pixel position is adjusted according to the weighted value of the region where the pixel is located to obtain the input matrix with the regional weight distribution;

[0075] First, at the input stage of the self-attention layer, the weight value of each region in the region weighted map is mapped to the input matrix of the self-attention layer. Specifically, each pixel position processed by the self-attention layer will be adjusted according to the weight value of the region in which it is located. Clear areas have higher weights, and when mapped to the input of the self-attention layer, their corresponding attention values ​​will be amplified, ensuring that the clear area features occupy a larger proportion in subsequent attention calculations; on the contrary, the low weight of the blurred area weakens the mapping, making the attention value of the area smaller. Through this weight mapping method, the region weighted map is decoded into the weight distribution of each region in the self-attention layer, and guides the model to pay more attention to the clear areas with high weights in the self-attention mechanism.

[0076] According to the input matrix with regional weight distribution, adjust the attention coefficient of the attention head to the clear and blurred areas, increase the attention to the clear area according to the weighted value, and reduce the attention to the blurred area. Integrate the attention matrix of each attention head to obtain the attention distribution map after weight adjustment;

[0077] Inside the self-attention layer, each attention head will pay different degrees of attention to each region according to the weight value in the regional weighted map. First, each attention head receives the mapped input matrix. For the parts with higher weight values ​​in the clear area, the attention head will assign a higher attention coefficient to increase its feature weight in the overall attention distribution; for the blurred area, each attention head assigns a lower coefficient to maintain low attention. After adjustment, each attention head generates a regional attention matrix, reflecting its attention weight in different regions. Then, the attention matrices generated by each attention head are merged and summarized to generate the attention distribution map after weight adjustment.

[0078] See also Figure 4 , the steps to obtain the similarity matrix between samples are as follows:

[0079] According to the attention distribution map after weight adjustment, the weight information of the image area is mapped to the image feature space using the formula:

[0080]

[0081] Calculate the similarity value of the feature vector of the labeled sample and the unlabeled sample , get the similarity information of feature vectors;

[0082] in, It is an indicator of the similarity between the model's unlabeled samples and the labeled samples. is a cumulative index representing each component in the eigenvector, It is the dimension of the feature vector, which refers to the total number of components of each feature vector. The feature vector usually contains image features such as brightness, color, and contrast. It is the feature vector of the labeled sample, which is composed of the eigenvalues ​​of each image region. The extraction process is formed by combining multiple features of the image (such as brightness, color change, and contrast) according to specific weights. It is usually calculated and standardized by feature extraction tools. is the first labeled sample The characteristic value, such as the characteristic component of brightness, color change or contrast, is obtained by summing the characteristic values ​​of the region. It is the feature vector of the unlabeled sample. The feature information (such as brightness, contrast, color, etc.) of each image area is extracted and combined according to the corresponding weights to obtain the feature vector. is the first unlabeled sample feature value, representing the unlabeled sample in The values ​​on the feature dimensions are obtained through image feature processing tools. is the weight adjustment factor, which indicates the weight value assigned to each feature dimension in the attention distribution map after weight adjustment. The weight value is obtained by the model's attention distribution to different regions. It is usually obtained through experimental tuning so that clear areas get higher weights and blurred areas get lower weights. For example, a set of labeled samples is selected and initial weights are set for the multi-dimensional features of the model (such as brightness, color, contrast, etc.). Assume that each feature weight is evenly distributed at the beginning (such as brightness weight is 0.33, color weight is 0.33, and contrast weight is 0.34), and record the initial classification accuracy of the model. Adjust the brightness feature weight to 0.5, and set the color and contrast weights to 0.25 respectively. Use the adjusted weight settings to perform a classification accuracy test on the model and record the new classification accuracy. Repeat the above steps to gradually increase or decrease the weights of each feature. For example, increase the color weight to 0.4, reduce the brightness weight to 0.3, keep the contrast at 0.3, and test the model again. Record the classification accuracy after each weight adjustment and compare the effects of different weight combinations. After multiple experiments, select the weight combination with the highest classification accuracy as the final weight adjustment factor setting.

[0083] If the feature vectors of labeled samples and unlabeled samples are and , the corresponding weight adjustment factor is , substitute these values ​​into the formula to calculate the similarity:

[0084]

[0085] The results show that the similarity value is 2.26.

[0086] The weight information of the image region is transferred to the feature space. In this process, the weight of each image region will be combined with the multidimensional feature data of the region so that the image region can be arranged in order in the feature space. Specifically, the weight information is fully reflected in the feature space by multiplying the weight value of each region by the feature vector component of the region (such as brightness, color, contrast, etc.). This step ensures that each feature dimension is adjusted according to the weight, thereby completing the conversion of the regional weight information to the feature space. After completing the weight mapping, the feature vectors of the image are gradually extracted. For the labeled samples, their feature vectors represent the image region features of known categories, and a set of reference feature vectors are formed after unified extraction by category; for the unlabeled samples, the weighted feature vectors of each region are extracted in turn to form a set of feature vectors. These vectors retain the spatial distribution characteristics of attention after weight adjustment. After the feature vectors of the labeled samples and the unlabeled samples are extracted, two sets of feature vectors are formed.

[0087] According to the similarity information of the feature vectors, the feature vectors of the unlabeled samples and the labeled samples are compared one by one, and the similarity values ​​of each pair of samples are filled into the corresponding matrix positions to construct the similarity matrix between the samples;

[0088] When clustering unlabeled samples, the similarity between unlabeled samples and labeled samples is first calculated using cosine similarity. For example, it represents an unlabeled sample With labeled samples When constructing the similarity matrix, the similarity values ​​calculated between the unlabeled samples and the labeled samples are filled into the corresponding positions of the matrix one by one. and Similarity , recorded in the matrix as matrix elements Repeat the above similarity calculation process, fill the similarity values ​​of all unlabeled samples and each labeled sample into the matrix, and build a complete similarity matrix. The rows of the matrix represent unlabeled samples, and the columns represent labeled samples. For example, if there are 3 unlabeled samples and 3 labeled samples , then the shape of the similarity matrix is , each element in the matrix Represents unlabeled samples And labeled samples The constructed similarity matrix is ​​applied to the hierarchical clustering method (such as the average linkage method) to cluster the unlabeled samples according to the similarity value. For example, if Represents unlabeled samples With labeled samples The similarity is higher, and , , then hierarchical clustering will convert unlabeled samples Prefer to be consistent with labeled samples Classify them into the same category. Through the hierarchical operation of clustering, each unlabeled sample is gradually merged into the most similar labeled sample category according to the size of the similarity value. The closest labeled sample The unlabeled samples are classified into the same cluster category, thus achieving preliminary classification. Each row of the matrix represents the similarity distribution between an unlabeled sample and each labeled sample, and each column represents the similarity distribution between all unlabeled samples and a labeled sample. This matrix structure clearly reflects the similarity between unlabeled samples and labeled samples, making it easy to apply clustering algorithms directly in the matrix to group unlabeled samples. The resulting similarity matrix not only provides a clear data basis for the clustering of unlabeled samples, but also makes the cluster distribution of similar samples more intuitive.

[0089] See also Figure 5 , the specific steps for obtaining the pseudo-label assignment results are:

[0090] Based on the data of the similarity matrix between samples, extract the similarity information of each unlabeled sample, compare the similarity values ​​of the unlabeled sample and the labeled sample, select the labeled sample with the highest similarity, assign the unlabeled sample to the corresponding category, and generate the initial classification result of the label;

[0091] In the process of assigning pseudo labels to unlabeled samples, we first refer to the similarity matrix between samples, which shows the similarity values ​​between each unlabeled sample and all labeled samples. The specific operation is as follows: extract the similarity row of each unlabeled sample from the similarity matrix in turn, and take out the similarity value between the unlabeled sample and all labeled samples. For example, for unlabeled samples , extract a whole row of values ​​in the similarity matrix, which correspond to Similarity with each labeled sample. For each unlabeled sample, traverse the similarity values ​​in its similarity row, compare the values, and find the highest similarity value. For example, if the unlabeled sample The similarity row contains the labeled sample , , The similarity values ​​of are 2.3, 1.8 and 2.1 respectively, so choose The similarity value of 2.3 is taken as the maximum value. The unlabeled sample is assigned to the category of the labeled sample with the highest similarity, and the category label of the labeled sample is recorded as the pseudo label of the unlabeled sample. Continuing with the previous example, the unlabeled sample With labeled samples The similarity is the highest, so Assign to The category to which it belongs and The category label is given As a pseudo label. Repeat the above process, compare the similarity and assign pseudo labels to all unlabeled samples in turn, ensuring that each unlabeled sample is assigned to the most similar labeled sample category. This process traverses all unlabeled sample rows in the similarity matrix, and finally obtains the initial division result for each unlabeled sample.

[0092] Based on the initial label division results, the unlabeled samples and labeled samples with pseudo labels are input into the Transformer model. The feature vectors of the unlabeled samples are gradually optimized through self-supervised learning. The pseudo labels are verified and fine-tuned in the training iterations to obtain the pseudo label assignment results.

[0093] Unlabeled samples with pseudo labels are combined with existing labeled samples to form a new training data set. The pseudo labels of unlabeled samples are used as preliminary category identifiers to provide more category information for the model's feature learning process. In the self-supervised learning stage of the model, unlabeled samples with pseudo labels are input into the model together with labeled samples. The model extracts features layer by layer through the self-attention mechanism and gradually updates the representation of the feature vector, so that the feature distribution of the unlabeled samples gradually fits the corresponding labeled category features. During the training process, the model continuously updates the feature vectors of unlabeled samples and compares them with the feature vectors of similar labeled samples to further verify the effectiveness of the pseudo labels. If the feature vectors of unlabeled samples show high category consistency during model training, the pseudo labels are kept unchanged; if deviations occur, the pseudo labels can be fine-tuned again in subsequent iterations. Through multiple iterations of training, the feature vectors of unlabeled samples gradually approach the category features corresponding to their pseudo labels in the feature space. At this time, the model adjusts the focus of the self-attention mechanism to strengthen the similarity between unlabeled samples and labeled samples, while suppressing the similarity between different categories to ensure that the feature representation of unlabeled samples is gradually clear and stable. After multiple rounds of training, the Transformer model completes the optimization of the unlabeled sample features, combines the pseudo-label assignment information with the final feature learning results, confirms the final position of the unlabeled samples in the feature space, and outputs the pseudo-label assignment results.

[0094] See also Figure 6 , the steps to obtain the context association matrix are as follows:

[0095] According to the similarity matrix information in the pseudo-label assignment result, the high-weight region is identified from the region weighted map, the color, texture and edge features in the region are extracted, and the overall image features are processed through the global pooling operation to obtain the local feature and global feature set;

[0096] Based on the similarity matrix information in the pseudo-label assignment result, the local image details of the high-weight region are extracted from the region weighted map. The specific operation process is as follows: the color distribution of each high-weight region is quantified by the color histogram method, and the color distribution in the region is obtained pixel by pixel. Specifically, by decomposing the RGB channels, the different color values ​​in each color channel are counted to generate a color histogram, so as to obtain the color feature data of the region, and the data is added to the local feature set. The texture information of each high-weight region is obtained by the frequency domain analysis method. First, Fourier transform is applied to the high-weight region to convert the image region into frequency information, and the main frequency components in the spectrum are extracted to reflect the texture density and directionality in the region. Then, the main frequency components are statistically processed to form a texture feature vector, the texture features of the region are recorded, and the results are added to the local feature set. An edge detection algorithm (such as the Sobel operator) is applied to each high-weight region to identify the edge density. The specific steps are: the Sobel operator is applied to the pixels in the region, the edge gradient value of each pixel is obtained and the edge pixel points are marked, and the total number and distribution density of the edge pixel points in the region are counted, which are used as the quantitative index of the edge feature. Finally, the edge feature data is recorded in the local feature set. After completing the color, texture and edge feature extraction of each high-weight area, the features are summarized to obtain the local feature set. Then, the average feature value of the entire image is obtained through the global pooling operation to reflect the structure and semantic background of the entire image. The specific operation is: weighted averaging of the color, texture and edge feature values ​​of all pixels in the entire image. For color features, the color average value of each channel of the entire image is obtained; for texture features, the average frequency value in the entire image spectrum is obtained; for edge features, the edge density mean of the image is obtained. The set of average feature values ​​generated by the global pooling constitutes the global feature set.

[0097] According to the local features and global feature sets, the formula is adopted:

[0098]

[0099] Calculate the normalized correlation , used for the matching degree between the local feature set and the global feature set, and constructing the context association matrix;

[0100] in, is an index variable used to indicate the components in local and global features. is the dimension of the feature vector, which indicates the number of components contained in the local and global feature vectors. This value depends on the dimension set during feature extraction. The specific extraction method is usually determined through experiments to capture the main feature information of the image. It is the feature vector of the local feature set, which represents the detail features of the high-weight area of ​​the image. is the local eigenvector A component reflects a certain feature of the high-weight area (such as color, texture or edge). The local detail feature set of the high-weight area is extracted region by region through image processing tools. For example, edge features are extracted by calculating the brightness difference of the image. After the image is divided into blocks, the detail features are extracted from the high-weight area. Each feature is extracted through the feature extraction method. If the color feature is extracted, the color distribution of the area is analyzed pixel by pixel and the color feature value is obtained. is the eigenvector of the global feature set, representing the average eigenvalue of the entire image. is the first feature in the global feature set Components, reflecting the overall structural information or semantic background of the image, obtain overall features through global pooling operations, capture the average feature value of the entire image, such as the mean or standard deviation, representing the average color or texture of the entire image, It is obtained by performing global pooling on the entire image, for example, averaging the pixels of the entire image and extracting the overall mean of each feature. If texture features are extracted, the mean of the frequency distribution of the entire image can be summed up through the frequency analysis tool. is the adjustment factor, The weight of each feature in the association between local and global features is used to highlight or reduce the matching importance of different features. The adjustment factor is adjusted according to the experiment to optimize the contribution of each feature to the overall association. By setting up experiments, it is determined according to the impact of the feature on the image structure. For example, local features (such as color, texture, and edges) are extracted from different image samples, and the correlation of these features is calculated separately in the context of global features to observe the impact of each feature on the overall matching of the image. Then, a higher adjustment factor weight is set for features with greater influence (such as color) to increase their contribution to the correlation to highlight their importance in the overall semantic matching; for features with less influence (such as edges), a lower adjustment factor weight is set to reduce their impact in the calculation. This process is optimized through successive experiments and adjustments, and finally the optimal weight configuration of each feature is determined, thereby obtaining a reasonable adjustment factor setting. Represents each component of the local feature set, Represents the components of the global feature set, is the first in the local feature set Quantity, is the first in the global feature set A quantity.

[0101] If the local eigenvector is , the global eigenvector is , adjustment factor , enter the formula to calculate the correlation :

[0102]

[0103] The results show that the correlation Indicates the matching degree between the local feature set and the global feature set.

[0104] See also Figure 7 , the specific steps for obtaining the cross-level feature association graph are:

[0105] According to the information of the context association matrix, referring to the association degree between each local feature and the global feature, the weight of the feature is dynamically adjusted according to the level of association, and an adjusted feature set is generated;

[0106] First, by analyzing each element of the context association matrix, we identify the correlation between local features and global features, and find features with high correlation, indicating that they occupy a significant position in the overall structure of the image. For example, the calculated correlation , this result can be used to judge the matching strength between local features and global features. If the matching standard threshold set by the system is , then the current correlation This indicates that the matching degree between the local feature and the global feature is low and is not enough to meet the set association standard. Based on this result, the weight of this feature can be reduced during the dynamic adjustment process to reduce its influence in the overall feature association; if the association degree is If it exceeds 0.3, its weight can be increased, thereby enhancing the impact of this feature on the overall structure of the image in the context association matrix. After the weight adjustment, all high-correlation features that meet the criteria will have a more positive impact on the overall image structure, while low-correlation features will have a weakened effect, ultimately generating an optimized dynamic feature weight distribution.

[0107] According to the adjusted feature set, a context association graph is constructed layer by layer in the multi-layer self-attention layer structure of the Transformer model, and the association information of each layer is accumulated and transmitted layer by layer to obtain a cross-layer feature association graph;

[0108] After obtaining the dynamically adjusted feature set, the set is input into each layer of the Transformer network layer by layer, and the cross-layer association of features is realized by constructing a multi-layer context association structure. The specific process is as follows: When entering the first layer of Transformer, each feature in the feature set already carries the dynamically adjusted weight value. The feature set is used as input and is assigned to each self-attention head so that each attention head independently pays attention to the relationship between different features. In each Transformer layer, the self-attention mechanism uses the weight information between features as a key reference and redistributes the relevance of features through the self-attention matrix. First, the correlation between each feature and other features in the current layer is calculated, and the degree of attention of each feature to the remaining features is calculated using the self-attention matrix, and the feature weights are proportionally adjusted so that features with high relevance are easier to be strengthened in the self-attention of this layer. The updated information in the self-attention matrix is ​​accumulated into the context association graph. This operation is repeated in each layer, and the context information of the features is gradually transmitted and expanded between layers by accumulating the updated results in each layer into the context association graph. In this accumulation process, the association results of each layer are gradually added to the association graph of the previous layer to ensure that the association weights of features are superimposed on multiple levels. After each layer of Transformer processing is completed, the feature set containing the accumulated context information is passed to the next layer. This cross-layer transfer operation makes the feature weights after each layer of processing carry the association information of the previous layers, ensuring that richer context information is obtained in higher layers. After the last layer completes the accumulation of features, the information of all layers is summarized into the final context association graph. This graph integrates the feature association and weight information of each layer to form a complete cross-level feature association graph. This association graph contains dynamic feature relationships from local to global, and has the integrity to be used as input in subsequent model processing.

[0109] See also Figure 8 , the specific steps for obtaining the multi-scale feature training results are:

[0110] Based on the cross-layer feature association map, the input image features are scaled, the large-scale and small-scale feature regions are identified and separated, the detail features in the small-scale region are extracted, and gradually embedded into the large-scale region through cross-layer information transfer to obtain the fused feature map;

[0111] First, the large-scale feature regions and small-scale feature regions in the image are identified and separated by the feature density, edge clarity, color and texture differences of each region in the association map. Specifically, the large-scale feature regions usually correspond to image regions with gentle color changes and low texture density, such as backgrounds or flat surfaces. Small-scale feature regions contain more detail information, such as parts with complex edges, obvious color transitions or dense textures. In this division process, the boundaries and contents of each feature region can be accurately located by referring to the distribution information in the cross-level feature association map. Next, the detail features are extracted from the divided small-scale regions, including the color, edge and texture information of each region. These features are subjected to feature extraction operations to generate a small-scale feature set. Then, the features of the small-scale region are gradually embedded in the large-scale region by cross-layer information transfer. The embedding method is the layer-by-layer superposition and matching of the small-scale feature set and the large-scale feature set. In this process, cross-layer information transfer is used to ensure that the detail features of the small-scale region are represented and retained in the large-scale region. The specific operations include pixel-by-pixel matching and superposition, so that small-scale features can be integrated into the details while maintaining the overall structure of the large-scale region, thereby improving the detail expression of the overall image and obtaining a fused feature map.

[0112] Based on the fused feature map, it is input into the subsequent Transformer layer, and the image features of each layer are dynamically updated through inter-layer information transmission, and the details and global structural features are accumulated layer by layer to generate multi-scale feature training results;

[0113] The fused feature map is input to the subsequent Transformer layer. This layer first preprocesses the fused multi-scale feature map to balance the feature map in the spatial dimension. Specifically, through the processing of the self-attention mechanism, the detailed features of the small-scale area are focused layer by layer, while the global structure of the large-scale area is maintained in each Transformer layer. The self-attention mechanism dynamically pays attention to the important feature areas in the image through the correlation information between the regions, and amplifies the information of specific areas in turn. In particular, when processing small-scale areas, the self-attention mechanism of each layer gradually accumulates the detailed performance of the features; while when processing large-scale areas, the stability of the overall structure is enhanced according to the inter-layer correlation. Through the information transfer between layers, the Transformer of each layer will gradually update and enhance the recognition of detailed features according to the feature map of the previous layer, while strengthening the continuity of the global structure. The accumulated feature information of each layer is combined to form a training result containing multi-scale features.

[0114] The above are only preferred embodiments of the present invention and are not intended to limit the present invention in other forms. Any technician familiar with the profession may use the technical contents disclosed above to change or modify them into equivalent embodiments with equivalent changes and apply them to other fields. However, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution of the present invention still falls within the protection scope of the technical solution of the present invention.

Claims

1. A Transformer-based visual large model training system, characterized in that: The system comprises: The fuzzy area selection module obtains the distribution information of fuzzy areas and clear areas based on the input image data, assigns weights according to the fuzzy areas and clear areas of the image, obtains the area weighted map, and uses the area weighted map in the Transformer self-attention layer to generate an attention distribution map after weight adjustment; The steps of obtaining the regional weighted map are specifically as follows: Based on the input image data, the image is divided into several sized blocks using the formula: Calculate the blur score , get the fuzziness evaluation value of the area; in, , is the combined weight coefficient used to adjust the relative influence of brightness, color change, contrast and edge clarity, and spatial frequency. is the brightness value of the area, is the upper limit of brightness normalization, that is, the maximum value of brightness, is the color change value, is the normalized upper limit of color variation, is the contrast difference value, is the normalized upper limit of contrast, is the edge definition, is the normalized upper limit of edge sharpness, is the spatial frequency, is the normalized upper limit of spatial frequency, , , , , is the feature weight coefficient; According to the blur evaluation value of the region, the blur degree of the region block is analyzed and the corresponding weight is set, and the weight data of each region block is filled into a weight matrix with the same size as the region, and the weight matrix is ​​spliced ​​and combined according to the position of the region in the image to obtain a region weighted map; The similarity pseudo-label assignment module refers to the weight of the image area in the attention distribution map after weight adjustment, constructs a similarity matrix between samples, assigns pseudo-labels to unlabeled samples, and assigns unlabeled samples to the category to which the most similar labeled samples belong, to obtain pseudo-label assignment results; The collaborative feature association module constructs a context association matrix based on the local detail areas of the image in the high-weighted areas in the region weighted map and the information of the similarity matrix between samples in the pseudo-label assignment result, dynamically adjusts the weights of local features and global features, and obtains a cross-level feature association map; The hierarchical multi-scale modeling module scales the input image features based on the cross-level feature association graph, embeds the small-scale region features into the large-scale region to achieve feature fusion through cross-layer information transmission, and inputs the fused features into the subsequent Transformer layer to generate multi-scale feature training results.

2. The Transformer-based visual large model training system according to claim 1, characterized in that: The steps for obtaining the attention distribution map after weight adjustment are specifically as follows: Based on the region weighted graph, the weighted value of each region is mapped to the input matrix of the Transformer self-attention layer, and the pixel position is adjusted according to the weighted value of the region where the pixel is located to obtain an input matrix with regional weight distribution; According to the input matrix with regional weight distribution, the attention coefficient of the attention head to the clear and blurred areas is adjusted, the attention to the clear area is increased according to the weighted value, and the attention to the blurred area is reduced. The attention matrix of each attention head is integrated to obtain the attention distribution map after weight adjustment.

3. The Transformer-based visual large model training system according to claim 2, characterized in that: The steps for obtaining the similarity matrix between the samples are specifically as follows: According to the weight-adjusted attention distribution map, the weight information of the image area is mapped to the image feature space using the formula: Calculate the similarity value of the feature vector of the labeled sample and the unlabeled sample , get the similarity information of feature vectors; in, represents each component in the eigenvector, is the dimension of the feature vector, is the feature vector of the labeled sample, is the first labeled sample eigenvalues, is the feature vector of the unlabeled sample, is the first unlabeled sample eigenvalues, is the weight adjustment factor; According to the similarity information of the feature vectors, the feature vectors of the unlabeled samples and the labeled samples are compared item by item, and the similarity values ​​of each pair of samples are filled into the corresponding matrix positions to construct a similarity matrix between the samples.

4. The Transformer-based visual large model training system according to claim 3, characterized in that: The steps for obtaining the pseudo label assignment result are specifically as follows: Based on the data of the similarity matrix between the samples, extract the similarity information of each unlabeled sample, compare the similarity values ​​of the unlabeled sample and the labeled sample, select the labeled sample with the highest similarity, assign the unlabeled sample to the corresponding category, and generate an initial classification result of the label; Based on the initial division results of the labels, unlabeled samples and labeled samples with pseudo labels are input into the Transformer model, the feature vectors of the unlabeled samples are gradually optimized through self-supervised learning, and the pseudo labels are verified and fine-tuned in training iterations to obtain pseudo label assignment results.

5. The Transformer-based visual large model training system according to claim 4, characterized in that: The steps of obtaining the context association matrix are specifically as follows: According to the similarity matrix information in the pseudo-label assignment result, a high-weight region is identified from the region weighted map, color, texture and edge features in the region are extracted, and the overall image features are processed through a global pooling operation to obtain a set of local features and a global feature; According to the local feature and global feature set, the formula is adopted: Calculate the normalized correlation , used for the matching degree between the local feature set and the global feature set, and constructing the context association matrix; in, Used to indicate the components of local and global features, represents the number of components contained in the local and global eigenvectors, is the first local eigenvector Quantity, is the first feature in the global feature set Quantity, is the adjustment factor, Represents each component of the local feature set, Represents the components of the global feature set, is the first in the local feature set Quantity, is the first in the global feature set A quantity.

6. The Transformer-based visual large model training system according to claim 5, characterized in that: The steps for obtaining the cross-level feature association graph are specifically as follows: According to the information of the context association matrix, referring to the association degree between each local feature and the global feature, dynamically adjusting the weight of the feature according to the level of association, and generating an adjusted feature set; According to the adjusted feature set, a context association graph is constructed layer by layer in the multi-layer self-attention layer structure of the Transformer model, and the association information of each layer is accumulated and transmitted layer by layer to obtain a cross-layer feature association graph.

7. The Transformer-based visual large model training system according to claim 6, characterized in that: The steps for obtaining the multi-scale feature training results are specifically as follows: Based on the cross-layer feature association map, the input image features are scaled, large-scale and small-scale feature regions are identified and separated, detail features in the small-scale region are extracted, and gradually embedded into the large-scale region through cross-layer information transfer to obtain a fused feature map; Based on the fused feature map, it is input into the subsequent Transformer layer, and the image features of each layer are dynamically updated through inter-layer information transmission, and the details and global structural features are accumulated layer by layer to generate multi-scale feature training results.

Citation Information

Patent Citations

  • Dynamic fuzzy processing algorithm for visual SLAM system

    CN110910332A

  • Weak supervision image segmentation method based on multi-scale pseudo label fusion

    CN117975002A