Method for representation understanding of a pointer based on data generation tuning and weight autonomous evolution

CN118918318BActive Publication Date: 2026-09-25NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411053464.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-02
Publication Date
2026-09-25
Estimated Expiration
2044-08-02

AI Technical Summary

Technical Problem

尽管通过调整视觉特征可以改善性能,但提取-调整的范式不可避免地包含大量具有固定权重的特征提取组件

Benefits of technology

[0029]1、本发明方法可以主动提取与表达相关的视觉特征,而无需手动修改视觉主干架构。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118918318B_ABST
    Figure CN118918318B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on data generation tuning and weight autonomous evolution's pointer representation understanding method, first part is initial pointer representation data generation module, second part is with negative example's context construction module, third part is context pointer representation data generation module.Fourth part is language main network, fifth part is language adaptive weight generator, sixth part is visual main network, seventh part is pointer representation understanding prediction module.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of language-image multimodal fusion technology, specifically involving a method for index expression understanding based on data generation optimization and weight autonomous evolution. Background Technology

[0002] Pointer expression understanding aims to detect specific objects based on given natural language descriptions. Compared to general object detection, which can only locate objects within a fixed set of categories, pointer expression understanding is more flexible and purposeful. Free-form language descriptions can specify specific visual attributes of the target object, such as category, attributes, relationship to other objects, relative / absolute position, etc. Due to its similarity to the detection task, previous pointer expression understanding methods typically follow a general object detection framework and emphasize the design of cross-modal interaction modules. Despite achieving impressive performance, visual backbone networks have not been well explored. Specifically, visual backbone networks extract visual features with fixed architectures and weights, without considering pointer expression. This passive feature extraction can lead to mismatches between the extracted visual features and the features required for various pointer expressions, such as missing or redundant features.

[0003] Meanwhile, a fixed visual backbone network has an inherent bias towards images, which may be unrelated to the text being referred to. Ideally, the visual backbone network should make full use of the referential expression, as the expression can provide information and biases about the desired visual features. Some methods have addressed this phenomenon and proposed corresponding solutions, which achieve visual feature extraction of the perceived meaning by inserting carefully designed interaction modules into the visual backbone network. Although performance can be improved by tuning visual features, the extraction-tuning paradigm inevitably contains a large number of feature extraction components with fixed weights. Summary of the Invention

[0004] To overcome the shortcomings of existing technologies, this invention provides an index expression understanding method based on data generation optimization and weight autonomous evolution. The method comprises seven parts: an initial index expression data generation module, a context construction module with negative examples, a context index expression data generation module, a language backbone network, a language adaptive weight generator, a visual backbone network, and an index expression understanding prediction module.

[0005] The technical solution adopted by this invention to solve its technical problem is as follows:

[0006] Step 1: Given an image dataset D, select one image from it. Using pointers to generate template T k Generate an initial expression x ikTraverse the dataset D to obtain the initial expression dataset X1;

[0007] Step 2: For a given image I i The model is given the prompt: "Please describe this image," and the model generates the corresponding description C. i ; regarding image I in X1 i The expression x ik Explicitly added to the context as context information for negative examples, while simultaneously providing the model with the hint: "Please avoid generating this type of problem," the context information known about negative examples is as follows: i C i ,x ik >;

[0008] Step 3: Using the context information obtained from the negative examples in Step 2, repeat Step 1 to generate the optimized expression data x. i1 The process is represented as:

[0009] I i2 =F θ (X i C i ,T k ,x i1 (1)

[0010] Step 4: Given a pointer to x i1 An N-layer language backbone network is used to segment the expression and add a [CLS] tag at the beginning to extract language features. Where L and d l Let F represent the number of tags and the dimension of linguistic features, respectively; then, the linguistic features F... l The input is fed into a language-adaptive weight generator to generate weights for the visual backbone network; next, given an image... Where W×H×3 represents the size of the image, expressing the perceived visual characteristics. The visual features are extracted via a visual backbone network, where C and s represent the number of channels and stride of the visual features, respectively; finally, the linguistic features, represented by [CLS] tags, are... Visual features are passed to the pointer representation understanding prediction module to predict the bounding box of the object pointed to by the pointer;

[0011] Step 5: Introduce a learnable layer-specific embedding. It operates on each layer of the visual backbone network to dynamically extract layer-related language features; for each group g, attention is... Assigned to and The normalized dot product is expressed as:

[0012] ​

[0013] Subsequently, actively perceived language features Through aggregation, we obtain: Finally, a fully connected layer is used to reduce the dimensionality of the aggregated language features of the i-th layer of the visual backbone network, as shown below:

[0014]

[0015] in Used to reduce the dimension to d h =d l / r, where r is the dimensionality reduction ratio; δ represents the eLU activation function;

[0016] Step 6: Generate adaptive weights for the language based on the given expression, which are then used to generate query X in the visual backbone network. q Key X k Sum X v , is represented as:

[0017] X q =θ(X; W) q ),X k =θ(X; W) k ),X v =θ(X; W) v (3)

[0018] Where θ(·;W) represents a linear mapping operation with parameter W, and X represents the input visual feature; It is a dynamic mapping weight used to generate queries, keys, and values, d in and d out These are feature X and the feature dimensions of the query / key / value;

[0019] The dynamic weights are generated according to the matrix factorization paradigm. For the i-th ViT block, this process is represented as follows:

[0020]

[0021] in, These are layer-specific static learnable weights; and These are statically learnable weights; It is a fully connected layer that aggregates language features As input, generate a shape of d w ×d w The dynamic matrix;

[0022] Step 7: Apply direct coordinate regression to predict the bounding box of the object being referred to;

[0023] First, visual features and language features Projected into a lower-dimensional space Then, attention weights are calculated using dot product similarity. Then, softmax normalization is performed; next, visual features are aggregated by weighted summation using attention weights A; finally, the aggregated visual features are input into a fully connected layer, and the sigmoid function is used to predict the bounding box.

[0024] Step 8: Given the predicted bounding box Given the ground truth bounding box b = (x, y, w, h), the detection loss function is defined as follows:

[0025]

[0026] in, and Let λ represent L1 loss and general IoU loss, respectively. L1 and λ giou This represents the correlation coefficient used to balance the loss.

[0027] Preferably, N = 6, L = 40, d l =512.

[0028] The beneficial effects of this invention are as follows:

[0029] 1. The method of the present invention can actively extract and express relevant visual features without manually modifying the visual backbone architecture.

[0030] 2. Extensive experiments have demonstrated the effectiveness of this invention, achieving advanced performance on widely used datasets. Attached Figure Description

[0031] Figure 1 This is a schematic diagram of the instruction expression method flow of the method of the present invention.

[0032] Figure 2 The result is understood by referring to the pointer expression using the method of the present invention. Detailed Implementation

[0033] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0034] This invention proposes a novel index expression comprehension method based on data generation optimization and autonomous weight evolution. This method considers that the index expression and model weights jointly determine the function of the visual backbone network, and adopts a simpler, finer-grained approach: using data generation optimization and autonomous weight evolution to better adapt to index expression comprehension.

[0035] The main modules of the technical solution of this invention are as follows: the first part is an initial indexing expression data generation module; the second part is a context construction module with negative examples; the third part is a context indexing expression data generation module; the fourth part is a language backbone network; the fifth part is a language adaptive weight generator; the sixth part is a visual backbone network; and the seventh part is an indexing expression comprehension and prediction module.

[0036] In the first part, initial image-related indexing representation data is generated to construct a context with negative examples in the subsequent second part. In the second part, the pre-trained model's preliminary visual-language alignment capabilities are utilized to construct the context with negative examples. In the third part, high-quality image-related indexing representation data is constructed using the constructed context information and the multimodal reasoning capabilities of the large model itself. In the fourth part, linguistic features are extracted from the high-quality indexing representation; in the fifth part, dynamic weights for the visual backbone network are generated based on the linguistic features. In the sixth part, visual features are extracted from the original image, and their behavior can be modified through language-adaptive weights. In the seventh part, the bounding box and mask of the pointed object are jointly predicted.

[0037] This method of expression involves the following main steps:

[0038] Step 1: Given an image dataset D, select one image from it. Using pointers to generate template T k Generate an initial expression x ik Traversing dataset D yields the initial expression dataset X1. In a multimodal large language model, when the language model part receives the expression generation template T... k At that time, because the model focused on language generation tasks during the pre-training stage and had not fully integrated visual understanding capabilities, the generated expression X1 lacked correlation with the image content and visual detail information, and X1 needed further optimization.

[0039] Step 2: For a given image I i The model is given the prompt: "Please describe this image.", and the model generates the corresponding description C. i Since the expression X1 generated in step 1 lacks relevance to the image content, inspired by self-criticism learning, the expression X1 related to image I is used... i The expression x ik Explicitly added to the context as context information for negative examples, while simultaneously providing the model with the hint: "Please avoid generating this type of problem.", the context information known about negative examples is as follows: i C i ,x ik > ​

[0040] Step 3: Using the context information obtained from the negative examples in Step 2, repeat Step 1 to generate the optimized expression data x. i1 This process can be represented as:

[0041] I i2 =Fx(X i C i ,T k ,x i1 (1)

[0042] Step 4: Given a pointer obtained in the above steps, reach x i1 An N-layer language backbone network is used to segment the expression, and a [CLS] tag is added at the beginning to extract language features. Where L and d l These represent the number of tags and the dimension of the linguistic features, respectively. Then, the linguistic features F... l The input is fed into a language-adaptive weight generator to generate weights for the visual backbone network. Next, given an image... Where W×H×3 represents the size of the image, expressing the perceived visual characteristics. It can be extracted using a visual backbone network, where C and s represent the number of channels and stride of the visual features, respectively. Finally, the language features represented by [CLS] tags are... Visual features are passed to the pointer interpretation prediction module to predict the bounding box of the object being pointed to.

[0043] Step 5: Considering that the representation corresponds to different numbers of linguistic tags, and that each layer of the visual backbone network may focus more on different linguistic tags, we attempt to independently aggregate linguistic features of a fixed size to each layer. Inspired by multi-head attention mechanisms, a learnable layer-specific embedding is introduced. This is applied to each layer of the visual backbone network to dynamically extract layer-related language features, which can improve the model's flexibility with almost no cost. The computation is performed across G groups. For each group g, attention is... Assigned to and The normalized dot product is expressed as:

[0044]

[0045] Subsequently, actively perceived language features It can be obtained through aggregation: Finally, a fully connected layer is used to reduce the dimensionality of the aggregated language features in the i-th layer of the visual backbone network, represented as:

[0046]

[0047] in Used to reduce the dimension to d h =d l / r, where r is the dimensionality reduction ratio. δ represents the eLU activation function.

[0048] Step 6: To guide the active perception of the visual backbone network, language adaptive weights are generated based on the index expression, which are used to generate query X in the visual backbone network. q Key X k Sum X v , represented as

[0049] X q =θ(X; W) q ),X k =θ(X; W) k ),X v =θ(X; W) v (3)

[0050] Where θ(·;W) represents a linear mapping operation with parameter W, and X represents the input visual feature. It is a dynamic mapping weight used to generate queries, keys, and values, d in and d out These are feature X and the feature dimensions of the query / key / value, respectively.

[0051] Considering the channel dimension (d) of dynamic weights out ×d in The computational burden of directly generating weights using fully connected layers is unbearable due to the large number of parameters. One solution is to alleviate this problem by weighted summation of K static convolutional kernels, but this increases the number of parameters by a factor of K and faces challenging joint optimization problems. Inspired by dynamic channel fusion, we attempt to generate dynamic weights according to the matrix factorization paradigm. Taking the i-th ViT block as an example, this process can be represented as:

[0052]

[0053] in, These are layer-specific static learnable weights. and They are also statically learnable weights, but are shared across all ViT blocks to reduce the number of parameters and prevent model overfitting. It is a fully connected layer that aggregates language features As input, generate a shape of d w ×d w The dynamic matrix.

[0054] Step 7: Unlike methods that require carefully designed cross-modal interaction modules, this method can acquire visual features related to expression through a language-aware visual backbone network without the need for additional cross-modal interaction modules. Through the proposed concise and efficient pointer expression comprehension prediction module, visual and linguistic features can be used to predict the bounding box of pointer expression comprehension. Specifically, direct coordinate regression is applied to predict the bounding box of the pointed object. To aggregate visual features along the spatial dimension, a language-adaptive aggregation module is proposed, which utilizes language-adaptive attention to aggregate visual features. First, the visual features are... and language features Projected into a lower-dimensional space Then, attention weights are calculated using dot product similarity. Softmax normalization is then performed. Next, visual features are aggregated by weighted summation using attention weights A. Finally, the aggregated visual features are input into a fully connected layer, and the sigmoid function is used to predict the bounding box.

[0055] Step 8: The framework of this invention can be optimized end-to-end for indicating expression understanding. Given the predicted bounding box... Given the ground truth bounding box b = (x, y, w, h), the detection loss function is defined as follows:

[0056]

[0057] in, and Let λ represent L1 loss and general IoU loss, respectively. L1 and λ giou This represents the correlation coefficient used to balance the loss.

[0058] Example:

[0059] This invention provides a novel indexing and interpretation method that better adapts to indexing and interpretation through data generation optimization and autonomous weight evolution. The specific process is as follows:

[0060] 1. Generation and optimization of pronoun expressions;

[0061] Given an image dataset of 10,000 images, select one image from it. Using pointers to generate template T k<Please describe this image.> Generate an initial expression. Iterate through 10,000 images to obtain the initial expression dataset. In a multimodal large language model, when the language model receives the expression generation template, because the model focuses on language generation during pre-training and has not fully integrated visual understanding, the generated expression lacks relevance and visual detail information to the image content, requiring further optimization. Since the expression lacks relevance to the image content at this point, inspired by self-criticism learning, explicitly add the expression to the context as negative example context information, while simultaneously providing the model with the prompt: "Please avoid generating this type of question.", thus obtaining the context information known from the negative example. Using the obtained context information known from the negative example, repeat step 1 to generate the optimized expression data x. i1 .

[0062] 2. Refers to the extraction of representative features;

[0063] Given an index obtained from the above steps, x i1 The expression is segmented using a 6-layer language backbone network based on a bidirectional encoder structure (BERT), and a [CLS] tag is added at the beginning to extract language features. Where L = 40 and d l =512 represents the number of tags and the dimension of the linguistic features, respectively. Then, the linguistic features F... l The input is fed into the language-adaptive weight generator to generate weights for the visual backbone network.

[0064] 3. Image feature extraction;

[0065] Given an image in a natural scene Where W×H×3 represents the image size, the entire image is resized to 448×448×3 and input into the feature extraction network for forward propagation. A Visual Transformer-B (ViT-B) structure is applied as the backbone visual network to acquire image features. The dimension size is C = 512, and the step size is s = 14.

[0066] 4. Refers to the prediction of the outcome;

[0067] Language features represented by [CLS] Visual features are passed to the pointer interpretation prediction module to predict the bounding box of the object being pointed to.

[0068] 5. Allocation of linguistic features;

[0069] Considering that each linguistic representation corresponds to a different number of linguistic tags, and that each layer of the visual backbone network may focus more on different linguistic tags, we attempt to independently aggregate linguistic features of a fixed size to each layer. Inspired by multi-head attention mechanisms, we introduce a learnable layer-specific embedding. It operates on each layer of the visual backbone network to dynamically extract layer-related language features. For each group g, attention is applied... Assigned to and The normalized dot product is expressed as:

[0070] 6. Aggregation of linguistic features;

[0071] Actively perceived language features It can be obtained through aggregation: Finally, a fully connected layer is used to reduce the dimensionality of the aggregated language features in the i-th layer of the visual backbone network, represented as: in Used to reduce the dimension to d h =32, where r=16 is the dimensionality reduction ratio. δ represents the eLU activation function.

[0072] 7. Generation of language-adaptive weights;

[0073] To guide the active perception of the visual backbone network, adaptive weights for language generation are generated based on the expression index, which are used to generate query X in the visual backbone network. q Key X k Sum X v , represented as X q =θ(X; W) q ),X k =θ(X; W) k ),X v =θ(X; W) v ), where θ(·;W) represents a linear mapping operation with parameter W, and X represents the input visual feature. It is a dynamic mapping weight used to generate queries, keys, and values, d in =512 and d out =512 represent the feature dimensions of feature X and query / key / value, respectively.

[0074] 8. Generation of dynamic weights;

[0075] We attempt to generate dynamic weights according to the matrix factorization paradigm. Taking the i-th ViT block as an example, this process can be represented as: in, These are layer-specific static learnable weights. and They are also statically learnable weights, but are shared across all ViT blocks to reduce the number of parameters and prevent model overfitting. It is a fully connected layer that aggregates language features As input, a dynamic matrix of shape 32×32 is generated.

[0076] 9. Refers to the prediction of the bounding box;

[0077] Unlike methods that require carefully designed cross-modal interaction modules, this method can acquire visual features related to expression through a language-aware visual backbone network without the need for additional cross-modal interaction modules. Through the proposed concise and efficient pointing expression comprehension prediction module, visual and linguistic features can be used to predict the bounding box of the pointing expression comprehension, and direct coordinate regression can be applied to predict the bounding box of the pointed object.

[0078] 10. Language adaptive aggregation;

[0079] To aggregate visual features along the spatial dimension, a language-adaptive aggregation module is proposed, which utilizes language-adaptive attention to aggregate visual features. First, the visual features... and language features Projected into a lower-dimensional space Then, attention weights are calculated using dot product similarity. Softmax normalization is then performed. Next, visual features are aggregated by weighted summation using attention weights A. Finally, the aggregated visual features are input into a fully connected layer, and the sigmoid function is used to predict the bounding box.

[0080] 11. Model loss;

[0081] This invention framework allows for end-to-end optimization for indicating expression understanding. Given a predicted bounding box... Given the ground truth bounding box b = (x, y, w, h), the detection loss function is defined as follows: in, and Let λ represent L1 loss and general IoU loss, respectively. L1 and λ giou This represents the correlation coefficient used to balance the loss.

[0082] 12. Model training;

[0083] The input image resolution was resized to 448×448. A ViT-Base (visual encoder) was used as the visual backbone, which was then adapted to higher resolution images. The visual backbone was pre-trained on the MS-COCO dataset, where overlapping images in the validation / test sets were excluded. The maximum length of the pronoun expressions was set to 40, and a six-layer BERT case-insensitive base model was used as the language backbone to generate linguistic features. λ L1 and λ giou Set to 1. λ focal and λ dice The initial learning rate was set to 4. The dimensionality reduction ratio r was set to 16. The initial learning rate for the visual and language backbones was 4e-5, and the initial learning rate for the remaining components was 4e-4. The model was optimized end-to-end using an optimizer (AdamW) for 90 epochs with a batch size of 256. The weight decay was set to 1e-4, and the learning rate was reduced by a factor of 10 after 60 epochs. Data augmentation operations included random horizontal flipping.

[0084] 13. Application of the model;

[0085] After the above training process, multiple models can be obtained. The optimal model (those with the best performance on the test set) is selected for application. For the input image and question, simply adjust the image to 448×448 pixels and normalize it, and perform word segmentation on the sentence before using it as input to the model. The parameters of the entire network model remain fixed; simply input the image and language data and propagate forward to directly obtain the prediction result. Actual results are shown in the image. Figure 2 As shown, this method enables efficient understanding of referential expressions.

Claims

1. A method for understanding index expression based on data generation optimization and autonomous weight evolution, characterized in that, Includes the following steps: Step 1: Given an image dataset D, select one image from it. Using pointers to generate template T k Generate an initial expression x ik Traverse the dataset D to obtain the initial expression dataset X1; Step 2: For a given image I i The model is given the prompt: "Please describe this image," and the model generates the corresponding description C. i ; regarding image I in X1 i The expression x ik Explicitly added to the context as context information for negative examples, while simultaneously providing the model with the hint: "Please avoid generating this type of problem," the context information known about negative examples is as follows: i C i ,x ik >;​ Step 3: Using the context information obtained from the negative examples in Step 2, repeat Step 1 to generate the optimized expression data x. i1 The process is represented as: I i2 =F θ (X i ,C i ,T k ,x i1 ) (1) Step 4: Given a pointer to x i1 An N-layer language backbone network is used to segment the expression and add a [CLS] tag at the beginning to extract language features. Where L and d l These represent the number of tags and the dimensions of language features, respectively. Then, the language features F l The input is fed into the language-adaptive weight generator to generate weights for the visual backbone network; Next, given an image Where W×H×3 represents the size of the image, expressing the perceived visual characteristics. The visual features are extracted via a visual backbone network, where C and s represent the number of channels and stride of the visual features, respectively; finally, the linguistic features, represented by [CLS] tags, are... Visual features are passed to the pointer representation understanding prediction module to predict the bounding box of the object pointed to by the pointer; Step 5: Introduce a learnable layer-specific embedding. It operates on each layer of the visual backbone network to dynamically extract layer-related language features; for each group g, attention is... Assigned to and F i g The normalized dot product is expressed as: Subsequently, actively perceived language features Through aggregation, we obtain: Finally, a fully connected layer is used to reduce the dimensionality of the aggregated language features of the i-th layer of the visual backbone network, as shown below: in Used to reduce the dimension to d h =d l / r, where r is the dimensionality reduction ratio; δ represents the eLU activation function; Step 6: Generate adaptive weights for the language based on the given expression, which are then used to generate query X in the visual backbone network. q Key X k Sum X v , is represented as: X q =θ(X;W q ),X k =θ(X;W k ),X v =θ(X;W v ) (3) Where θ(·;W) represents a linear mapping operation with parameter W, and X represents the input visual feature; It is a dynamic mapping weight used to generate queries, keys, and values, d in and d out These are feature X and the feature dimensions of the query / key / value; The dynamic weights are generated according to the matrix factorization paradigm. For the i-th ViT block, this process is represented as follows: in, These are layer-specific static learnable weights; and These are statically learnable weights; It is a fully connected layer that aggregates language features As input, generate a shape of d w ×d w The dynamic matrix; Step 7: Apply direct coordinate regression to predict the bounding box of the object being referred to; First, visual features and language features Projected into a lower-dimensional space Then, attention weights are calculated using dot product similarity. Then, softmax normalization is performed; next, visual features are aggregated by weighted summation using attention weights A; finally, the aggregated visual features are input into a fully connected layer, and the sigmoid function is used to predict the bounding box. Step 8: Given the predicted bounding box Given the ground truth bounding box b = (x, y, w, h), the detection loss function is defined as follows: in, and Let λ represent L1 loss and general IoU loss, respectively. L1 and λ giou This represents the correlation coefficient used to balance the loss.

2. The index expression method based on data generation optimization and weight autonomous evolution according to claim 1, characterized in that, The N=6, L=40, d l =512.

Citation Information

Patent Citations

  • Method, device, equipment and medium for processing visual task by using universal model

    CN115830330A

  • ViT and sliding window attention fusion-based visual pointer understanding method and system

    CN116258931A