Visual feature and text feature fusion location encoding method, system and device
By combining image and text features through convolutional neural networks and image domain location encoding methods, the computational complexity and local feature extraction efficiency of large image-text modal models in high-resolution image processing are solved, thereby improving the model's visual understanding and generation efficiency.
Patent Information
- Application Number
- CN202411353187.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-26
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2044-09-26
AI Technical Summary
Existing large image-text modal models have high computational complexity and low efficiency in local feature extraction when processing high-resolution images, which limits their performance on fine vision tasks.
A convolutional neural network is used to obtain multi-scale features of the image, and a fused positional encoding vector is generated by image domain positional encoding and text embedding mapping. The image features and text features are then combined for unified encoding.
It improves the model's visual understanding and generation efficiency, enhances its perception of global information, and improves the accuracy of image information position in text sequences.
Smart Images

Figure CN119379787B_ABST
Abstract
Description
Technical Field
[0001] This document relates to the field of image and text feature fusion technology, and in particular to a location coding method, system and device for fusing visual features and text features. Background Technology
[0002] With the rapid development of artificial intelligence technology, especially the widespread application of deep learning in multimedia processing, multi-modal image-text large models have become an important bridge connecting vision and language, greatly promoting the development of cross-modal understanding and generation technologies. These models can parse visual information in images and generate related natural language descriptions, or generate corresponding images based on text content, demonstrating enormous application potential.
[0003] However, current large image-text modal models primarily rely on the Visual Transformer architecture for processing visual information. While this architecture demonstrates advantages in capturing global dependencies and modeling complex visual patterns, its computational complexity is high and its processing efficiency is relatively low, especially when processing high-resolution images. Furthermore, the Transformer architecture still falls short compared to traditional convolutional neural networks in handling local details and texture features, which limits the model's performance on fine-grained visual tasks.
[0004] Therefore, effectively integrating convolutional features to compensate for the shortcomings of visual Transformers in efficiency and local feature extraction has become a key bottleneck for the further development of large-scale image-text modal models. Convolutional neural networks, through their hierarchical convolutional kernel design, can efficiently extract multi-scale local features from images, which are crucial for understanding and describing details within images. Therefore, exploring a method to effectively integrate convolutional features into the generation process of large-scale image-text modal models can not only improve the model's visual understanding ability but also significantly enhance its generation efficiency and effectiveness, possessing significant research significance and application value. Summary of the Invention
[0005] This invention provides a location encoding method, system, and apparatus that fuses visual features and text features, aiming to solve the above-mentioned problems.
[0006] This invention provides a location encoding method that fuses visual features and text features, including:
[0007] S1. Obtain the image and text sequences from the data to be processed;
[0008] S2. Use a convolutional neural network to obtain multi-scale features of the image, and perform multi-scale feature fusion operation on the multi-scale features to obtain a compressed feature map;
[0009] S3. Image domain position coding is used to perform image domain position coding on the compressed feature map to obtain image feature embedding;
[0010] S4. Perform text embedding mapping on the text sequence to obtain text feature embeddings;
[0011] S5. The image feature embedding and the text feature embedding are uniformly implemented by positional encoding to generate a fused positional encoding vector.
[0012] This invention provides a location coding system that fuses visual features and text features, comprising:
[0013] The data acquisition module is used to acquire image and text sequences from the data to be processed;
[0014] The multi-scale feature fusion module is used to obtain multi-scale features of the image by using a convolutional neural network, and to perform multi-scale feature fusion operation on the multi-scale features to obtain a compressed feature map.
[0015] The image location encoding module uses image domain location encoding to perform image domain location encoding on the compressed feature map to obtain image feature embedding;
[0016] The text mapping module is used to perform text embedding mapping on the text sequence to obtain text feature embeddings;
[0017] The image-text fusion module is used to uniformly implement image feature encoding and text feature encoding by using positional encoding for the image feature embedding and the text feature embedding, and generate a fused positional encoding vector.
[0018] This specification provides an electronic device comprising: one or more embodiments thereof.
[0019] Processor; and,
[0020] A memory is configured to store computer-executable instructions, which, when executed, cause the processor to perform the steps of the location encoding method described above, which fuses scaled visual features and text features.
[0021] This specification provides one or more embodiments of a storage medium for storing computer-executable instructions that, when executed, implement the steps of the location encoding method described above, which fuses scaled visual features and text features.
[0022] By employing embodiments of the present invention, convolutional networks are used to improve the efficiency of image feature extraction. Simultaneously, image features are encoded in the image domain to ensure that the features are contained in the image, thereby enhancing the perception of global information in subsequent generative large models. After unified encoding through the fusion of image and text features, the position of image features in the text sequence is determined, thus forming a two-level encoding form and improving the accuracy of image information location. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in one or more embodiments of this specification or in the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 This is a flowchart of a location encoding method for fusing scaled visual features and text features according to an embodiment of the present invention;
[0025] Figure 2 This is a schematic diagram illustrating a specific implementation of the location encoding method for fusing scaled visual features and text features according to an embodiment of the present invention;
[0026] Figure 3 This is a schematic diagram of multi-scale convolutional feature fusion in this specific embodiment;
[0027] Figure 4 This is a schematic diagram of image domain position encoding according to an embodiment of the present invention;
[0028] Figure 5 This is a schematic diagram of a location coding system that fuses scaled visual features and text features according to an embodiment of the present invention. Detailed Implementation
[0029] To enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the technical solutions in one or more embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of this document.
[0030] Method Implementation Examples
[0031] According to embodiments of the present invention, a location encoding method for fusing scaled visual features and text features is provided. Figure 1This is a flowchart of a location encoding method for fusing scaled visual features and text features according to an embodiment of the present invention. Figure 1 As shown, the location encoding method for fusing scaled visual features and text features according to an embodiment of the present invention specifically includes:
[0032] S1. Obtain the image and text sequences from the data to be processed;
[0033] S2. Use a convolutional neural network to obtain multi-scale features of the image, and perform multi-scale feature fusion operation on the multi-scale features to obtain a compressed feature map;
[0034] Figure 3 This is a schematic diagram illustrating a specific implementation of the location encoding method for fusing scaled visual features and text features according to an embodiment of the present invention. Figure 2 As can be seen, the multi-scale feature fusion operation performed on the multi-scale features in this embodiment of the invention specifically includes:
[0035] The system consists of three parts: horizontal semi-global feature fusion, vertical semi-global feature fusion, and bidirectional global feature fusion. Horizontal semi-global feature fusion compresses features into [1, height, channel]; vertical semi-global feature fusion compresses features into [width, 1, channel]; and bidirectional global feature fusion fuses and compresses the horizontal and vertical features into a vector of dimension [1, 1, channel]. Compression methods can include unidirectional average pooling and unidirectional convolution.
[0036] The convolutional features are the results of the intermediate layers of the convolutional network. The specific number of layers can be selected, and the dimensions are [widht, height, channel]. After feature fusion, the feature dimensions are [widht+1, height+1, channel].
[0037] S3. Image domain position coding is used to perform image domain position coding on the compressed feature map to obtain image feature embedding;
[0038] Figure 4 This is a schematic diagram of image domain position encoding according to an embodiment of the present invention. Figure 4 It can be seen that image domain position coding includes two dimensions: horizontal coding and vertical coding, where:
[0039] The formula for calculating the horizontal encoding is: sin(x / (2^i));
[0040] The formula for calculating the vertical encoding is: cos(x / (2^j));
[0041] Where x is a trainable parameter, i is the horizontal position of the feature in the image, and j is the vertical position.
[0042] The encoding dimension is [widht+1, height+1, 2]. After concatenation, the feature is generated as [widht+1, height+1, channel+2]. The feature dimension of channel+2 needs to be consistent with the text feature (commonly 768, so the channel dimension can be 766, which can be varied). After stretching, a feature vector with dimension [(widht+1)*(height+1), channel+2] can be obtained.
[0043] S4. Perform text embedding mapping on the text sequence to obtain text feature embeddings;
[0044] For the input text, a vector of length [seq_length, emb_dim] can be obtained using a common word embedding model. Here, emb_dim is usually 786, and seq_length is related to the length of the text input.
[0045] S5. The image feature embedding and the text feature embedding are uniformly implemented by positional encoding to generate a fused positional encoding vector.
[0046] By concatenating a text vector of dimension [seq_length, emb_dim] with an image feature vector of dimension [(widht+1)*(height+1), channel+2], as explained above, channel+2 must be the same as emb_dim, resulting in a vector of dimension [seq_length+(widht+1)*(height+1), emb_dim].
[0047] Positional encoding is used to uniformly encode image and text features, generating a fused positional encoding vector. This vector can then be processed according to a generative text model.
[0048] Figure 2 This is a schematic diagram illustrating a specific implementation of the location encoding method for fusing scaled visual features and text features according to an embodiment of the present invention. Image domain location encoding is the core of this patent. This method generates feature vectors with image domain encoding based on convolutional features, and can use convolutional features in conventional generative large model location encoding methods, such as rotation location encoding.
[0049] Image multi-scale convolutional feature fusion, oriented towards the characteristics of multi-layer positional coding features, utilizes local image information to generate semi-global and global information, and combines it with image domain positional coding to improve the perception of global information in subsequent generative large models.
[0050] Image-text feature fusion coding mainly describes how generated image features are fused with regular text embedding features.
[0051] The embodiments of the present invention have the following beneficial effects:
[0052] By employing embodiments of the present invention, convolutional networks are used to improve the efficiency of image feature extraction. Simultaneously, image features are encoded in the image domain to ensure that the features are contained in the image, thereby enhancing the perception of global information in subsequent generative large models. After unified encoding through the fusion of image and text features, the position of image features in the text sequence is determined, thus forming a two-level encoding form and improving the accuracy of image information location.
[0053] System Implementation Examples
[0054] According to embodiments of the present invention, a location coding system that fuses scaled visual features and textual features is provided. Figure 5 This is a schematic diagram of a location coding system that fuses scaled visual features and text features according to an embodiment of the present invention. Figure 5 As shown, embodiment 5 of the present invention is a location coding system for the fusion of scaled visual features and text features, specifically including:
[0055] Data acquisition module 50 is used to acquire image and text sequences from the data to be processed;
[0056] The image multi-scale convolutional feature fusion module 52 is used to obtain multi-scale features of the image by using a convolutional neural network, and to perform multi-scale feature fusion operation on the multi-scale features to obtain a compressed feature map.
[0057] Image domain position encoding module 54 performs image domain position encoding on the compressed feature map to obtain image feature embedding;
[0058] Text mapping module 56 is used to perform text embedding mapping on the text sequence to obtain text feature embeddings;
[0059] The image and text feature fusion encoding module 58 is used to uniformly implement image feature encoding and text feature encoding by using position encoding for the image feature embedding and the text feature embedding, and generate a fused position encoding vector.
[0060] The image multi-scale convolutional feature fusion module 52 performs multi-scale feature fusion operations on the multi-scale features, specifically including:
[0061] The multi-scale features of the image are subjected to horizontal semi-global feature fusion, vertical semi-global feature fusion, and bidirectional global feature fusion, respectively.
[0062] The horizontal semi-global feature fusion compresses the image into a vector of dimension [1, height, channel], the vertical semi-global feature fusion compresses the image into a vector of dimension [width, 1, channel], and the bidirectional global feature fusion compresses the horizontal and vertical semi-global feature fusions into a vector of dimension [1, 1, channel], where height represents the height of the image, width represents the width of the image, and channel represents the number of channels of the image features.
[0063] The image domain position encoding module 54 is specifically used for:
[0064] The compressed feature map is encoded laterally and vertically, wherein:
[0065] Horizontal encoding is obtained using the following formula:
[0066] sin(x / (2^i)) (1);
[0067] Vertical encoding is obtained using the following formula:
[0068] cos(x / (2^j)(2);
[0069] Where x is a trainable parameter, i is the horizontal position of the feature in the image, and j is the vertical position.
[0070] Device Example 1
[0071] An electronic device is provided according to an embodiment of the present invention, comprising:
[0072] Processor; and,
[0073] A memory is configured to store computer-executable instructions, which, when executed, cause the processor to perform the steps described in the above method embodiments.
[0074] Device Example 2
[0075] According to an embodiment of the present invention, a storage medium is provided for storing computer-executable instructions, which, when executed, perform the steps described in the above method embodiments.
[0076] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A location encoding method that fuses visual features and textual features, characterized in that... include: S1. Obtain the image and text sequences from the data to be processed; S2. Use a convolutional neural network to obtain multi-scale features of the image, and perform multi-scale feature fusion operation on the multi-scale features to obtain a compressed feature map; S3. Image domain position coding is used to perform image domain position coding on the compressed feature map to obtain image feature embedding; S4. Perform text embedding mapping on the text sequence to obtain text feature embeddings; S5. The image feature embedding and the text feature embedding are uniformly implemented by positional encoding to generate a fused positional encoding vector; The step of performing multi-scale feature fusion on the multi-scale features to obtain a compressed feature map specifically includes: The multi-scale features of the image are respectively subjected to horizontal semi-global feature fusion, vertical semi-global feature fusion, and bidirectional global feature fusion. The horizontal semi-global feature fusion compresses the image into a vector of dimension [1, height, channel], the vertical semi-global feature fusion compresses the image into a vector of dimension [width, 1, channel], and the bidirectional global feature fusion compresses the horizontal and vertical semi-global feature fusions into a vector of dimension [1, 1, channel], where height represents the height of the image, width represents the width of the image, and channel represents the number of channels of the image features.
2. The method according to claim 1, characterized in that, The image domain position encoding specifically includes two dimensions: horizontal encoding and vertical encoding, wherein: Horizontal encoding is obtained using the following formula: (1); Vertical encoding is obtained using the following formula: (2); in, x For trainable parameters, i The feature is located at its horizontal position in the image. j This refers to the vertical position.
3. A location coding system that fuses visual features and textual features, characterized in that, include: The data acquisition module is used to acquire image and text sequences from the data to be processed; The image multi-scale convolutional feature fusion module is used to obtain multi-scale features of the image using a convolutional neural network, and to perform multi-scale feature fusion operation on the multi-scale features to obtain a compressed feature map. The image domain position encoding module performs image domain position encoding on the compressed feature map to obtain image feature embedding. The text mapping module is used to perform text embedding mapping on the text sequence to obtain text feature embeddings; The image-text feature fusion encoding module is used to uniformly implement image feature encoding and text feature encoding by using positional encoding to generate a fused positional encoding vector; The image multi-scale convolutional feature fusion module performs multi-scale feature fusion operations on the multi-scale features, specifically including: The multi-scale features of the image are respectively subjected to horizontal semi-global feature fusion, vertical semi-global feature fusion, and bidirectional global feature fusion. The horizontal semi-global feature fusion compresses the image into a vector of dimension [1, height, channel], the vertical semi-global feature fusion compresses the image into a vector of dimension [width, 1, channel], and the bidirectional global feature fusion compresses the horizontal and vertical semi-global feature fusions into a vector of dimension [1, 1, channel], where height represents the height of the image, width represents the width of the image, and channel represents the number of channels of the image features.
4. The system according to claim 3, characterized in that, The image domain position encoding module is specifically used for: The compressed feature map is encoded laterally and vertically, wherein: Horizontal encoding is obtained using the following formula: (1); Vertical encoding is obtained using the following formula: (2); in, x For trainable parameters, i The feature is located at its horizontal position in the image. j This refers to the vertical position.
5. An electronic device, comprising: processor; as well as, A memory configured to store computer-executable instructions, which, when executed, cause the processor to perform the steps of the location encoding method for fusing visual and textual features as described in any one of claims 1-2.
6. A storage medium for storing computer-executable instructions, which, when executed, implement the steps of the location encoding method for fusing visual features and text features as described in any one of claims 1-2.
Citation Information
Patent Citations
Multi-modal data fusion method based on image-text interaction
CN115659279A
Financial violation detection method and device based on multi-modal fusion, equipment and medium
CN117421631A