Multi-modal knowledge retrieval method, system and device based on adaptive feature expression

By adopting adaptive feature expression and cross-modal interaction mechanisms in multimodal knowledge retrieval, combined with image segmentation and graphic encoder, the problems of low accuracy and time-consuming cross-modal search in the existing technology are solved, and more efficient and accurate multimodal knowledge retrieval is achieved.

CN120067405APending Publication Date: 2025-05-30DATA SPACE RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510177212.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing multimodal knowledge retrieval methods have shortcomings in cross-modal feature interaction and computing efficiency, resulting in low retrieval accuracy and long time.

Method used

The multimodal knowledge retrieval method based on adaptive feature expression is adopted to fuse the semantic features of images and text through cross-modal interaction and self-attention mechanism, combine image segmentation and graphic encoder, extract global and local features, and perform feature stitching and interaction.

Benefits of technology

Improve the accuracy and efficiency of cross-modal retrieval, better capture the correlation between different modalities, and enhance understanding of image content, especially when dealing with complex images and specific area queries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067405A_ABST
    Figure CN120067405A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, in particular to a multi-modal knowledge retrieval method, system and device based on adaptive feature expression. According to the multi-modal knowledge retrieval method based on adaptive feature expression provided by the invention, a new cross-modal interaction mode is provided, semantic features of query objects and semantic features of query conditions are spliced, and self-attention features are extracted; and segmenting the self-attention signs according to the dimensions of the semantic features of the queried object and the dimensions of the semantic features of the query conditions to obtain cross-modal semantic features of the queried object and the query conditions. The invention overcomes the defects of low cross-modal retrieval precision and long consumed time in the prior art, provides a multi-modal knowledge retrieval method based on adaptive feature expression, can better capture association among different modals, and greatly improves the cross-modal retrieval precision and efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a multi-modal knowledge retrieval method, system and device based on adaptive feature expression. Background Art

[0002] Multi-modal knowledge retrieval refers to retrieving data with similar semantics in another modality according to the input data of one modality. The current mainstream methods include the method based on dual encoders and the method based on cross encoders. The method based on dual encoders uses an image encoder and a text encoder to encode images and texts respectively. The method based on cross encoders uses an image-text encoder to jointly encode the spliced text and image for cross-modal feature interaction.

[0003] However, there are certain problems with the above methods. 1. Most of the image feature extraction of dual encoders and cross encoders cuts the pictures into image blocks with uniform sizes for extracting local image features. Since the image blocks are not independent and complete in content, problems occur in local image feature extraction. 2. The method based on dual encoders ignores the cross-modal feature interaction between images and texts and cannot capture rich semantic representation features. 3. Although the method based on cross encoders is superior to the dual encoder in performance metrics, it needs to encode texts and images simultaneously, consuming a great deal of computing time. Summary of the Invention

[0004] In order to overcome the defects of low cross-modal retrieval accuracy and long time consumption in the above-mentioned prior art, the present invention proposes a multi-modal knowledge retrieval method based on adaptive feature expression, which can better capture the association between different modalities and greatly improve the accuracy and efficiency of cross-modal retrieval.

[0005] A multi-modal knowledge retrieval method based on adaptive feature expression proposed by the present invention first obtains the semantic features of the object to be queried and the semantic features of the query condition in the database, and then performs cross-modal interaction on the semantic features of the object to be queried and the semantic features of the query condition to obtain the cross-modal semantic features of the object to be queried and the cross-modal semantic features of the query condition; the object to be queried is an image or text; the query condition is text or image;

[0006] The cross-modal semantic features of the query condition are successively matched with the cross-modal semantic features of the object to be queried for similarity, and the object to be queried with the highest matching degree is obtained as the query result;

[0007] The way of performing cross-modal interaction on the semantic features of the object to be queried and the semantic features of the query condition is: splicing the two and extracting self-attention features, and dividing the self-attention features according to the dimensions of the semantic features of the object to be queried and the dimensions of the semantic features of the query condition to obtain the cross-modal semantic features of the object to be queried and the query condition.

[0008] Preferably, the extraction method of the semantic features of the image is as follows: First, the original image is segmented according to semantics, and the semantic features of the original image and the segmented image patches are extracted respectively, and then dimensional interaction is performed.

[0009] Preferably, after the original image is segmented, the global image semantic features of the original image and the local image semantic features of the segmented image patches are extracted; the global image semantic features and the local image semantic features are dimensionally concatenated and then self-attention features are extracted, and then the self-attention features are mapped to a specified feature dimension to obtain the image semantic features of the original image.

[0010] Preferably, the original image is combined with a prompt template and segmented by a large model, and the large model uses Qwen-VL, MiniGPT-v2 or LLaVA-1.5.

[0011] Preferably, the semantic features of the original image and the segmented image patches are obtained through a text-image encoder.

[0012] Preferably, the text is input into the text-image encoder to extract the text semantic features.

[0013] Preferably, the text-image encoder uses a variational autoencoder, a ResNet model, a Transformer model or a CLIP model.

[0014] A multi-modal knowledge retrieval device based on adaptive feature expression proposed by the present invention includes:

[0015] A picture database for storing the original images annotated with image semantic features;

[0016] A receiving module for receiving a query condition in text form and extracting the text semantic features of the query condition;

[0017] A similarity matching module is respectively connected to the receiving module and the picture database; the similarity matching module adopts the multi-modal knowledge retrieval method based on adaptive feature expression to find the image semantic feature with the highest similarity to the text semantic feature as the query target;

[0018] An output module is connected to the similarity matching module for outputting the query target.

[0019] A multi-modal knowledge retrieval system based on adaptive feature expression proposed by the present invention includes a memory and a processor. A computer program is stored in the memory, and the processor is connected to the memory. The processor is used to execute the computer program to implement the multi-modal knowledge retrieval method based on adaptive feature expression.

[0020] A storage medium provided by the present invention stores a computer program, which is used to implement the multi-modal knowledge retrieval method based on adaptive feature representation when executed.

[0021] The advantages of the present invention are as follows:

[0022] (1) Cross-modal interaction and self-attention mechanism are used to fuse semantic features of different modalities, enhancing the adaptability of feature representation, enabling dynamic association of image and text features in the shared semantic space, and better capturing the associations between different modalities. The divide-and-conquer strategy of segmentation after feature concatenation effectively preserves the original modality feature dimension information, avoids feature confusion, and improves cross-modal alignment accuracy. Then, the most relevant object is found through similarity matching, thus improving the accuracy of retrieval.

[0023] (2) For image processing, image regions are segmented, global and local features are extracted, and dimension interaction is performed to achieve dual-channel extraction of global and local features, realizing multi-level expression of image semantics. In this way, the details of the image can be captured more carefully, enhancing the understanding of the image content. Especially when the query text involves a specific region, the recognition of local features is more important. For example, when a user searches for "red car", the local features can help identify the location and color of the car.

[0024] (3) Using a large model for image segmentation improves the accuracy and efficiency of segmentation, especially when dealing with complex images. Using a pre-trained large model may reduce the training time and computational resource requirements. Combining the segmentation strategy of the object detection model can make the local features focus on the key semantic regions. For example, in the retrieval of product images, specific product components can be accurately located, greatly improving the retrieval efficiency.

[0025] (4) Applying a text-image encoder for semantic extraction performs excellently in text-image matching, can extract higher-quality semantic features, improve the effect of cross-modal retrieval, and generate more robust feature representations.

[0026] (5) Through the innovative cross-modal interaction mechanism and adaptive feature representation strategy, the present invention has achieved triple breakthroughs in accuracy, efficiency, and applicability in the field of multi-modal knowledge retrieval, and has important theoretical value and industrial application prospects.

[0027] (6) The device and system proposed by the present invention can be deployed on communication devices such as computers and servers, which is conducive to the popularization and application of this method and improves the processing speed of retrieval tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 It is a flowchart of a multi-modal knowledge retrieval method based on adaptive feature representation.

[0029] Figure 2It is a flowchart of another multi-modal knowledge retrieval method based on adaptive feature expression. Specific implementation manner

[0030] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0031] A multi-modal knowledge retrieval method based on adaptive feature expression proposed in this implementation manner includes the following steps:

[0032] S1. Regionally segment the original image based on semantics to obtain a plurality of segmented image blocks.

[0033] In specific implementation, a pre-set prompt template can be used in combination with a large model to regionally segment the original image, that is, the prompt template and the original image are input into the large model, and the large model outputs the segmented image blocks.

[0034] Specifically, the large model can be Qwen-VL, MiniGPT-v2, LLaVA-1.5, etc.

[0035] The hint template is set as Template = "Please perform semantic segmentation on the following image and meet the following requirements: 1. **Segmentation target**: Segment the image into multiple regions, and the text and images (if any) within each region should form a semantically independent unit. For example, a region may contain a title and related paragraphs, or a table and its description, or an object. 2. **Output format**: - **Bounding box information**: For each segmented region, provide the coordinate information of its bounding box. The coordinate information should contain four values (x_min, y_min, x_max, y_max), representing the horizontal and vertical coordinates of the upper left and lower right corners of the region respectively. - **Region content**: Although it is not necessary to directly extract the text content, please briefly describe the main information type contained in each region (such as title, paragraph, image, table, etc.) in the description. 3. **Processing details**: - Please ensure that each segmented region is semantically independent, that is, the information within each region can form a complete and independent semantic unit. - Considering that the image may contain text and images of different sizes and positions, flexibly adjust the segmentation strategy to adapt to these changes. - When segmenting, try to keep the integrity of each region and avoid splitting the same semantic unit into multiple regions. 4. **Example**: - Suppose the image contains a title 'Introduction to Baidu Encyclopedia', followed by a paragraph about Baidu Encyclopedia, and a table about the historical development of Baidu Encyclopedia. In this case, you can segment the title and the paragraph into one region and the table into another region. - For each region, you may output bounding box information similar to the following: 'Region 1: (10, 20, 300, 150), containing the title and the paragraph', 'Region 2: (310, 20, 600, 300), containing the table'. Now, please segment the following image and output the bounding box information and brief description of each region. {image i}”

[0036] Perform semantic segmentation on the image image i based on the large model to obtain segmented image patches containing multiple locally independent semantics

[0037] where, image i represents the original image, respectively represent the 1st, 2nd, kth, and mth segmented image patches; m represents the number of segmented image patches; 1 ≤ k ≤ m.

[0038] S2. Extract semantic features from the original image and the segmented image patches to obtain global image semantic features and local image semantic features respectively.

[0039] Semantic feature extraction can specifically adopt image - text encoders such as variational auto - encoder (VAE), ResNet model, Transformer model, CLIP model, etc. Taking the Transformer model as an example, it is expressed by the formula:

[0040]

[0041] Among them, H i represents the set of global image semantic features and local image semantic features;

[0042] Let image i ∈R 1×d represent the global image semantic features after encoding the original image image i , d represents the semantic feature dimension, represents the local image semantic features after encoding the segmented image patch ;

[0043] Then

[0044] S3. Concatenate the global image semantic features and local image semantic features according to the feature dimension, and use the self - attention mechanism to perform interaction between the global image semantic features and local image semantic features in the feature dimension to obtain the self - attention features imagech i of the original image;

[0045]

[0046] Concatenate(·) represents dimension concatenation, Attention(·) represents self - attention operation; imagech i ∈R (d*(1+m))×1 represents the output after the interaction between the global image semantic features and local image semantic features using the self - attention mechanism.

[0047] In this step, the interaction between the global image semantic features and local image semantic features in the feature dimension is performed to obtain image semantic features that contain both global information and local details; and further use the self - attention mechanism to perform cross - modal interaction between the image semantic features and text semantic features in the feature dimension, so as to obtain richer semantic features on the premise of consuming less computing time.

[0048] S4. Map the feature dimension of the self - attention features imagech i of the original image to d to obtain the final image semantic features imagefh i ∈R 1×d ;

[0049] Specifically, the projection layer Projection can be used to map the feature dimension to d, and the formula is expressed as:

[0050] imagefh i = Projection(imagech i )

[0051] S5. Extract the semantic features of the text q j to obtain the text semantic feature qh j ∈R 1×r , where r represents the feature dimension;

[0052] Specifically, a text-image encoder can be used for text semantic feature extraction. For example:

[0053] qh j = Transformer(q j )

[0054] S6. Perform cross-modal interaction on the text semantic feature qh j and the image semantic feature imagefh i . First, concatenate the two dimensions and then perform self-attention operation, and then perform dimension splitting to obtain the image semantic feature imagefh' after cross-modal interaction i and the text semantic feature qh' after cross-modal interaction j .

[0055] Specifically, the formula is expressed as:

[0056] [imagefh', qh'] = Attention(Concatenate(imagefh i , qh j )) i , qh j ))

[0057] where Concatenate(·) represents the concatenation operation of feature dimensions, Attention(·) represents the self-attention operation, [imagefh', qh'] ∈ R i , qh' j ∈R 1×(d+r) , imagefh' i ∈R 1×d , qh' j ∈R 1×r .

[0058] S7. Then, for the image semantic feature imagefh' after cross-modal interaction i and the text semantic feature qh' after cross-modal interaction jPerform matching to obtain the matching text and the original image, thereby completing cross-modal knowledge query.

[0059] Using this method, if the query condition is text q j , then the image semantic features imagefh of each original image in the database are obtained through steps S1 - S4 i , and then steps S6 - S7 are used to match the text with each original image one by one based on the text semantic feature qh j , so as to obtain the most matching original image.

[0060] If the query condition is the original image, then the image semantic features imagefh of the original image are obtained through steps S1 - S4 i , and the text semantic features qh of each text in the database are obtained through step S5 j , and then steps S6 - S7 are used to match the original image with each text one by one, so as to obtain the most matching text.

[0061] A multi-modal knowledge retrieval device based on adaptive feature expression proposed by the present invention includes a picture database, a receiving module, a similarity matching module, and an output module.

[0062] The picture database is used to store the original images marked with image semantic features;

[0063] The receiving module is used to receive the query condition in text form and extract the text semantic features of the query condition;

[0064] The similarity matching module is respectively connected to the receiving module and the picture database; the similarity matching module uses the multi-modal knowledge retrieval method based on adaptive feature expression to find the image semantic feature with the highest similarity to the text semantic feature as the query target;

[0065] The output module is connected to the similarity matching module and is used to output the query target.

[0066] This device can be applied to mall management. Through the mall picture database, pictures of specified commodities are searched according to the text input, so as to realize the rapid positioning of commodity locations;

[0067] It can also be used to support natural language-driven multimedia content retrieval. Through video frame decomposition and the method of the present invention, rapid query of multimedia content according to text input is realized, improving the efficiency of information processing;

[0068] It can also be used in an intelligent security system. The surveillance video is decomposed into video frames and stored in the picture database, and then the method of the present invention is used to search for surveillance content through text input, such as "a 5-year-old little girl with double ponytails wearing a yellow top".

[0069] Of course, for those skilled in the art, the present invention is not limited to the details of the above-described exemplary embodiments, but also includes the same or similar structures that can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.

[0070] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative manner of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

[0071] The technologies, shapes, and structures not described in detail in the present invention are all well-known technologies.

Claims

1. A multimodal knowledge retrieval method based on adaptive feature expression, characterized in that: First, the semantic features of the queried object and the semantic features of the query conditions in the database are obtained, and then the semantic features of the queried object and the semantic features of the query conditions are cross-modally interacted to obtain the cross-modal semantic features of the queried object and the cross-modal semantic features of the query conditions; The queried object is an image or text; The query condition is text or image; The cross-modal semantic features of the query conditions are matched with the cross-modal semantic features of the queried objects one by one in terms of similarity, and the queried object with the highest matching degree is obtained as the query result; The semantic features of the queried object and the semantic features of the query conditions interact cross-modally by concatenating the two and extracting self-attention features, and segmenting the self-attention features according to the dimensions of the semantic features of the queried object and the dimensions of the semantic features of the query conditions to obtain cross-modal semantic features of the queried object and the query conditions.

2. The multimodal knowledge retrieval method based on adaptive feature expression according to claim 1, characterized in that: The method for extracting the semantic features of an image is as follows: first, the original image is segmented according to semantics, the semantic features of the original image and the semantic features of the segmented image blocks are extracted respectively, and then dimensional interaction is performed.

3. The multimodal knowledge retrieval method based on adaptive feature expression as claimed in claim 2, characterized in that: After performing regional segmentation on the original image, the global image semantic features of the original image and the local image semantic features of the segmented image blocks are extracted; the global image semantic features and the local image semantic features are dimensionally spliced ​​and then the self-attention features are extracted. The self-attention features are then mapped to the specified feature dimensions to obtain the image semantic features of the original image.

4. The multimodal knowledge retrieval method based on adaptive feature expression as claimed in claim 2, characterized in that: The original image is combined with the prompt template for segmentation through a large model, which uses Qwen-VL, MiniGPT-v2 or LLaVA-1.

5.

5. The multimodal knowledge retrieval method based on adaptive feature expression as claimed in claim 2, characterized in that: The semantic features of the original image and segmented image blocks are obtained through the image-text encoder.

6. The multimodal knowledge retrieval method based on adaptive feature expression according to any one of claims 1 to 5, characterized in that: The text is input into the graph-text encoder to extract text semantic features.

7. The multimodal knowledge retrieval method based on adaptive feature expression according to claim 6, characterized in that: The image and text encoder uses a variational autoencoder, ResNet model, Transformer model, or CLIP model.

8. A multimodal knowledge retrieval device based on adaptive feature expression, characterized in that: include: Image database, used to store original images annotated with image semantic features; A receiving module, used for receiving a query condition in text form and extracting text semantic features of the query condition; A similarity matching module is connected to the receiving module and the image database respectively; the similarity matching module uses the multimodal knowledge retrieval method based on adaptive feature expression as claimed in claim 1 to find the image semantic feature with the highest similarity to the text semantic feature as the query target; The output module is connected to the similarity matching module and is used to output the query target.

9. A multimodal knowledge retrieval system based on adaptive feature expression, characterized in that: It includes a memory and a processor, the memory stores a computer program, the processor is connected to the memory, and the processor is used to execute the computer program to implement the multimodal knowledge retrieval method based on adaptive feature expression as described in any one of claims 1 to 5.

10. A storage medium, characterized in that: A computer program is stored, and when the computer program is executed, it is used to implement the multimodal knowledge retrieval method based on adaptive feature expression as described in any one of claims 1 to 5.