Knowledge retrieval system

By using multimodal large models and CLIP model-based image-text semantic alignment technology, the problems of structural information loss and insufficient cross-modal retrieval capabilities in traditional RAG technology are solved, achieving efficient and accurate knowledge retrieval and meeting users' needs for comprehensive and timely knowledge acquisition.

CN120873247APending Publication Date: 2025-10-31INSPUR ENTERPRISE CLOUD TECHNOLOGY (SHANDONG) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510906892.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Traditional RAG technology cannot preserve structural information during text retrieval, and the semantic integrity and information granularity of individual segments cannot be configured. In cross-modal retrieval, it lacks strong semantic differentiation and cross-modal semantic representation capabilities, making it difficult to achieve efficient and accurate knowledge retrieval.

Method used

Employing the multimodal understanding capabilities of a multimodal large model, graphs are interpreted as structured and semantically rich text descriptions. The text is represented as a highly discriminative embedding form using LLM2VEC technology, and the CLIP model is used for graph-text semantic alignment. Redis is then combined for efficient vector database management and similarity calculation.

Benefits of technology

It significantly improves semantic discrimination accuracy, enables precise cross-modal knowledge retrieval, and enhances the comprehensiveness and efficiency of knowledge retrieval, providing users with relevant knowledge quickly and comprehensively from massive multimodal data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120873247A_ABST
    Figure CN120873247A_ABST
Patent Text Reader

Abstract

The invention provides a knowledge retrieval system, which belongs to the technical field of AI large models and information retrieval, and comprises a multi-modal data analysis module, a cross-modal representation learning module, a knowledge storage module, a query processing module and a knowledge recall module. The method comprises the steps of knowledge material modal separation and recombination, cross-modal representation learning, knowledge storage, query processing and knowledge recall. Semantic association of different modal data is achieved through ASR, LLM2VEC, LLM2CLIP and a multi-modal large language model technology. The method can effectively solve the problem that multi-modal data association information cannot be efficiently utilized in the prior art, the comprehensiveness, accuracy and efficiency of knowledge recall are remarkably improved, and the method can be widely applied to various scenes needing comprehensive multi-modal knowledge retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of AI large-scale models and information retrieval technology, and in particular to a knowledge retrieval system. Background Technology

[0002] In today's information-saturated era, knowledge emerges in massive quantities in multiple modalities, such as language, graphs, audio, and video. The effective analytical components of this modal information flowing into a system can be broken down into text, graphs, and structures. While graphs can describe and convey certain aspects, they ultimately need to be represented in the form of text structures. Therefore, a complete knowledge representation system, with its logically accurate expression, ultimately relies on language. Language conveys what can be conveyed, and what cannot be accurately conveyed by language cannot be accurately represented. Traditional RAG (Research and Retrieval of Information) technology faces numerous challenges due to the loss of information elements during processing. First, in text retrieval, structural information cannot be preserved after knowledge is segmented. Second, the semantic integrity and information granularity of individual fragments cannot be configured after knowledge segmentation. Third, when recalling results, the top-K strategy of recommendation systems is adopted, but this inherently contradicts the goals of knowledge retrieval. Fourth, in cross-modal retrieval, existing methods cannot effectively utilize the inherent connections between different modalities. For example, if a user wants to find textual records and corresponding images related to a historical event, traditional systems struggle to accurately and comprehensively retrieve relevant knowledge from both text and image databases simultaneously. This is mainly due to the fact that traditional models lack strong semantic differentiation and cross-modal semantic representation capabilities, making it difficult to achieve efficient and accurate knowledge retrieval in massive multimodal data. Summary of the Invention

[0003] To address the aforementioned technical issues, this invention provides a knowledge retrieval system that utilizes the multimodal understanding capabilities of a multimodal large model to interpret graphs into structured and semantically rich text descriptions, forming original graph-text pairs. Then, LLM2VEC technology is applied to represent the text as a highly discriminative embedding form. Finally, through training the CLIP model, graph-text semantic alignment is performed.

[0004] The technical solution of this invention is:

[0005] A knowledge retrieval system, comprising:

[0006] Multimodal data acquisition, separation, and preprocessing module: responsible for collecting text and image data from various data sources; cleaning text data to remove low-quality semantic data; and extracting and analyzing image data, as well as deduplicating images.

[0007] Training data generation module: LLM2VEC dataset preparation: using the separated images, and after understanding and analyzing each image, a set of titles is generated, thus forming a corpus; LLM2CLIP dataset preparation: still using the separated images, and after understanding and analyzing each image, a summary description and detailed information are generated, thus forming a corpus;

[0008] CLIP Text Encoder Module: Introduces a large language model and fine-tunes it to replace the text representation in traditional CLIP as a semantic representation model;

[0009] CLIP image encoder module: Employs LLM2CLIP technology to achieve cross-modal semantic alignment; fine-tuning corpora are derived from the dataset generation module and open datasets;

[0010] Knowledge storage module: Applying the trained CLIP model, the deduplicated image data to be retrieved in the system is used as knowledge pieces, and stored in the vector database along with relevant metadata. The specific vector database selection is configurable. At the same time, Redis is used as an auxiliary cache and data management tool to store metadata and the index information of the vector database in Redis. Redis's high-speed read and write characteristics are used to improve the overall system's data access efficiency.

[0011] The query processing and knowledge retrieval module receives user-input queries, parses the query semantics (the query can be text, image, or speech modality), preprocesses and parses the queries, and applies the CLIP model for vectorization representation to retrieve data from the vector database. In the knowledge storage module's vector database, it retrieves multimodal knowledge related to the query by calculating the similarity between the query vector and stored vectors. The results are then filtered according to a similarity threshold, sorted in descending order, and returned to the user.

[0012] Furthermore,

[0013] The data in the multimodal data acquisition, separation, and preprocessing module comes from user-uploaded knowledge material files, which belong to private domain knowledge data. The knowledge materials are then standardized and transcribed at the file level using knowledge material modality separation and recombination methods.

[0014] After transcribing, the file metadata information is preserved, and the various modal information in the file content is extracted in an orderly manner and assembled into a unified format according to the original order; at the same time, the mapping relationship is stored in the database.

[0015] Each file can correspond to several records. Each record should include a unique file identifier, the modality type of the data, a unique header identifier, a unique mapping identifier for the modality semantic content of the data, and time information.

[0016] Furthermore,

[0017] The fine-tuning corpus of the CLIP text encoder module comes from the dataset generation module and the open dataset. Following the LLM2VEC training method, the output features of LLM are fine-tuned by applying title contrast, treating different titles of the same image as positive samples and the rest as negative samples.

[0018] Furthermore,

[0019] The CLIP image encoder module employs the LLM2CLIP technique to achieve cross-modal semantic alignment. The fine-tuning corpus is derived from the dataset generation module and an open dataset. During training, the gradients of the LLM are frozen to preserve its inherent capabilities, and to compensate for the frozen LLM, several new linear layers are introduced as adapters after the LLM. These layers serve as learnable parameters.

[0020] Following the original design, CLIP also uses a projection layer to align the dimensions of the two encoders, making it easier to train using CLIP loss.

[0021] Furthermore,

[0022] The voice modality of the query processing and knowledge retrieval module is converted into text modality using the ASR model.

[0023] The query processing and knowledge retrieval module is configured to return a maximum number of results by default. If the maximum number of results is exceeded, the user will be prompted to load more results.

[0024] The beneficial effects of this invention are

[0025] Significantly improves semantic discrimination accuracy: LLM2VEC leverages the extensive world knowledge and powerful semantic understanding capabilities of large language models to accurately distinguish text semantics and demonstrates significant discriminative power in semantic relevance. The relative stability of the relevance or similarity threshold provides a strong basis for recall.

[0026] Achieve accurate cross-modal knowledge retrieval: Through innovative LLM2VEC and LLM2CLIP technologies, efficient alignment between data features of different modalities is achieved, and the limitations of simple semantic alignment in traditional cross-modal retrieval are broken. Users can quickly obtain related graph modal knowledge through highly summarized or dense and rich language descriptions, and can also obtain text descriptions with different information densities related to graphs.

[0027] Enhancing the comprehensiveness and efficiency of knowledge retrieval: The use of a dedicated vector database in conjunction with Redis, along with efficient similarity calculation and indexing structures, enables the system to perform fast and comprehensive knowledge retrieval from massive amounts of multimodal data. It can retrieve as much relevant knowledge as possible for users in a short time, fully meeting their needs for timely and comprehensive knowledge acquisition. Attached Figure Description

[0028] Figure 1 This is a schematic diagram illustrating an example of a document tree visualization structure;

[0029] Figure 2 This is a diagram illustrating the knowledge parsing and storage process;

[0030] Figure 3 This is a diagram illustrating the knowledge retrieval and recall process. Detailed Implementation

[0031] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0032] This invention aims to construct a knowledge retrieval system based on the complex semantic understanding capabilities of a large language model and the cross-modal semantic alignment capabilities of the CLIP model. It utilizes the multimodal understanding capabilities of a large multimodal model to interpret graphs into structured and semantically rich text descriptions, forming original graph-text pairs. Then, using LLM2VEC technology, the text is represented as a highly discriminative embedding. Finally, through training the CLIP model, graph-text semantic alignment is achieved. The resulting visual model possesses structural knowledge anchors, semantic richness, and discriminability, significantly improving the effectiveness of the visual model. The model, trained with domain-specific data, overcomes the challenges of existing knowledge retrieval technologies in graph semantic differentiation and cross-modal retrieval, achieving accurate and efficient retrieval of multimodal knowledge and meeting the complex and multifaceted semantic retrieval needs of professional users. This invention solves the fourth problem mentioned in the background section and lays the groundwork for solving the first and third problems.

[0033] This invention mainly includes the following modules:

[0034] The multimodal data acquisition, separation, and preprocessing module is responsible for collecting multimodal data, including text and images, from various data sources. For text data, it cleans the data, removing low-quality semantic data. For image data, it extracts and analyzes the data and removes duplicates. The data primarily originates from user-uploaded knowledge material files, belonging to private domain knowledge data. The knowledge materials are standardized and transcribed at the file level using knowledge material modality separation and recombination methods. After transcribing, the file metadata information is retained, and the various modal information in the file content is extracted in an orderly manner and assembled into a unified format, such as Markdown, according to the original order. Simultaneously, the mapping relationships are stored in a database. Each file can correspond to multiple records, and each record should include a unique file identifier, the modality type of the data, a unique identifier for the header data, a unique mapping identifier for the modal semantic content of the data, and time information. The unique mapping identifier for modal semantic content is explained as follows: due to the large amount of duplicate semantic content in the actual material set, and considering factors such as system performance and model training, this unique semantic record information is retained after deduplication. This information is associated with multiple semantically consistent material information.

[0035] Training data generation module: LLM2VEC dataset preparation: Using the separated images, and after understanding and analyzing each image, a set of two titles is generated, thus forming a corpus. LLM2CLIP dataset preparation: Again, using the separated images, and after understanding and analyzing each image, a summary description and detailed information are generated, thus forming a corpus.

[0036] The CLIP text encoder module introduces and fine-tunes an advanced large language model, replacing the traditional text representation in CLIP as a semantic representation model. The fine-tuning corpus comes from the dataset generation module and an open dataset. Following the LLM2VEC (BehnamGhader et al., 2024) training method, title contrast (CC) is applied to the LLM output features for fine-tuning, treating different titles of the same image as positive samples and the rest as negative samples. After training with the CLIP text encoder, the language model encoder module can more rationally allocate the semantic space and optimize text representation, making it a highly suitable super text encoder for CLIP training.

[0037] The CLIP image encoder module employs the LLM2CLIP technique to achieve cross-modal semantic alignment. Fine-tuning corpora are derived from the dataset generation module and open datasets. During training, the gradients of the LLM are frozen to preserve its inherent capabilities, and to compensate for the frozen LLM, several new linear layers are introduced as adapters after the LLM. These layers serve as learnable parameters to improve the alignment between the LLM and the CLIP visual encoder. Following the original design, CLIP also uses a projection layer to align the dimensions of the two encoders, facilitating training with CLIP loss. With this powerful LLM-based supertext encoder, CLIP achieves a substantial leap in language understanding capabilities. The open-world knowledge of LLM enables the CLIP visual encoder to learn more structured and global visual representations consistent with human knowledge. Furthermore, this approach can fully utilize high-quality, long, and dense image interpretation datasets without requiring any special architectural tweaks.

[0038] The knowledge storage module utilizes the trained CLIP model to store deduplicated image data as knowledge pieces, along with relevant metadata (such as data source, timestamp, and data category), in a vector database. The specific vector database selection is configurable, supporting Milvus, Faiss, and others. These vector databases possess efficient vector index structures, facilitating rapid vector retrieval. Simultaneously, Redis is used as an auxiliary caching and data management tool, storing frequently used or critical metadata and vector database index information in Redis. Leveraging Redis's high-speed read / write capabilities improves the overall system's data access efficiency.

[0039] The query processing and knowledge retrieval module receives user-input queries, parses the query semantics (queries can be in various modalities such as text, image, and speech), preprocesses and parses the queries (speech modalities are converted to text modalities using an ASR model), and applies the CLIP model for vectorization representation before retrieving data from the vector database. In the knowledge storage module's vector database, it retrieves multimodal knowledge related to the query by calculating the similarity (e.g., cosine similarity) between the query vector and stored vectors. It filters the results according to a similarity threshold, sorts the results in descending order, and returns the sorted results to the user. A default maximum number of results can be configured; for example, 10 or fewer results are returned by default, while more than 10 results prompt the user to load more.

[0040] The above description is merely a preferred embodiment of the present invention and is used only to illustrate the technical solution of the present invention, and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.

Claims

1. A knowledge retrieval system, characterized in that, include: Multimodal data acquisition, separation, and preprocessing module: responsible for collecting text and image data from various data sources; cleaning text data to remove low-quality semantic data; and extracting and analyzing image data, as well as deduplicating images. Training data generation module: LLM2VEC dataset preparation: using the separated images, and after understanding and analyzing each image, a set of titles is generated, thus forming a corpus; LLM2CLIP dataset preparation: still using the separated images, and after understanding and analyzing each image, a summary description and detailed information are generated, thus forming a corpus; CLIP Text Encoder Module: Introduces a large language model and fine-tunes it to replace the text representation in traditional CLIP as a semantic representation model; CLIP image encoder module: Employs LLM2CLIP technology to achieve cross-modal semantic alignment; fine-tuning corpora are derived from the dataset generation module and open datasets; Knowledge storage module: Applying the trained CLIP model, the deduplicated image data to be retrieved in the system is used as knowledge pieces, and stored in the vector database along with relevant metadata. The specific vector database selection is configurable. At the same time, Redis is used as an auxiliary cache and data management tool to store metadata and the index information of the vector database in Redis. Redis's high-speed read and write characteristics are used to improve the overall system's data access efficiency. The query processing and knowledge retrieval module receives user-input queries, parses the query semantics (the query can be text, image, or speech modality), preprocesses and parses the queries, applies the CLIP model for vectorization, and retrieves vector data from the vector database. In the knowledge storage module's vector database, it calculates the similarity between the query vector and stored vectors to retrieve multimodal knowledge related to the query. It then filters the results according to a similarity threshold, sorts the filtered results in descending order, and returns the sorted results to the user.

2. The system according to claim 1, characterized in that, The data in the multimodal data acquisition, separation and preprocessing module comes from knowledge material files uploaded by users and belongs to private domain knowledge data; The knowledge materials are standardized and transcribed at the file level using knowledge material modality separation and recombination methods. After transcribing, the file metadata information is preserved, and the various modal information in the file content is extracted in an orderly manner and assembled into a unified format according to the original order; at the same time, the mapping relationship is stored in the database.

3. The system according to claim 2, characterized in that, Each file can correspond to several records. Each record should include a unique file identifier, the modality type of the data, a unique header identifier, a unique mapping identifier for the modality semantic content of the data, and time information.

4. The system according to claim 1, characterized in that, The fine-tuning corpus of the CLIP text encoder module comes from the dataset generation module and the open dataset. Following the LLM2VEC training method, the output features of LLM are fine-tuned by applying title contrast, treating different titles of the same image as positive samples and the rest as negative samples.

5. The system according to claim 1, characterized in that, The CLIP image encoder module employs LLM2CLIP technology to achieve cross-modal semantic alignment; the fine-tuning corpus comes from the dataset generation module and the open dataset; during the training phase, the gradients of LLM are frozen to preserve their inherent capabilities, and new linear layers are introduced after LLM as adapters; these layers serve as learnable parameters.

6. The system according to claim 5, characterized in that, Following the original design, CLIP also uses a projection layer to align the dimensions of the two encoders, making it easier to train using CLIP loss.

7. The system according to claim 1, characterized in that, The voice modality of the query processing and knowledge retrieval module is converted into text modality using the ASR model.

8. The system according to claim 1, characterized in that, The query processing and knowledge retrieval module is configured to return a maximum number of results by default. If the maximum number of results is exceeded, the user will be prompted to load more results.