Dynamic vector knowledge base construction and retrieval method based on multi-modal large model

By constructing a dynamic vector knowledge base of multimodal large models, the semantic gap problem of traditional knowledge retrieval system in multi-source heterogeneous data processing is solved, and an efficient and personalized knowledge retrieval experience is achieved, which improves the comprehensiveness of the search results and user satisfaction.

CN120277223AActive Publication Date: 2025-07-08南京迅集科技有限公司

Patent Information

Application Number
CN202510765415.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-07-08
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

Traditional knowledge retrieval systems are difficult to process multi-source heterogeneous data, and lack a unified semantic representation framework, which leads to one-sided and incomplete search results, which cannot meet the complex and diverse query needs of users, and lacks personalized retrieval experience.

Method used

A dynamic vector knowledge base is constructed based on multimodal large models, and structured vector knowledge graphs are constructed through preprocessing, feature extraction, semantic correlation analysis and hierarchical clustering, and personalized search and optimization are carried out in combination with user feedback.

Benefits of technology

It realizes unified representation and efficient retrieval of multimodal data, improves the comprehensiveness and accuracy of search results, provides a personalized user experience, and significantly improves user satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277223A_ABST
    Figure CN120277223A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of knowledge retrieval, and discloses a multi-modal large model-based dynamic vector knowledge base construction and retrieval method, which comprises the following steps of: obtaining a multi-source heterogeneous modal data set, and carrying out preprocessing and modal standardization processing on the multi-source heterogeneous modal data set to obtain a standardized multi-modal data set; performing feature extraction and semantic vector representation generation by using the pre-trained multi-modal large model, and constructing a multi-modal knowledge vector set; semantic association analysis and hierarchical clustering are carried out on the multi-modal knowledge vector set, and a structured vector knowledge base is constructed; performing semantic similarity calculation and relation modeling on the vector knowledge base to form a vector relation network; intention analysis and vector representation are performed based on mixed modal query information input by a user, and efficient similarity retrieval is realized in combination with a vector relation network; dynamic optimization is carried out through user feedback, personalized retrieval result adjustment is achieved, and the problem of limitation of a traditional retrieval system during multi-modal data processing is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of knowledge retrieval, and more specifically, to a method for constructing and retrieving a dynamic vector knowledge base based on a multimodal large model. Background Art

[0002] In the field of knowledge retrieval and data management, multimodal information processing, as a core technical ability, is of great significance for achieving comprehensive and accurate knowledge acquisition. With the rapid development of information technology, users' requirements for the intelligence, personalization, and efficiency of knowledge retrieval systems are increasing day by day, and many problems have gradually emerged in the traditional knowledge retrieval and management mode.

[0003] Compared with the prior art, traditional knowledge retrieval systems mainly rely on single-modal data processing. This method is difficult to process multi-source data forms such as text, images, audio, and video, resulting in one-sided and incomplete retrieval effects. Moreover, different modal data are stored dispersedly and have inconsistent representation forms, making it difficult to perform effective integration and correlation analysis. In addition, due to the lack of a unified semantic representation framework, cross-modal information retrieval faces the problem of semantic gap, resulting in low retrieval accuracy and inability to meet the complex and diverse query needs of users. Traditional methods model the semantic relationships of knowledge relatively simply, lacking deep semantic understanding and vector representation, and it is difficult to capture the implicit associations between knowledge. As a result, the retrieval system can only perform simple keyword matching, unable to understand the true intention of users' queries, and at the same time, it also affects the relevance and comprehensiveness of retrieval results. In terms of knowledge organization structure, traditional methods are difficult to effectively process large-scale heterogeneous data, lacking an efficient indexing mechanism and optimized storage strategy, resulting in low retrieval efficiency and increasing the system burden. It is impossible to achieve a personalized retrieval experience and difficult to dynamically adjust the retrieval strategy according to user feedback, leading to low user satisfaction.

[0004] In view of this, the present invention proposes a method for constructing and retrieving a dynamic vector knowledge base based on a multimodal large model to solve the above problems. Summary of the Invention

[0005] In order to overcome the above-mentioned defects of the prior art and to achieve the above object, the present invention provides the following technical solutions: A method for constructing and retrieving a dynamic vector knowledge base based on a multimodal large model, comprising: Step 1: Obtain a multi-source heterogeneous modal data set, and perform preprocessing and modal normalization processing on the obtained multi-source heterogeneous modal data set to obtain a corresponding standardized multimodal data set; Step 2: Based on a pre-constructed multimodal large model, perform feature extraction and semantic vector representation generation on the standardized multimodal data set to obtain a corresponding multimodal knowledge vector set; Step 3: Conduct semantic association analysis and hierarchical clustering on the obtained multi-modal knowledge vector set, construct a vector knowledge graph structure, and build a structured vector knowledge base based on the vector knowledge graph structure; Step 4: Calculate the semantic similarity and model the relationships between vectors in the obtained vector knowledge base to obtain the corresponding vector relationship network; Step 5: Conduct multi-modal query understanding and intention analysis on the mixed-modal query information input by the user to obtain the corresponding query vector representation; and perform efficient similarity retrieval in combination with the vector relationship network to obtain the corresponding initial retrieval result set; Step 6: Dynamically feedback and optimize the obtained initial retrieval result set to obtain an optimized retrieval result set.

[0006] Furthermore, a data acquisition unit is set up. The data acquisition unit is used to collect and integrate different modal data sets from multiple data sources to obtain the corresponding multi-source heterogeneous modal data set; the multi-source heterogeneous modal data set includes text data, image data, audio data, and video data.

[0007] Furthermore, the process of obtaining the standardized multi-modal data set includes: Conduct modal recognition and format conversion on the obtained multi-source heterogeneous modal data set to obtain the corresponding initial modal data set; Conduct data cleaning and outlier detection on the obtained initial modal data set to obtain the corresponding purified modal data set; Conduct data augmentation and normalization on the obtained purified modal data set to obtain the corresponding enhanced modal data set; Conduct cross-modal alignment and temporal synchronization on the obtained enhanced modal data set to obtain the corresponding multi-modal aligned data set; Conduct feature standardization and dimension consistency processing on the obtained multi-modal aligned data set to obtain the corresponding standardized multi-modal data set.

[0008] Furthermore, the process of obtaining the multi-modal knowledge vector set includes: Select pre-trained large model components suitable for different modal data based on the known modal types; integrate the selected pre-trained large model components and build the corresponding multi-modal large model based on them; Extract features from the corresponding standardized multi-modal data set based on the constructed multi-modal large model to obtain the corresponding modal feature matrix; Conduct cross-modal attention fusion and semantic alignment on the modal feature matrix to obtain the fusion feature representation in the unified semantic space; Conduct knowledge distillation and feature compression on the obtained fusion feature representation to obtain the corresponding efficient dense vectors; Perform semantic similarity verification and quality assessment on the obtained high-efficiency dense vectors to obtain the corresponding multi-modal knowledge vector set.

[0009] Furthermore, the construction process of the vector knowledge graph structure includes: Perform similarity calculation and neighbor analysis on the obtained multi-modal knowledge vector set, and construct the corresponding initial semantic association network based on it; Based on the community detection algorithm, identify semantic clustering clusters for the constructed initial semantic association network to obtain the corresponding semantic community structure; Based on the obtained semantic community structure, construct a multi-level semantic classification system by applying the hierarchical clustering algorithm to obtain the corresponding hierarchical semantic tree; enhance the node relationships of the obtained hierarchical semantic tree, and add cross-level and cross-branch semantic associations to obtain the corresponding vector knowledge graph structure.

[0010] Furthermore, the construction process of the vector knowledge base includes: Adopt a multi-level indexing mechanism to organize the high-efficiency dense vectors, and design the corresponding hierarchical indexing structure based on it; the multi-level indexing mechanism can include inverted index, tree index; Perform compression encoding on the obtained high-efficiency dense vectors to obtain quantization vectors with low storage cost; and construct the corresponding quantization-reduction conversion model based on it; Integrate the hierarchical indexing structure and the quantization-reduction conversion model between vectors based on the vector knowledge graph structure, and construct a multi-level structure vector knowledge base based on it.

[0011] Furthermore, the construction process of the vector relationship network includes: Perform semantic similarity calculation on the obtained vector knowledge base to obtain the semantic similarity between different vectors in the vector knowledge base; and construct the corresponding similarity matrix based on it; Based on the obtained similarity matrix, identify high-density semantic association regions by applying threshold filtering and density clustering to obtain the corresponding semantic association clusters; and perform relationship type identification and attribute annotation on the obtained semantic association clusters to obtain the corresponding semantic relationship triples; Construct a relationship graph structure between vectors based on the semantic relationship triples, and perform graph embedding optimization to obtain the corresponding initial vector relationship network; and perform relationship reasoning and link prediction on the initial vector relationship network to obtain the corresponding enhanced vector relationship network; Perform weight optimization and path analysis on the obtained enhanced vector relationship network to obtain the corresponding vector relationship network.

[0012] Furthermore, the process of obtaining the query vector representation includes: Set up a user interaction unit and obtain the hybrid-modal query information input by the target user based on it; Perform modal separation and recognition on the input hybrid-modal query information; obtain corresponding modality-specific query components; perform normalization processing on the obtained modality-specific query components to obtain corresponding standardized query representations; Use a multi-modal large model to perform vector mapping on the obtained standardized query representation and convert it into a vector representation in a unified semantic space to obtain a corresponding initial query vector; Based on the constructed vector relationship network, perform intent understanding and expansion on the input hybrid-modal query information to obtain corresponding intent-enhanced descriptions; and based on this, re-weight and adjust the obtained initial query vector to obtain an intent-reinforced query vector; Perform context correlation analysis on the intent-reinforced query vector and combine the user's historical query behavior to obtain a corresponding query vector representation.

[0013] Furthermore, the process of obtaining the initial retrieval result set includes: Design a multi-level retrieval strategy based on the constructed hierarchical index structure and perform coarse-grained retrieval in combination with the obtained query vector representation to obtain a corresponding candidate result set; and filter the candidate result set based on pre-set metadata filtering conditions to obtain a refined candidate set; the metadata filtering conditions include time range, source limit, and modality type; Perform fine-grained similarity calculation on the refined candidate set to obtain accurate similarity scores; and sort the refined candidate set based on this to obtain a corresponding ordered result list; Based on the ordered result list, obtain the distribution and correlation of data corresponding to different modality types in the retrieval results, and perform multi-modal fusion and complementary enhancement based on this to obtain an initial retrieval result set with modal balance.

[0014] Furthermore, the process of obtaining the optimized retrieval result set includes: Feed the obtained initial retrieval results back to the target user, and collect the feedback information and user interaction behavior of the user on the initial retrieval results based on the user interaction unit; Mark the corresponding initial retrieval results based on the feedback information of the target user to obtain the retrieved content selected by the user; and construct a corresponding query-result association graph by analyzing the implicit association pattern between it and the input hybrid-modal query information; Integrate the obtained historical feedback information, historical user interaction behavior, and user historical query behavior; and perform cleaning and standardization processing on it to form structured user preference data; Based on user preference data and query-result association graphs, and using reinforcement learning methods to adjust the designed multi-level retrieval strategy to form a personalized retrieval model; Obtain the initial retrieval results and user preference data for threshold parameter optimization; obtain the corresponding adaptive threshold strategy; Combine the personalized retrieval model and the adaptive threshold strategy to re-rank and optimize the initial retrieval result set to obtain the corresponding optimized retrieval result set.

[0015] Technical effects and advantages of the method for constructing and retrieving a dynamic vector knowledge base based on a multi-modal large model of the present invention: 1. Through processes such as pre-training a multi-modal large model and unified semantic space mapping for preprocessing, feature extraction, and semantic vector generation of the collected multi-source heterogeneous modal data sets, realizing the unified representation and semantic interoperability of different modal data, and obtaining a high-quality multi-modal knowledge vector set; it can not only accurately capture cross-modal semantic associations, but also effectively eliminate the semantic gap problem, providing strong support for complex queries, breaking through the limitations of traditional single-modal retrieval, and improving the comprehensiveness and accuracy of retrieval results; 2. By using algorithms such as hierarchical clustering and graph structure modeling to perform semantic association analysis and structured organization on the multi-modal knowledge vector set; constructing a vector knowledge graph with rich semantics, realizing the design of a multi-level indexing mechanism, providing a product quantization compression storage strategy, and forming an efficient structured vector knowledge base; identifying potential semantic associations through vector relationship network modeling, providing key support for efficient retrieval, effectively solving the storage and retrieval efficiency problems of traditional systems in processing massive heterogeneous data, significantly reducing system resource consumption, and improving retrieval response speed; 3. By combining user feedback and a reinforcement learning mechanism, realizing the dynamic optimization of personalized retrieval; forming a highly personalized retrieval experience; it can not only accurately understand the user's true query intention, but also continuously learn the user's preferences and dynamically optimize the retrieval results, solving the problem of the lack of personalization ability in traditional retrieval systems, and significantly improving user satisfaction and system usage value. Brief Description of the Drawings

[0016] Figure 1 It is a schematic diagram of the method for constructing and retrieving a dynamic vector knowledge base based on a multi-modal large model of the present invention; Figure 2 It is a schematic diagram of the system for constructing and retrieving a dynamic vector knowledge base based on a multi-modal large model of the present invention. Detailed Embodiments

[0017] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0018] Embodiment 1 Please refer to Figure 1 As shown, the method for constructing and retrieving a dynamic vector knowledge base based on a multi-modal large model in this embodiment includes: Step 1: Obtain a multi-source heterogeneous modal data set, and perform preprocessing and modal normalization processing on the obtained multi-source heterogeneous modal data set to obtain a corresponding standardized multi-modal data set; Step 2: Based on a pre-constructed multi-modal large model, perform feature extraction and semantic vector representation generation on the standardized multi-modal data set to obtain a corresponding multi-modal knowledge vector set; Step 3: Perform semantic association analysis and hierarchical clustering on the obtained multi-modal knowledge vector set, construct a vector knowledge graph structure, and construct a structured vector knowledge base based on the vector knowledge graph structure; Step 4: Perform semantic similarity calculation and vector relationship modeling on the obtained vector knowledge base to obtain a corresponding vector relationship network; Step 5: Perform multi-modal query understanding and intent analysis on the mixed-modal query information input by the user to obtain a corresponding query vector representation; and perform efficient similarity retrieval in combination with the vector relationship network to obtain a corresponding initial retrieval result set; Step 6: Perform dynamic feedback optimization on the obtained initial retrieval result set to obtain an optimized retrieval result set.

[0019] It should be further noted that in the specific implementation process, the process of obtaining the standardized multi-modal data set includes: Set up a data collection unit, which is used to collect and integrate different modal data sets from multiple data sources to obtain a corresponding multi-source heterogeneous modal data set; the multi-source heterogeneous modal data set includes text data, image data, audio data, and video data; the data sources include publicly available databases and web pages, etc.; for example: collect text corpora from public academic databases, collect high-quality image data from open data set platforms, obtain various sound records from audio libraries, and grab diverse video samples from video content platforms; at the same time, use web crawler technology to obtain multi-modal content from Internet web pages and integrate with third-party services through API interfaces to achieve automated collection of multi-source data; Perform modal recognition and format conversion on the obtained multi-source heterogeneous modal dataset to obtain the corresponding initial modal dataset; modal recognition refers to extracting content features of the modal data in the corresponding multi-source heterogeneous modal dataset through a content feature recognition algorithm, and combining format metadata such as file extensions and MIME types for modal type recognition; modal types include visual modality, text modality, auditory modality, and other composite modalities; format conversion refers to converting the collected multi-source heterogeneous modal data into a preset format requirement based on the recognized modal type; for example: for image data under the visual modality, it can be converted into a unified PNG or JPEG format, and the image resolution can be unified; for text data within the text modality, it can be converted into a TXT or JSON format encoded in UTF-8. Perform data cleaning and outlier detection on the obtained initial modal dataset to obtain the corresponding purified modal dataset; data cleaning includes processes such as removing duplicate data, repairing damaged data, and filling in missing values; that is, data cleaning is used to remove or repair data containing errors, incompleteness, or irrelevance; for example: perform spelling correction on text data, remove HTML tags, and delete meaningless symbols; perform blurriness detection on image data and remove overly blurred images; perform noise detection on audio data and filter out audio segments with too low signal-to-noise ratio; outlier detection identifies and processes data samples that deviate from the normal distribution through statistical methods; the statistical methods used include the Z-score and IQR methods; furthermore, through data cleaning and anomaly detection, a purified modal dataset with higher data quality is obtained. Perform data augmentation and normalization on the obtained purified modal dataset to obtain the corresponding augmented modal dataset; data augmentation refers to performing augmentation processing on the modal data in the purified modal dataset using different augmentation methods based on the modal type; for example: use methods such as synonym replacement, syntactic tree transformation, and backtranslation to augment text data; use methods such as rotation, flipping, scaling, cropping, and color transformation to augment image data; normalization processing refers to unifying the purified modal dataset after data augmentation to the corresponding numerical range to reduce the deviation of data distribution and the heterogeneity within the modality, and obtain the corresponding augmented modal dataset; for example: normalize the image pixel values to the interval [0, 1]. Perform cross-modal alignment and temporal synchronization on the obtained augmented modal dataset to obtain the corresponding multi-modal alignment dataset; cross-modal alignment refers to establishing the corresponding relationship between different modal data within the augmented modal dataset; for example: the corresponding relationship between video frames and corresponding audio segments, images and descriptive texts; temporal synchronization is used to ensure the unity of modal data with time attributes in the time dimension; among them, the temporal synchronization process is achieved through the use of timestamp matching and interpolation algorithms; furthermore, through cross-modal alignment and temporal synchronization, the semantic consistency of different modal data is ensured. Furthermore, perform feature standardization processing and dimension consistency processing on the obtained multi-modal alignment dataset to obtain the corresponding standardized multi-modal dataset. Feature standardization processing refers to scaling different modal data within a unified magnitude. Commonly used methods include Z-score standardization, Min-Max scaling, etc. Dimension consistency processing is used to ensure that different modal data have the same or compatible dimension structures during the data processing, providing a unified interface for subsequent model processing.

[0020] It should be further noted that in the specific implementation process, the acquisition process of the multi-modal knowledge vector set includes: Select pre-trained large model components suitable for different modal data based on the known modal types. The pre-trained large model components include pre-trained language models, audio models, temporal visual models, etc. For example, for text data, pre-trained language models such as BERT or RoBERTa can be selected; for image data, pre-trained visual models such as ResNet, ViT, CLIP, etc. can be selected; for audio data, pre-trained audio models such as Wav2Vec, HuBERT, etc. can be selected; for video data, temporal visual models such as VideoSwin, TimeSformer, etc. can be selected. Furthermore, integrate the selected pre-trained large model components and build the corresponding multi-modal large model based on them. The multi-modal large model can be used to extract features from modal data under different modal types. Among them, the training process of the corresponding multi-modal large model needs to ensure the interface compatibility between each pre-trained large model component and the fusibility of the output representations. Furthermore, based on the constructed multi-modal large model, extract features from the corresponding standardized multi-modal dataset to obtain the corresponding modal feature matrix. The modal feature matrix refers to the set of high-dimensional feature representations obtained after each modal data is extracted by the corresponding pre-trained large model component. It should be further noted that during the feature extraction process, corresponding modal preprocessing strategies need to be adopted for different modal types. For example, text data is converted into the model input format through steps such as word segmentation and encoding; image data is prepared as the input of the visual model through steps such as scaling and normalization; audio data is converted into a form that can be processed by the audio model through steps such as spectral analysis and MFCC feature extraction. Perform cross-modal attention fusion and semantic alignment on the modal feature matrix to obtain a fused feature representation in a unified semantic space; the fused feature representation is a comprehensive feature representation containing multi-modal feature information in the unified semantic space; cross-modal attention fusion refers to using a multi-modal attention mechanism or a Transformer architecture to make the features between different modal data pay attention to each other to capture the semantic associations between cross-modal data; semantic alignment means mapping the features of different modal data into the same semantic space through algorithms such as contrastive learning or shared projection space to ensure that the distances of semantically similar content in the corresponding semantic space are relatively close; Furthermore, perform knowledge distillation and feature compression on the obtained fused feature representation to obtain corresponding efficient dense vectors; knowledge distillation refers to migrating the extracted complex fused feature representation to a more concise lightweight feature representation; feature compression reduces the vector dimension and computational complexity while retaining key semantic information and integrity by using methods such as principal component analysis (PCA), autoencoders, or low-rank approximation; the dimension of the vector after feature compression is usually set within a fixed interval to balance the expressive ability and computational efficiency; Perform semantic similarity verification and quality assessment on the obtained efficient dense vectors to obtain a corresponding multi-modal knowledge vector set; the multi-modal knowledge vector set includes information such as the identifier, vector value, and metadata of each efficient dense vector; Semantic similarity verification refers to testing the discrimination ability of the obtained efficient dense vectors through a preset semantic similarity task; commonly used semantic similarity tasks include similar document retrieval, image-text matching, etc.; quality assessment evaluates the distribution quality and effectiveness of the corresponding efficient dense vectors in the semantic space by using methods such as the silhouette coefficient and semantic preservation metric; the silhouette coefficient represents the average Euclidean distance between an efficient dense vector and other vectors in the corresponding semantic space; the semantic preservation metric is used to quantify the degree of retention of semantic information during the vector representation of the corresponding modal data in the semantic space; Furthermore, perform high-quality screening of efficient dense vectors through semantic similarity verification and quality assessment to form a corresponding multi-modal knowledge vector set.

[0021] It should be further noted that in the specific implementation process, the construction process of the structured vector knowledge base includes: Calculate the similarity and perform neighbor analysis on the obtained multi-modal knowledge vector set, and construct the corresponding initial semantic association network based on it; the similarity calculation obtains the semantic similarity between different high-efficiency dense vectors in the multi-modal knowledge vector set by using measurement methods such as cosine similarity, Euclidean distance or inner product; the neighbor analysis is to identify other high-efficiency dense vectors that are most similar to the corresponding high-efficiency dense vector through the k-nearest neighbor algorithm; divide them into a vector pair; and synchronously construct the corresponding initial semantic association network based on it; the initial semantic association network is a weighted graph structure, composed of nodes and edges; the nodes represent the corresponding high-efficiency dense vectors; the edges represent semantic relationships, and the edge weights are used to reflect the similarity strength of semantic similarity; Identify semantic clustering clusters for the constructed initial semantic association network based on the community detection algorithm to obtain the corresponding semantic community structure; the community detection algorithm includes the Louvain algorithm, the InfoMap algorithm or the label propagation algorithm, which is used to discover tightly connected node groups in the graph structure; the semantic clustering cluster refers to a set of vectors with high semantic similarity and tight connection; the semantic community structure contains multiple semantic clustering clusters and their associated relationships; Based on the obtained semantic community structure, and through the application of the hierarchical clustering algorithm to construct a multi-level semantic classification system, obtain the corresponding hierarchical semantic tree; the hierarchical clustering algorithm adopts a bottom-up agglomerative clustering or a top-down divisive clustering method, and gradually merges or divides the clusters according to the semantic distance; thus organizing the corresponding semantic community structure into a tree structure; the multi-level semantic classification system refers to semantic categories containing multiple abstract levels; forming a hierarchical relationship from specific concepts to abstract-level concepts; Enhance the node relationships of the obtained hierarchical semantic tree, and add cross-level and cross-branch semantic associations to obtain the corresponding vector knowledge graph structure; the vector knowledge graph structure is a complex network containing various semantic relationships; it not only retains the classification information of the hierarchical structure but also contains rich horizontal semantic associations; node relationship enhancement means adding non-hierarchical association relationships based on information such as semantic similarity, co-occurrence frequency between vector representations or external knowledge bases on the basis of the hierarchical semantic tree; Adopt a multi-level index mechanism to organize the high-efficiency dense vectors, and design the corresponding hierarchical index structure based on it; the multi-level index mechanism can include inverted index, tree index, etc.; the design process of the hierarchical index structure can be realized by algorithms such as HNSW or IVF, divide the semantic space where the high-efficiency dense vectors are located into multiple levels of sub-spaces, and construct a multi-level index hierarchy from coarse-grained to fine-grained; at the same time, on each index level, design a suitable index structure according to the distribution characteristics of the high-efficiency dense vectors, such as tree index, graph index or hybrid index, to form a vector index system that can be quickly retrieved; Meanwhile, the product quantization algorithm technology is used to compress and encode the high-efficiency dense vectors to reduce the storage space requirement and obtain quantized vectors with low storage costs. The product quantization algorithm decomposes the high-dimensional vectors into multiple low-dimensional subspaces and performs independent quantization in each subspace, significantly reducing the storage space requirement. For example, a 256-dimensional vector can be divided into 16 16-dimensional sub-vectors, and each sub-vector is independently quantized into an index of one of the 256 codebook vectors, compressing the original vector to 1 / 8 of its original size. Furthermore, the mapping relationship before and after quantization is retained, and a codebook and reconstruction parameters are established based on it to support the rapid conversion between the quantized vectors and the original vectors, obtaining the corresponding quantization-reduction conversion model. Based on the vector knowledge graph structure, the hierarchical index structure and the quantization-reduction conversion model between vectors are integrated, and a vector knowledge base with a multi-level structure is constructed based on them. The vector knowledge base contains the semantic relationship structure of high-efficiency dense vectors, an efficient retrieval mechanism, and a storage optimization strategy, and can support the storage, retrieval, and semantic analysis of large-scale vectors. Moreover, the vector knowledge base adopts a distributed storage architecture to support the efficient read-write and retrieval operations of vector data.

[0022] It should be further noted that in the specific implementation process, the construction process of the vector relationship network includes: Calculate the semantic similarity of the obtained vector knowledge base to obtain the semantic similarity between different vectors in the vector knowledge base, and construct the corresponding similarity matrix based on it. The matrix elements in the similarity matrix represent the similarity values between the corresponding vector pairs. The semantic similarity calculation is realized by a combination of multiple distance measurement methods to capture the semantic relationships from different angles. The distance measurement methods used include calculating the geometric distance between vectors using the Euclidean distance, calculating the angular similarity between vectors using the cosine similarity, calculating the absolute difference between vectors using the L1 distance, and considering the correlation between features using the L1 distance. Moreover, the block calculation and parallel processing strategies are adopted in the semantic similarity calculation process to process large-scale vector sets. Based on the obtained similarity matrix, identify the high-density semantic association regions by applying threshold filtering and density clustering to obtain the corresponding semantic association clusters. Threshold filtering refers to screening vectors based on a pre-set similarity threshold, only retaining the vector pair associations higher than the similarity threshold. Density clustering uses algorithms such as DBSCAN to identify high-density semantic association regions based on the similarity relationship between vector pairs. The high-density semantic association region refers to a group of vectors with high similarity and close connection, and the semantic association cluster is a subset of vectors with clear boundaries. Moreover, there is a strong semantic association between different vector pairs within the corresponding subset of vectors. Perform relationship type recognition and attribute annotation on the obtained semantic association clusters to obtain corresponding semantic relationship triples; relationship type recognition refers to inferring the possible relationship types between vector pairs through rule inference algorithms and combining the relative positions, distances, and distribution characteristics between vectors; for example: "contains", "similar", "premise", "causality", etc.; attribute annotation is used to add attribute information such as weights, confidence levels, and timeliness to the inferred relationship types; the semantic relationship triples are expressed in the form of (head entity, relationship type, tail entity); which can be used to reflect the semantic relationship between vectors. Furthermore, construct a relationship graph structure between vectors based on the semantic relationship triples and perform graph embedding optimization to obtain the corresponding initial vector relationship network; the relationship graph structure is a graph network structure with vectors as nodes and semantic relationships as edges, and the edge attribute information includes relationship types, etc.; graph embedding optimization uses knowledge graph embedding methods such as TransE, RotatE, or ComplEx to learn the vector representations of semantic relationships, so that the head entity vector is close to the tail entity vector after semantic relationship transformation. Perform relationship reasoning and link prediction on the initial vector relationship network to supplement potential association relationships and obtain the corresponding enhanced vector relationship network; relationship reasoning refers to inferring relationships that are not explicitly expressed based on existing semantic relationship triples and transitivity rules through rule reasoning or statistical reasoning methods; for example: "if A contains B and B contains C, then A may contain C"; link prediction refers to predicting potential association relationships that may exist but have not been labeled in the corresponding initial vector relationship network through graph neural networks or path ranking algorithms; potential association relationships are semantic relationships that may exist but are not directly reflected. Perform weight optimization and path analysis on the obtained enhanced vector relationship network to obtain the corresponding vector relationship network; weight optimization refers to adjusting the weights of the edges based on factors such as the importance and credibility of semantic relationships; path analysis is used to identify relationship paths and central nodes in the enhanced vector relationship network to optimize the network structure; relationship paths include the shortest path and the critical path; the vector relationship network contains rich semantic relationship types, weights, and path information to support complex semantic queries.

[0023] It should be further noted that in the specific implementation process, the process of obtaining the query vector representation includes: Set up a user interaction unit and obtain the mixed-modal query information input by the target user based on it; among them, the user interaction unit provides multiple input interfaces and supports multiple query methods such as text input, voice input, image upload, and video upload; the mixed-modal query information is composed of a single modality or a combination of multiple modalities, for example: text + image, voice + text, etc. Perform modality separation and recognition on the input mixed-modal query information to distinguish different modality components such as text, images, and audio within the mixed-modal query information; modality separation refers to separating the input mixed-modal query information into different modality query components such as text, images, and audio based on methods such as content feature recognition, MIME type analysis, or feature detection; for example: identifying the type of text queries (keywords, natural language questions, commands, etc.), the type of image queries (photos, screenshots, sketches, etc.), and the type of audio queries (voice instructions, music segments, etc.); the recognition process refers to the process of determining the modality type corresponding to the query component; obtain the corresponding modality-specific query components through modality classification and recognition; Furthermore, perform normalization processing on the obtained modality-specific query components to obtain the corresponding standardized query representation; normalization processing includes operations such as text standardization, image preprocessing, and audio feature extraction; for example: performing text standardization processing operations such as word segmentation, stemming, and stop word filtering on text queries; performing image preprocessing operations such as resizing, normalization, enhancement, and feature extraction on image queries; performing audio feature extraction operations such as noise reduction, segmentation, and feature transformation on audio queries; convert query components of different modalities into a standard format through the normalization processing process; Use a multi-modal large model to perform vector mapping on the obtained standardized query representation and convert it into a vector representation in a unified semantic space to obtain the corresponding initial query vector; among them, the vector mapping is implemented based on a pre-constructed multi-modal large model to ensure that the obtained initial query vector is in the same semantic space as the obtained efficient dense vector; the initial query vector is a set of vector representations of each modality query component; for example: semantic vectors can be extracted for text queries; visual features can be extracted for image queries; and map the feature vectors of each modality to the same semantic space to ensure the semantic consistency of different modality queries; Furthermore, based on the constructed vector relationship network, perform intent understanding and expansion on the input mixed-modal query information. Intent understanding refers to parsing the potential meaning within the corresponding mixed-modal query information through context understanding and the initial query vector, and performing processing such as ambiguity elimination and implicit information supplementation; based on the intent understanding result, rewrite and expand the original mixed-modal query information to generate supplementary information such as synonymous expressions, related concepts, or hypernym and hyponym concepts, and extract its corresponding intent type (such as information search, recommendation request, operation instruction, etc.) and focus of attention; integrate based on the intent understanding result and the expansion result to form a more comprehensive query semantic representation and obtain the corresponding intent-enhanced description; Re-weight and adjust the obtained initial query vector based on the intention enhancement description to obtain an intention-strengthened query vector. The re-weighting process refers to adjusting the weights of the corresponding query components of the initial query vector based on the intention enhancement description and in combination with the attention mechanism to strengthen the key semantic features and suppress the secondary features. The adjustment process is to shift the position of the query vector in the semantic space to make it closer to the vector position corresponding to the user's true intention. Conduct context correlation analysis on the intention-strengthened query vector and combine the user's historical query behavior to obtain the corresponding query vector representation. Context correlation analysis refers to identifying potential query sequence relationships and providing context support for the current query by obtaining the historical mixed-modal query information in the current session and the semantic similarity between the query result and the current mixed-modal query information. The user's historical query behavior includes historical query content, click records, and dwell time, etc., which are used to infer the user's interest preferences. The query vector representation is a comprehensive vector representation that integrates multi-modal information, intention understanding, and context correlation. It should be further noted that in the specific implementation process, the process of obtaining the initial retrieval result set includes: Design a multi-level retrieval strategy based on the constructed hierarchical index structure and conduct a coarse-grained retrieval in combination with the obtained query vector representation to obtain the corresponding candidate result set. The stacked retrieval strategy obtains the corresponding candidate result set by quickly screening in the corresponding vector relationship network based on the query vector representation in a phased manner of coarse-grained retrieval and fine-grained retrieval. Among them, the coarse-grained retrieval is implemented by using the approximate nearest neighbor algorithm to quickly identify the vector set similar to the query vector expression. The fine-grained retrieval conducts exploratory retrieval based on semantic path reasoning on the corresponding vector set to ensure that potential relevant results are not missed. Among them, the semantic path reasoning process is implemented based on the breadth-first search algorithm. Screen the candidate result set based on the pre-set metadata filtering conditions to obtain a refined candidate set. The metadata filtering conditions include attribute restrictions such as time range, source limit, and modality type, etc., which are used to narrow the retrieval range. The screening process realizes the precise matching of complex query requirements by combining multiple metadata filtering conditions using boolean logic. The filtering process realizes efficient filtering by using the Bitmap Index. Use a cross-encoder model to calculate the fine-grained similarity of the refined candidate set to obtain the exact similarity score, and sort the refined candidate set based on it to obtain the corresponding ordered result list. The cross-encoder model is an encoder based on the Transformer structure that can receive both the query and the candidate item as inputs at the same time, and can be used to capture more complex semantic interactions. Multiple dimensions such as semantic relevance, structural matching degree, and context adaptability need to be considered in the fine-grained similarity calculation process. Furthermore, based on the ordered result list, the distribution and correlation of data corresponding to different modality types in the retrieval results are obtained, and multi-modal fusion and complementary enhancement are performed based on them to obtain an initial retrieval result set with modal balance; the distribution and correlation of data corresponding to different modality types can be used to evaluate the coverage and diversity of content of different modality types; multi-modal fusion refers to integrating retrieval results of different modality types to improve the comprehensiveness and diversity of retrieval results; complementary enhancement refers to using the correlation between data corresponding to different modality types to improve the quality and richness of retrieval results; through multi-modal fusion and complementary enhancement, a reasonable distribution of retrieval results among different modality types can be ensured, avoiding a single modality dominating the final results.

[0024] It should be further noted that in the specific implementation process, the process of optimizing the retrieval result set includes: The obtained initial retrieval results are fed back to the target user, and feedback information and user interaction behaviors of the user on the initial retrieval results are collected based on the user interaction unit; explicit feedback includes information such as user clicks, dwell time, favorites, ratings, etc. Based on the feedback information corresponding to the target user, the corresponding initial retrieval results are marked to obtain the retrieved content selected by the user; and the implicit association pattern between it and the input mixed-modal query information is analyzed to obtain the corresponding query-result association graph; among them, the implicit association pattern refers to the potential relationship between the query and the result inferred from the user interaction behavior, such as the user's tendency to select results of a specific modality or a specific topic; the query-result association graph is a bipartite graph structure representing the association strength between the query and the relevant results. Furthermore, the obtained historical feedback information, historical user interaction behaviors, and historical user query behaviors are integrated; and they are cleaned and standardized to form structured user preference data. Based on the user preference data and the query-result association graph, a reinforcement learning method is used to adjust the designed multi-level retrieval strategy to form a personalized retrieval model; among them, the corresponding adjustment process includes: modeling the retrieval task formed based on the multi-level retrieval strategy as a Markov decision process (MDP), the state is the current query and result set, the action is the result ranking adjustment, and the reward is the user satisfaction index; using a deep Q-network (DQN) or a policy gradient method to learn the optimal ranking strategy; the learning goal is to maximize the cumulative user satisfaction; furthermore, the ranking strategy is continuously optimized through reinforcement learning to form a personalized retrieval model. Further, obtain the initial retrieval results and user preference data for threshold parameter optimization; obtain the corresponding adaptive threshold strategy; the adaptive threshold strategy refers to dynamically adjusting the original similarity threshold according to the distribution and correlation of data corresponding to different modality types in the initial retrieval results, the query difficulty, and historical feedback information; among them, the adaptive threshold strategy includes a multi-level structure combining global and local thresholds to adapt to different retrieval scenario requirements; the dynamic adjustment process is implemented through the Bayesian optimization method; Re-rank and optimize the initial retrieval result set by combining the personalized retrieval model and the adaptive threshold strategy to obtain the corresponding optimized retrieval result set; the re-ranking process is re-ranked by comprehensively considering multiple factors such as the semantic similarity between vectors, user preferences, result diversity, and timeliness.

[0025] The present invention comprehensively preprocesses and standardizes the collected multi-source heterogeneous modality data set to establish a standardized multi-modal data set; uses a pre-trained multi-modal large model for deep feature extraction to generate a knowledge vector representation with unified semantic expression ability; constructs a vector knowledge graph through semantic association analysis and hierarchical clustering to form a structured vector knowledge base; finely models the relationship between vectors to construct a vector relationship network with rich semantics; realizes in-depth understanding and intention analysis of user queries, and obtains initial retrieval results by combining efficient retrieval strategies; dynamically optimizes based on user feedback to improve the relevance of retrieval results and user satisfaction. This method combines the semantic understanding ability of the multi-modal large model and the efficient characteristics of vector retrieval to construct an intelligent, efficient, and sustainable evolving multi-modal knowledge retrieval system.

[0026] Embodiment 2 Please refer to Figure 2 As shown, for the parts not described in detail in this embodiment, refer to the description content of Embodiment 1. Provide a dynamic vector knowledge base construction and retrieval system based on a multi-modal large model, including: A data preprocessing module, configured to obtain a multi-source heterogeneous modality data set, and perform preprocessing and modality normalization processing on the obtained multi-source heterogeneous modality data set to obtain the corresponding standardized multi-modal data set; A semantic vector generation module, based on a pre-constructed multi-modal large model, performs feature extraction and semantic vector representation generation on the standardized multi-modal data set to obtain the corresponding multi-modal knowledge vector set; A knowledge base construction module, performs semantic association analysis and hierarchical clustering on the obtained multi-modal knowledge vector set, constructs a vector knowledge graph structure, and constructs a structured vector knowledge base based on the vector knowledge graph structure; A network modeling module: configured to perform semantic similarity calculation and vector relationship modeling on the obtained vector knowledge base to obtain the corresponding vector relationship network; A query and retrieval module for performing multi-modal query understanding and intent analysis on the mixed-modal query information input by the user to obtain the corresponding query vector representation; and performing efficient similarity retrieval in combination with a vector relation network to obtain the corresponding initial retrieval result set; A dynamic optimization module for dynamically feedback-optimizing the obtained initial retrieval result set to obtain an optimized retrieval result set.

[0027] Each module is connected by wired and / or wireless means to achieve data transmission between modules.

[0028] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

[0029] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "including one..." does not exclude the presence of additional identical elements in the process, method, article or device including the said element.

[0030] In the description of the present invention, it should be understood that the terms "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0031] In the description of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more.

[0032] In the description of the present invention, the meaning of "several" is one or more, and the meaning of "a large number" is two or more.

[0033] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0034] For the formulas in this specification, only the numerical values are calculated after dimensionlessization. The formulas are obtained by collecting a large amount of data and performing software simulations to obtain a formula that is closest to the actual situation. The preset parameters and threshold values in the formulas are set by those skilled in the art according to the actual situation.

[0035] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.

Claims

1. A method for constructing and retrieving a dynamic vector knowledge base based on a multimodal large model, characterized in that Including: Step 1: Obtain a multi-source heterogeneous modality dataset, and perform preprocessing and modality normalization on the obtained multi-source heterogeneous modality dataset to obtain a corresponding standardized multi-modal dataset; Step 2: Based on a pre-constructed multi-modal large model, perform feature extraction and semantic vector representation generation on the standardized multi-modal dataset to obtain a corresponding multi-modal knowledge vector set; Step 3: Perform semantic association analysis and hierarchical clustering on the obtained multi-modal knowledge vector set, construct a vector knowledge graph structure, and construct a structured vector knowledge base based on the vector knowledge graph structure; Step 4: Perform semantic similarity calculation and vector relationship modeling on the obtained vector knowledge base to obtain a corresponding vector relationship network; Step 5: Perform multi-modal query understanding and intention analysis on the mixed-modal query information input by the user to obtain a corresponding query vector representation; and perform efficient similarity retrieval in combination with the vector relationship network to obtain a corresponding initial retrieval result set; Step 6: Perform dynamic feedback optimization on the obtained initial retrieval result set to obtain an optimized retrieval result set.

2. The method for constructing and retrieving a dynamic vector knowledge base based on a multimodal large model according to claim 1, wherein Set up a data acquisition unit, and the data acquisition unit is used to collect and integrate different modality data sets from multiple data sources to obtain a corresponding multi-source heterogeneous modality dataset; the multi-source heterogeneous modality dataset includes text data, image data, audio data and video data.

3. The method for constructing and retrieving a dynamic vector knowledge base based on a multimodal large model according to claim 2, wherein The process of obtaining the standardized multi-modal dataset includes: Perform modality recognition and format conversion on the obtained multi-source heterogeneous modality dataset to obtain a corresponding initial modality dataset; Perform data cleaning and outlier detection on the obtained initial modality dataset to obtain a corresponding purified modality dataset; Perform data augmentation and normalization on the obtained purified modality dataset to obtain a corresponding augmented modality dataset; Perform cross-modal alignment and temporal synchronization on the obtained augmented modality dataset to obtain a corresponding multi-modal alignment dataset; Perform feature standardization processing and dimension consistency processing on the obtained multi-modal alignment dataset to obtain a corresponding standardized multi-modal dataset.

4. The method for constructing and retrieving a dynamic vector knowledge base based on a multimodal large model according to claim 3, wherein The process of obtaining the multi-modal knowledge vector set includes: Select pre-trained large model components applicable to different modality data based on known modality types; integrate the selected pre-trained large model components and construct a corresponding multi-modal large model based on them; Based on the constructed multi-modal large model, perform feature extraction on the corresponding standardized multi-modal dataset to obtain a corresponding modality feature matrix; Perform cross-modal attention fusion and semantic alignment on the modality feature matrix to obtain a fusion feature representation in a unified semantic space; Perform knowledge distillation and feature compression on the obtained fusion feature representation to obtain a corresponding efficient dense vector; Perform semantic similarity verification and quality assessment on the obtained efficient dense vector to obtain a corresponding multi-modal knowledge vector set.

5. The method for constructing and retrieving a dynamic vector knowledge base based on a multimodal large model according to claim 4, wherein The process of constructing the vector knowledge graph structure includes: Perform similarity calculation and nearest neighbor analysis on the obtained multi-modal knowledge vector set, and construct a corresponding initial semantic association network based on it; Based on the community detection algorithm, perform semantic clustering cluster recognition on the constructed initial semantic association network to obtain a corresponding semantic community structure; Based on the obtained semantic community structure, a multi-level semantic classification system is constructed by applying a hierarchical clustering algorithm to obtain a corresponding hierarchical semantic tree; the node relationships of the obtained hierarchical semantic tree are enhanced, and semantic associations across levels and branches are added to obtain a corresponding vector knowledge graph structure.

6. The method for constructing and retrieving a dynamic vector knowledge base based on a multimodal large model according to claim 5, wherein The construction process of the vector knowledge base includes: A multi-level indexing mechanism is used to organize the efficient dense vectors, and a corresponding hierarchical indexing structure is designed based on it; the multi-level indexing mechanism can include an inverted index and a tree index; The obtained efficient dense vectors are compressed and encoded to obtain quantized vectors with low storage costs; and a corresponding quantization-reduction conversion model is constructed based on them; Based on the vector knowledge graph structure, the hierarchical indexing structure and the quantization-reduction conversion model among vectors are integrated, and a vector knowledge base with a multi-level structure is constructed based on them.

7. The method for constructing and retrieving a dynamic vector knowledge base based on a multi-modal large model according to claim 6, characterized in that The construction process of the vector relationship network includes: Semantic similarity calculation is performed on the obtained vector knowledge base to obtain the semantic similarity between different vectors in the vector knowledge base; and a corresponding similarity matrix is constructed based on it; Based on the obtained similarity matrix, high-density semantic association regions are identified by applying threshold filtering and density clustering to obtain corresponding semantic association clusters; and the relationship types and attribute annotations of the obtained semantic association clusters are identified to obtain corresponding semantic relationship triples; Based on the semantic relationship triples, a relationship graph structure between vectors is constructed and graph embedding optimization is performed to obtain a corresponding initial vector relationship network; and relationship reasoning and link prediction are performed on the initial vector relationship network to obtain a corresponding enhanced vector relationship network; Weight optimization and path analysis are performed on the obtained enhanced vector relationship network to obtain a corresponding vector relationship network.

8. The method for constructing and retrieving a dynamic vector knowledge base based on a multi-modal large model according to claim 7, wherein The acquisition process of the query vector representation includes: A user interaction unit is set, and based on it, the mixed-modal query information input by the target user is obtained; The input mixed-modal query information is subjected to modal separation and recognition; corresponding modality-specific query components are obtained; the obtained modality-specific query components are normalized to obtain corresponding standardized query representations; The obtained standardized query representation is vectorized and mapped by using a multi-modal large model and converted into a vector representation in a unified semantic space to obtain a corresponding query initial vector; Based on the constructed vector relationship network, the intent of the input mixed-modal query information is understood and extended to obtain a corresponding intent-enhanced description; and based on it, the obtained query initial vector is re-weighted and adjusted to obtain an intent-strengthened query vector; Contextual association analysis is performed on the intent-strengthened query vector, and combined with the user's historical query behavior, a corresponding query vector representation is obtained.

9. The method for constructing and retrieving a dynamic vector knowledge base based on a multimodal large model according to claim 8, wherein, The acquisition process of the initial retrieval result set includes: Based on the constructed hierarchical indexing structure, a multi-level retrieval strategy is designed, and rough-grained retrieval is performed in combination with the obtained query vector representation to obtain a corresponding candidate result set; and the candidate result set is filtered based on pre-set metadata filtering conditions to obtain a refined candidate set; the metadata filtering conditions include time range, source limit, and modality type; Perform fine-grained similarity calculation on the refined candidate set to obtain the exact similarity score; and sort the refined candidate set based on it to obtain the corresponding ordered result list; Obtain the distribution and correlation of data corresponding to different modality types in the retrieval results based on the ordered result list, and perform multimodal fusion and complementary enhancement based on it to obtain an initial retrieval result set with modal balance.

10. The method for constructing and retrieving a dynamic vector knowledge base based on a multimodal large model according to claim 9, wherein The process of optimizing the retrieval result set includes: Feed the obtained initial retrieval results back to the target user, and collect the feedback information of the user on the initial retrieval results and the user interaction behavior based on the user interaction unit; Mark the corresponding initial retrieval results based on the feedback information of the target user to obtain the retrieved content selected by the user; and construct the corresponding query-result association graph by analyzing the implicit association pattern between it and the input mixed-modal query information; Integrate the obtained historical feedback information, historical user interaction behavior, and user historical query behavior; and perform cleaning and standardization processing on them to form structured user preference data; Based on the user preference data and the query-result association graph, and adopt the reinforcement learning method to adjust the designed multi-level retrieval strategy to form a personalized retrieval model; Obtain the initial retrieval results and user preference data for threshold parameter optimization; obtain the corresponding adaptive threshold strategy; Re-rank and optimize the initial retrieval result set by combining the personalized retrieval model and the adaptive threshold strategy to obtain the corresponding optimized retrieval result set.

Citation Information

Patent Citations

  • Multi-modal data distributed retrieval method and system based on mapping knowledge domain and vector matching

    CN118551086A

  • Equipment operation manual vector knowledge base construction method

    CN119128219A

Cited By

  • Heat supply system pipe network anomaly detection method based on multi-mode AI large model

    CN120430214A

  • Intelligent Agent collaborative decision-making system for environment monitoring

    CN120509614A

  • Emotion analysis method and system based on multi-modal enhanced retrieval generation technology

    CN120524444A

  • Document type recommendation method based on large model

    CN120561382A

  • Cross-regional data acquisition method and system based on intelligent decision

    CN120578658A