Multi-modal data retrieval method and device, equipment, medium and product

By introducing a collaborative mechanism of semantic understanding, routing decision-making, and knowledge enhancement, the accuracy and scalability issues of multimodal information retrieval methods in electronic signature and video dual recording scenarios are solved. This enables efficient and controllable multi-level business information extraction and compliance review, and improves the accuracy and transparency of cross-modal semantic understanding and information retrieval.

CN122019749APending Publication Date: 2026-05-12中移信息技术有限公司 +1
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
中移信息技术有限公司
Filing Date
2026-01-21
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing multimodal information retrieval methods suffer from insufficient accuracy, timeliness, and scalability in business scenarios such as electronic signatures and video recording. They struggle to support efficient, controllable, and multi-layered business information extraction and compliance review. In particular, they fail to effectively address issues such as modality selection and routing when facing users' natural language queries, unified retrieval and structured output of multimodal data, knowledge-driven semantic enhancement and retrieval accuracy improvement, and adaptive retrieval paths in complex query scenarios.

Method used

By introducing a collaborative mechanism of semantic understanding, routing decision-making, and knowledge enhancement, the system obtains the user's query intent in natural language questions, selects the most suitable modality type for retrieval, and verifies the results using a multimodal business knowledge graph. It also constructs a multi-head routing parallel strategy and a structured semantic alignment mechanism to achieve high efficiency and accuracy in cross-modal semantic understanding and information retrieval.

Benefits of technology

It significantly improves cross-modal semantic understanding capabilities and information retrieval performance, solves the problems of "semantic misjudgment" and "cross-modal ambiguous matching" in existing technologies, realizes multi-channel and multi-source collaborative semantic integration and ranking, and improves retrieval accuracy and result interpretability under complex business semantics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122019749A_ABST
    Figure CN122019749A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a multi-modal data retrieval method and device, equipment, a medium and a product, and the method comprises the steps: obtaining a natural language problem of a user, carrying out the intention understanding of the natural language problem, and obtaining the query intention of the natural language problem; performing retrieval path selection on the query intention, and determining at least one target modal type matched with the natural language question; the natural language question is distributed to vector databases corresponding to the target modal type to be retrieved, and candidate multi-modal retrieval results of the natural language question are obtained; and verifying the candidate multi-modal retrieval result according to a pre-constructed multi-modal business knowledge graph to obtain a target multi-modal retrieval result of the natural language problem. By utilizing the method, a collaborative mechanism of semantic comprehension, routing decision and knowledge enhancement is introduced, so that the cross-modal semantic comprehension capability and the information retrieval performance are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a multimodal data retrieval method, apparatus, device, medium, and product. Background Technology

[0002] In recent years, with the development of artificial intelligence technology in the multimodal field, research on multimodal information retrieval, reasoning, and question answering systems has become a focus. Currently, multimodal information retrieval and reasoning have gradually matured in fields such as intelligent question answering, video understanding, and e-government. Common approaches include semantic expansion and matching based on knowledge graphs, using visual / language models to process multimodal inputs, and fusing retrieval enhancement generation mechanisms to improve generation quality. Some technologies also introduce structures such as autoregressive large models, graph neural networks, and knowledge adapters for feature fusion and reasoning enhancement.

[0003] However, in business scenarios such as electronic signatures and video recording, the data exhibits complex characteristics such as unstructured, heterogeneous modalities, and large semantic spans, resulting in insufficient accuracy, timeliness, and scalability of existing methods. In particular, it is difficult to support efficient, controllable, and multi-layered business information extraction and compliance review. Summary of the Invention

[0004] This invention provides a multimodal data retrieval method, apparatus, device, medium, and product. By introducing a collaborative mechanism of semantic understanding, routing decision, and knowledge enhancement, it constructs a highly adaptable and intelligent multimodal content retrieval method, fundamentally solving the bottlenecks of existing technologies and significantly improving cross-modal semantic understanding capabilities and information retrieval performance.

[0005] Firstly, this embodiment provides a multimodal data retrieval method, which includes:

[0006] Obtain the user's natural language question, perform intent understanding on the natural language question, and obtain the query intent of the natural language question;

[0007] The query intent is used to select a retrieval path and determine at least one target modality type that matches the natural language question;

[0008] The natural language question is distributed to each vector database corresponding to the target modality type for retrieval, and candidate multimodal retrieval results for the natural language question are obtained.

[0009] The candidate multimodal retrieval results are verified based on a pre-constructed multimodal business knowledge graph to obtain the target multimodal retrieval result for the natural language question.

[0010] Secondly, this embodiment provides a multimodal data retrieval device, which includes:

[0011] The intent determination module is used to obtain the user's natural language question, perform intent understanding on the natural language question, and obtain the query intent of the natural language question.

[0012] The modality determination module is used to select a retrieval path for the query intent and determine at least one target modality type that matches the natural language question;

[0013] The initial determination module is used to distribute the natural language question to each vector database corresponding to the target modality type for retrieval, and obtain candidate multimodal retrieval results for the natural language question.

[0014] The retrieval determination module is used to verify the candidate multimodal retrieval results based on a pre-constructed multimodal business knowledge graph, and obtain the target multimodal retrieval result for the natural language question.

[0015] Thirdly, this embodiment provides an electronic device, including:

[0016] At least one processor; and

[0017] A memory communicatively connected to the at least one processor; wherein,

[0018] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the multimodal data retrieval method according to any embodiment of the present invention.

[0019] Fourthly, this embodiment provides a computer-readable storage medium storing computer instructions that, when executed by a processor, implement the multimodal data retrieval method as described in any embodiment of the present invention.

[0020] Fifthly, embodiments of the present invention also provide a computer program product, the computer program product including a computer program, which, when executed by a processor, implements the multimodal data retrieval method as described in any embodiment of the present invention.

[0021] This invention provides a multimodal data retrieval method, apparatus, device, medium, and product. The method includes: acquiring a user's natural language question; performing intent understanding on the natural language question to obtain the query intent of the natural language question; selecting a retrieval path based on the query intent to determine at least one target modality type matching the natural language question; distributing the natural language question to various vector databases corresponding to the target modality type for retrieval to obtain candidate multimodal retrieval results for the natural language question; and verifying the candidate multimodal retrieval results based on a pre-constructed multimodal business knowledge graph to obtain the target multimodal retrieval result for the natural language question. This technical solution, by introducing a collaborative mechanism of semantic understanding, routing decision, and knowledge enhancement, constructs a highly adaptable and intelligent multimodal content retrieval method, fundamentally solving the bottlenecks of existing technologies and significantly improving cross-modal semantic understanding capabilities and information retrieval performance. The dynamic modality routing mechanism based on semantic understanding can adaptively select the optimal modality retrieval path according to the query content, supporting efficient mapping from query intent to modality path. To address complex queries, a "parallel modal fusion + routing fusion" mechanism is introduced. This mechanism constructs a multi-head routing parallel strategy to support complex queries, performing searches separately across text, image, and video modalities, followed by multimodal feature alignment and fusion. The final ranking output is driven by the fusion result. This solves the problem of existing solutions' inability to effectively handle multimodal mixed semantics, achieving multi-channel, multi-source collaborative semantic integration and ranking, and improving robustness in cross-modal and multi-intent queries. Simultaneously, a business knowledge graph enhancement step is introduced to achieve "structured semantic alignment + retrieval reasoning" capabilities. A knowledge graph integrating multimodal business elements is constructed, supporting the parsing of user queries into structured semantics and graph-driven reliable ranking of preliminary search results. This effectively solves problems such as "semantic misjudgment" and "cross-modal ambiguous matching" in existing technologies, improving retrieval accuracy and result interpretability under complex business semantics.

[0022] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1This is a flowchart illustrating a multimodal data retrieval method provided in Embodiment 1 of the present invention;

[0025] Figure 2 This is a flowchart illustrating a multimodal data retrieval method provided in Embodiment 1 of the present invention.

[0026] Figure 3 This is a flowchart illustrating another multimodal data retrieval method provided in Embodiment 2 of the present invention;

[0027] Figure 4 This is a schematic diagram of the structure of a multimodal data retrieval device provided in Embodiment 3 of the present invention;

[0028] Figure 5 This is a schematic diagram of the structure of an electronic device provided in Embodiment 4 of the present invention. Detailed Implementation

[0029] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0030] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0031] It's important to understand that in multimodal information retrieval tasks, a common approach is to combine knowledge graphs, pre-trained models, and multimodal embedding representations to enhance the understanding of relationships between heterogeneous modalities such as text, images, and videos. While some current technical solutions attempt to incorporate knowledge graphs into the retrieval process to improve semantic modeling capabilities and contextual consistency, significant limitations remain in areas such as retrieval accuracy, modality fusion mechanisms, and system scalability. The following is an introduction to typical representative technical solutions:

[0032] The Contrastive Language-Image Pre-Training (CLIP) series of models, by jointly training image and text encoders, enables image-text pairs to have similar distributions in a shared vector space, thus supporting cross-modal retrieval from text to image. Models such as the VIOLET video language learning model introduce frame-level modeling and temporal location encoding, expanding the ability to model temporal information and achieving significant progress in video retrieval tasks. The multimodal video understanding foundation model (InternVideo2) further integrates image, video, and audio multi-source modalities, achieving even more powerful multimodal modeling capabilities through a unified visual coding backbone.

[0033] On the other hand, a number of multimodal semantic enhancement schemes integrating knowledge graph structures have emerged in existing technologies. Some studies (such as the graph convolutional network-based model MMGCN and the multimodal model VQA-KG that combines knowledge graphs and visual question answering) attempt to introduce graph entities or concepts into the graph-text joint coding structure, and model the semantic relationships between visual regions and knowledge entities through graph neural networks to improve the representation and reasoning capabilities of complex semantics. However, most of these methods are limited to "static graph-text understanding" tasks and lack the ability to model structured retrieval paths for natural language queries, making it difficult to directly adapt to the paragraph-level multimodal retrieval needs in real-world scenarios.

[0034] The existing technology has the following problems:

[0035] (1) Modality selection and routing problem under natural language query: When faced with users' natural language queries, the existing system has difficulty in accurately identifying the query intent and selecting the most matching modality index (image, video, text), and lacks a flexible and efficient modality and granular routing mechanism.

[0036] (2) The problem of unified retrieval and structured output of multimodal data: Multimodal data (text, images and videos) present challenges such as heterogeneous information structure and large differences in granularity, making it difficult for traditional index retrieval methods to achieve cross-modal, high-precision paragraph-level content location and return.

[0037] (3) Knowledge-driven semantic enhancement and retrieval accuracy improvement: Existing multimodal retrieval methods generally lack deep semantic reasoning capabilities and cannot make full use of external knowledge graphs for semantic completion, context association and result reordering, resulting in insufficient relevance and accuracy of retrieval results.

[0038] (4) Adaptive retrieval path problem in complex query scenarios: When the user input is semantically ambiguous or the query target is complex, the system lacks an effective mechanism to dynamically adjust the retrieval path, which makes it impossible to achieve modal collaboration and strategy optimization, thus reducing the intelligence and availability of the system.

[0039] Therefore, a method is needed to solve the above problems.

[0040] Example 1

[0041] Figure 1 This is a flowchart illustrating a multimodal data retrieval method provided in Embodiment 1 of the present invention. This method is applicable to situations that support high-precision retrieval tasks at the image, text, and video segment level under complex business semantics. This method can be executed by a multimodal data retrieval device, which can be implemented in hardware and / or software and is generally integrated into an electronic device.

[0042] like Figure 1 As shown, the multimodal data retrieval method provided in this embodiment may specifically include the following steps:

[0043] S101. Obtain the user's natural language question, perform intent understanding on the natural language question, and obtain the query intent of the natural language question.

[0044] In this embodiment, the user's natural language question can be specifically understood as a query question described in natural language. Upon receiving the user's natural language question, intent understanding is performed on the natural language question to further understand the user's query intent. For example, an intent classifier and a named entity recognition module are used to perform intent understanding on the natural language question from different perspectives to obtain the query intent of the natural language question.

[0045] S102. Select a retrieval path for the query intent and determine at least one target modality type that matches the natural language question.

[0046] In this embodiment, vector databases corresponding to different modalities are pre-built, such as vector databases for text modalities, image modalities, and video modalities. This embodiment aims to automatically select the appropriate modal (e.g., text / image / video) database for retrieval based on the user's query intent, constructing a query understanding and dynamic routing system. Through deep semantic analysis and multi-level decision-making mechanisms, it achieves accurate identification of query intent and adaptive selection of the optimal retrieval path. Each query is routed using a "modality + granularity" approach, sending it to the most suitable vector database to find the content required for the answer.

[0047] Then, through a multi-level routing system that combines modality awareness and granularity awareness, refined classification and targeted processing of query requests are achieved. The router uses deep semantic analysis technology to extract multi-dimensional features from the user's query intent, intelligently predicting the most suitable modality type combination, denoted as the target modality type. Target modality types include text modality, image modality, and video modality. That is, the final modality type combination may include pure text information queries (such as text information like subtitles, titles, and optical character recognition text), image content queries (such as scene recognition, object detection, and person recognition), video content queries (such as queries focusing on dynamic features like action sequences, temporal changes, and motion trajectories), and multi-modal fusion queries (such as queries requiring complex semantic understanding integrating text, visual, and audio information). Each query type will automatically trigger the corresponding modality vector database for matching.

[0048] S103. Distribute the natural language question to each vector database corresponding to the target modality type for retrieval to obtain candidate multimodal retrieval results for the natural language question.

[0049] In this embodiment, the natural language question is distributed to the vector databases corresponding to each target modality type for retrieval, and the retrieval results in each vector database are denoted as single retrieval results corresponding to the target modality type. Alternatively, the single retrieval results of each modality type can be fused to obtain fused retrieval results. The single retrieval results and fused retrieval results of the multimodal search are then used together as candidate multimodal retrieval results for the natural language question.

[0050] For complex multimodal fusion query scenarios, an advanced multi-head parallel routing strategy is deployed. First, parallel retrieval is performed simultaneously in three independent modal spaces: text, image, and video. Then, a fusion mechanism is used to perform deep feature fusion on the multimodal retrieval results. Finally, an intelligent re-ranking algorithm optimizes the result quality and returns the most relevant search content. The single multimodal retrieval results and the fused retrieval results are combined as candidate multimodal retrieval results for natural language questions.

[0051] S104. Verify the candidate multimodal retrieval results based on the pre-constructed multimodal business knowledge graph to obtain the target multimodal retrieval result for the natural language question.

[0052] After the initial retrieval, the rich semantic relationships of the knowledge graph are used to perform secondary verification and re-ranking of the candidate multimodal results. By calculating the semantic consistency score between the candidate content and the graph path, non-compliant, irrelevant, or semantically conflicting results are excluded, and the Top-K paragraph results with the highest credibility are output as the target multimodal retrieval results for the natural language question.

[0053] Specifically, if the query intent of the natural language question includes an image, the target multimodal retrieval result includes the image related to the natural language question and the image index location; if the query intent of the natural language question includes a video, the target multimodal retrieval result includes the video name, segment start and end timestamps, keyframes, and corresponding recognized text of the video related to the natural language question; if the query intent of the natural language question includes text, the target multimodal retrieval result includes the text related to the natural language question and the text index location.

[0054] In this embodiment, the paragraphs in the sorting results are returned in a structured form. If the user's query intent is an image or text paragraph, the index position of the relevant image or text paragraph is returned; if the target is a video, auxiliary information such as the video name, paragraph start and end timestamps, keyframes, and recognized text from Optical Character Recognition 8 / Automatic Speech Recognition (OCR / ASR) is returned. For example, the output format includes: {video_id, start_time, end_time, frame_thumbnail, matched_entities, matched_text}. The graph path and matching nodes called during this search are also recorded to support subsequent interpretative review and compliance retrospective analysis.

[0055] It's important to understand that by including structured information such as video ID, timestamp, keyframes, transcribed text, and matching paths in the output results, the system provides a record and explanation of each search path, achieving structured and auditable output, and enhancing transparency and compliance. This differs from traditional "black box" searches that only return content indexes, improving traceability and accountability in high-trust scenarios. The traceable and auditable structured output format includes rich information such as video ID, segment start and end times, keyframes, OCR / ASR transcribed text, and knowledge graph matching paths in the search results, supporting highly reliable retrieval traceability and compliance review. This clearly distinguishes it from existing technologies that only return coarse-grained search results, improving usability in scenarios requiring high transparency.

[0056] It should be noted that this embodiment employs a paragraph-level indexing structure and a cross-modal time alignment mechanism. In the video modality, the system supports a feature vector-based timestamp binding mechanism based on paragraph segmentation, achieving precise paragraph-level alignment between video content and semantic intent. Compared to traditional methods that can only return the entire video or full-image information, this solution supports locating semantic paragraph units accurate to the second, optimizing the user interaction experience.

[0057] The aforementioned technical solution, by introducing a collaborative mechanism of semantic understanding, routing decision-making, and knowledge enhancement, constructs a highly adaptable and intelligent multimodal content retrieval method, fundamentally solving the bottlenecks of existing technologies and significantly improving cross-modal semantic understanding capabilities and information retrieval performance. The dynamic modal routing mechanism based on semantic understanding can adaptively select the optimal modal retrieval path according to the query content, supporting efficient mapping from query intent to modal path. For complex queries, a "parallel modal fusion + routing fusion" mechanism is introduced, constructing a multi-head routing parallel strategy to support complex queries. Retrieval is performed separately across text, image, and video modalities, followed by multimodal feature alignment and fusion, with the final ranking output driven by the fusion result. This solves the problem that existing solutions cannot effectively handle multimodal mixed semantics, achieving multi-channel, multi-source collaborative semantic integration and ranking, improving robustness in cross-modal, multi-intent queries. Simultaneously, a business knowledge graph enhancement step is introduced to achieve "structured semantic alignment + retrieval reasoning" capabilities, constructing a knowledge graph that integrates multimodal business elements, supporting the parsing of user queries into structured semantics, and performing graph-driven reliable ranking of preliminary retrieval results. It effectively solves problems such as "semantic misjudgment" and "cross-modal ambiguous matching" in existing technologies, and improves the retrieval accuracy and interpretability of results under complex business semantics.

[0058] To more clearly illustrate the multimodal retrieval method provided in the embodiments of the present invention, exemplarily, Figure 2 This is a flowchart illustrating a multimodal retrieval method provided in Embodiment 1 of the present invention, as shown below. Figure 2 As shown, the system receives user queries (i.e., the user's natural language question); processes the user query through a text encoder to obtain word embeddings (i.e., query intent); routes the queries to text index / vector databases, image index / vector databases, and video index / vector databases respectively, and performs retrieval and ranking for each; the retrieval and ranking results are then augmented with a knowledge graph to output paragraph retrieval results. Each index / vector database is pre-processed based on business data through a text / visual encoder to obtain the text index / vector database, image index / vector database, and video index / vector database respectively, for subsequent retrieval.

[0059] As an optional embodiment of the present invention, based on the above embodiments, the method can be optimized to further include the following step before obtaining the user's natural language question:

[0060] a1) A hierarchical feature extraction method is used to convert the original multimodal data into feature vectors of the corresponding modality type.

[0061] In this embodiment, the original multimodal data can be understood as data including multiple modal types, such as text, images, and videos. Specifically, semantically rich feature representations are extracted from the original multimodal data. A hierarchical feature extraction method is used to convert the original multimodal data into a standardized vector representation, denoted as a feature vector.

[0062] As a specific implementation method, the step of converting the original multimodal data into feature vectors of the corresponding modality type using a hierarchical feature extraction method can be optimized, including:

[0063] a11) Extract image modal data from the original multimodal data, and use a visual encoder to extract features from the image modal data to generate a visual feature vector.

[0064] In this embodiment, for image modalities, image modal data is extracted from the original multimodal data, and the feature vector of the image modal data is directly extracted using a visual encoder, denoted as the visual feature vector.

[0065] a12) Extract video modal data from the original multimodal data and divide the video modal data into multiple video segments, so as to use the visual encoder to extract features from each video segment and generate a visual feature vector.

[0066] In this embodiment, for video modalities, video modal data is extracted from the original multimodal data. The video modal data is then divided into multiple video segments of fixed length using video preprocessing techniques, for example, dividing the video modal data into segments of 10 seconds each, with each segment serving as a basic retrieval unit. Subsequently, a visual encoder is used to extract image frame features from each segment to generate corresponding visual feature vectors.

[0067] a13) Extract text modal data from the original multimodal data, and use a pre-trained language model or text encoder to perform semantic encoding on the text modal data to generate a text feature vector containing contextual information.

[0068] In this embodiment, for the text modality, text modality data is extracted from the original multimodal data, and the speech transcribed from the video is extracted using high-precision speech recognition technology to form a semantically complete text corpus. Subsequently, a pre-trained language model or multimodal text encoder is used for semantic encoding to generate text feature vectors containing contextual information.

[0069] b1) Represent the feature vectors of each modality type according to the paragraph index, and store the feature vectors marked with the modality type into the corresponding vector database according to the modality type.

[0070] In this embodiment, after obtaining the feature vectors of each modality type, the modality labeling and index library construction are finally carried out. The extracted visual feature vectors and text feature vectors are labeled with modality types respectively, and the labeled feature vectors are classified and stored in the corresponding vector database according to modality type.

[0071] As a specific implementation, the step of representing the feature vectors of each modality type according to paragraph indexes and storing the feature vectors marked with the modality type into the corresponding vector database according to the modality type can be optimized, including:

[0072] b11) Divide the original multimodal data into segments to obtain data segments of the corresponding modality type.

[0073] In this embodiment, the obtained multimodal feature vectors are used to construct a paragraph-level index representation to support vector matching and paragraph localization for subsequent natural language queries. Specifically, the original multimodal data is divided into paragraphs to obtain data paragraphs of corresponding modality types.

[0074] b12) For each modality type, bind each feature vector to the corresponding data segment, and store the feature vector marked with the modality type in the corresponding vector database according to the modality type.

[0075] In this embodiment, a corresponding vector database is constructed for each modality type. For example, a vector database for image modality is constructed. A vector database for video modality is constructed. A vector database for text modality is constructed. For each modality type, each feature vector is bound to the corresponding data segment, and the modality type is labeled. Then, the feature vector labeled with the modality type is stored in the corresponding vector database.

[0076] For example, for video feature vectors, the video feature vectors are bound to the start and end times of video segments, and a video segment index is built using the Facebook AI Similarity Search (FAISS) approximate nearest neighbor indexing method. Specifically, during a query, after mapping to their respective vectors, a fast nearest neighbor search is performed in the index, returning the item with the highest similarity. This ultimately achieves a structured representation of video semantic segments. In the video modality, a segmentation processing and timestamp binding mechanism is designed, and an approximate nearest neighbor index is built using FAISS, binding video segment features to the temporal structure to improve the granularity of query results.

[0077] Unlike existing technologies that use coarse-grained processing of entire video segments or full-frame images, the above-mentioned technical solution improves the accuracy of query results by designing a segmented processing and timestamp binding mechanism to bind feature vectors to segments and build a vector database. It supports segment-level indexing and fine-grained retrieval, optimizes retrieval granularity and time-based positioning capabilities, and achieves retrieval accuracy down to the second level for time segments, meeting the time-sensitive review needs of fields such as signature verification and risk control.

[0078] Example 2

[0079] Figure 3 This is a flowchart illustrating another multimodal data retrieval method provided in Embodiment 2 of the present invention. This embodiment is a further optimization of the above embodiment. In this embodiment, the following optimizations are made: "performing intent understanding on the natural language question to obtain the query intent of the natural language question"; "selecting a retrieval path for the query intent to determine at least one target modality type matching the natural language question"; "distributing the natural language question to each vector database corresponding to the target modality type for retrieval to obtain candidate multimodal retrieval results of the natural language question"; and "verifying the candidate multimodal retrieval results based on a pre-constructed multimodal business knowledge graph to obtain the target multimodal retrieval results of the natural language question".

[0080] like Figure 3 As shown in the figure, this embodiment 2 provides a multimodal data retrieval method, which specifically includes the following steps:

[0081] S201. Obtain the user's natural language question, use an intent classifier to identify the intent of the natural language question, and obtain the intent type of the natural language question.

[0082] In this embodiment, an intent classifier and a named entity recognition module are first used to construct a comprehensive query understanding framework. The intent classifier, based on a pre-trained language model, can accurately identify the core intent type of the query. By using the intent classifier to identify the intent of the natural language question, the purpose of the user's query is obtained, which serves as the intent type of the natural language question.

[0083] S202. The named entity recognition module is used to extract entity information from the natural language problem to obtain the business elements of the natural language problem.

[0084] Specifically, the named entity recognition module focuses on extracting key entity information from natural language problems, including business elements such as names, contract numbers, and specific actions.

[0085] S203. The intent type and business elements of the natural language are used as the query intent of the natural language question.

[0086] The two modules described above work together to provide rich semantic features and structured information for subsequent routing decisions. Specifically, they combine the intent type of the natural language query with business elements to form the query intent of the natural language question. These steps essentially analyze the user's natural language question from different perspectives, collaboratively improving the accuracy of the query.

[0087] S204. Add the query intent to the prompt words and input the prompt words into the basic large model to output at least one target modality type that matches the natural language question; or input the query intent into the fine-tuning large model to output at least one target modality type that matches the natural language question.

[0088] In this embodiment, two optional router implementation strategies are provided, allowing for flexible selection of the most suitable router implementation method based on actual deployment needs, computing resource constraints, and accuracy requirements.

[0089] One implementation approach is a training-free routing method, which involves adding the query intent to prompt words and inputting these prompt words into a base model to output at least one target modality type that matches the natural language question. Specifically, it leverages the powerful reasoning capabilities of large-scale language models by using designed prompt word templates, such as "You are now a router, and your task is to determine the type of knowledge to search for," and adding examples like "Question 1 → You should search for video clips; Question 2 → You should search for documents; Question 3 → You should search for images..." This achieves plug-and-play routing prediction functionality, offering advantages such as simple deployment and high adaptability.

[0090] Another approach is trainable optimization routing, which involves inputting the query intent into a fine-tuned large model to output at least one target modality type that matches the natural language question. Specifically, pseudo-labels are constructed using a standard dataset (benchmark), and a lightweight model is trained as a router. Each question is automatically labeled using the "default modality / granularity" preferences of the existing dataset. Through supervised learning and fine-tuning on specific business data, higher routing accuracy and faster inference speeds are achieved, making it suitable for performance-critical production environments.

[0091] S205. Perform structured parsing on the natural language problem to obtain the structured parsing result.

[0092] In this embodiment, an intelligent semantic understanding and reasoning mechanism is constructed. A multimodal business knowledge graph is used to achieve deep semantic parsing and context enhancement of queries, improving the accuracy and compliance of retrieval. First, the natural language question is parsed in a structured manner to obtain semantic triples, and the parsing result is used as the structured parsing result. For example, the query "Zhang San signed the first page of the contract" is precisely parsed into the semantic triple: <Zhang San, signed, first page of the contract>.

[0093] S206. Align and reason with the structured parsing results and the pre-built multimodal business knowledge graph to obtain the enhanced retrieval instructions corresponding to the natural language question.

[0094] In this embodiment, the structured parsing results obtained above are aligned and reasoned with the pre-constructed multimodal business knowledge graph to obtain enhanced retrieval instructions.

[0095] As a specific implementation, the step of aligning and reasoning the structured parsing results with a pre-built multimodal business knowledge graph to obtain the enhanced retrieval instructions corresponding to the natural language question can be optimized, including:

[0096] a1) Match the structured parsing results with the entity relationship paths in the multimodal business knowledge graph.

[0097] In this embodiment, the multimodal business knowledge graph is pre-built and stores multiple entity relationship paths. Specifically, the structured parsing results are used to perform a path search algorithm in the multimodal business knowledge graph to find entity relationship paths that match the query.

[0098] b1) If a match is successful, the structured parsing result is used as the core semantics of the query intent, and the associated information of the structured parsing result is used as the auxiliary semantics of the query intent.

[0099] Specifically, if a complete matching path is found, it is determined as the core semantics of the query intent. For the complex dynamic relationships between entities in the graph, dynamic relationship reasoning is performed, and the association information from the structured parsing results is used as auxiliary semantics of the query intent.

[0100] For example, the query "Zhang San signed the first page of the contract" is precisely parsed into a semantic triple: <Zhang San, sign, first page of the contract>. A path search algorithm is then performed in the graph to find entity relationship paths that match the query triples. If a complete matching path is found, it is determined as the core semantic of the query intent.

[0101] For the complex dynamic relationships between entities in the graph, dynamic relationship reasoning is implemented. For example, when it is detected that "Zhang San" signed the first page of the "purchase contract" and that the document page appears in the 15th-20th second of "Video A", the structured information is converted into "directional cues" (the spatiotemporal information of "Video A 15th-20th second") as auxiliary input for the vector retrieval machine.

[0102] c1) The core semantics and auxiliary semantics of the query intent are used as the enhanced retrieval instructions corresponding to the natural language question.

[0103] Specifically, the core semantics and auxiliary semantics of the query intent are used as enhanced retrieval instructions corresponding to natural language questions.

[0104] It should be noted that when encountering situations such as missing knowledge graph paths, semantic ambiguity, or incomplete information, an intelligent completion mechanism is employed. By analyzing the semantic features of the query intent and the type of missing information, the routing strategy in step S204 is dynamically invoked, and the optimal combination of retrieval modalities (such as image / text / video / multimodal fusion) is adaptively selected to perform the retrieval.

[0105] The aforementioned technical solution combines multimodal knowledge graphs for semantic alignment and enhanced reasoning, constructing a multimodal business knowledge graph strongly correlated with video / image / text. It also introduces a graph-driven reasoning mechanism to perform structured parsing of user query intent, entity path retrieval, and result reordering. It supports mapping queries such as "the first page of a contract signed by someone" to verifiable knowledge graph triples and transforming the reasoning results into paragraph-level vector query intents. Compared to existing technologies that rely solely on surface-level semantic similarity matching, this technical solution significantly improves the interpretability and compliant retrieval capabilities of the results.

[0106] S207. Distribute the enhanced search instruction to each vector database corresponding to the target modality type for retrieval, and obtain a single search result corresponding to each target modality type.

[0107] Specifically, the enhanced search command is distributed to the vector databases corresponding to each target modality type for retrieval, and the search results in each vector database are denoted as single search results corresponding to the target modality type. For example, assuming the target modality types include text, image, and video modalities, the enhanced search command is distributed to the vector database corresponding to the text modality for retrieval, obtaining a single search result for the text modality. Similarly, the enhanced search command is distributed to the vector database corresponding to the image modality for retrieval, obtaining a single search result for the image modality. And the enhanced search command is distributed to the vector database corresponding to the video modality for retrieval, obtaining a single search result for the video modality. In summary, a single search result for multiple modalities is obtained.

[0108] S208. The individual search results are merged to obtain a merged search result.

[0109] Subsequently, a fusion mechanism based on a multimodal large model (BridgeTower) architecture is used to perform deep feature fusion on the single search results of multimodal search. Finally, the most relevant search content is returned as the fused search result by optimizing the result quality through an intelligent re-ranking algorithm.

[0110] S209. The single search results and the fused search results are used as the candidate multimodal search results.

[0111] Specifically, single and fused multimodal search results are used together as candidate multimodal search results for natural language questions.

[0112] S210. Determine the semantic consistency score between the candidate multimodal retrieval results and the entity relationship paths in the multimodal business knowledge graph.

[0113] In this embodiment, after the initial retrieval is completed, the rich semantic relationships of the multimodal business knowledge graph are used to perform secondary verification and re-ranking of the candidate multimodal retrieval results. By calculating the semantic consistency score between the candidate content and the graph path, non-compliant, irrelevant, or semantically conflicting results are excluded, and the Top-K paragraph results with the highest credibility are output. This step is used to calculate the semantic consistency score between each candidate multimodal retrieval result and the entity relationship path in the multimodal business knowledge graph.

[0114] S211. Sort the candidate multimodal retrieval results from high to low according to their semantic consistency scores, and obtain a predetermined number of candidate multimodal retrieval results as the target multimodal retrieval results for the natural language problem.

[0115] The set number can be set according to the actual situation, and no specific limit is set here. Specifically, the candidate multimodal search results are sorted from high to low according to their calculated semantic consistency scores, and the first set number of candidate multimodal search results are taken as the most relevant search results for the natural language question, and are recorded as the target multimodal search results.

[0116] The above technical solution specifies the steps for understanding the intent of natural language queries, selecting a retrieval path based on the query intent, distributing the natural language query to the vector databases corresponding to the target modality type for retrieval, and verifying the candidate multimodal retrieval results based on a pre-constructed multimodal business knowledge graph. It implements a multimodal routing mechanism to address the problems of fixed modality selection and poor adaptability in existing technologies. By constructing a query understanding and modality routing mechanism, it introduces a semantic classifier, entity recognizer, and multi-level modality selection strategies to achieve automatic recognition and routing of user natural language queries. Furthermore, it introduces two strategies: training-free language model routing and trainable lightweight model routing, ensuring flexible use under different deployment conditions. It improves the system's ability to perceive complex query intents, significantly enhancing the accuracy and resource utilization efficiency of modality retrieval. It introduces a multimodal perception routing mechanism based on natural language queries and designs a dynamic modality routing mechanism based on semantic understanding, which can adaptively select the optimal modality retrieval path (image / text / video) based on the query content, supporting efficient mapping from query intent to modality path. The routing mechanism supports both training-free and trainingable implementation schemes, significantly different from the fixed modality matching strategy in traditional retrieval systems. This mechanism is the foundation for ensuring the accuracy, flexibility, and computational efficiency of multimodal retrieval. Furthermore, a multi-head routing fusion strategy is employed to support complex query types. For complex queries requiring joint judgment of text, images, videos, and other information, this technical solution designs a parallel modality routing strategy and a fusion retrieval mechanism to achieve feature fusion and unified ranking. Unlike single-modality or fixed fusion structures, this solution supports on-demand combination of modalities, possessing stronger generalization ability and task adaptability. Simultaneously, a business knowledge graph enhancement step is introduced to achieve "structured semantic alignment + retrieval reasoning" capabilities, constructing a knowledge graph that integrates multimodal business elements. This supports parsing user queries into structured semantic triples and uses path search and semantic alignment to achieve assisted identification and re-ranking of query targets. Combined with a semantic consistency scoring mechanism, the initial retrieval results are subjected to graph-driven reliable ranking. This effectively solves problems such as "semantic misjudgment" and "cross-modal ambiguous matching" existing technologies, improving retrieval accuracy and result interpretability under complex business semantics.

[0117] Example 3

[0118] Figure 4 This is a schematic diagram of a multimodal data retrieval device provided in Embodiment 3 of the present invention. This device is suitable for supporting high-precision retrieval tasks at the image, text, and video segment level under complex business semantics. This multimodal data retrieval device can be implemented in hardware and / or software and is generally integrated into electronic devices. Figure 4As shown, the device includes: an intent determination module 31, a modality determination module 32, an initial determination module 33, and a retrieval determination module 34, wherein,

[0119] The intent determination module 31 is used to obtain the user's natural language question, perform intent understanding on the natural language question, and obtain the query intent of the natural language question.

[0120] Modality determination module 32 is used to select a retrieval path for the query intent and determine at least one target modality type that matches the natural language question;

[0121] The initial determination module 33 is used to distribute the natural language question to each vector database corresponding to the target modality type for retrieval, and obtain candidate multimodal retrieval results for the natural language question;

[0122] The retrieval determination module 34 is used to verify the candidate multimodal retrieval results based on a pre-constructed multimodal business knowledge graph, and obtain the target multimodal retrieval result for the natural language question.

[0123] The aforementioned technical solution, by introducing a collaborative mechanism of semantic understanding, routing decision-making, and knowledge enhancement, constructs a highly adaptable and intelligent multimodal content retrieval method, fundamentally solving the bottlenecks of existing technologies and significantly improving cross-modal semantic understanding capabilities and information retrieval performance. The dynamic modal routing mechanism based on semantic understanding can adaptively select the optimal modal retrieval path according to the query content, supporting efficient mapping from query intent to modal path. For complex queries, a "parallel modal fusion + routing fusion" mechanism is introduced, constructing a multi-head routing parallel strategy to support complex queries. Retrieval is performed separately across text, image, and video modalities, followed by multimodal feature alignment and fusion, with the final ranking output driven by the fusion result. This solves the problem that existing solutions cannot effectively handle multimodal mixed semantics, achieving multi-channel, multi-source collaborative semantic integration and ranking, improving robustness in cross-modal, multi-intent queries. Simultaneously, a business knowledge graph enhancement step is introduced to achieve "structured semantic alignment + retrieval reasoning" capabilities, constructing a knowledge graph that integrates multimodal business elements, supporting the parsing of user queries into structured semantics, and performing graph-driven reliable ranking of preliminary retrieval results. It effectively solves problems such as "semantic misjudgment" and "cross-modal ambiguous matching" in existing technologies, and improves the retrieval accuracy and interpretability of results under complex business semantics.

[0124] Optionally, the intent determination module 31 is specifically used for:

[0125] An intent classifier is used to identify the intent of the natural language question to obtain the intent type of the natural language question.

[0126] A named entity recognition module is used to extract entity information from the natural language question to obtain the business elements of the natural language question.

[0127] The intent type and business elements of the natural language are used as the query intent of the natural language question.

[0128] Optionally, the modal determination module 32 is used for:

[0129] Add the query intent to the prompt words and input the prompt words into the base model to output at least one target modality type that matches the natural language question; or

[0130] The query intent is input into a fine-tuned large model to output at least one target modality type that matches the natural language question.

[0131] Optionally, the initial determination module 33 is used for:

[0132] The natural language problem is then parsed in a structured manner to obtain the structured parsing results;

[0133] The structured parsing results are aligned and reasoned with a pre-built multimodal business knowledge graph to obtain enhanced retrieval instructions corresponding to the natural language question;

[0134] The enhanced search command is distributed to each vector database corresponding to the target modality type for separate retrieval, and a single search result corresponding to each target modality type is obtained;

[0135] The individual search results are merged to obtain a fused search result.

[0136] The individual search results and the fused search results are used as the candidate multimodal search results.

[0137] Optionally, the initial determination module 33 is used to align and reason with the structured parsing result and a pre-built multimodal business knowledge graph to obtain the enhanced retrieval instruction corresponding to the natural language question, including:

[0138] The structured parsing results are matched with the entity relationship paths in the multimodal business knowledge graph;

[0139] If a match is successful, the structured parsing result will be used as the core semantics of the query intent, and the association information of the structured parsing result will be used as the auxiliary semantics of the query intent.

[0140] The core semantics and auxiliary semantics of the query intent are used as enhanced retrieval instructions corresponding to the natural language question.

[0141] Optionally, the retrieval determination module 34 is specifically used for:

[0142] Determine the semantic consistency score between the candidate multimodal retrieval results and the entity relationship paths in the multimodal business knowledge graph;

[0143] The candidate multimodal retrieval results are sorted from high to low according to their semantic consistency scores, and a predetermined number of candidate multimodal retrieval results are obtained as the target multimodal retrieval results for the natural language problem.

[0144] Optionally, the device includes an index building module, comprising:

[0145] The vector extraction unit is used to convert the original multimodal data into feature vectors of the corresponding modality type using a hierarchical feature extraction method.

[0146] A vector storage unit is used to represent the feature vectors of each modality type according to the paragraph index, and to store the feature vectors marked with the modality type into the corresponding vector database according to the modality type.

[0147] Optionally, the vector extraction unit is used for:

[0148] Image modal data is extracted from the original multimodal data, and a visual encoder is used to extract features from the image modal data to generate a visual feature vector.

[0149] Video modal data is extracted from the original multimodal data, and the video modal data is divided into multiple video segments. The visual encoder is used to extract features from each video segment to generate a visual feature vector.

[0150] Text modal data is extracted from the original multimodal data, and semantically encoded using a pre-trained language model or text encoder to generate text feature vectors containing contextual information.

[0151] Optionally, the vector storage unit is used for:

[0152] The original multimodal data is divided into segments to obtain data segments of corresponding modality types;

[0153] For each modality type, each feature vector is bound to the corresponding data segment, and the feature vector marked with the modality type is stored in the corresponding vector database according to the modality type.

[0154] Optionally,

[0155] If the query intent of the natural language question includes an image, then the target multimodal retrieval result includes the image related to the natural language question and the image index location;

[0156] If the query intent of the natural language question includes video, then the target multimodal retrieval result includes the video name, paragraph start and end timestamps, keyframes, and corresponding recognized text of the video related to the natural language question.

[0157] If the query intent of the natural language question includes text, then the target multimodal retrieval result includes the text related to the natural language question and the text index position.

[0158] The multimodal data retrieval device provided in the embodiments of the present invention can execute the multimodal data retrieval method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.

[0159] Example 4

[0160] Figure 5 This is a schematic diagram of an electronic device provided in Embodiment 4 of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0161] like Figure 5 As shown, the electronic device 40 includes at least one processor 41 and a memory, such as a read-only memory (ROM) 42 or a random access memory (RAM) 43, communicatively connected to the at least one processor 41. The memory stores computer programs executable by the at least one processor. The processor 41 can perform various appropriate actions and processes based on the computer program stored in the ROM 42 or loaded from storage unit 48 into the RAM 43. The RAM 43 may also store various programs and data required for the operation of the electronic device 40. The processor 41, ROM 42, and RAM 43 are interconnected via a bus 44. An input / output (I / O) interface 45 is also connected to the bus 44.

[0162] Multiple components in electronic device 40 are connected to I / O interface 45, including: input unit 46, such as keyboard, mouse, etc.; output unit 47, such as various types of monitors, speakers, etc.; storage unit 48, such as disk, optical disk, etc.; and communication unit 49, such as network card, modem, wireless transceiver, etc. Communication unit 49 allows electronic device 40 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0163] Processor 41 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 41 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 41 performs the various methods and processes described above, such as multimodal data retrieval methods.

[0164] In some embodiments, the multimodal data retrieval method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 48. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 40 via ROM 42 and / or communication unit 49. When the computer program is loaded into RAM 43 and executed by processor 41, one or more steps of the multimodal data retrieval method described above may be performed. Alternatively, in other embodiments, processor 41 may be configured to perform the multimodal data retrieval method by any other suitable means (e.g., by means of firmware).

[0165] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0166] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0167] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0168] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0169] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0170] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0171] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements the multimodal data retrieval method provided in any embodiment of this invention.

[0172] In implementing a computer program product, computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof. These programming languages ​​include, but are not limited to, object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0173] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0174] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A multimodal data retrieval method, characterized in that, include: Obtain the user's natural language question, perform intent understanding on the natural language question, and obtain the query intent of the natural language question; The query intent is used to select a retrieval path and determine at least one target modality type that matches the natural language question; The natural language question is distributed to each vector database corresponding to the target modality type for retrieval, and candidate multimodal retrieval results for the natural language question are obtained. The candidate multimodal retrieval results are verified based on a pre-constructed multimodal business knowledge graph to obtain the target multimodal retrieval result for the natural language question.

2. The method according to claim 1, characterized in that, The process of performing intent understanding on the natural language question to obtain the query intent of the natural language question includes: An intent classifier is used to identify the intent of the natural language question to obtain the intent type of the natural language question. A named entity recognition module is used to extract entity information from the natural language question to obtain the business elements of the natural language question. The intent type and business elements of the natural language are used as the query intent of the natural language question.

3. The method according to claim 1, characterized in that, The step of selecting a retrieval path for the query intent and determining at least one target modality type that matches the natural language question includes: Add the query intent to the prompt words and input the prompt words into the base model to output at least one target modality type that matches the natural language question; or The query intent is input into a fine-tuned large model to output at least one target modality type that matches the natural language question.

4. The method according to claim 1, characterized in that, The step of distributing the natural language question to the vector databases corresponding to the target modality type for retrieval, and obtaining candidate multimodal retrieval results for the natural language question, includes: The natural language problem is then parsed in a structured manner to obtain the structured parsing results; The structured parsing results are aligned and reasoned with a pre-built multimodal business knowledge graph to obtain enhanced retrieval instructions corresponding to the natural language question; The enhanced search command is distributed to each vector database corresponding to the target modality type for separate retrieval, and a single search result corresponding to each target modality type is obtained; The individual search results are merged to obtain a fused search result. The individual search results and the fused search results are used as the candidate multimodal search results.

5. The method according to claim 4, characterized in that, The step of aligning and reasoning the structured parsing results with a pre-constructed multimodal business knowledge graph to obtain enhanced retrieval instructions corresponding to the natural language question includes: The structured parsing results are matched with the entity relationship paths in the multimodal business knowledge graph; If a match is successful, the structured parsing result will be used as the core semantics of the query intent, and the association information of the structured parsing result will be used as the auxiliary semantics of the query intent. The core semantics and auxiliary semantics of the query intent are used as enhanced retrieval instructions corresponding to the natural language question.

6. The method according to claim 5, characterized in that, The step of verifying the candidate multimodal retrieval results based on a pre-constructed multimodal business knowledge graph to obtain the target multimodal retrieval result for the natural language question includes: Determine the semantic consistency score between the candidate multimodal retrieval results and the entity relationship paths in the multimodal business knowledge graph; The candidate multimodal retrieval results are sorted from high to low according to their semantic consistency scores, and a predetermined number of candidate multimodal retrieval results are obtained as the target multimodal retrieval results for the natural language problem.

7. The method according to claim 1, characterized in that, Before obtaining the user's natural language question, the following is also included: A hierarchical feature extraction method is used to convert the original multimodal data into feature vectors of the corresponding modality type; The feature vectors of each modality type are represented by paragraph indexes, and the feature vectors marked with the modality type are stored in the corresponding vector database according to the modality type.

8. The method according to claim 7, characterized in that, The method of hierarchical feature extraction converts the original multimodal data into feature vectors of corresponding modality types, including: Image modal data is extracted from the original multimodal data, and a visual encoder is used to extract features from the image modal data to generate a visual feature vector. Video modal data is extracted from the original multimodal data, and the video modal data is divided into multiple video segments. The visual encoder is used to extract features from each video segment to generate a visual feature vector. Text modal data is extracted from the original multimodal data, and semantically encoded using a pre-trained language model or text encoder to generate text feature vectors containing contextual information.

9. The method according to claim 7, characterized in that, The step of representing the feature vectors of each modality type according to paragraph indexes and storing the feature vectors marked with the modality type into the corresponding vector database according to the modality type includes: The original multimodal data is divided into segments to obtain data segments of corresponding modality types; For each modality type, each feature vector is bound to the corresponding data segment, and the feature vector marked with the modality type is stored in the corresponding vector database according to the modality type.

10. The method according to claim 1, characterized in that, If the query intent of the natural language question includes an image, then the target multimodal retrieval result includes the image related to the natural language question and the image index location; If the query intent of the natural language question includes video, then the target multimodal retrieval result includes the video name, paragraph start and end timestamps, keyframes, and corresponding recognized text of the video related to the natural language question. If the query intent of the natural language question includes text, then the target multimodal retrieval result includes the text related to the natural language question and the text index position.

11. A multimodal data retrieval device, characterized in that, include: The intent determination module is used to obtain the user's natural language question, perform intent understanding on the natural language question, and obtain the query intent of the natural language question. The modality determination module is used to select a retrieval path for the query intent and determine at least one target modality type that matches the natural language question; The initial determination module is used to distribute the natural language question to each vector database corresponding to the target modality type for retrieval, and obtain candidate multimodal retrieval results for the natural language question. The retrieval determination module is used to verify the candidate multimodal retrieval results based on a pre-constructed multimodal business knowledge graph, and obtain the target multimodal retrieval result for the natural language question.

12. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the multimodal data retrieval method as described in any one of claims 1-10.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the multimodal data retrieval method as described in any one of claims 1-10.

14. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the multimodal data retrieval method as described in any one of claims 1-10.