Multi-modal retrieval method and device
By constructing a multi-level content block structure and vector encoding, the problem of extracting non-textual modal information in multimodal retrieval is solved, enabling accurate matching and logical retrieval of multimodal tables, thus improving retrieval accuracy and flexibility.
Patent Information
- Application Number
- CN202511404026.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-28
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-09-28
AI Technical Summary
Existing technologies cannot effectively extract and locate non-textual modal information in tables during multimodal retrieval, resulting in low retrieval accuracy and efficiency.
A multi-layered content block structure is constructed, including an atomic layer, a semantic layer, and a relational layer. Vector encoding is used to achieve accurate matching and logical retrieval of multimodal tables. Combining vector databases and relational databases improves retrieval accuracy and flexibility.
It achieves complete coverage and multi-dimensional retrieval of every cell in a multimodal table, reduces missed and false detections, improves retrieval accuracy and flexibility, and adapts to user query intent in complex scenarios.
Smart Images

Figure CN120910110A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of multi-modal retrieval, in particular to a multi-modal retrieval method and device. BACKGROUND
[0002] Retrieval-Augmented Generation (RAG) technology effectively reduces the "hallucination" phenomenon in generated answers by combining information retrieval with large language models. Traditional RAG is based on embedding technology, which retrieves unstructured text information through semantic similarity matching, and has good performance.
[0003] However, the table data in actual enterprise applications often contains structured text, images, formulas and other multi-modal content. Therefore, the above-mentioned traditional RAG method focuses on the extraction of text-type cell content, and thus ignores the extraction of other modal information, resulting in low retrieval accuracy. At the same time, the structured data in the table often organizes content according to a specific structure and hierarchy, and does not have a fixed order or common rules in natural language. Therefore, subsequent queries cannot be quickly retrieved through structural positioning, but can only be fuzzy queries, which is inefficient and prone to data omission.
[0004] Therefore, how to improve the accuracy of multi-modal retrieval has become a problem to be solved. SUMMARY
[0005] The present application provides a multi-modal retrieval method and device to at least solve the problem of low multi-modal retrieval accuracy in related technologies.
[0006] The present application provides a multi-modal retrieval method, comprising: receiving a user query text; vector encoding the user query text to obtain a query vector corresponding to the user query text; querying in a vector database according to the query vector to generate a query result; the vector database stores vector encodings corresponding to multi-level content block structures of a plurality of multi-modal table documents; Wherein, the multi-level content block structure of any of the multi-modal table documents includes: an atomic layer composed of a plurality of atomic blocks, a semantic layer composed of a plurality of semantic blocks, and a relationship layer composed of a plurality of relationship blocks; the atomic block includes target element content and structure information of any cell in the multi-modal table document; the semantic block includes aggregated content corresponding to any row, or any column, or any entire table in the multi-modal table document; the relationship block includes text describing the association between a plurality of atomic blocks and a plurality of semantic blocks; the structure information includes row and column indexes corresponding to each cell, and cross-modal association between multi-modal elements.
[0007] The application also provides a multi-modal retrieval device, comprising: a receiving unit configured to receive a user query text; an encoding unit configured to perform vector encoding on the user query text to obtain a query vector corresponding to the user query text; a querying unit configured to perform querying in a vector database according to the query vector to generate a query result; the vector database stores vector encodings corresponding to multi-level content block structures of a plurality of multi-modal table documents; Any multi-modal table document includes a plurality of atomic blocks, a plurality of semantic blocks, and a plurality of relationship blocks in the multi-level content block structure; the atomic block includes target element content and structure information of any cell in the multi-modal table document; the semantic block includes aggregated content corresponding to any row, any column, or any whole table in the multi-modal table document; the relationship block includes text describing the association between a plurality of atomic blocks and a plurality of semantic blocks; and the structure information includes row and column indexes corresponding to each cell and cross-modal association between multi-modal elements.
[0008] The application also provides an electronic device, comprising a memory configured to store a computer program and a processor configured to execute the computer program to implement the steps of any of the multi-modal retrieval methods.
[0009] The application also provides a computer-readable storage medium, which stores a computer program, wherein the computer program is executed by a processor to implement the steps of any of the multi-modal retrieval methods.
[0010] The application also provides a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the steps of any of the multi-modal retrieval methods.
[0011] The application can completely cover the content of each cell, the aggregation semantics of the row and column of the table and the cross-element association logic in the multi-modal table by constructing a multi-level content block structure including an atomic layer (corresponding to a single cell multi-modal element and structure information), a semantic layer (corresponding to the aggregation content of a row, a column or a whole table) and a relationship layer (describing the association relationship between atomic blocks and semantic blocks), realizing integrated retrieval of multi-modal elements and structures and associations to adapt to complex scenarios. With the representation of the content blocks at each level by means of vector coding, the user query vector can be accurately matched with the vectors of the atomic blocks, the semantic blocks and the relationship blocks in the vector database, which not only covers multi-dimensional retrieval requirements such as cell details and row / column aggregation semantics, but also breaks through the limitation of traditional table retrieval which only focuses on the content itself by separately setting the “relationship layer”, supports deep retrieval based on association logic, reduces missed and false detections, and significantly improves the semantic matching accuracy. At the same time, the combination of the multi-level structure and the vector retrieval can also flexibly adapt to the full-scenario user query intention from fine-grained details to medium-grained aggregation and then to association logic, greatly improving the retrieval flexibility and user experience. BRIEF DESCRIPTION OF DRAWINGS
[0012] In order to more clearly illustrate the embodiments of the present application, the drawings needed in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0013] Figure 1 One of the flowcharts of a multi-modal retrieval method provided by an embodiment of the present application; Figure 2 The second flowchart of a multi-modal retrieval method provided by an embodiment of the present application; Figure 3 The third flowchart of a multi-modal retrieval method provided by an embodiment of the present application; Figure 4 The fourth flowchart of a multi-modal retrieval method provided by an embodiment of the present application; Figure 5 The structural diagram of a multi-modal retrieval device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0014] The technical solutions in the embodiments of the present application will be described clearly and completely in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present application.
[0015] It should be noted that in the description of the present application, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that processes, methods, articles or devices including a series of elements not only include those elements, but also include other elements not explicitly listed, or further include elements inherent to such processes, methods, articles or devices. The terms "first", "second" and the like in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence.
[0016] In order to enable those skilled in the art to better understand the present application, the present application will be further described in detail below in combination with the drawings and specific embodiments.
[0017] The embodiments of the present application provide a multi-modal retrieval method. The method is described in detail in combination with the execution process of the multi-modal retrieval method, with reference to Figure 1 The flowchart of the multi-modal retrieval method provided by the present application is shown in FIG. 1. The specific steps include the following: S101, receiving a user query text.
[0018] In some embodiments, the received user query text can be a natural language query text input by a user (for example, "What is the profit of product A in Q3 of 2024?"), and then the query result is generated by further querying the vector database and / or the relational database in combination with the actual demand corresponding to the user query text.
[0019] S102, vector encoding the user query text to obtain a query vector corresponding to the user query text.
[0020] Further, the user query text needs to be vector encoded to convert the query intention in the form of text into a vector form that can be operated by a computer, i.e., a query vector, so that the matching degree between vectors can be measured by vector similarity in the vector database in the subsequent process; finally, the result is queried and generated in the vector database according to the query vector.
[0021] S103, querying in the vector database according to the query vector to generate a query result.
[0022] The vector database stores the vector encoding corresponding to the multi-level content block structure of the plurality of multi-modal table documents.
[0023] Specifically, the multi-level content block structure of any multi-modal table document includes: an atomic layer composed of a plurality of atomic blocks, a semantic layer composed of a plurality of semantic blocks, and a relationship layer composed of a plurality of relationship blocks; the atomic block includes target element content and structure information of any cell in the multi-modal table document; the semantic block includes aggregated content corresponding to any row, or any column, or any whole table in the multi-modal table document; the relationship block includes text describing the association between the plurality of atomic blocks and the plurality of semantic blocks; the structure information includes the row and column indexes corresponding to each cell, and the cross-modal association relationship between the multi-modal elements.
[0024] It should be noted that the vector database in the embodiment of the present application pre-stores a large number of vector encodings of multi-level content block structures of multi-modal tables that may be used, such as atomic layer (single cell content and structure), semantic layer (row / column / whole table aggregated content), and relationship layer (association between content blocks). Therefore, based on the three dimensions of the content of each cell, the aggregated semantics of the row, column, and whole table, and the cross-element association logic, the vector can be accurately matched, and the content most suitable for the user's query intention can be searched, and finally a comprehensive and accurate retrieval result can be generated.
[0025] Specifically, the multi-level content block structure includes three layers, specifically including the following: (1) Atomic layer The atomic layer includes a plurality of atomic blocks. The atomic block can be understood as the smallest carrier of fine-grained information. In units of single cells, the target element content and structure information of each cell are associated and bound to form an independent atomic block.
[0026] For example: Atomic block 1 = [target element content (Markdown format "Product A") + structure information (row 3 column 2 index + no cross-modal association)].
[0027] For example: Atomic block 2 = [target element content (LaTeX format "Profit = Sales × (Unit Price - Cost)") + structure information (row 5 column 4 index + associated row 5 column 2 "Sales", row 5 column 3 "Unit Price" atomic block)].
[0028] By constructing an atomic block for each cell, an atomic layer is formed, thereby preserving the information integrity of the table at the finest granularity. Each atomic block is an independent unit that can be positioned and includes its semantic information, which is the basis for subsequent aggregation and association. At the same time, when querying in the subsequent query, if the information of a specific cell needs to be accurately positioned, the atomic block can be directly matched quickly.
[0029] (2) Semantic layer The semantic layer includes a plurality of semantic blocks, which are formed by aggregating atomic blocks in three dimensions of row, column and table based on row indexes and column indexes in the structure information, and retaining the association relationship between the corresponding atomic blocks.
[0030] It should be noted that aggregation refers to the process of integrating and summarizing a plurality of dispersed and fine-grained information units (e.g., atomic blocks) into an information set (i.e., semantic block) with more overall semantics according to a specific rule (e.g., the structural logic of row, column and table).
[0031] Among them, the row-level semantic block is aggregated with all atomic blocks in the same row, for example: row 3 semantic block = atomic block set of row 3 column 1-column 5, and thus can describe the complete semantics of a row. When the data of a certain row is for a plurality of cells of product A, the semantic information corresponding to the row-level semantic block corresponding to the row can include "name, unit price, sales, cost and profit of product A".
[0032] The column-level semantic block is aggregated with all atomic blocks in the same column, for example: column 2 semantic block = atomic block set of column 2 row 1-row 10, and thus can describe the dimension semantics of a column, such as the unit price data of all products.
[0033] The table-level semantic block is aggregated with the row-level and column-level semantic blocks of the whole table, and records the total range of the rows and columns of the table, for example, "row 1-row 10, column 1-column 5", which can describe the global semantics of the whole table, such as the product sales data table of Q3 (third quarter) of 2024.
[0034] (3) Relationship layer The relationship layer includes a plurality of relationship blocks, which are formed by describing the association logic between atomic blocks and semantic blocks with natural language text based on the cross-modal association relationship in the structure information.
[0035] For example: relationship block 1 = [atomic block (row 5 column 4 formula) references atomic block (row 5 column 2 sales), atomic block (row 5 column 3 unit price) - association type: data dependency].
[0036] For another example: relationship block 2 = [row 3 semantic block (product A data) is associated with column 5 semantic block (all product profit data) - association type: data attribution].
[0037] Through the relationship blocks in the relationship layer, the association logic of the multi-modal elements is established, so that the retrieval can extend from a single block to an associated block (for example, when searching for a formula, the original data cell referenced by the formula can be traced back).
[0038] Specifically, the construction process of the multi-level content block structure includes the following steps 1 to 3: Step 1, parse a plurality of multi-modal table documents to obtain target element content and structure information corresponding to each cell in each multi-modal table document.
[0039] The structure information includes the row and column indexes corresponding to each cell, and the cross-modal association relationship between multi-modal elements.
[0040] Specifically, the multi-modal table document refers to a table whose multiple cells contain multiple data types (modalities), for example, the multi-modal table document can be an Excel format table document; the multi-modal can be different modalities including text, pictures, formulas, etc. The text element can be understood as ordinary numbers and character content; the picture element is an icon, schematic diagram, photo, etc. embedded in the cell; the formula element is a mathematical formula, chemical equation, etc. For example, the multi-modal table document can be a data table in a scientific research report, a corporate financial statement, a product parameter table, etc.
[0041] It should be noted that, due to the non-uniform format of multi-modal elements in the multi-modal table document, for example, the text element can carry differentiated formats such as bold, line break, and special symbol coding, the picture element can be embedded in different formats such as JPG and PNG without textual description, and the formula element can be a "picture formula" in the form of a screenshot or rely on a specific software (such as Math Type) for exclusive coding. Directly preserving the original format will cause difficulties in identifying the same content and associating non-text elements during subsequent parsing; therefore, the elements in the multi-modal table document need to be converted in format, and ultimately the target element content format is unified and semantically, laying a foundation for subsequent construction of multi-level content blocks.
[0042] Therefore, the data format is unified based on steps 11-14 in the embodiments of the present application to obtain the target element content corresponding to each cell: Step 11, parse a plurality of multi-modal table documents to extract initial element content in each cell in each multi-modal table document.
[0043] The initial element content includes at least one of a text element, a picture element, and a formula element.
[0044] Specifically, the initial element content in each cell, i.e. the text element, the picture element, and the formula element, can be extracted from the plurality of multi-modal table documents by a table parsing tool (such as openpyxl for Excel, Tabula for PDF, etc.).
[0045] Step 12, when the initial element content is a text element, converting the format of the text element into a target lightweight markup language format to generate corresponding target element content.
[0046] In some embodiments, the target lightweight markup language format can be Markdown, which converts raw text from different software and formats into a unified, concise, and parsable standard format, eliminating format interference while preserving key semantics and necessary format information. For example, use "\n" to represent line breaks, thereby eliminating format redundancy and differences, focusing on core semantics, and facilitating subsequent vector encoding and semantic matching.
[0047] Step 13, when the initial element content is a picture element, performing semantic understanding on the picture element to generate a textual description corresponding to the picture element, thereby generating the corresponding target element content.
[0048] Specifically, through an image understanding model, such as a CLIP-based image-text association model, a textual description is generated for icons, photos, etc. within the cell, for example, the textual description of the picture can be "male, dark short hair, wearing a blue shirt, background is white", converting visual information that cannot directly participate in semantic calculation into text information, enabling the picture to establish semantic association with other elements (such as corresponding text data).
[0049] Step 14, when the initial element content is a formula element, converting the formula element into a text in the target syntax format, thereby generating the corresponding target element content.
[0050] Wherein, the target syntax format can be LaTeX structured syntax text, thereby uniformly converting the formula into LaTeX structured syntax format, preserving the mathematical logic of the formula and achieving cross-platform compatibility, facilitating subsequent parsing of formula meaning and associated cell data.
[0051] The embodiments of the present application convert text into a unified lightweight markup language format and formulas into standardized target syntax text. At the same time, the semantic associability of non-text elements such as pictures and formulas is activated, laying a foundation for high-quality data for subsequent extraction of structural information, construction of multi-level content blocks, and database retrieval.
[0052] It should be further noted that the structural information is information that defines the spatial position of elements and the association relationship between elements in the multi-modal table document, and integrates scattered multi-modal elements into a logical and traceable whole.
[0053] Specifically, the generation process of the structural information can refer to the following steps: For any one of the plurality of multi-modal table documents, extract the row and column indexes of each cell in the multi-modal table document in the multi-modal table document file, and the cross-modal association relationship between the text elements, picture elements and formula elements corresponding to each cell, to generate the structural information.
[0054] Wherein, the row and column indexes can explicitly specify the specific coordinates of each cell in the table (for example, row 3 column 2), providing a unique spatial identifier for all multi-modal elements; the cross-modal association relationship between the text elements, picture elements and formula elements corresponding to each cell, that is, can be understood as the logical reference relationship between different modal elements, and then the semantic association logic of the multi-modal elements is established. This lays a foundation for subsequent sorting of the association logic between atomic blocks and semantic blocks, and for associated traceability query when querying.
[0055] Step 2, for each multi-modal table document, based on the target element content and structure information corresponding to each cell, a multi-level content block structure corresponding to the multi-modal table document is constructed.
[0056] Further, according to the target element content and structure information corresponding to each cell of each multi-modal table document obtained in step 1, a multi-level content block structure corresponding to the multi-modal table document is constructed.
[0057] Specifically, the explanation of the constructed multi-level content block structure can refer to the explanation in the above step S103, which will not be repeated here.
[0058] Step 3, vector encoding is performed on the multi-level content blocks, and the vector database is stored. In this step, the atomic blocks, semantic blocks, and relationship blocks, i.e., the multi-level content block structure constructed in the foregoing, are vector encoded. Specifically, an embedded model (such as BERT, CLIP, etc.) can be used to convert the “textual information” (such as Markdown text of atomic blocks, textual description of pictures, LaTeX text of formulas, and association description of relationship blocks) in the atomic blocks, semantic blocks, and relationship blocks into high-dimensional feature vectors. These vectors can accurately capture the semantic features of the content (for example, the vectors of “product A sales” and “product A sales” will be highly similar); the generated vectors and corresponding “block identifiers” (such as “atomic block-table 1-row 3-column 2”, “row-level semantic block-table 1-row 5”) are stored in a vector database (such as Milvus, Chroma, etc.).
[0059] Further, when the user inputs a natural language query text (for example, “find product A sales data”), the system can also convert the query text into a vector, calculate the similarity between the query vector and the content block vector stored in the vector database, and quickly locate the semantically associated content block without relying on accurate keywords or formats.
[0060] The present application can completely cover the content of each cell, the aggregation semantics of the row and column of the whole table, and the cross-element association logic in the multi-modal table by constructing a multi-level content block structure including an atomic layer (corresponding to a single cell multi-modal element and structure information), a semantic layer (corresponding to the aggregation content of a row, a column or a whole table), and a relationship layer (describing the association relationship between atomic blocks and semantic blocks), realizing integrated retrieval of multi-modal elements and structures, and association to adapt to complex scenarios. With the representation of the content blocks at each level by vector encoding, the user query vector can be accurately matched with the vectors of the atomic blocks, the semantic blocks and the relationship blocks in the vector database, covering multi-dimensional retrieval requirements such as cell details and row / column aggregation semantics, breaking through the limitation of traditional table retrieval which only focuses on the content itself through the separately set “relationship layer”, supporting deep retrieval based on association logic, reducing missed and false detections, and significantly improving the semantic matching accuracy. At the same time, the combination of the multi-level structure and vector retrieval can also flexibly adapt to the full-scenario user query intent from fine-grained details to medium-grained aggregation and then to association logic, greatly improving the retrieval flexibility and user experience.
[0061] As an extension and refinement of the above-mentioned embodiments, referring to FIG. 8, the present application also provides another multi-modal retrieval method, and the specific steps include the following: Figure 2 S201, receiving a user query text.
[0062] S202, vector encoding the user query text to obtain a query vector corresponding to the user query text.
[0063] S203, inputting the user query text into a target semantic recognition model to perform intent recognition on the semantics of the user query text by the target semantic recognition model, and obtaining an intent recognition result.
[0064] Specifically, the user query text (for example, “find cost accounting formula” “statistical product AQ3 profit”) can be input into the target semantic recognition model (usually a text classification / intent recognition model based on BERT / LLaMA), and then the intent recognition result of the user query text output by the model is obtained. Specifically, the intent recognition result of the user can be divided into two categories, namely, a fuzzy semantic intent and a refined semantic intent. The fuzzy semantic intent refers to no explicit data condition / calculation requirement, and focuses on semantic association matching (for example, “which content and sales trend graph are related” “what are the formulas for cost accounting”); the refined semantic intent refers to containing explicit data conditions, calculation targets or location information, and focusing on “precise data extraction / operation” (for example, “products with profit > 100,000 in table 1” “total sales of product A in Q3 of 2024”).
[0065] S204, querying the vector database and the relational database according to the intent recognition result and the query vector to generate a query result.
[0066] Among them, the relational database stores the structured content of multiple multimodal table documents.
[0067] In this embodiment of the application, it is also necessary to extract information with clear structured features from the multimodal table. These structured features may include: basic table information, namely table ID, table name, and total number of rows / columns; cell structured data, namely row / column index, target element content (e.g., standardized plain text, formula calculation results, image description keywords), and the table ID to which it belongs; content block association information, namely the affiliation relationship between atomic blocks and semantic blocks (e.g., "atomic block-table1-row3column2 belongs to row-level semantic block-table1-row3"), and the identifiers and types of the associated parties in the relationship block records.
[0068] Then, the structured information is stored in a relational database (such as MySQL or PostgreSQL) according to the structure of "table-row-column-cell-content block".
[0069] By using SQL statements to perform conditional filtering (e.g., "filter products in Table 1 with a unit price > 20 yuan"), "data aggregation" (e.g., "statistics the total sales of all products in Table 1"), and "relational queries" (e.g., "query all atomic blocks under the semantic block in row 3"), we can ensure an efficient response to "precise data manipulation needs".
[0070] Therefore, we can then combine the characteristics of both vector databases and relational databases. Relational databases can compensate for the insufficient semantic matching accuracy of vector databases, while vector databases can overcome the limitation of traditional retrieval requiring precise keywords for matching. This combination can cover a variety of user needs, from vague queries to precise data calculations.
[0071] Specifically, refer to Figure 3 As shown, the detailed steps of step S204 (based on the intent recognition result and query vector, combined with the vector database and relational database to perform a query and generate query results) include the following: S2041. Determine whether the intent recognition result contains the target requirement.
[0072] The target requirement is any one of the following: data aggregation, conditional filtering, or multi-table join requirements, which requires a high degree of precision in data querying.
[0073] In S2041 above, if the intent recognition result contains the target requirement, then S2042 is executed as follows; if the intent recognition result does not contain the target requirement, then S2043 is executed as follows: S2042, query the vector database based on the query vector to obtain a first data query result, and obtain a structured query statement corresponding to the user query text, query the relational database through the structured query statement to obtain a second data query result; and integrate the first data query result and the second data query result to generate a query result.
[0074] When the intent recognition result is a query with both fuzzy semantic requirements and precise requirements, for example, "How much is the profit of product A in Q3 of 2024?", it indicates that the current user requirements contain explicit precise data extraction, condition filtering or calculation, and therefore, the profit data needs to be extracted precisely and the structured query formula information needs to be obtained. In order to improve the accuracy when dealing with precise requirements, at this time, the vector database also needs to be queried, and then the two databases are queried to further ensure the accuracy of the query result.
[0075] Specifically, the user query text needs to be vector encoded first to generate a query vector consistent with the vector dimension of the content block in the vector database; by calculating the similarity of the query vector and the atomic block, semantic block, and relationship block vector, the top N content blocks with the highest semantic correlation (such as "product AQ3 profit related row level semantic block" "profit calculation formula atomic block") are selected and integrated into the first data query result, i.e. the first data query result is to lock the semantic range related to the precise requirements, and then by combining the second data query result obtained by querying the relational database, the query result is more comprehensive and accurate.
[0076] Again, for the precise requirements in the query, a large language model (such as GPT-3.5 / 4) is called to convert the natural language description of the precise requirements (such as "extract product A profit data in Q3 of 2024" "filter products with unit price > 200 yuan and sales") into structured query statements (such as SQL) executable by the relational database; through the statement, the structured data of the matching conditions (such as "product A profit in Q3 of 2024: 5000 yuan" "products with unit price > 200 yuan: product B (sales 120 pieces), product C (sales 90 pieces)") is accurately extracted in the relational database as the second data query result.
[0077] Finally, the first data query result (semantically associated content block description, such as formula source, data association) and the second data query result (precise structured data) are integrated according to the logic of "core data first, then associated background". For example: first show the precise data: "Product A's profit in Q3 2024 is 5000 yuan (data source: table 1 row 5 column 5)"; then supplement the semantically associated content: "the corresponding profit calculation formula is atomic block-table 1 row 5 column 4 (LaTeX text: \text{profit}=\text{sales}\times(\text{unit price}-\text{cost}) ), which references table 1 row 5 column 2 (sales 100 pieces) and row 5 column 3 (unit price 80 yuan) data".
[0078] S2043, based on the query vector, query in the vector database, obtain the first data query result, and take the first data query result as the query result.
[0079] When the intent recognition result is that the query is only a fuzzy semantic requirement without precise data filtering / calculating requirement, for example, "find the formula and data source description related to cost accounting", only the associated content needs to be obtained through semantic similarity matching.
[0080] The first step of S2042 can be referred to for vector encoding of the user query text to generate a query vector, calculate the similarity in the vector database, and filter out the top N content blocks with the highest semantic association degree (such as "atomic block of cost accounting formula", "picture description block corresponding to sales trend chart", and "relationship block of profit calculation logic"), and integrate them into the first data query result.
[0081] Since there is no precise data requirement for the query, there is no need to trigger the large language model to generate a structured query statement and a relational database query process, and the first data query result is directly sorted according to "similarity from high to low" and output. For example: "1. Atomic block-table 1 row 3 column 4: cost accounting formula, LaTeX text: \text{cost}=\text{raw material cost}+\text{labor cost}; 2. Relationship block-table 1-002: this formula references table 1 row 3 column 2 (raw material cost) and row 3 column 3 (labor cost) atomic blocks; 3. Semantic block-table 2 row 6: product sales trend chart description-2024Q1-Q4 product A sales line chart, Q3 reaches the peak".
[0082] Through the subdivision of the query intent, differentiated retrieval strategies are adopted to find the optimal solution between precision, efficiency, and result integrity, and finally output the query result that meets the user's real needs.
[0083] In the embodiments of the present application, for a query containing target requirements, the collaborative strategy of combining vector library semantic positioning and relational database accurate extraction is used to ensure that the results are associated with the query semantics and the data is accurate; for a query not containing target requirements, only vector library retrieval is used to simplify the process, reduce redundant calculations, and realize resource on-demand allocation. At the same time, in the accuracy requirement scenario, the associated content of the vector library and the accurate data of the relational database are integrated to avoid the problems of traditional retrieval that accurate data lacks background and associated content lacks accurate values; in the fuzzy requirement scenario, semantic matching is focused on to ensure that the results are highly related to the query intent.
[0084] Further, the processes of obtaining the first data query result and the second data query result involved in S2042 and S2043 can refer to the following steps A-E: Step A: Convert the user query text into a corresponding query vector.
[0085] The user query text, i.e., the natural language text, is converted into a computer-calculable vector form, so that the query requirements can be compared with the content block vectors in the vector database at the semantic level.
[0086] Specifically, the same embedded model (e.g., BERT, CLIP) as S203 (content block vector encoding) can be used to convert the user query text (e.g., "Product A's profit in Q3 of 2024 and calculation formula") into a query vector with the same dimension. This ensures that the query vector and the content block vector are consistent in the semantic space, and thus the vectors generated based on the same model have reference value in subsequent similarity calculation.
[0087] Step B: Calculate the similarity scores between the query vector and the plurality of content block vectors stored in the vector database.
[0088] This step is to quantify the semantic association degree between the user query requirements and the content blocks in the vector database, and to find the content block that best fits the query intent.
[0089] Specifically, a vector similarity calculation algorithm (e.g., cosine similarity, Euclidean distance) can be used to calculate the similarity scores between the query vector and all atomic block vectors, semantic block vectors, and relational block vectors in the vector database (the score range is usually 0-1, and the closer to 1, the more similar the semantics).
[0090] For example, the similarity score between the query vector "Product A's profit in Q3 of 2024" and the "row-level semantic block - Table 1 - Row 5 (Product A Q3 data)" is 0.92, and the similarity score between the query vector and the "column-level semantic block - Table 1 - Column 5 (all product profits)" is 0.85.
[0091] Step C, rank the plurality of content blocks according to the similarity scores, select the top pre-set threshold number of content blocks with the highest similarity scores, and generate a first data query result based on the pre-set threshold number of content blocks.
[0092] Further, the most semantically relevant content is filtered from a large number of content blocks to form the first data query result. That is, all content blocks can be ranked from high to low according to the similarity scores; by setting a pre-set threshold (for example, the top 5), the top N content blocks with the highest scores are selected; the key information (such as block identifier, textual content, and association description) of these content blocks is extracted and integrated into the first data query result.
[0093] For example, when the pre-set threshold is 3, the final selection is "row-level semantic block-table 1-row 5", "atomic block-table 1-row 5-column 5 (profit data)", and "relationship block-table 1-001 (profit formula association description)", as the first data query result.
[0094] Step D, input the user query text into a large language model to generate a structured query statement corresponding to the user query text.
[0095] Specifically, this step converts the precise requirements of the user query text into a structured query statement (such as SQL) that can be directly executed by a relational database. The user query text can be input into a large language model with structured query statement conversion function, so that the model generates a corresponding structured query statement according to the precise requirements in the query (such as "extract the profit of product A in Q3 2024" and "query the cost accounting formula").
[0096] For example, when the user query text is "What is the profit of product A in Q3 2024", the SQL statement generated by the model can be represented as: SELECT target element content FROM cell table WHERE table ID = 1 AND row index = 5 AND column index = 5 AND target element content LIKE '% product A in Q3 2024 profit %'.
[0097] Step E, based on the structured query statement, query in the relational database to obtain a second data query result.
[0098] Further, the structured query statement is used to accurately extract the quantitative data and specific content required by the user from the relational database to form the second data query result.
[0099] Specifically, the structured query statement generated in step D is submitted to the relational database, and the database returns the results that completely match the conditions after executing the statement as the second data query result.
[0100] For example, after executing the above SQL statement, the "target element content: product A profit in Q3 2024 - 5000 yuan" is returned as the second data query result.
[0101] The embodiment of the application converts the user query into a vector and calculates the similarity with the content block vector, which can break through the limitations of keyword matching, accurately identify semantically related content, and further filter the pre-set threshold number of content blocks to ensure that the first data query result focuses on core associated information and avoids irrelevant content interference.
[0102] At the same time, the user's natural language query is converted into a structured query statement by a large language model, and based on this statement, the structured query statement is retrieved in a relational database to directly obtain accurate data that meets the conditions; further, the first data query result provides semantically associated background information (such as data sources and associated relationships), and the second data query result provides accurate structured data, which improves the accuracy of the query result.
[0103] As an extension and refinement of the above embodiment, referring to Figure 4 The application also provides a construction process of a multi-level content block structure corresponding to a multi-modal table document in a multi-modal retrieval method, which can refer to the following steps: S401, for each cell in any multi-modal table document, associate the target element content of the cell with the row and column indexes of the cell to generate a plurality of atomic blocks.
[0104] Specifically, for each cell of the multi-modal table, the target element content of each cell is associated and bound with the row index and column index of the cell to generate a plurality of independent atomic blocks.
[0105] Further, each cell information is given a spatial identifier to solve the problem of traditional table information only storing content without positioning, and further subsequent retrieval of individual cell content or tracing of content sources can quickly locate the corresponding atomic block through the row index and column index to ensure traceability of fine-grained information.
[0106] It should be noted that the refinement steps of S401 (for each cell in any multi-modal table document, associate the target element content of the cell with the row and column indexes of the cell to generate a plurality of atomic blocks) include the following steps a1-a3: Step a1, traverse each cell in the multi-modal table document to obtain the target element content corresponding to each cell.
[0107] This step traverses each cell in the multi-modal table document one by one without missing any cell; for the multi-modal elements (text, pictures, formulas, etc.) in the cell, it is uniformly processed as standardized target element content. Avoid the defect of traditional table processing that only extracts text and ignores non-text elements, ensure that multi-modal information is completely collected, provide full and standardized content basis for atomic blocks, and eliminate the problem of ignoring non-text elements in subsequent retrieval.
[0108] Step a2, extract the row and column indexes corresponding to each cell, and associate the table identifier of the multi-modal table document corresponding to the cell.
[0109] For each traversed cell, extract its row index and column index in the table to determine its position in the table; at the same time, it is necessary to additionally associate the table identifier of the multi-modal table document to which the cell belongs. The table identifier is the unique identifier of each multi-modal table, which is used to distinguish tables in different documents.
[0110] In some embodiments, only the row index and column index cannot uniquely locate the cell (for example, "row 3 column 2" may correspond to different contents in "table ID_001" and "table ID_002"), and the present application assigns a globally unique position identifier to each cell by combining the table identifier, row index, and column index, solving the problem of repeated cell positions across tables and inability to accurately trace the source.
[0111] Step a3, bind the target element content, row index, column index, and table identifier of each cell to generate an atomic block.
[0112] Force the binding of the target element content, row index, column index, and "table identifier" of each cell to form an independent atomic block.
[0113] For example, atomic block 1 = [table identifier: table ID_001; row index: row 3; column index: column 2; target element content: Markdown text "Product A, unit price 200 yuan"].
[0114] The present application ensures that all types of cell content such as text, pictures, and formulas can be converted into the target element content of the atomic block, avoiding the exclusion of non-text elements from the retrieval range, allowing subsequent retrieval to cover all information dimensions of the table. At the same time, by associating the index and the table identifier, each atomic block can be accurately located, and when the user queries the atomic block content, the position in the original table can be directly traced back; at the same time, this identifier also provides a key basis for subsequent row / column aggregation to generate semantic blocks.
[0115] S402, based on the row and column indexes of the cells, aggregating the atomic blocks corresponding to the cells by row, column, and the whole table to generate a plurality of semantic blocks corresponding to the row level, the column level, and the table level respectively.
[0116] Further, based on the atomic blocks generated in S401, the plurality of semantic blocks corresponding to the row level, the column level, and the table level respectively are generated according to the row and column indexes of the cells.
[0117] It should be noted that the detailed steps of S402 (based on the row and column indexes of the cells, aggregating the atomic blocks corresponding to the cells by row, column, and the whole table to generate a plurality of semantic blocks corresponding to the row level, the column level, and the table level respectively) include the following steps b1-b3: Step b1, based on the row index of the cell, aggregating the atomic blocks in the same row to generate a row-level semantic block.
[0118] Based on the row index of the cell, all atomic blocks (i.e. information of all cells in a row) in the same table with the same row index are packaged and integrated to form an independent row-level semantic block.
[0119] For example, if the atomic block of the table "row 5" contains "product B (column 1), 180 yuan per piece (column 2), 200 pieces (column 3), 36000 yuan (column 4)", the "row 5 row-level semantic block" after aggregation will uniformly carry the complete single industry business information of "product B's unit price 180 yuan per piece, sales 200 pieces, and revenue 36000 yuan".
[0120] Further, the scattered cell information is linked into a logically associated single row, meeting the user's demand for querying full-dimensional data of a subject (such as a product) (without the need to search each atomic block in the row one by one).
[0121] Step b2, based on the column index of the cell, aggregating the atomic blocks in the same column to generate a column-level semantic block.
[0122] Based on the column index of the cell, all atomic blocks (i.e. information of all cells in a column) in the same table with the same column index are packaged and integrated to form an independent column-level semantic block.
[0123] For example, if the atomic block of the table "column 3" (sales column) contains "150 pieces (row 4), 200 pieces (row 5), 120 pieces (row 6)", the "column 3 column-level semantic block" after aggregation will uniformly carry the summary information of the single column dimension "all product sales data: product A 150 pieces, product B 200 pieces, product C 120 pieces".
[0124] Integrate atomic blocks of the same dimension (such as all unit prices, all sales) into an ordered data sequence to meet the user's global summary demand for a certain dimension (such as all product unit prices) without traversing each row of the corresponding atomic block.
[0125] Step b3, integrate all row-level semantic blocks and column-level semantic blocks in the same table to generate a table-level semantic block.
[0126] This step takes table ownership as the aggregation basis to further integrate all row-level semantic blocks and all column-level semantic blocks within the same multi-modal table, extract the "core global information" of the table, and form a "table-level semantic block".
[0127] For example, a certain "2024Q4 product sales table" contains 6 row-level semantic blocks (corresponding to 6 products) and 4 column-level semantic blocks (product name, unit price, sales, and revenue). The aggregated "table-level semantic block" will carry the global overview of "2024Q4 product sales table: a total of 6 products, covering unit price 150-220 yuan / piece, sales 100-250 pieces, and revenue 15000-55000 yuan, with the core data dimensions of product name, unit price, sales, and revenue".
[0128] The embodiments of the present application integrate row and column semantics into a global correlation body through row-level semantic blocks, column-level semantic blocks, and table-level semantic blocks. This correlation allows the subsequent vector encoding to capture contextual semantics and improve semantic matching accuracy.
[0129] S403, based on the cross-modal association relationship, extracting the association relationship between atomic blocks and semantic blocks to generate a plurality of relationship blocks.
[0130] Based on the cross-modal association relationship parsed in the early stage, the association logic between atomic blocks and semantic blocks is extracted and described to generate relationship blocks. The association logic of multi-modal elements is established, and then during subsequent retrieval, the user can trace the original data block referenced by the formula, and associate the text block that explains the picture.
[0131] It should be noted that the detailed steps of the above S403 (based on the cross-modal association relationship, extracting the association relationship between atomic blocks and semantic blocks to generate a plurality of relationship blocks) include the following step c1: Step c1, based on the cross-modal association relationship in the structure information, extracting the association relationship between different atomic blocks and the association relationship between different semantic blocks, and describing the association relationship in text form to construct a relationship block.
[0132] Based on the cross-modal association relationship in the previously parsed structural information, these associations are the implicit logic that naturally exists in multi-modal tables (such as "the text cell 'data source' points to the sales trend chart of the picture cell B4" "the formula cell C5 references the sales data cell A5 and the unit price data cell B5" "the row-level semantic block 'product A data' corresponds to the third data of the column-level semantic block 'profit column'").
[0133] Specifically, focusing on the cross-modal / cross-location association at the level of individual cells, such as "the text atomic block (row 2 column 1:'sales trend chart') and the picture atomic block (row 2 column 2: '2024Q1-Q4 line chart') are in a 'description and described' relationship" "the formula atomic block (row 5 column 4: 'profit = sales × (unit price - cost)') and the data atomic block (row 5 column 2: sales, row 5 column 3: unit price) are in a 'data reference' relationship".
[0134] And focusing on the association at the aggregation level, such as "the row-level semantic block (row 3: product A data) is subordinate to the column-level semantic block (column 5: profit column)" "the table-level semantic block (table 1: 2024Q3 sales data) contains the row-level semantic block (rows 3-8: 6 product data) and the column-level semantic block (columns 1-5: 5 indicator dimensions)".
[0135] Furthermore, the extracted association relationships are converted into standardized text (such as natural language descriptions or structured statements) to form independent "relationship blocks".
[0136] For example, a relationship block representing the relationship between atomic blocks can be represented as: [association type: data reference; association parties: formula atomic block (table 1-row 5-column 4) → data atomic block (table 1-row 5-column 2, table 1-row 5-column 3); corresponding text form description: the formula calculates profit by referencing sales and unit price data].
[0137] For another example, a relationship block representing the relationship between semantic blocks can be represented as: [association type: subordinate inclusion; association parties: row-level semantic block (table 1-row 3) → column-level semantic block (table 1-column 5); description: the profit data of product A is the third item of data in the profit column].
[0138] Furthermore, in traditional multi-modal tables, the correspondence between pictures and text, the reference between formulas and data, and other associations exist only in the table layout and are not explicitly recorded, so users cannot directly query which data a formula references or what the textual explanation of a picture is. The relationship blocks of the present application make these associations explicit through textual descriptions, and in subsequent retrieval, the system can locate the relevant relationship blocks through vector matching to achieve multi-dimensional information tracing.
[0139] S404, constructing a multi-level content block structure corresponding to the multi-modal table document according to the plurality of atomic blocks, the plurality of semantic blocks, and the plurality of relationship blocks.
[0140] Further, the atomic blocks generated in S401, the semantic blocks generated in S402, and the relationship blocks generated in S403 are integrated to form a three-layer structure of an atomic layer, a semantic layer, and a relationship layer, that is, a multi-level content block structure corresponding to the multi-modal table.
[0141] The three kinds of dispersed information units are converted into a hierarchical, logical, and linkable three-dimensional framework. The atomic layer is the foundation, the semantic layer is the aggregation and extension, and the relationship layer is the association link. The three together support the subsequent needs of granular retrieval and association tracing.
[0142] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software and a necessary general hardware platform, of course, it can also be implemented by hardware, but in many cases the former is a better embodiment.
[0143] The embodiments of the present application also provide a multi-modal retrieval device, which corresponds to the method. Figure 5 A structural schematic diagram of a multi-modal retrieval device 500 provided by the present disclosure is provided. The device 500 of the present embodiment includes: A receiving unit 51 is configured to receive a user query text. An encoding unit 52 is configured to perform vector encoding on the user query text to obtain a query vector corresponding to the user query text. A query unit 53 is configured to perform a query in a vector database according to the query vector to generate a query result. The vector database stores vector encodings corresponding to multi-level content block structures corresponding to a plurality of multi-modal table documents. The multi-level content block structure of any of the multi-modal table documents includes an atomic layer composed of a plurality of atomic blocks, a semantic layer composed of a plurality of semantic blocks, and a relationship layer composed of a plurality of relationship blocks. The atomic block includes target element content and structure information of any cell in the multi-modal table document. The semantic block includes aggregated content corresponding to any row, or any column, or any entire table in the multi-modal table document. The relationship block includes text describing the association between a plurality of atomic blocks and a plurality of semantic blocks. The structure information includes row and column indexes corresponding to each cell, and cross-modal association relationships between multi-modal elements.
[0144] As an optional implementation of the embodiment of the present application, the multi-modal retrieval device further comprises a construction unit, specifically configured to parse a plurality of multi-modal table documents, obtain target element content and structure information corresponding to each cell in each multi-modal table document, and construct a multi-level content block structure corresponding to the multi-modal table document based on the target element content and structure information corresponding to each cell for each multi-modal table document.
[0145] As an optional implementation of the embodiment of the present application, the query unit 53 is specifically configured to input the user query text into a target semantic recognition model, perform intent recognition on the semantics of the user query text through the target semantic recognition model, obtain an intent recognition result, query the vector database and the relational database according to the intent recognition result and the query vector, and generate the query result. The relational database stores structured content of a plurality of multi-modal table documents.
[0146] As an optional implementation of the embodiment of the present application, the query unit 53 is specifically configured to determine whether the intent recognition result contains a target requirement. The target requirement is any one of data aggregation, condition filtering, and multi-table association requirement. If yes, the query unit 53 performs a query in the vector database based on the query vector to obtain a first data query result, obtains a structured query statement corresponding to the user query text, performs a query in the relational database through the structured query statement to obtain a second data query result, integrates the first data query result and the second data query result, and generates the query result. If no, the query unit 53 performs a query in the vector database based on the query vector to obtain the first data query result, and takes the first data query result as the query result.
[0147] As an optional implementation of the embodiment of the present application, the query unit 53 is specifically configured to input the user query text into a large language model to generate a structured query statement corresponding to the user query text, and perform a query in the relational database based on the structured query statement to obtain the second data query result.
[0148] As an optional implementation of the embodiment of the present application, the query unit 53 is specifically configured to perform vector conversion on the user query text to generate a corresponding query vector, calculate a similarity score between the query vector and a plurality of content block vectors stored in the vector database, sort the plurality of content blocks according to the similarity score, select a preset threshold number of content blocks with the highest similarity, and generate the first data query result based on the preset threshold number of content blocks.
[0149] As an optional implementation of the embodiment of the present application, the construction unit is specifically configured to parse a plurality of multi-modal table documents, and extract initial element content in each cell in each of the multi-modal table documents. The initial element content includes at least one of a text element, a picture element, and a formula element. When the initial element content is the text element, the format of the text element is converted into a target lightweight markup language format to generate corresponding target element content. When the initial element content is the picture element, semantic understanding is performed on the picture element to generate a textual description corresponding to the picture element to generate corresponding target element content. When the initial element content is the formula element, the formula element is converted into a text in a target syntax format to generate corresponding target element content.
[0150] As an optional implementation of the embodiment of the present application, the construction unit is specifically configured to, for any one of the plurality of multi-modal table documents, extract a row and column index of each cell in the multi-modal table document in a multi-modal table document file, and a cross-modal association relationship between a text element, a picture element, and a formula element corresponding to the each cell to generate the structure information.
[0151] As an optional implementation of the embodiment of the present application, the construction unit is specifically configured to, for each cell in any of the multi-modal table documents, associate the target element content of the cell with the row and column index of the cell to generate a plurality of atomic blocks. The atomic blocks corresponding to the cell are aggregated according to the row and column index of the cell to generate a plurality of semantic blocks corresponding to rows, columns, and tables, respectively. The association relationship between the atomic blocks and the semantic blocks is extracted based on the cross-modal association relationship to generate a plurality of relationship blocks. The multi-level content block structure corresponding to the multi-modal table document is constructed according to the plurality of atomic blocks, the plurality of semantic blocks, and the plurality of relationship blocks.
[0152] The features of the embodiments of the multi-modal retrieval device can be referred to the related descriptions of the embodiments of the multi-modal retrieval method, which will not be repeated here.
[0153] Embodiments of the present application also provide an electronic device, comprising a memory and a processor, the memory storing a computer program, and the processor is configured to run the computer program to perform the steps in any of the above multi-modal retrieval method embodiments.
[0154] Embodiments of the present application also provide a computer readable storage medium storing a computer program, wherein the computer program is configured to perform the steps in any of the above multi-modal retrieval method embodiments when running.
[0155] In an example embodiment, the above computer readable storage medium can include, but is not limited to, a U disk, a Read-Only Memory (ROM), a Random Access Memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.
[0156] Embodiments of the present application also provide a computer program product comprising a computer program, wherein the computer program is executed by a processor to implement the steps in any of the above multi-modal retrieval method embodiments.
[0157] Embodiments of the present application also provide another computer program product comprising a non-volatile computer readable storage medium storing a computer program, wherein the computer program is executed by a processor to implement the steps in any of the above multi-modal retrieval method embodiments.
[0158] The skilled person can further realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be realized in electronic hardware, computer software or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been described in general terms in the above description. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0159] The above provides a detailed description of a multi-modal retrieval method and device. The principles and implementation methods of the present application are described by applying specific examples. The above example descriptions are only used to help understand the method and its core idea. It should be noted that for those skilled in the art, without departing from the principles of the present application, some improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A multi-modal retrieval method, characterized by, The method comprises the following steps: receiving a user query text; vector encoding the user query text to obtain a query vector corresponding to the user query text; querying in a vector database according to the query vector to generate a query result; the vector database stores vector encodings corresponding to multi-level content block structures of a plurality of multi-modal table documents; any multi-level content block structure of the multi-modal table document comprises an atomic layer composed of a plurality of atomic blocks, a semantic layer composed of a plurality of semantic blocks, and a relationship layer composed of a plurality of relationship blocks; the atomic block comprises target element content and structure information of any cell in the multi-modal table document; the semantic block comprises aggregated content corresponding to any row, any column, or any entire table in the multi-modal table document; the relationship block comprises text describing the association between a plurality of atomic blocks and a plurality of semantic blocks; the structure information comprises row and column indexes corresponding to each cell, and cross-modal association between multi-modal elements.
2. The method of claim 1, wherein, Before receiving the user query text, the method further comprises: parsing a plurality of multi-modal table documents to obtain target element content and structure information corresponding to each cell in each multi-modal table document; for each multi-modal table document, constructing a multi-level content block structure corresponding to the multi-modal table document based on the target element content and structure information corresponding to each cell; vector encoding the multi-level content block and storing it in the vector database.
3. The method of claim 1, wherein, After vector encoding the user query text to obtain a query vector corresponding to the user query text, the method further comprises: inputting the user query text into a target semantic recognition model to perform intent recognition on the semantics of the user query text through the target semantic recognition model to obtain an intent recognition result; querying in the vector database and a relational database according to the intent recognition result and the query vector to generate the query result; the relational database stores structured content of a plurality of multi-modal table documents.
4. The method of claim 3, wherein, The querying in the vector database and the relational database according to the intent recognition result and the query vector to generate the query result comprises: determining whether the intent recognition result contains a target requirement; the target requirement is any one of data aggregation, condition filtering, and multi-table association requirement; if yes, querying in the vector database based on the query vector to obtain a first data query result, obtaining a structured query statement corresponding to the user query text, querying in the relational database based on the structured query statement to obtain a second data query result, and integrating the first data query result and the second data query result to generate the query result; if no, querying in the vector database based on the query vector to obtain the first data query result, and taking the first data query result as the query result.
5. The method of claim 4, wherein, The query vector is queried in the vector database based on the query vector, and a first data query result is obtained, including: The user query text is converted into a vector to generate a corresponding query vector; Calculate the similarity score between the query vector and the plurality of content block vectors stored in the vector database; According to the similarity score, the plurality of content blocks are sorted, the top preset threshold of the content blocks with the highest similarity is selected, and the first data query result is generated based on the preset threshold of the content blocks.
6. The method of claim 4, wherein, The second data query result is obtained by querying the relational database through the structured query statement, including: The user query text is input into a large language model to generate a structured query statement corresponding to the user query text; Based on the structured query statement, the relational database is queried to obtain the second data query result.
7. The method of claim 2, wherein, The plurality of multi-modal table documents are parsed, including: Parse a plurality of multi-modal table documents to extract initial element content in each cell of each multi-modal table document; The initial element content includes at least one of a text element, a picture element, and a formula element; When the initial element content is the text element, the format of the text element is converted into a target lightweight markup language format to generate corresponding target element content; When the initial element content is the picture element, the picture element is semantically understood to generate a textual description corresponding to the picture element to generate corresponding target element content; When the initial element content is the formula element, the formula element is converted into a text in a target syntax format to generate corresponding target element content.
8. The method of claim 2, wherein, The plurality of multi-modal table documents are parsed, including: For any one of the plurality of multi-modal table documents, extract the row and column indexes of each cell in the multi-modal table document file, and the cross-modal association relationship between the text element, picture element and formula element corresponding to the each cell to generate the structure information.
9. The method of claim 2, wherein, For each multi-modal table document, the multi-level content block structure corresponding to the multi-modal table document is constructed based on the target element content and structure information corresponding to each cell, including: For each cell in any multi-modal table document, associate the target element content of the cell with the row and column indexes of the cell to generate a plurality of atomic blocks; Based on the row and column indexes of the cell, the atomic blocks corresponding to the cell are aggregated by row, column and table to generate a plurality of semantic blocks corresponding to the row level, column level and table level respectively; Based on the cross-modal association relationship, the association relationship between the atomic blocks and the semantic blocks is extracted to generate a plurality of relationship blocks; According to a plurality of atomic blocks, a plurality of semantic blocks and a plurality of relationship blocks, a multi-level content block structure corresponding to the multi-modal table document is constructed.
10. An electronic device, comprising: It includes: A memory for storing a computer program; a processor for implementing the multi-modal search method according to any one of claims 1-9 when executing the computer program.
Citation Information
Patent Citations
Multi-modal document retrieval enhancement generation method based on large model
CN119988588A
Multi-source heterogeneous data knowledge base system construction method, equipment and medium
CN120386896A
Data analysis question and answer platform based on large model and knowledge vector library
CN120448508A
Browsing knowledge on the basis of semantic relations
US20090070322A1
Cited By
Cross-modal retrieval method and device based on building specification document and medium
CN121524338A
Retrieval enhanced question and answer method and system for mixed mode document content
CN121636682A