Multi-mode fusion data processing method and device, medium, electronic equipment and program product

By receiving multimode data search statements in a multimode fusion database, parsing and translating them into singlemode data search statements, and performing data retrieval and feature fusion processing, the problem of low processing efficiency of multimode databases during data integration is solved, and efficient multimode data processing is achieved.

CN120086273APending Publication Date: 2025-06-03BEIJING VOLCANO ENGINE TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510246518.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

When multi-mode fusion database responds to application services to integrate multiple modal data, the data processing complexity is high and the data processing efficiency is low, making it difficult to meet the service requirements of high timeliness.

Method used

A multimode fusion data processing method is provided, including receiving multimode data search statements, parsing and translating them into singlemode data search statements corresponding to each of multiple modals, respectively, data search in the corresponding singlemode database in the multimode database, obtaining target data of different modalities, and generating a unified multimode data representation through feature fusion and modal alignment processing.

Benefits of technology

It reduces the complexity of data processing, improves data processing efficiency, and can meet the service needs of high-time efficiency requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086273A_ABST
    Figure CN120086273A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-mode fusion data processing method and device, a medium, electronic equipment and a program product, belongs to the technical field of computers, and can reduce the data processing complexity and meet the service requirement of high timeliness when responding to application service to integrate multi-mode data. The multi-mode fusion data processing method comprises the following steps: receiving a multi-mode data retrieval statement, analyzing the multi-mode data retrieval statement, and translating the multi-mode data retrieval statement into single-mode data retrieval statements corresponding to a plurality of modes respectively; according to each single-mode data retrieval statement, performing data retrieval in a corresponding single-mode database in the multi-mode database to obtain target data of different modalities; wherein single-mode databases corresponding to a plurality of modes are maintained in the multi-mode database, and independent data management is carried out on each single-mode database; and performing feature fusion and modal alignment processing on the retrieved target data of different modalities to generate a unified multi-modal data representation, and feeding back the unified multi-modal data representation as a retrieval result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a multi-modal fusion data processing method, apparatus, medium, electronic device, and program product. Background Art

[0002] A multi-modal fusion database is an integrated intelligent database that can perform various tasks such as storing, accessing, and processing different types of data, and support diverse application services.

[0003] Currently, when a multi-modal fusion database integrates multiple-modal data in response to an application service, the data processing complexity is high and the data processing efficiency is low, making it difficult to meet the service requirements with high timeliness requirements. Summary of the Invention

[0004] This Summary of the Invention section is provided to introduce concepts in a brief form that will be described in detail in the subsequent Detailed Description section. This Summary of the Invention section is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to be used to limit the scope of the claimed technical solution.

[0005] In a first aspect, the present disclosure provides a multi-modal fusion data processing method, including: Receiving a multi-modal data retrieval statement, parsing the multi-modal data retrieval statement, and translating it into single-modal data retrieval statements corresponding to each modality; Performing data retrieval in the single-modal databases corresponding to each single-modal data retrieval statement in a multi-modal database to obtain target data of different modalities; wherein, multiple single-modal databases corresponding to each modality are maintained in the multi-modal database, and independent data management is performed for each single-modal database; Performing feature fusion and modality alignment processing on the retrieved target data of different modalities to generate a unified multi-modal data representation, and feeding back the unified multi-modal data representation as a retrieval result.

[0006] In a second aspect, the present disclosure provides a multi-modal fusion data processing apparatus, including: A parsing and translation module, configured to receive a multi-modal data retrieval statement, parse the multi-modal data retrieval statement, and translate it into single-modal data retrieval statements corresponding to each modality; A retrieval module, configured to perform data retrieval in the single-modal databases corresponding to each single-modal data retrieval statement in a multi-modal database to obtain target data of different modalities; wherein, multiple single-modal databases corresponding to each modality are maintained in the multi-modal database, and independent data management is performed for each single-modal database; A feedback module, configured to perform feature fusion and modality alignment processing on the retrieved target data of different modalities to generate a unified multi-modal data representation, and feed back the unified multi-modal data representation as a retrieval result.

[0007] In a third aspect, the present disclosure provides a computer-readable medium, on which a computer program is stored, and when the computer program is executed by a processing device, the steps of any one of the methods in the first aspect are implemented.

[0008] In a fourth aspect, the present disclosure provides an electronic device, including: A storage device, on which a computer program is stored; A processing device, configured to execute the computer program in the storage device to implement the steps of any one of the methods in the first aspect.

[0009] In a fifth aspect, the present disclosure provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of any one of the methods in the first aspect are implemented.

[0010] By adopting the above technical solution, it is possible to receive a multi-modal data retrieval statement, parse the multi-modal data retrieval statement and translate it into single-modal data retrieval statements respectively corresponding to multiple modalities, perform data retrieval in the single-modal databases respectively corresponding to each single-modal data retrieval statement in a multi-modal database to obtain target data of different modalities, wherein the multi-modal database maintains single-modal databases respectively corresponding to multiple modalities and performs independent data management for each single-modal database, and then perform feature fusion and modality alignment processing on the retrieved target data of different modalities to generate a unified multi-modal data representation, and feed back the unified multi-modal data representation as a retrieval result. In this way, when responding to an application service to integrate multiple-modal data, the data processing complexity can be reduced, the data processing efficiency can be improved, and the service requirements with high timeliness requirements can be met.

[0011] Other features and advantages of the present disclosure will be described in detail in the subsequent specific implementation part. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Combined with the drawings and referring to the following specific implementation manners, the above and other features, advantages and aspects of the embodiments of the present disclosure will become more obvious. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic, and the original elements and elements are not necessarily drawn to scale. In the drawings: Figure 1 is a flowchart of a multi-modal fusion data processing method according to an embodiment of the present disclosure.

[0013] Figure 2Shows a schematic diagram of the definition of multimodal SQL retrieval syntax according to an embodiment of the present disclosure and semantic understanding based on a large language model and a multimodal RAG model.

[0014] Figure 3 Shows a schematic diagram of translating an SQL multimodal retrieval statement into an InfluxDB time series retrieval statement and an ElasticSearch retrieval statement through semantic understanding and translation according to an embodiment of the present disclosure.

[0015] Figure 4 Is an architecture diagram of a multimodal fusion data processing architecture according to an embodiment of the present disclosure.

[0016] Figure 5 Is a schematic block diagram of a multimodal fusion data processing device according to an embodiment of the present disclosure.

[0017] Figure 6 Shows a schematic diagram of the structure of an electronic device suitable for implementing an embodiment of the present disclosure. Detailed implementation manners

[0018] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0019] It should be understood that the various steps recorded in the method embodiments of the present disclosure can be executed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.

[0020] As used herein, the term "including" and its variations are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0021] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependence.

[0022] It should be noted that the modifications of "one" and "multiple" mentioned in this disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".

[0023] The names of the messages or information exchanged between multiple devices in the embodiments of this disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0024] It can be understood that before using the technical solutions disclosed in the embodiments of this disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0025] For example, when responding to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that performs the operations of the technical solutions of this disclosure based on the prompt message.

[0026] As an optional but non-limiting implementation manner, the way of sending a prompt message to the user in response to receiving an active request from the user can be, for example, in the form of a pop-up window. The prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0027] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not constitute a limitation on the implementation manner of this disclosure. Other ways that meet relevant laws and regulations can also be applied to the implementation manner of this disclosure.

[0028] At the same time, it can be understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws, regulations and related provisions.

[0029] In the related art, a multi-modal database is usually implemented by adopting an architecture with a single storage engine supporting multiple data models. In this architecture, a general storage engine is designed at the bottom layer of the multi-modal database, and this storage engine can store data of different data models in a unified manner. For example, by abstracting data into basic units (such as byte streams or specific object structures), and then classifying and storing them according to the type tags of the data (the type tags are used to represent whether it is relational data, document data, graph data, etc.). Taking graph data as an example, in the storage engine, both nodes and edges can be regarded as special objects for storage, while the table structure of relational data is transformed into a storage form of object-relational, and document data is stored according to its own hierarchical structure. At the same time, there will be some metadata at the storage level to identify the type and structural characteristics of the data. When performing data queries, the native multi-modal database will design a unified query language or support multiple query grammars under a query framework. For example, develop a new query language that can be syntactically compatible with some functions of the Structured Query Language (SQL) of relational databases, and at the same time can handle similar JSON path queries for document data and path traversal queries for graph data, etc. In terms of data consistency guarantee, when performing operations such as data update and deletion, the native multi-modal database ensures data consistency between different data models through a transaction management mechanism. For example, when a record in a relational data table is associated with a certain document in a document data (such as through a foreign key or a logical relationship), during the transaction processing, the multi-modal database will ensure that the operations on both are either successful at the same time or fail at the same time, which may be achieved by setting logging records, lock mechanisms, etc. at the storage engine level.

[0030] When this architecture responds to the integration of multiple-modal data in application services, the data processing complexity is high and the data processing efficiency is low, making it difficult to meet the service requirements of high-timeliness requirements.

[0031] Figure 1 It is a flowchart of a multi-modal fusion data processing method according to an embodiment of the present disclosure. This multi-modal fusion data processing method can be applied to scenarios of performing various operations on a multi-modal database, such as retrieval, storage, etc. As Figure 1 shown, this multi-modal fusion data processing method may include the following steps S11 to S13.

[0032] In step S11, receive a multi-modal data retrieval statement, parse the multi-modal data retrieval statement and translate it into single-modal data retrieval statements corresponding to each modality.

[0033] The multi-modal data retrieval statement may be an SQL retrieval statement.

[0034] In some embodiments, the multi-modal data retrieval statement can be parsed in the following manner.

[0035] First, call the Large Language Model (LLM) to decompose the multi-modal data retrieval statement into words and grammar units, and construct a syntax tree based on the words and grammar units. For example, the large language model can use lexical analysis and syntactic analysis techniques to decompose the multi-modal data retrieval statement into individual words and grammar units, and construct a syntax tree based on the words and grammar units. Additionally, before decomposing the multi-modal data retrieval statement into words and grammar units, the large language model can also check whether the multi-modal data retrieval statement conforms to SQL syntax rules. If there are syntax errors, it can return an error message to the user.

[0036] Then, call the Retrieval Augmented Generation (RAG) model to retrieve a preliminary semantic understanding result from the corresponding knowledge base based on the syntax tree, and call the large language model to process the preliminary semantic understanding result to obtain the parsing of the multi-modal data retrieval statement. That is, based on the constructed syntax tree, further analyze the semantics of the multi-modal data retrieval statement to determine the user's data processing intent, operation object, condition constraints, etc. For example, determine whether the user wants to perform data query, insertion, update, or deletion operations, and which tables, columns, and conditional expressions in the multi-modal database are involved. When performing semantic understanding, a pre-constructed knowledge graph can be referred to. For example, the multi-modal RAG model can be called to retrieve a preliminary semantic understanding result from the pre-constructed knowledge graph based on the constructed syntax tree, and then the large language model can be called to process the preliminary semantic understanding result to obtain the parsing of the multi-modal data retrieval statement.

[0037] In some embodiments, translating the parsing result of the multi-modal data retrieval statement into single-modal data retrieval statements corresponding to each modality can be achieved in the following manner. First, according to the modality information in the multi-modal database, match the table names, column names, etc. in the parsing result with the actual data modalities, and translate the parsing result into corresponding operations for the matched modality data. For example, for a joint query involving different data modalities (such as relational data modality, document data modality, graph data modality, key-value pair data modality, etc.), the parsing result can be decomposed into sub-queries for each data modality, and relevant conversions and coordinations can be performed. For example, it can be translated into operation instructions for the corresponding single-modal database that the multi-modal database can understand and process.

[0038] The relational data modality stores data in the form of tables and manages and queries data through relationships (such as primary key - foreign key relationships). The document - type data modality stores data in a JSON - like document form, and each document can have a different structure, which is suitable for storing semi - structured data. The graph - type data modality represents entities and the relationships between entities with nodes and edges, and is suitable for representing complex relational data such as social networks and knowledge graphs. The key - value pair data modality is a simple key - value store, suitable for quickly looking up the simple value corresponding to a specific key, such as in cache scenarios.

[0039] Figure 2 Taking the multi - modal data retrieval statement as an SQL retrieval statement as an example, the multi - modal SQL retrieval syntax definition and the parsing schematic diagram of the multi - modal data retrieval statement based on the large - language model and the multi - modal RAG model are shown. From Figure 2 It can be seen that through the parsing and translation processing based on the large - language model and the multi - modal RAG model, the multi - modal data retrieval statement is parsed and translated into InfluxDB time - series retrieval statements, MongoDB retrieval statements, large - language model multi - modal retrieval statements, etc. Figure 3 Taking the multi - modal data retrieval statement as an SQL retrieval statement as an example, the schematic diagram of parsing and translating the multi - modal data retrieval statement into an InfluxDB time - series retrieval statement and an ElasticSearch retrieval statement through parsing and translation according to the embodiments of the present disclosure is shown. Among them, InfluxDB time - series retrieval refers to using the InfluxDB database for time - based sequence data retrieval.

[0040] InfluxDB is a time - series database used to store time - series data, such as the time - series data collected by sensors (e.g., the changes of temperature and pressure over time). These data are indexed by timestamps, which facilitates quickly querying data within a specific time period. MongoDB is a document database used to store document data, including various formats of text files (such as PDF, Word documents), image description documents, audio transcription texts, etc. The document database can store data in a flexible JSON format, which is convenient for processing complex document structures and nested data. In addition, the modal database can also include wide - table databases and relational databases. Wide - table databases use data warehouse technologies (such as Hive or Snowflake) to store wide - table data, which is suitable for storing data sets with a large number of columns and sparse data, such as user behavior data, product attribute data, etc. Wide - tables can facilitate multi - dimensional analysis and aggregation operations. Relational databases (such as MySQL or PostgreSQL) can be used to store data with a clear relational structure, such as the associated data between user tables, order tables, and product tables, and perform efficient relational queries and transaction processing through SQL.

[0041] In some embodiments, after translation, the single-mode data retrieval statements obtained by translation can be optimized according to the data distribution, statistical information of the multi-mode database, and the complexity of the single-mode data retrieval statements obtained by translation. Among them, the data distribution refers to the probability distribution of the data stored in each single-mode database, which can generally be represented by a probability density function. The statistical information refers to the statistical information related to the data stored in each single-mode database. For example, statistical information such as retrieval frequency, data volume, and retrieval method. When optimizing, at least one of the connection method, connection order, index construction method, and view creation method of the single-mode data retrieval statements obtained by translation can be adjusted according to the data distribution, statistical information of the multi-mode database, and the complexity of the single-mode data retrieval statements obtained by translation. For example, the query plan can be adjusted, appropriate indexes can be selected, operation steps can be merged or simplified, etc. For example, to retrieve a complete logistics information of a purchased commodity, it is necessary to retrieve relevant retrieval data from the commodity transaction relationship database and the logistics transfer process time series database respectively and then splice them to obtain a data set for return. Then, the process of parsing and translating a multi-mode data retrieval statement into two or more different-mode database retrievals and then merging can be achieved by identifying frequently accessed data and storing it uniformly in the relational database or the time series database, so as to reduce the retrieval quantity or frequency. That is to say, through optimization, an optimal retrieval plan can be generated, the retrieval efficiency and performance can be improved, and unnecessary calculations and data access can be reduced.

[0042] In step S12, data retrieval is performed in the single-mode databases corresponding to each single-mode data retrieval statement in the multi-mode database to obtain target data of different modalities; among them, multiple single-mode databases corresponding to each modality are maintained in the multi-mode database, and independent data management is performed for each single-mode database.

[0043] In step S13, the target data of different modalities retrieved are subjected to feature fusion and modality alignment processing to generate a unified multi-modal data representation, and the unified multi-modal data representation is fed back as the retrieval result.

[0044] By adopting the above technical solution, it is possible to receive multimodal data retrieval statements, parse the multimodal data retrieval statements and translate them into unimodal data retrieval statements corresponding to each modality, perform data retrieval in the unimodal databases corresponding to each unimodal data retrieval statement in the multimodal database respectively to obtain target data of different modalities. Among them, in the multimodal database, unimodal databases corresponding to each modality are maintained, and independent data management is performed for each unimodal database. Then, the target data of different modalities retrieved are subjected to feature fusion and modality alignment processing to generate a unified multimodal data representation, and the unified multimodal data representation is fed back as the retrieval result. In this way, when responding to the application service for integrating multiple modality data, the data processing complexity can be reduced, the data processing efficiency can be improved, and the service requirements of high timeliness can be met.

[0045] In some embodiments, for the unimodal database of each modality, the data therein is vectorized, and indexes are established for all the vectorized data. For example, for text data, text indexing techniques such as inverted indexes can be used to establish indexes; for data such as images and audio, index structures based on feature vectors, such as KD-Tree, hash tables, etc., can be used. In addition, a hybrid index structure can be constructed in combination with the characteristics of multimodal data to enable efficient retrieval of data of different modalities. In this way, after obtaining the unimodal data retrieval statements for the unimodal databases of each modality through translation, the data targeted by the unimodal data retrieval statements can be determined according to the corresponding indexes, improving the retrieval efficiency. Moreover, during retrieval, the retrieval operation can be performed based on the corresponding adaptation interfaces and operation methods provided for different data modalities. For example, for the relational data modality, operation methods such as table structure definition, index creation, data insertion, and query are provided; for the document data modality, functions such as document creation, update, and full-text search are provided, and the retrieval can be performed based on these provided operation methods.

[0046] In some embodiments, the step of performing data retrieval in the unimodal databases corresponding to each unimodal data retrieval statement in the multimodal database respectively according to each unimodal data retrieval statement in step S13 to obtain target data of different modalities may include: calling each of the multiple multimodal agents to perform data retrieval from the unimodal database corresponding to each according to the unimodal data retrieval statement for each unimodal database, and vectorizing the retrieved data to obtain target data of different modalities; wherein, the multimodal agents and the unimodal databases are in one-to-one correspondence, one multimodal agent is responsible for retrieving one unimodal database, and each multimodal agent is independent of each other.

[0047] When performing retrieval, the multimodal agent can adopt appropriate retrieval strategies based on each unimodal data retrieval statement. For example, the multimodal agent can select the corresponding retrieval method according to the modality type to be retrieved. If it is text retrieval, text retrieval techniques can be used, such as using an inverted index for keyword matching retrieval; if it is image retrieval, an image feature matching retrieval method can be used; for image and audio data, feature vector indexing can be used for similarity retrieval; for relational data and wide table data, retrieval can be performed in relational databases and data warehouses according to the traditional SQL query execution process. In addition, information from multiple modalities can be combined for joint retrieval to improve the accuracy and comprehensiveness of retrieval.

[0048] When vectorizing the retrieved data, the following method can be used to achieve it. For data of different modalities, such as text, images, audio, etc., corresponding technologies can be used to extract features from the retrieved data, thereby obtaining the vectors of the retrieved data. For example, for text, word vectors and text embedding models (such as BERT, Word2Vec, etc.) can be used to convert the retrieved data into numerical vector representations for the computer to process and understand. For images, convolutional neural networks can be used to extract the feature vectors of the images, and these feature vectors can capture information such as the color, texture, and shape of the images. For audio data, it can be encoded through technologies such as acoustic models and converted into audio feature vectors.

[0049] By adopting the above technical solutions, the retrieval efficiency can be improved, the data processing complexity can be reduced, and the service requirements of high timeliness can be met.

[0050] In some embodiments, the feature fusion and modality alignment processing of the retrieved target data of different modalities in step S13 includes: invoking a large language model to map the target data of different modalities into the same embedding space, performing modality alignment processing on the target data of different modalities in this embedding space, and performing feature fusion on the target data of different modalities after the modality alignment processing. Mapping the target data of different modalities into the same embedding space can make the target data of different modalities comparable in this same embedding space, which facilitates subsequent alignment and feature fusion operations. For example, a multimodal encoder can be trained to encode the target data of modalities such as text, images, and audio into the same feature space, so that the similarity between these target data can be measured by the distance between vectors. In addition, there are many methods for feature fusion, such as simple concatenation, weighted summation, fusion based on the attention mechanism, etc. The attention mechanism can automatically assign different weights according to the importance of different modality data, so as to better fuse the features of multimodal data and obtain a fusion result in a specified format (for example, a two-dimensional table, JSON, XML, etc.). Through feature fusion, a unified multimodal data representation can be generated. After feature fusion, further processing (such as aggregation, sorting, filtering, etc.) can be performed on the fused data according to the specific retrieval requirements of the user, such as aggregation, sorting, filtering, etc.

[0051] Performing modality alignment on the target data of each modality before feature fusion can ensure that the target data of different modalities are aligned in terms of time, space, and / or semantics, which can ensure the accuracy of the feature fusion result. For example, in multimodal data of video and audio, it is necessary to align the video frames and audio segments to better understand and fuse the information between them. For multimodal data of text and images, semantic alignment is also required, such as matching the objects described in the text with the objects in the image.

[0052] After obtaining the unified multimodal data representation, it can be formatted and visualized according to the requirements of the user interaction layer as the retrieval result, and then returned to the user interface for display.

[0053] In some embodiments, the multimodal fusion data processing method according to the embodiments of the present disclosure can also implement the storage of multimodal data. For example, the storage medium for the data targeted by the storage operation can be determined according to the access frequency, importance, and storage cost of the data targeted by the storage operation. For example, frequently accessed data is stored in memory or a solid-state drive to improve data access speed; infrequently accessed data is stored on a disk to reduce storage costs. In this way, efficient storage and access of data can be achieved.

[0054] In some embodiments, large language models, multimodal RAG models, and multimodal agents can be trained in the following manner. The following is described by taking SQL statements as an example.

[0055] First, pre-train and fine-tune the large language model, multimodal RAG model, and multimodal agent. In the pre-training stage, a large amount of text data containing SQL statements can be incorporated into the training corpus to enable the large language model, multimodal RAG model, and multimodal agent to learn the syntax, structure, and common patterns of SQL. In the fine-tuning stage, a specific labeled SQL dataset is used to further train the pre-trained large language model, multimodal RAG model, and multimodal agent, enabling the large language model, multimodal RAG model, and multimodal agent to generate and understand SQL statements that meet the requirements more accurately. The annotation information can include the function of the SQL statement, the corresponding database table structure, query conditions, etc., to help the large language model, multimodal RAG model, and multimodal agent better understand and generate correct SQL.

[0056] In addition, a clear input format can be defined. For example, specific prompts or markers are used to indicate the start and end of the SQL statement, enabling the large language model, multimodal RAG model, and multimodal agent to clearly know that the content to be processed is related to SQL. For the output, the large language model, multimodal RAG model, and multimodal agent can be required to generate SQL statements in a predetermined format, such as following specific indentation rules and capitalizing keywords, to make it easier to parse and use the generated SQL statements.

[0057] In addition, a knowledge graph related to the multimodal database can be created, which contains information such as database architecture, table relationships, and field meanings, and the RAG framework is used to combine it with the LLM. In this way, when generating SQL, the information in the knowledge graph can be referred to to better understand the associations between data and query requirements, thereby generating more accurate and logical SQL statements.

[0058] In addition, the large language model, multi-modal RAG model, and multi-modal agent can be guided to generate correct SQL through multiple rounds of interaction. Users can gradually provide more information and constraints, allowing the model to continuously adjust and improve the generated SQL statements based on the feedback until the requirements are met. When errors exist in the SQL generated by the large language model, multi-modal RAG model, and multi-modal agent, clear error prompts and corrective suggestions should be given in a timely manner, enabling the large language model, multi-modal RAG model, and multi-modal agent to learn from the errors and improve their SQL recognition and generation capabilities. For example, feedback on the retrieval accuracy of the retrieval results can be obtained from the user, the retrieval accuracy of the retrieval results can be evaluated based on the feedback, and the parameters of the large language model, multi-modal RAG model, and multi-modal agent can be adjusted according to the evaluation results. For example, the translation strategy, retrieval strategy, data embedding strategy, index construction strategy, fusion method, and the parameters of the large language model, multi-modal RAG model, and multi-modal agent can be adjusted based on the evaluation results to improve the performance and accuracy of the system.

[0059] In addition, the SQL parser can be integrated with the LLM. After the large language model, multi-modal RAG model, and multi-modal agent generate SQL statements, the SQL parser is immediately used to perform syntax checking and semantic analysis on them, and the parsing results are fed back to the large language model, multi-modal RAG model, and multi-modal agent, enabling the large language model, multi-modal RAG model, and multi-modal agent to understand whether the generated SQL is correct and what problems exist, so as to make self-corrections. Further combined with the SQL executor, the generated SQL statements are executed in an actual database environment, and the execution results are returned to the large language model, multi-modal RAG model, and multi-modal agent, allowing the large language model, multi-modal RAG model, and multi-modal agent to further optimize and adjust the SQL based on the results to better meet the query requirements.

[0060] By adopting the above technical solutions, the training of the large language model, multi-modal RAG model, and multi-modal agent can be achieved.

[0061] Figure 4 It is an architecture diagram of a multi-modal fusion data processing architecture according to an embodiment of the present disclosure.

[0062] As Figure 4 shown, the multi-modal fusion data processing architecture includes a semantic understanding and translation layer, a transaction management layer, a multi-modal data management layer, and a multi-modal data storage layer. External application systems can access the multi-modal fusion data processing architecture according to an embodiment of the present disclosure through the network to achieve the processing requirements of external application systems, such as querying, data storage, etc.

[0063] The multi-modal fusion data processing architecture is implemented based on multi-modal RAG models, multi-modal agents, and large language models. For multi-modal data such as time series, text, and images, the multi-modal data storage layer can be compatible with existing single-modal databases such as document databases, time series databases, wide table databases, and relational databases, as well as object storage, avoiding the solution of unified fusion storage of all data with multiple copies retained. Moreover, the multi-modal fusion data processing architecture encapsulates data query capabilities, result generation capabilities, and extraction vector embedding RAG capabilities with decoupled and independent multi-modal agents, and provides unified and consistent capabilities to the multi-modal semantic understanding translation layer and the multi-modal data fusion retriever based on the LLM model.

[0064] The semantic understanding translation layer can be used to implement operations such as parsing, translation, and semantic understanding described above.

[0065] The transaction management layer can be used to manage and coordinate transaction operations in the multi-modal database, ensuring the consistency, atomicity, isolation, and durability of the data in the multi-modal database.

[0066] The transaction management layer can be used to manage the start and end of transactions, for example, responsible for identifying the start and end of transactions. When a user executes a set of related database operations and wishes to commit or roll them back as a whole, the transaction management layer records the starting point of the transaction and, after all operations are completed, decides whether to commit the transaction to make all operations effective or roll back the transaction to restore the data to the state before the transaction started, based on the execution situation.

[0067] The transaction management layer can perform concurrent control of transactions. In a multi-user environment, the transaction management layer can coordinate the concurrent execution of multiple transactions, preventing data inconsistencies and conflicts. By adopting techniques such as lock mechanisms, timestamp ordering, and multi-version concurrent control, it ensures that the operations between different transactions are scheduled and isolated according to certain rules, avoiding the visibility of the uncompleted operations of one transaction to other transactions, thus guaranteeing data consistency and isolation.

[0068] The transaction management layer can perform transaction recovery. In case of system failures or anomalies, the transaction management layer is responsible for recovering uncompleted transactions, ensuring data durability. By recording transaction log information, including operation records, values before and after data modification, etc., the transaction management layer can perform redo or undo operations based on the log when the system restarts or recovers, restoring the multi-modal database to a consistent state.

[0069] The transaction management layer can also handle nested transactions. The transaction management layer can support the execution and management of nested transactions, where a transaction can contain other sub-transactions. In this case, the transaction management layer needs to correctly handle the relationship between sub-transactions and the parent transaction, including transaction commit, rollback, and resource allocation and release, etc., to ensure the correctness and consistency of the entire transaction hierarchy.

[0070] The multi-modal data management layer is the core part of the multi-modal fusion data processing architecture, responsible for the unified management and operation of data in multiple data modalities, including data storage, retrieval, update, deletion, etc., and at the same time providing data modality conversion and fusion functions.

[0071] The multi-modal data management layer can perform data modality adaptation. For example, for different data modalities (such as relational data modality, document data modality, graph data modality, key-value pair data modality, etc.), corresponding adaptation interfaces and operation methods are provided, enabling the effective storage and management of various modal data. For example, for relational data, operations such as table structure definition, index creation, data insertion, and query are provided; for document data, functions such as document creation, update, and full-text search are supported.

[0072] The multi-modal data management layer can also perform data fusion and conversion, that is, achieve data conversion and fusion between different data modalities, and support cross-data modality queries and operations. For example, the multi-modal data management layer can perform associative queries on relational data and document data, or convert node and edge information in graph data into relational data for storage and analysis. By defining mapping rules and conversion functions between data modalities, the multi-modal data management layer can achieve seamless data flow and collaborative processing.

[0073] The multi-modal data management layer can achieve data consistency maintenance, ensuring data consistency during the operation of multi-modal data. When updating or modifying related data in different data modalities, the multi-modal data management layer needs to ensure data consistency and integrity, and avoid data inconsistency. For example, when updating entity data that exists in both a relational table and a document collection, the multi-modal data management layer needs to update the data in both places to ensure data consistency.

[0074] The multi-modal data management layer can also perform data indexing and query optimization. For example, according to the characteristics of different data modalities and query requirements, various index structures can be created and managed to improve the efficiency of data retrieval. In addition, the query of multi-modal data can be optimized, analyzing the data modalities and operation types involved in the query statement, and selecting appropriate query execution plans and indexes to accelerate the query and access speed of data.

[0075] The multi-modal data storage layer is responsible for storing multi-modal data on the underlying storage medium in a suitable physical storage manner and providing efficient data reading, writing, and access functions. According to requirements, the multi-modal data storage layer can be constructed based on the data retrieval requirements of existing databases or application systems.

[0076] The multi-modal data storage layer can use multi-modal agents to encapsulate and decouple the reading, writing, caching, embedding, etc. capabilities of each single-modal database. That is, each single-modal database uses an independent multi-modal agent for reading, writing, caching, embedding, etc.

[0077] The multi-modal data storage layer can be responsible for managing multiple storage media, such as disks, memory, solid-state drives, etc. Appropriate storage media can be selected to store data based on factors such as data access frequency, importance, and storage cost. For example, frequently accessed data can be stored in memory or solid-state drives to improve data access speed; infrequently accessed data can be stored on disks to reduce storage costs.

[0078] The multi-modal data storage layer can perform data organization and layout. For example, reasonable data organization and layout methods can be designed according to different data modalities and data characteristics. For example, for relational data, it can be stored in a tabular form and corresponding storage structures can be established based on the primary key and index; for document-type data, it can be stored according to the document's structure and attributes, and an inverted index can be established to support full-text search. By optimizing data organization and layout, data storage efficiency and access performance can be improved.

[0079] The multi-modal data storage layer can also select the data storage format. That is, choose a suitable data storage format to store multi-modal data, such as binary format, text format, JSON format, etc. Different data storage formats have different advantages and disadvantages, and need to be selected according to factors such as data type, application scenario, and performance requirements. For example, for structured relational data, a binary relational database storage format can be selected; for semi-structured or unstructured document-type data, text formats such as JSON or XML can be used for storage.

[0080] The multi-modal data storage layer can also adopt data reading, writing, and caching mechanisms. For example, it can provide efficient data reading and writing interfaces and caching mechanisms to speed up data access. By using caching technology, frequently used data is cached in memory, reducing the number of accesses to the underlying storage medium and improving data reading and writing performance. In addition, the multi-modal data storage layer can also optimize data reading and writing operations, adopting batch reading and writing, asynchronous reading and writing, etc. to improve data reading and writing efficiency and the system's concurrent processing ability.

[0081] The multi-modal fusion data processing architecture may also include a metadata manager for uniformly maintaining the metadata of different modal databases and providing a consistent data access view for multi-modal data fusion retrieval. For example, when performing multi-modal fusion data retrieval, the metadata of each single-modal database can be obtained from the metadata manager.

[0082] In addition, in the multi-modal fusion data processing architecture, the embedding model can convert the retrieved data into vectors. The converted vectors can be re-constructed and indexed. The in-library vector aggregation data pipeline can then transmit the converted vectors to the multi-modal data fusion retriever. The multi-modal data fusion retriever can fuse the vectors transmitted by the in-library vector aggregation data pipeline to obtain a fusion result, which will be fed back to the user. In addition, the multi-modal data fusion retriever can store the vectors transmitted by the in-library vector aggregation data pipeline into the vector database. When needed, the vectors in the vector database can be analyzed based on the large language model.

[0083] In addition, the multi-modal fusion data processing architecture may also include a multi-modal agent manager for managing multi-modal agents. The multi-modal fusion data processing architecture may also include an out-of-library data import engine for importing external data sources into the corresponding single-modal databases.

[0084] By adopting the above multi-modal fusion data processing architecture, the following beneficial effects can be achieved: (1) It can fully explore the associations and complementarities between different modal data, avoid the limitations of single-modal data queries, and improve the overall utilization efficiency of data; (2) The extended query syntax and multi-modal query functions enable users to flexibly construct multi-modal data retrieval statements according to specific needs, meet diverse query scenarios, and enhance the flexibility of queries; (3) The application of multi-modal agents and multi-modal RAG models enables the provision of intelligent query services, which can not only accurately return data but also provide value-added services such as data analysis, decision-making suggestions, and knowledge expansion, improving the user experience and realizing a simple and efficient service mode; (4) Utilizing the semantic understanding ability of the large language model can enhance the understanding and compatibility of statements, have a certain tolerance for local input errors in grammar, table names, column names, etc., and give accurate semantic understanding through error correction or fuzzy queries to return the expected query results, improving the query tolerance and compatibility ability.

[0085] Figure 5 It is a schematic block diagram of a multi-modal fusion data processing device according to an embodiment of the present disclosure. As Figure 5As shown, the multi-modal fusion data processing device can be applied to scenarios where various operations are performed on a multi-modal database, such as retrieval, storage, etc. The multi-modal fusion data processing device 50 may include: a parsing and translation module 51, configured to receive a multi-modal data retrieval statement, parse the multi-modal data retrieval statement and translate it into single-modal data retrieval statements corresponding to each modality; a retrieval module 52, configured to perform data retrieval in the corresponding single-modal databases in the multi-modal database according to each single-modal data retrieval statement to obtain target data of different modalities; wherein, multiple single-modal databases corresponding to each modality are maintained in the multi-modal database, and independent data management is performed for each single-modal database; a feedback module 53, configured to perform feature fusion and modality alignment processing on the retrieved target data of different modalities to generate a unified multi-modal data representation, and feedback the unified multi-modal data representation as a retrieval result.

[0086] By adopting the above technical solution, it is possible to receive a multi-modal data retrieval statement, parse the multi-modal data retrieval statement and translate it into single-modal data retrieval statements corresponding to each modality, perform data retrieval in the corresponding single-modal databases in the multi-modal database according to each single-modal data retrieval statement to obtain target data of different modalities, wherein multiple single-modal databases corresponding to each modality are maintained in the multi-modal database, and independent data management is performed for each single-modal database, and then perform feature fusion and modality alignment processing on the retrieved target data of different modalities to generate a unified multi-modal data representation, and feedback the unified multi-modal data representation as a retrieval result. In this way, when integrating multi-modal data in response to an application service, the data processing complexity can be reduced, the data processing efficiency can be improved, and the service requirements of high timeliness can be met.

[0087] Optionally, the parsing and translation module 51 parses the multi-modal data retrieval statement, including: Invoking a large language model to decompose the multi-modal data retrieval statement into words and grammar units and constructing a syntax tree based on the words and the grammar units; Invoking a multi-modal retrieval enhanced generation model to retrieve a preliminary semantic understanding result from the corresponding knowledge base based on the syntax tree, and invoking the large language model to process the preliminary semantic understanding result to obtain the parsing of the multi-modal data retrieval statement.

[0088] Optionally, the multi-modal data fusion processing device further includes an optimization module, configured to: Optimize the translated single-modal data retrieval statements according to the data distribution, statistical information of the multi-modal database, and the complexity of the translated single-modal data retrieval statements, wherein the data distribution refers to the probability distribution of the data stored in each single-modal database, and the statistical information refers to the statistical information related to the data stored in each single-modal database.

[0089] Optionally, the optimization module optimizes the single-modal data retrieval statement obtained by translation according to the data distribution, statistical information of the multi-modal database, and the complexity of the single-modal data retrieval statement obtained by translation, including: Adjust at least one of the connection method, connection order, index construction method, and view creation method of the single-modal data retrieval statement obtained by translation according to the data distribution, statistical information of the multi-modal database, and the complexity of the single-modal data retrieval statement obtained by translation.

[0090] Optionally, the retrieval module 52 retrieves target data of different modalities in the single-modal databases corresponding to each single-modal data retrieval statement in the multi-modal database respectively, including: Call each multi-modal agent among a plurality of multi-modal agents to respectively retrieve retrieval data from their corresponding single-modal databases according to the single-modal data retrieval statement for each single-modal database and vectorize the retrieval data, so as to obtain target data of different modalities; wherein, the multi-modal agents correspond to the single-modal databases one by one, one multi-modal agent is responsible for retrieving one single-modal database, and each multi-modal agent is independent of each other.

[0091] Optionally, the feedback module 53 performs feature fusion and modality alignment processing on the retrieved target data of different modalities, including: Call a large language model to map the target data of different modalities to the same embedding space, perform modality alignment processing on the target data of different modalities in the embedding space, and perform feature fusion on the target data of different modalities after modality alignment processing.

[0092] Optionally, the multi-modal fusion data processing device further includes an adjustment module for: Obtain the user's feedback on the retrieval accuracy of the retrieval result; Evaluate the retrieval accuracy of the retrieval result according to the user's feedback; Adjust the parameters of the relevant model according to the evaluation result.

[0093] The specific implementation manners of the operations performed by each module in the multi-modal fusion data processing device according to the embodiments of the present disclosure have been described in detail in the related methods and will not be elaborated herein.

[0094] The present disclosure also provides a computer-readable medium, on which a computer program is stored, and when the computer program is executed by a processing device, the steps of any method in the present disclosure are implemented.

[0095] The present disclosure also provides an electronic device, including: A storage device on which a computer program is stored; A processing device for executing the computer program in the storage device to implement the steps of any one of the methods in the present disclosure.

[0096] The present disclosure also provides a computer program product, including a computer program which, when executed by a processor, implements the steps of any one of the methods in the present disclosure.

[0097] Reference is made below to Figure 6 , which shows a schematic structural diagram of an electronic device 600 suitable for implementing the embodiments of the present disclosure. The terminal devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The electronic device shown is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0098] As Figure 6 shown, the electronic device 600 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage device 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.

[0099] Generally, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or wirelesly to exchange data. Although Figure 6 the electronic device 600 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices may be alternatively implemented or had.

[0100] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowchart can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program code for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above functions defined in the method of the embodiment of the present disclosure are performed.

[0101] It should be noted that the computer-readable medium in the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in combination with an instruction execution system, apparatus, or device. And in the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0102] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0103] The above computer-readable medium can be included in the above electronic device; or can exist separately without being assembled into the electronic device.

[0104] The above computer-readable medium carries one or more programs. When the above one or more programs are executed by the electronic device, the electronic device is caused to: receive a multi-modal data retrieval statement, parse the multi-modal data retrieval statement and translate it into single-modal data retrieval statements corresponding to each modality; perform data retrieval in the single-modal databases corresponding to each single-modal data retrieval statement in the multi-modal database to obtain target data of different modalities; wherein, the multi-modal database maintains single-modal databases corresponding to each modality, and performs independent data management for each single-modal database; perform feature fusion and modality alignment processing on the retrieved target data of different modalities to generate a unified multi-modal data representation, and feedback the unified multi-modal data representation as the retrieval result.

[0105] Computer program code for performing the operations of the present disclosure can be written in one or more programming languages or combinations thereof. The above programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).

[0106] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that, in some alternative implementations, the functions noted in the blocks may occur in a different order than that noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.

[0107] The modules described in the embodiments of the present disclosure can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation on the module itself in some cases. For example, the parsing and translation module can also be described as "a module that receives a multimodal data retrieval statement, parses the multimodal data retrieval statement, and translates it into unimodal data retrieval statements corresponding to each modality".

[0108] The functions described above herein can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGA), Application Specific Integrated Circuits (ASIC), Application Specific Standard Products (ASSP), System on Chip (SOC), Complex Programmable Logic Devices (CPLD), and so on.

[0109] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a Random Access Memory (RAM), a Read Only Memory (ROM), an Erasable Programmable Read Only Memory (EPROM or Flash Memory), an optical fiber, a portable Compact Disc Read Only Memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0110] According to one or more embodiments of the present disclosure, Example 1 provides a multimodal fusion data processing method, including: Receiving a multimodal data retrieval statement, parsing the multimodal data retrieval statement, and translating it into unimodal data retrieval statements corresponding to each modality; Performing data retrieval in the unimodal databases corresponding to each unimodal data retrieval statement in the multimodal database to obtain target data of different modalities; wherein, the multimodal database maintains unimodal databases corresponding to each modality, and performs independent data management for each unimodal database; Performing feature fusion and modality alignment processing on the retrieved target data of different modalities to generate a unified multimodal data representation, and feeding back the unified multimodal data representation as a retrieval result.

[0111] According to one or more embodiments of the present disclosure, Example 2 provides the method of Example 1, wherein, parsing the multimodal data retrieval statement includes: Invoking a large language model to decompose the multimodal data retrieval statement into words and grammar units and constructing a syntax tree based on the words and the grammar units; Invoking a multimodal retrieval enhanced generation model to retrieve a preliminary semantic understanding result from the corresponding knowledge base based on the syntax tree, and invoking the large language model to process the preliminary semantic understanding result to obtain the parsing of the multimodal data retrieval statement.

[0112] According to one or more embodiments of the present disclosure, Example 3 provides the method of Example 1, wherein, the method further includes: Optimizing the translated unimodal data retrieval statements according to the data distribution, statistical information of the multimodal database, and the complexity of the translated unimodal data retrieval statements, wherein the data distribution refers to the probability distribution of the data stored in each unimodal database, and the statistical information refers to the statistical information related to the data stored in each unimodal database.

[0113] According to one or more embodiments of the present disclosure, Example 4 provides the method of Example 3, wherein, optimizing the translated unimodal data retrieval statements according to the data distribution, statistical information of the multimodal database, and the complexity of the translated unimodal data retrieval statements includes: Adjusting at least one of the connection mode, connection order, index construction mode, and view creation mode of the translated unimodal data retrieval statements according to the data distribution, statistical information of the multimodal database, and the complexity of the translated unimodal data retrieval statements.

[0114] According to one or more embodiments of the present disclosure, Example 5 provides the method of Example 1, wherein the data retrieval for obtaining target data of different modalities in the corresponding unimodal database in the multimodal database according to each unimodal data retrieval statement includes: Invoking each of the multiple multimodal agents to respectively perform data retrieval from the corresponding unimodal database according to the unimodal data retrieval statement for each of the unimodal databases, vectorizing the retrieved data, and obtaining target data of different modalities; wherein, the multimodal agents correspond one-to-one with the unimodal databases, one multimodal agent is responsible for retrieving one unimodal database, and each multimodal agent is independent of each other.

[0115] According to one or more embodiments of the present disclosure, Example 6 provides the method of Example 1, wherein the feature fusion and modality alignment processing of the retrieved target data of different modalities includes: Invoking a large language model to map the target data of different modalities to the same embedding space, performing modality alignment processing on the target data of different modalities in the embedding space, and performing feature fusion on the target data of different modalities after the modality alignment processing.

[0116] According to one or more embodiments of the present disclosure, Example 7 provides the method of Example 2, 5 or 6, wherein the method further includes: Obtaining the user's feedback on the retrieval accuracy of the retrieval result; Evaluating the retrieval accuracy of the retrieval result according to the user's feedback; Adjusting the parameters of the relevant model according to the evaluation result.

[0117] According to one or more embodiments of the present disclosure, Example 8 provides a multimodal fusion data processing device, including: A parsing and translation module, configured to receive a multimodal data retrieval statement, parse the multimodal data retrieval statement and translate it into unimodal data retrieval statements respectively corresponding to multiple modalities; A retrieval module, configured to perform data retrieval in the corresponding unimodal database in the multimodal database according to each unimodal data retrieval statement to obtain target data of different modalities; wherein, multiple unimodal databases respectively corresponding to multiple modalities are maintained in the multimodal database, and independent data management is performed for each unimodal database; A feedback module, configured to perform feature fusion and modality alignment processing on the retrieved target data of different modalities to generate a unified multimodal data representation, and feedback the unified multimodal data representation as the retrieval result.

[0118] According to one or more embodiments of the present disclosure, Example 9 provides a computer-readable medium having a computer program stored thereon, and when the computer program is executed by a processing device, the steps of the method described in any one of Examples 1-7 are implemented.

[0119] According to one or more embodiments of the present disclosure, Example 10 provides an electronic device, including: A storage device having a computer program stored thereon; A processing device configured to execute the computer program in the storage device to implement the steps of the method described in any one of Examples 1-7.

[0120] According to one or more embodiments of the present disclosure, Example 11 provides a computer program product including a computer program, and when the computer program is executed by a processor, the steps of the method described in any one of Examples 1-7 are implemented.

[0121] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, a technical solution formed by mutually replacing the above features with (but not limited to) technical features having similar functions disclosed in the present disclosure.

[0122] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although a number of specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented combinatorially in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0123] Although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims. Regarding the devices in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated herein.

Claims

1. A multi-mode fusion data processing method, characterized in that: include: Receiving a multi-mode data search statement, parsing the multi-mode data search statement and translating it into single-mode data search statements corresponding to each of the multiple modes; According to each single-mode data search statement, data is searched in the corresponding single-mode database in the multi-mode database to obtain target data of different modes; wherein the multi-mode database maintains single-mode databases corresponding to multiple modes, and independent data management is performed for each single-mode database; The retrieved target data of different modalities are subjected to feature fusion and modality alignment processing to generate a unified multimodal data representation, and the unified multimodal data representation is fed back as a retrieval result.

2. The method according to claim 1, characterized in that The parsing of the multi-mode data search statement includes: Calling a large language model to decompose the multimodal data retrieval sentence into words and grammatical units and constructing a grammatical tree based on the words and the grammatical units; The multimodal retrieval enhancement generation model is called to retrieve a preliminary semantic understanding result from the corresponding knowledge base based on the syntax tree, and the large language model is called to process the preliminary semantic understanding result to obtain the analysis of the multimodal data retrieval sentence.

3. The method according to claim 1, characterized in that The method further comprises: The translated single-mode data retrieval statement is optimized based on the data distribution, statistical information of the multi-mode database and the complexity of the translated single-mode data retrieval statement, wherein the data distribution refers to the probability distribution of the data stored in each single-mode database, and the statistical information refers to the statistical information related to the data stored in each single-mode database.

4. The method according to claim 3, characterized in that The step of optimizing the translated single-mode data search statement according to the data distribution and statistical information of the multi-mode database and the complexity of the translated single-mode data search statement comprises: According to the data distribution and statistical information of the multi-mode database and the complexity of the translated single-mode data retrieval statement, at least one of the connection method, connection order, index construction method and view creation method of the translated single-mode data retrieval statement is adjusted.

5. The method according to claim 1, characterized in that The step of performing data retrieval in the corresponding single-mode database in the multi-mode database according to each single-mode data retrieval statement to obtain target data of different modes includes: Call each of the multiple multimodal agents to perform data search from the corresponding single-mode database according to the single-mode data search statement for each of the single-mode databases, obtain search data and vectorize the search data to obtain target data of different modalities; wherein the multimodal agents correspond to the single-mode databases one-to-one, one multimodal agent is responsible for searching one single-mode database, and each multimodal agent is independent of each other.

6. The method according to claim 1, characterized in that The retrieved target data of different modalities are subjected to feature fusion and modality alignment processing, including: The large language model is called to map target data of different modalities to the same embedding space, modality alignment is performed on the target data of different modalities in the embedding space, and feature fusion is performed on the target data of different modalities after the modality alignment.

7. The method according to claim 2, 5 or 6, characterized in that: The method further comprises: Obtaining user feedback on the search accuracy of the search results; According to the feedback from the user, evaluating the retrieval accuracy of the retrieval result; According to the evaluation results, the parameters of the relevant models are adjusted.

8. A multi-mode fusion data processing device, characterized in that: include: A parsing and translation module, used for receiving a multi-mode data search statement, parsing the multi-mode data search statement and translating it into single-mode data search statements corresponding to each of the multiple modes; A retrieval module, used to perform data retrieval in the corresponding single-mode database in the multi-mode database according to each single-mode data retrieval statement to obtain target data of different modes; wherein the multi-mode database maintains single-mode databases corresponding to multiple modes, and independently manages data for each single-mode database; The feedback module is used to perform feature fusion and modality alignment processing on the retrieved target data of different modalities to generate a unified multimodal data representation, and to feed back the unified multimodal data representation as a retrieval result.

9. A computer readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processing device, the steps of the method according to any one of claims 1 to 7 are implemented.

10. An electronic device, characterized in that: include: a storage device having a computer program stored thereon; A processing device, configured to execute the computer program in the storage device to implement the steps of the method according to any one of claims 1 to 7.

11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • View optimization query method and system oriented to multimode database

    CN120723789A

  • A view optimization query method and system for a multi-mode database

    CN120723789B

  • Multi-modal grain data fusion and decision-making method based on artificial intelligence

    CN120743941A