Cross-modal data query method

Through the triple registration format and dynamic path selection, combined with multi-level caching and comparison models, the problems of data heterogeneity, high latency and insufficient accuracy in cross-modal data queries are solved, and efficient and flexible cross-modal data queries are achieved.

CN120632128AActive Publication Date: 2025-09-12INSPUR GENERSOFT CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511135485.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-14
Publication Date
2025-09-12
Estimated Expiration
2045-08-14

Smart Images

  • Figure CN120632128A_ABST
    Figure CN120632128A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-modal data query method, which belongs to the technical field of data query processing, and comprises the following steps: extracting a plurality of descriptor information of a data source in advance according to a preset triple registration format to generate an access interface for accessing data to the data source; the method further comprises the steps of responding to query content, analyzing to obtain modal types, selecting data sources of corresponding types according to the modal types, selecting access paths of the data sources according to historical delay data of the data sources, and setting a cache mode according to the number of the modal types; and accessing a corresponding database according to the access path, obtaining an access result of the corresponding modal type through a preset comparison model based on the relevance between the query content and the data corresponding to the modal type, and caching based on the caching mode. According to the method, the interface is automatically generated through the triple descriptor, and dynamic adaptation, low-delay access and high-precision query of a cross-modal data source are realized in combination with modal analysis and a multi-level cache mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data query processing, and in particular relates to a cross-modal data query method. Background Art

[0002] With the rapid development of artificial intelligence (AI), particularly breakthroughs in large language models (LLMs) and multimodal learning, the demand for cross-modal data interaction is growing exponentially. In areas such as intelligent customer service, cross-modal search engines, digital content management, medical image analysis, and autonomous driving, systems must simultaneously process heterogeneous data from multiple sources, including text, images, audio, and video. For example, media companies need to semantically correlate and retrieve massive amounts of multimodal content, such as news images, documentary videos, and soundtracks.

[0003] While multimodal data processing technology has made some progress, the storage formats and access protocols for different modal data vary significantly. Traditional approaches require writing independent adaptation code for each data source. For example, accessing an image database requires developing an image encoder interface, while accessing an audio database requires developing an audio decoder interface. Furthermore, adding new data sources (such as a new video database) requires redeveloping the adaptation logic, which is time-consuming and error-prone.

[0004] Furthermore, cross-modal queries require multiple round trips to different data sources, resulting in long response times and high resource consumption due to the accumulated delays. Furthermore, multimodal feature vectors reside in different semantic spaces, making it difficult to establish correlations using traditional methods. For example, the matching accuracy between musical and poetic emotions is low, resulting in irrelevant query results.

[0005] In general, existing technologies generally have data heterogeneity obstacles, query delays, resource waste, and insufficient query result accuracy when querying cross-modal data. Summary of the Invention

[0006] The present invention provides a cross-modal data query method to solve the problems of data heterogeneity, query delay, resource waste, and insufficient query result accuracy.

[0007] The technical solution adopted in the present invention is: A cross-modal data query method includes: extracting multiple descriptor information of a data source according to a preset triple registration format in advance to generate an access interface corresponding to the data source for accessing data from the data source; The method further includes, in response to the query content, parsing to obtain a modality type, selecting a data source of a corresponding type according to the modality type, selecting an access path to the data source according to historical latency data of the data source, and setting a cache mode according to the number of the modality types; According to the access path, the corresponding database is accessed, and through a preset comparison model, based on the query content and the correlation between the data in the data source corresponding to the modality type, the access result of the corresponding modality type is obtained, and cached based on the cache mode.

[0008] The cross-modal data query method disclosed in the present invention also has the following additional technical features: The triplet registration format is specifically: The triplet registration format includes modality type, data format, and access protocol. The modality types include text type, image type, and audio type.

[0009] In response to the query content, the modal type is parsed, specifically: Based on the query content, keywords are identified through grammatical analysis, so as to parse and obtain the modality type according to a mapping relationship table between the keywords and the modality type.

[0010] Select an access path for the data source based on the historical latency data of the data source, specifically: Based on the latency data obtained from historical queries and the load status of the current data source access path, a learning model is used to predict the latency of the current data source on different paths. If the local GPU latency is lower than the threshold, the data source of the local resource path is called; Otherwise, call the data source of the cloud API path.

[0011] The threshold is specifically: According to the modality type, initial thresholds are set one by one, with the initial thresholds of the text type, the image type, and the audio type decreasing in sequence; According to the complexity of the query content, the initial threshold is adjusted to obtain the threshold, and the threshold is negatively correlated with the complexity.

[0012] According to the number of modal types, set the cache mode, specifically: When the modality type obtained by parsing the query content is 1, the access result cache mode is set to short-term cache, and a first pre-cache time is set accordingly; otherwise, the access result cache mode is set to long-term cache, and a second pre-cache time is set accordingly, wherein the first pre-cache time is shorter than the second pre-cache time; Adjusting the first pre-caching time and the second pre-caching time according to the modality type obtained by parsing and the hit rate of the query content, thereby obtaining a first cache time and a second cache time accordingly; The first cache time and the second cache time are positively correlated with the hit rate.

[0013] Based on the query content and the correlation between the data in the data source corresponding to the modality type, an access result of the corresponding modality type is obtained, including: When the parsed modality type is 1 in response to the query content, An embedding vector is obtained according to the query content, and the embedding vector is compared with a feature vector of data in a corresponding data source. An access result is returned according to the comparison result.

[0014] Obtaining access results corresponding to the modality type based on the correlation between the query content and the data in the data source corresponding to the modality type, further comprising: When multiple modal types are parsed in response to the query content, Obtaining a query modality and a target modality according to the description completeness of the query content, wherein the description completeness of the query modality is greater than that of the target modality; Comparing the embedding vector of the query modality with the feature vector of the data in the corresponding data source to obtain a query vector; According to the correlation between the query vector corresponding to the query modality and the feature vector corresponding to the target modality, the target vector is obtained; An access result is obtained according to the query vector and the target vector.

[0015] The comparison model is specifically: According to the text type, image type, and audio type, a trimodal training set is obtained; According to the trimodal training set, corresponding feature vectors are obtained through corresponding encoders; According to the similarity of the feature vectors between the modalities, the feature vectors between the modalities are aligned and associated; In response to user query feedback, a reward value is set and the alignment association of feature vectors between modalities is adjusted.

[0016] The present invention also discloses a processing device, comprising: memory for storing computer programs; A processor is configured to implement the steps of the cross-modal data query method when executing the computer program.

[0017] Due to the adoption of the above technical solution, the beneficial effects achieved by the present invention are as follows: 1. This invention pre-extracts multiple data source descriptors based on a preset triple registration format to generate an access interface corresponding to that data source, resolving the data heterogeneity barrier in existing technologies. By standardizing triple descriptors, interfaces are automatically generated, eliminating the need for manual adaptation. This eliminates the need to write separate adapter code for each data source, reducing development costs. Adding new data sources requires only registering triples, eliminating the need to modify core code, improving system scalability.

[0018] In response to the query content, the modality type is parsed and the corresponding data source is selected based on the modality type, solving the problem of fragmented adaptation in existing technologies and improving system flexibility. Keywords are identified through grammatical analysis and dynamically selected using a modality mapping table. This improves modality parsing accuracy and avoids invalid queries caused by misjudgments. By selecting the corresponding data source based on the modality type, full-modality queries are avoided, improving system throughput.

[0019] The access path to the data source is selected based on the data source's historical latency data, resolving the high query latency issue in existing technologies and significantly reducing response time. By selecting the optimal path based on historical latency data, average query latency is reduced, throughput is improved in high-concurrency scenarios, and recalculation is avoided using a caching mechanism.

[0020] The cache mode is set according to the number of modal types. The present invention optimizes resource utilization through a multi-level cache mechanism, solves the problem of repeated calculations in the prior art, and significantly improves system efficiency. The short-term cache is used to store original query results (such as text retrieval results) to avoid repeated requests. The long-term cache is used to store cross-modal feature vectors for multiple query reuse. The cache mode is set according to the number of modal types (single modality or cross-modality). Single modality uses short-term cache, and cross-modality uses long-term cache. The cache hit rate is improved and resource utilization is improved.

[0021] In addition, according to the access path, the corresponding database is accessed, and through a preset comparison model, based on the query content and the correlation between the data in the data source corresponding to the modality type, the access result of the corresponding modality type is obtained, and cached based on the cache mode, which solves the semantic gap problem in the existing technology, significantly improves the accuracy of cross-modal queries, supports complex queries, and improves the accuracy of cross-modal queries. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings: Figure 1 The figure is a flowchart of the cross-modal data query method according to one embodiment of the present invention. DETAILED DESCRIPTION

[0023] In order to more clearly illustrate the overall concept of the present invention, a detailed description is given below in an exemplary manner in conjunction with the accompanying drawings.

[0024] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the scope of protection of the present invention is not limited to the specific embodiments disclosed below.

[0025] like Figure 1 As shown, a cross-modal data query method includes: S100: extracting multiple descriptor information of a data source in advance according to a preset triple registration format to generate an access interface corresponding to the data source for accessing data from the data source.

[0026] The core goal of this step is to solve the data heterogeneity barrier problem in existing technologies, uniformly describe multi-source heterogeneous data through a preset triple registration format, and automatically generate standardized access interfaces, thereby reducing development complexity and improving system scalability.

[0027] The data source must be<modality_type, data_format, access_protocol> The modality_type is a triplet for registration, where modality_format represents the modality type, data_format represents the data format, and access_protocol represents the access protocol.

[0028] The data source submits a descriptor (a triplet) to the MCP agent, specifying its type, format, and access protocol. The MCP agent automatically generates the corresponding data source access interface code based on the descriptor (for example, encapsulating gRPC as a REST API), eliminating the need for developers to manually write adaptation logic.

[0029] In addition, when adding a new data source, you only need to submit a new triplet descriptor, and the system will automatically adapt and generate an interface without modifying the core code.

[0030] Understandably, traditional methods require writing independent adaptation code for each data source (such as text, image, and audio) (for example, an image database requires developing an image encoder interface, and an audio database requires developing an audio decoder interface), which results in high development costs and difficult maintenance.

[0031] The present invention automatically generates interfaces by standardizing triplet descriptors without manual adaptation. For example, the image data source is registered as<image, jpeg, http> Afterward, the system automatically encapsulates the data into a unified REST API, freeing developers from worrying about underlying implementation details. This eliminates the need to write separate adapter code for each data source, reducing development costs. Adding a new data source requires only registering a triplet, without modifying the core code, improving system scalability.

[0032] Furthermore, unified access to multimodal data is achieved by uniformly describing all data sources in a triple format and dynamically generating interfaces by parsing triples. Standardized descriptors are used to uniformly manage heterogeneous data from multiple sources, simplifying the system architecture. Developers no longer need to write adaptation logic for different modal data, improving development efficiency.

[0033] This step systematically solves the data heterogeneity barrier in cross-modal data queries through triple registration format and automated interface generation, which not only reduces development costs but also improves the system's scalability and dynamic adaptability.

[0034] The method also includes, S200: in response to the query content, parsing to obtain the modality type, selecting a data source of a corresponding type according to the modality type, selecting an access path of the data source according to historical delay data of the data source, and setting a cache mode according to the number of the modality types.

[0035] The core goal of this step is to solve the problems of low query efficiency and resource waste in existing technologies. By dynamically parsing modality types, optimizing data source access paths, and setting cache modes, the latency of cross-modal queries can be significantly reduced, thereby improving system throughput and resource utilization.

[0036] Syntax analysis of the query content identifies keywords and parses the modality type. For example, if a user enters "Find documents related to this image," the system parses the modality types: text (text type) and image (image type). Based on the parsed modality type, the corresponding type is matched from registered data sources (e.g., image → image database, audio → audio database). Combined with the MCP protocol's triple descriptor, the adaptation code is dynamically loaded, eliminating the need for manual interface development. Syntax analysis automatically identifies modalities and dynamically parses modality types, avoiding full-modality queries.

[0037] It should be noted that data source access paths generally include local resource access and cloud API access. The latency of the current data source on the local resource access path is estimated to determine whether to call the local resource access or cloud API access path. The optimal path is selected based on historical latency data, reducing response time and average query latency.

[0038] The cache mode is set based on the number of modal types. For unimodal queries (e.g., image only), a short-term cache is set. For cross-modal queries (e.g., image + text), a long-term cache is set. The short-term cache stores the original query results (e.g., text retrieval results) to avoid repeated requests; the long-term cache stores cross-modal feature vectors for reuse in multiple queries. It is understandable that the computational complexity of cross-modal queries is much greater than that of unimodal queries. Therefore, a long-term cache is set for cross-modal queries, with a longer cache time than for unimodal queries, to balance redundant computation and resource utilization.

[0039] This step systematically solves the problems of high query latency and resource waste in cross-modal data queries through dynamic parsing of modality types, intelligent routing algorithms, and multi-level caching mechanisms, significantly improving query efficiency and resource utilization.

[0040] S300: According to the access path, the corresponding database is accessed, and through a preset comparison model, based on the query content and the correlation between the data in the data source corresponding to the modality type, an access result of the corresponding modality type is obtained, and cached based on the cache mode.

[0041] The core goal of this step is to address the semantic gap problem in existing technologies. By aligning the multimodal feature space through a preset comparison model and combining it with a caching mechanism to reduce redundant calculations, the accuracy and efficiency of cross-modal queries can be significantly improved.

[0042] Based on the selected access path (such as local GPU or cloud API), the corresponding data source interface (such as REST API) is called to extract data related to the query content (such as image feature vectors, audio embeddings) from the database.

[0043] Using a pre-set comparison model, the query content and data source data are semantically correlated to generate access results. This supports complex queries and improves cross-modal accuracy. Access results are stored based on the configured caching mode (short-term or long-term).

[0044] This step systematically addresses the semantic gap and resource waste problems in cross-modal data queries by comparative learning to align multimodal feature spaces and multi-level caching mechanisms, significantly improving query accuracy and resource utilization.

[0045] As a preferred embodiment of the present invention, the triplet registration format is specifically: The triplet registration format includes modality type, data format, and access protocol. The modality types include text type, image type, and audio type.

[0046] The core goal of this implementation is to solve the fragmented adaptation problem in the existing technology, uniformly describe multi-source heterogeneous data through a standardized triple registration format, and automatically generate a standardized access interface, thereby reducing development complexity and improving system scalability.

[0047] The triple format is as follows:<modality_type, data_format, access_protocol> modality_type indicates the modality type, including text, image, and audio. data_format indicates the data format, such as json (text), jpeg (image), and mp3 (audio). access_protocol indicates the access protocol, such as http (HTTP interface), grpc (gRPC interface), and s3 (object storage protocol).

[0048] For example, the image data source is registered as<image, jpeg, http> ;Audio data source is registered as<audio, mp3,grpc> .

[0049] It is understandable that the access methods of different data sources can be automatically adapted according to the triple content. For example,<image, jpeg, http> The HTTP interface call code will be generated.<audio, mp3, grpc> gRPC interface call code is generated. Regardless of the protocol or format of the data source, the system encapsulates it into a standardized interface (such as a REST API) through a triple descriptor. Developers can access different data sources by simply calling the unified interface, without manual adaptation. Developers do not need to worry about the underlying implementation details, reducing development costs and improving system scalability.

[0050] In this implementation, the "fragmented adaptation" problem in cross-modal data queries is systematically solved through standardized descriptors in triple registration format.

[0051] As a preferred embodiment of the present invention, in response to the query content, the modality type is parsed and obtained, specifically: Based on the query content, keywords are identified through grammatical analysis, so as to parse and obtain the modality type according to a mapping relationship table between the keywords and the modality type.

[0052] The core goal of this implementation is to identify keywords through grammatical analysis and parse modal types in combination with a mapping relationship table, thereby dynamically adapting data sources and improving system flexibility and query efficiency.

[0053] Based on the grammatical structure of the query, keywords (such as "image," "audio," and "document") are extracted. Specifically, natural language processing (NLP) techniques (such as word segmentation and dependency parsing) are used to identify entities or function words in the query. For example, if a user enters "Find documents related to this image," the system extracts the keywords "image" and "document."

[0054] Keywords are converted to their corresponding modal types based on a pre-defined keyword-modal type mapping table (e.g., "image" → image, "audio" → audio, "document" → text). Modalities are dynamically identified through syntax analysis and the mapping table, eliminating the need for manual intervention. Furthermore, adding new data sources requires only updating the mapping table, without modifying the core code.

[0055] According to the parsed modal type (such as image or text), the corresponding type is matched from the registered data source (such as<image, jpeg, http> or<text, json, elasticsearch> ), and call its interface to reduce resource waste by reducing full-modal queries and improve system throughput.

[0056] This embodiment identifies keywords through grammatical analysis and parses modal types in combination with a mapping relationship table, thereby reducing development complexity and improving system flexibility and query efficiency.

[0057] As a preferred embodiment of the present invention, the access path of the data source is selected according to the historical delay data of the data source, specifically: Based on the latency data obtained from historical queries and the load status of the current data source access path, a learning model is used to predict the latency of the current data source on different paths. If the local GPU latency is lower than the threshold, the data source of the local resource path is called; Otherwise, call the data source of the cloud API path.

[0058] The core goal of this implementation is to address the high query latency problem in existing technologies. By combining a dynamic routing algorithm with a learning model to predict latency and selecting the optimal access path (local GPU or cloud API) based on a threshold, the response time of cross-modal queries can be significantly reduced, thereby improving system throughput.

[0059] Collect historical data source latency data (such as average latency of local GPUs and network latency of cloud APIs). Combined with the load status of the current data source access path (such as GPU utilization and network bandwidth), use learning models (such as linear regression and LSTM) to predict the current data source latency along different paths.

[0060] Specifically, historical data is used to train the model and learn the relationship between latency and load status; real-time load status is obtained through monitoring tools (such as Prometheus) and input into the model for dynamic prediction.

[0061] Based on the prediction results, a threshold-based access path selection strategy is used. If the local GPU latency is below the threshold, the local GPU is prioritized for processing; if the local GPU latency is above or equal to the threshold, the cloud API is called. This implementation dynamically selects the optimal path based on historical latency data and real-time load status.

[0062] This implementation solves the problems of high query latency and resource waste in cross-modal data queries by integrating historical delay data with load status, learning models to predict delays, and making dynamic routing decisions, thus balancing query efficiency and resource utilization.

[0063] As a preferred embodiment of this implementation, the threshold is specifically: According to the modality type, initial thresholds are set one by one, with the initial thresholds of the text type, the image type, and the audio type decreasing in sequence; According to the complexity of the query content, the initial threshold is adjusted to obtain the threshold, and the threshold is negatively correlated with the complexity.

[0064] This embodiment further optimizes the access path selection between the local GPU and the cloud API by dynamically adjusting the threshold (based on the modality type and query complexity), thereby significantly reducing the response time of cross-modal queries and improving system throughput and resource utilization.

[0065] Set the initial threshold according to the modality type (text, image, audio): For text types, the initial threshold is higher (such as 200ms) because it has a higher tolerance for delay; for image types, the initial threshold is medium (such as 150ms) because it needs to balance calculation and delay; for audio types, the initial threshold is the lowest (such as 100ms) because it has a higher requirement for real-time performance.

[0066] In addition, the initial threshold is adjusted based on the complexity of the query content, and complexity is assessed through feature extraction (such as query length and embedding vector dimension). For example, long text queries or high-dimensional audio features have higher complexity.

[0067] Complexity is negatively correlated with threshold (the higher the complexity, the lower the threshold). For example: When the complexity is 1 (low), the threshold = initial threshold × 1.0; when the complexity is 3 (high), the threshold = initial threshold × 0.5.

[0068] The access path is selected based on the adjusted threshold. If the predicted latency is lower than the adjusted threshold, the local GPU is called; otherwise, the cloud API is called.

[0069] For example, if the image query complexity is 2, the initial threshold is 150ms, the adjusted threshold is 120ms, and the predicted local GPU latency is 110ms, the local path is selected.

[0070] It’s important to note that static thresholds cannot adapt to the needs of different modalities and query complexity. For example, a static threshold for text queries might be too low, leading to unnecessary cloud calls, while a static threshold for audio queries might be too high, overloading the local GPU.

[0071] Dynamically adjust the threshold based on modality type and query complexity to ensure path selection is more tailored to actual needs. For example, high-complexity audio queries have lower thresholds, prioritizing the use of the local GPU to reduce latency. This reduces average query latency and improves throughput in high-concurrency scenarios by reducing redundant requests and optimizing path selection.

[0072] Furthermore, the threshold is dynamically adjusted based on complexity to balance the load between the local GPU and the cloud API. For example, low-complexity text queries allow for higher thresholds to fully utilize local resources, while high-complexity audio queries have strict threshold limits to avoid overloading the local GPU.

[0073] This embodiment sets an initial threshold based on the modality type, dynamically adjusts the threshold based on query complexity, and selects an access path based on the adjusted threshold. This solves the problems of high query latency and resource waste in cross-modal data queries and significantly improves query efficiency and resource utilization.

[0074] As a preferred embodiment of the present invention, a cache mode is set according to the number of modality types, specifically: When the modality type obtained by parsing the query content is 1, the access result cache mode is set to short-term cache, and a first pre-cache time is set accordingly; otherwise, the access result cache mode is set to long-term cache, and a second pre-cache time is set accordingly, wherein the first pre-cache time is shorter than the second pre-cache time; Adjusting the first pre-caching time and the second pre-caching time according to the modality type obtained by parsing and the hit rate of the query content, thereby obtaining a first cache time and a second cache time accordingly; The first cache time and the second cache time are positively correlated with the hit rate.

[0075] This implementation significantly improves system resource utilization and query efficiency, and reduces redundant calculations and data access delays by dynamically setting the cache mode (short-term cache or long-term cache) according to the number of modal types and adjusting the cache time based on the hit rate of the query content.

[0076] Set the cache mode based on the number of parsed modal types (single modal or cross modal): For single-modal queries (such as image-only queries), set up a short-term cache to store the original query results (such as text retrieval results); For cross-modal queries (such as image+text), set up long-term cache to store cross-modal feature vectors (such as CLIP image embedding).

[0077] Adjust the pre-caching time based on the parsed modality type and query hit rate. Count the number of times the same query is repeated within a specific time period. For example, a high-frequency query like "dog pictures" has a high hit rate, while a low-frequency query like "rare bird audio" has a low hit rate. When the hit rate is high, extend the short-term cache time; when the hit rate is low, extend the long-term cache time.

[0078] Additionally, set a hit rate coefficient based on the modality type to adjust the cache duration. Note that the hit rate coefficients for text, image, and audio types increase in descending order. Ensure that the first and second cache times are positively correlated with the hit rate: short-term cache time = first pre-cache time × (1 + hit rate coefficient); long-term cache time = second pre-cache time × (1 + hit rate coefficient).

[0079] Specifically, if the hit rate of a cross-modal query increases from 20% to 80%, the long-term cache time is extended from 72 hours to 30 days.

[0080] Single-modal queries use short-term caching to avoid repeated requests; cross-modal queries use long-term caching and reuse feature vectors. The cache duration for high-frequency queries is extended based on the hit rate, reducing repeated calculations, improving cache hit rates, and enhancing resource utilization. Furthermore, the cache duration for high-hit-rate queries is extended, while that for low-hit-rate queries is shortened, optimizing cache space allocation.

[0081] This implementation solves the problem of resource waste in cross-modal data queries by setting the cache mode (short-term cache or long-term cache) according to the number of modal types and dynamically adjusting the cache time based on the query hit rate, thereby significantly improving the cache hit rate and resource utilization.

[0082] As a preferred implementation of the present invention, based on the query content and the correlation between the data in the data source corresponding to the modality type, an access result of the corresponding modality type is obtained.

[0083] In one embodiment, when the modality type obtained by parsing in response to the query content is 1, An embedding vector is obtained according to the query content, and the embedding vector is compared with a feature vector of data in a corresponding data source. An access result is returned according to the comparison result.

[0084] The core goal of this embodiment is to accurately match the query content with the data source data by comparing the embedding vector and the feature vector under unimodal query, thereby significantly improving the accuracy and efficiency of unimodal query.

[0085] Based on the query content (such as "find pictures similar to this picture"), the embedding vector is extracted and the image is converted into the embedding vector through encoders such as CLIP and ResNet.

[0086] The query embedding vector is compared with the feature vectors stored in the data source for similarity. This embodiment does not limit the comparison method; algorithms such as cosine similarity and Euclidean distance can be used to calculate the degree of match between vectors. It is understood that the data in the data source has been converted into feature vectors and stored in advance using the same encoder.

[0087] Specifically, the embedding vector of the query image is compared with the CLIP image feature vector stored in the data source, and the similarity is calculated and the Top-K matching results are returned.

[0088] Based on the comparison results (such as similarity sorting), the access results that best match the query content are returned. For single-modal matching, the data with the highest matching degree (such as "dog pictures") are directly returned.

[0089] In another embodiment, obtaining access results of the corresponding modality type based on the correlation between the query content and the data in the data source corresponding to the modality type further includes: When multiple modal types are parsed in response to the query content, Obtaining a query modality and a target modality according to the description completeness of the query content, wherein the description completeness of the query modality is greater than that of the target modality; Comparing the embedding vector of the query modality with the feature vector of the data in the corresponding data source to obtain a query vector; According to the correlation between the query vector corresponding to the query modality and the feature vector corresponding to the target modality, the target vector is obtained; An access result is obtained according to the query vector and the target vector.

[0090] The core goal of this embodiment is to accurately match cross-modal data (such as image → text, audio → text) through description completeness evaluation and feature vector association analysis under multimodal queries, significantly improving the accuracy of complex queries and system scalability.

[0091] The query modality and target modality are distinguished based on the completeness of the description in the query content. Description completeness is defined as the query modality being more specific (e.g., "dog pictures"), while the target modality being more vague (e.g., "dog barking"). Specifically, the query content is analyzed for entity completeness through natural language processing (NLP). (e.g., "dog pictures" contains both a specific entity and modality type, while "dog barking" only contains an entity and a vague modality).

[0092] The query embedding vector is compared with the feature vector of the data in the corresponding data source. The query embedding vector is then compared with the feature vector stored in the data source (e.g., image feature vector) by performing a similarity calculation (e.g., cosine similarity) to generate a query vector.

[0093] A target vector is generated based on the correlation between the query vector and the feature vector of the target modality. Based on the query vector and the target vector, the query results are returned. By weightedly fusing the query vector and the target vector, the data that best matches the query content is returned (for example, matching a "dog picture" with a "dog barking" audio file).

[0094] This embodiment evaluates the completeness of the query description, prioritizing high-completeness modalities (such as images) and then generating results for low-completeness modalities (such as audio) through feature alignment. The query vector is derived from the embedding vector describing the complete query modality, and the target vector is further derived from the query vector. This significantly improves the accuracy of access results compared to directly deriving the target vector from the embedding vector describing the incomplete target modality.

[0095] This embodiment significantly improves the accuracy and resource utilization of complex queries through description completeness evaluation, feature vector comparison, and cross-modal alignment under multimodal queries.

[0096] As a preferred embodiment of the present invention, the comparison model is specifically: According to the text type, image type, and audio type, a trimodal training set is obtained; According to the trimodal training set, corresponding feature vectors are obtained through corresponding encoders; According to the similarity of the feature vectors between the modalities, the feature vectors between the modalities are aligned and associated; In response to user query feedback, a reward value is set and the alignment association of feature vectors between modalities is adjusted.

[0097] The core goal of this implementation is to significantly improve the accuracy and system adaptability of cross-modal queries by constructing a trimodal training set, aligning feature vectors, and dynamically optimizing alignment associations based on user feedback, while reducing reliance on manual labeling.

[0098] Based on three types of modal data: text, image, and audio, a trimodal training set is constructed. By automatically matching natural language descriptions with multimodal data (such as using the CLIP model to generate text-image pairs), the cost of manual labeling is reduced.

[0099] The trimodal feature vector is extracted through the corresponding encoder, where the text encoder uses BERT or Transformer model to extract semantic features; the image encoder uses CLIP or ResNet to extract visual features; and the audio encoder uses AST or Wav2Vec to extract audio features.

[0100] By contrastive learning to align the feature vectors between modalities, InfoNCE loss or Triplet Loss is used to maximize the similarity of matching modality pairs and minimize the similarity of unmatched pairs.

[0101] Specifically, text-image alignment (TV) uses the CLIP dual-tower structure to narrow the feature distance between semantically similar text and images; text-audio alignment (TA) uses the VALOR model to jointly optimize text and audio features; image-audio alignment (VA) aligns visual and auditory features through the cross-modal attention mechanism.

[0102] In response to user query feedback, reward values ​​are set and alignment associations are adjusted. The user clicks on a positive example (e.g., "correctly matched audio") or an error tag (e.g., "irrelevant image") as a reward signal. Reward values ​​are set based on the feedback type (positive / negative) (e.g., +1 for positive examples, -0.5 for negative examples). Feature alignment parameters are adjusted using reinforcement learning (e.g., Policy Gradient) or online learning (e.g., incremental model updates).

[0103] Specifically, through the user points incentive system and dopamine feedback mechanism, user feedback is transformed into the driving force for model optimization.

[0104] This implementation uses comparative learning to align text, image, and audio feature vectors (e.g., TV, TA, and VA) to establish a unified semantic space and implement trimodal alignment training. Alignment relationships are dynamically adjusted through a reward mechanism (e.g., reinforcement learning to update model parameters), with user feedback driving optimization.

[0105] This implementation improves query accuracy by constructing a trimodal training set, aligning feature vectors, and dynamically optimizing alignment associations based on user feedback.

[0106] The present invention also provides a processing device, comprising: Memory for storing computer programs; A processor is configured to implement the steps of the cross-modal data query method when executing the computer program.

[0107] Therefore, any effect in the cross-modal data query method can be achieved, which will not be elaborated here.

[0108] Anything not described in the present invention can be achieved by adopting or drawing on existing technologies.

[0109] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.

[0110] The foregoing is merely an embodiment of the present invention and is not intended to limit the present invention. It will be apparent to those skilled in the art that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention are intended to be included within the scope of the claims of the present invention.

Claims

1. A cross-modal data query method, characterized in that: include: Extracting multiple descriptor information of a data source in advance according to a preset triple registration format to generate an access interface corresponding to the data source for accessing data from the data source; The method further includes, in response to the query content, parsing to obtain a modality type, selecting a data source of a corresponding type according to the modality type, selecting an access path to the data source according to historical latency data of the data source, and setting a cache mode according to the number of the modality types; According to the access path, the corresponding database is accessed, and through a preset comparison model, based on the query content and the correlation between the data in the data source corresponding to the modality type, the access result of the corresponding modality type is obtained, and cached based on the cache mode.

2. The cross-modal data query method according to claim 1, characterized in that: The triplet registration format is specifically: The triplet registration format includes modality type, data format, and access protocol. The modality types include text type, image type, and audio type.

3. The cross-modal data query method according to claim 1, characterized in that: In response to the query content, the modal type is parsed, specifically: Based on the query content, keywords are identified through grammatical analysis, so as to parse and obtain the modality type according to a mapping relationship table between the keywords and the modality type.

4. The cross-modal data query method according to claim 2, characterized in that: Based on the historical delay data of the data source, select the access path of the data source, specifically: Based on the latency data obtained from historical queries and the load status of the current data source access path, a learning model is used to predict the latency of the current data source on different paths. If the local GPU latency is lower than the threshold, the data source of the local resource path is called; Otherwise, call the data source of the cloud API path.

5. The cross-modal data query method according to claim 4, characterized in that: The threshold is specifically: According to the modality type, initial thresholds are set one by one, with the initial thresholds of the text type, the image type, and the audio type decreasing in sequence; According to the complexity of the query content, the initial threshold is adjusted to obtain the threshold, and the threshold is negatively correlated with the complexity.

6. The cross-modal data query method according to claim 1, characterized in that: According to the number of modal types, set the cache mode, specifically: When the modality type obtained by parsing the query content is 1, the access result cache mode is set to short-term cache, and a first pre-cache time is set accordingly; otherwise, the access result cache mode is set to long-term cache, and a second pre-cache time is set accordingly, wherein the first pre-cache time is shorter than the second pre-cache time; Adjusting the first pre-caching time and the second pre-caching time according to the modality type obtained by parsing and the hit rate of the query content, thereby obtaining a first cache time and a second cache time accordingly; The first cache time and the second cache time are positively correlated with the hit rate.

7. The cross-modal data query method according to claim 1, characterized in that: Based on the query content and the correlation between the data in the data source corresponding to the modality type, an access result of the corresponding modality type is obtained, including: When the parsed modality type is 1 in response to the query content, An embedding vector is obtained according to the query content, and the embedding vector is compared with a feature vector of data in a corresponding data source. An access result is returned according to the comparison result.

8. The cross-modal data query method according to claim 1, characterized in that: Obtaining access results corresponding to the modality type based on the correlation between the query content and the data in the data source corresponding to the modality type, further comprising: When multiple modal types are parsed in response to the query content, Obtaining a query modality and a target modality according to the description completeness of the query content, wherein the description completeness of the query modality is greater than that of the target modality; Comparing the embedding vector of the query modality with the feature vector of the data in the corresponding data source to obtain a query vector; According to the correlation between the query vector corresponding to the query modality and the feature vector corresponding to the target modality, the target vector is obtained; An access result is obtained according to the query vector and the target vector.

9. The cross-modal data query method according to claim 7 or 8, characterized in that: The comparison model is specifically: According to the text type, image type, and audio type, a trimodal training set is obtained; According to the trimodal training set, corresponding feature vectors are obtained through corresponding encoders; According to the similarity of the feature vectors between the modalities, the feature vectors between the modalities are aligned and associated; In response to user query feedback, a reward value is set and the alignment association of feature vectors between modalities is adjusted.

10. A processing device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the cross-modal data query method according to any one of claims 1 to 9 when executing the computer program.

Citation Information

Patent Citations

  • Parameter alignment method and device for visual language large model and storage medium

    CN119558379A

  • Cross-modal image-text retrieval processing method and system

    CN119988664A

  • Online teaching interaction method based on multi-modal knowledge graph, medium and equipment

    CN120339011A

  • Method and system for federated querying of data sources

    US20050234889A1

  • Cross-modal data processing method and device, storage medium, and electronic device

    WO2022068196A1