Multimodal database query engine and method based on unified semantic representation
By using a multimodal database query engine based on unified semantic representation, the challenges of semantic understanding and cross-modal association in multimodal data processing of traditional databases are solved, enabling efficient querying and intelligent retrieval of multimodal data, and improving query efficiency and semantic relevance.
Patent Information
- Application Number
- CN202511097961.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2045-08-06
AI Technical Summary
Traditional database query engines cannot effectively perform semantic understanding and cross-modal association when processing multimodal data, resulting in a lack of coordinated optimization of semantic relevance and execution efficiency. Existing multimodal data processing solutions cannot simultaneously optimize semantic relevance and physical execution efficiency.
A multimodal database query engine based on unified semantic representation is adopted. Through the query parsing and semantic encoding module, the query execution plan generation module, the multimodal query execution module, and the query result fusion module, unified semantic representation and query processing of multimodal data are realized. The query execution plan is optimized by combining dynamic weight adaptive algorithm.
It realizes intelligent and optimized cross-modal query operations, improves the semantic understanding capability and query efficiency of multimodal data, solves the problems of superficial semantic understanding and difficulty in cross-modal association in traditional multimodal data query, and provides intelligent retrieval and application solutions for massive heterogeneous data.
Smart Images

Figure CN120929654B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of database, in particular to a multi-modal database query engine and method based on unified semantic representation. BACKGROUND
[0002] With the deepening of digital transformation, enterprises and organizations have accumulated massive multi-modal data, including text, image, audio and video, etc. How to effectively manage and query these heterogeneous data has become a major challenge in the current database technology field. Traditional database query engines are mainly designed for structured data and have limited support for multi-modal data queries, especially in the aspect of semantic understanding and association. There is a technical gap between current multi-modal AI research and database engineering practice. Multi-modal AI models focus on optimizing semantic accuracy, but lack awareness of underlying data access cost and execution efficiency. Traditional database query optimizers have mature cost estimation and execution plan generation capabilities, but lack semantic understanding capabilities and cannot handle cross-modal semantic association. Existing multi-modal data processing solutions cannot optimize semantic relevance and physical execution efficiency simultaneously, which is manifested in the lack of a unified semantic representation mechanism for heterogeneous data types, the inability of semantic association to be optimized in coordination with execution plans, and the inability of traditional query optimization techniques to perceive semantic features for intelligent decision-making. Therefore, a unified optimization mechanism that can simultaneously perceive semantic features and execution features is needed to achieve the coordinated improvement of semantic accuracy and execution efficiency. SUMMARY
[0003] To solve the above technical problems, the technical scheme adopted by the present application is as follows:
[0004] According to the first aspect of the present application, a multi-modal database query engine based on unified semantic representation is provided, the query engine is associated with a vector library, the vector library stores unified semantic vectors of a plurality of entities, and the query engine comprises:
[0005] A query analysis and semantic encoding module is configured to perform semantic analysis on a user input query request to obtain corresponding semantic structure features, and to perform unified semantic encoding on the semantic structure features to generate corresponding encoding results. If the semantic structure features contain single-modal data, the corresponding semantic vector of the single-modal data is directly generated as the encoding result. If the semantic structure features contain multi-modal data, the semantic vector of each modality is generated and the generated semantic vectors are fused to generate a fused semantic vector as the encoding result. The query request includes single-modal query and multi-modal query. The semantic structure features include query intent, key entity and modality information.
[0006] A query execution plan generation module is configured to generate an optimal query execution plan based on semantic features in the encoding result, data distribution characteristics of the vector library, query complexity, and query engine resource status, wherein the optimal query execution plan comprises a vector retrieval strategy, a modality weight allocation, and a computing resource scheduling scheme.
[0007] A multi-modal query execution module is configured to perform a query operation on the vector library according to the optimal query execution plan, and obtain candidate results similar to the semantic of the query request.
[0008] A query result fusion module is configured to perform a weighted fusion operation on the multi-path results, and generate a weighted score of each candidate result.
[0009] An output module is configured to output each candidate result in a descending order of the weighted score. According to a second aspect of the present application, a multi-modal database query method based on unified semantic representation is provided, and the method is implemented based on the query engine provided in the first aspect of the present application, and the method comprises the following steps:
[0010] Performing semantic analysis on a query request input by a user to obtain corresponding semantic structure features; the query request comprises a single-modal query and a multi-modal query; the semantic structure features comprise a query intention, key entities, and modality information.
[0011] Performing unified semantic encoding on the semantic structure features to generate corresponding encoding results, wherein if the semantic structure features comprise single-modal data, a semantic vector of the single-modal data is directly generated as the encoding result; if the semantic structure features comprise multi-modal data, semantic vectors of each modality are generated, and the generated semantic vectors are fused to generate a fused semantic vector as the encoding result.
[0012] Generating an optimal query execution plan based on semantic features in the encoding result, data distribution characteristics of the vector library, query complexity, and query engine resource status, wherein the optimal query execution plan comprises a vector retrieval strategy, a modality weight allocation, and a computing resource scheduling scheme.
[0013] Performing a query operation on the vector library according to the optimal query execution plan, and obtaining candidate results similar to the semantic of the query request.
[0014] Performing a weighted fusion operation on the multi-path results, and generating a weighted score of each candidate result.
[0015] Outputting each candidate result in a descending order of the weighted score.
[0016] The present application has at least the following beneficial effects:
[0017] The multi-modal database query engine and method based on unified semantic representation provided by the embodiment of the application map heterogeneous data to a unified semantic space through a fusion device with encoder and execution plan perception, wherein the fusion device adopts a three-factor dynamic weight adaptive algorithm to combine database execution plan theory and multi-modal semantic fusion; a semantic perception-based query optimization framework incorporates semantic characteristics and execution plan features into an optimization decision process based on weight calculation results of the fusion device; and a multi-modal query execution and result fusion mechanism supports cross-modal query operations based on dynamic weight configuration of the fusion device. The query engine effectively solves key technical problems such as shallow semantic understanding, cross-modal association difficulty and low query efficiency in traditional multi-modal data query through the dynamic weight adaptive algorithm with execution plan perception, and achieves important breakthroughs in the intelligentization of weight allocation and the precision of execution optimization, thereby providing a complete technical solution for intelligent retrieval and application of massive heterogeneous data.
[0018] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the application, nor is it used to limit the scope of the application. Other features of the application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.
[0020] Figure 1 The structural block diagram of the multi-modal database query engine based on unified semantic representation provided by the embodiment of the application is shown in the figure.
[0021] Figure 2 The flowchart of the multi-modal database query method based on unified semantic representation provided by the embodiment of the application is shown in the figure. DETAILED DESCRIPTION
[0022] The technical solutions in the embodiments of the application will be described clearly and completely in combination with the drawings in the embodiments of the application. Obviously, the described embodiments are only some of the embodiments of the application, not all. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the application.
[0023] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the description herein is for describing particular embodiments only and is not intended to be limiting of the application. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0024] It is to be understood that some of the example embodiments are described in terms of a process or method being performed at a location, such as an electronic processing device. Although processes or methods can be described as being performed at a location, such as an electronic processing device, it is understood that such processes or methods can be performed at any location, such as an electronic processing device, and that the processes or methods include any suitable steps, which can be performed at any suitable location.
[0025] The present application aims to provide a multi-modal semantic understanding capability into the query engine architecture, to achieve a unified semantic representation of a variety of modal data and query processing based on unified semantic representation of multi-modal database query engine.
[0026] The query engine in the embodiments of the present application is associated with a vector library, and the vector library stores unified semantic vectors of a plurality of entities.
[0027] In the embodiments of the present application, the vector library is constructed based on a Unified Semantic Representation (USR) model. The USR model includes a plurality of encoders and a fusioner. The encoders are used to convert data of different forms into unified semantic vector representation, so that the query engine can process data of multiple modalities under the same semantic framework. In the embodiments of the present application, a special encoder implementation is designed for each main data modality to ensure that each modality data can be accurately mapped into a unified semantic space, which can include, for example, a semantic encoder, an image encoder, a video encoder, and an audio encoder, etc.
[0028] The text encoder implements the Etext(xtext) function defined in the USR model, and its implementation process follows strict technical specifications to ensure the accuracy and consistency of semantic encoding. The text preprocessing stage adopts standardized word segmentation and cleaning processes, including removing special characters, unifying encoding formats, and sentence processing, to ensure the consistency of the input text format. The semantic encoding stage uses a pre-trained language model based on the Transformer architecture to extract deep semantic features of the text through multiple layers of self-attention mechanisms, generating 768-dimensional semantic vector representations. The vector normalization stage performs L2 normalization on the generated semantic vectors to ensure that the vector length is 1, allowing semantic vectors of different texts to be compared on a unified unit sphere. The quality verification stage verifies the encoding quality through semantic similarity tests to ensure that texts with similar semantics are close in vector space, and texts with large semantic differences are far apart in vector space.
[0029] The image encoder implements the Eimage(ximage) function, which has higher complexity in technical implementation than the text encoder, and needs to process multi-level semantic abstraction of visual information. The image preprocessing stage includes size standardization, color space conversion, and data enhancement operations to unify the input image to 224x224 pixels in RGB format. The feature extraction stage uses a pre-trained visual model based on the ResNet or VisionTransformer architecture to extract visual features of the image through convolutional neural networks or self-attention mechanisms. The semantic mapping stage maps the visual features to a 768-dimensional unified semantic space through a fully connected layer, ensuring consistency with the output dimension of the text SEncoder. The cross-modal alignment stage optimizes the semantic consistency between the image vector and the corresponding text description vector through a contrast learning method, so that images and texts describing the same content are close in semantic space.
[0030] The audio encoder implements the Eaudio(xaudio) function, which uses an audio encoding model based on Wav2Vec or similar architecture to extract semantic features of the audio through time-frequency analysis and deep learning methods. The video encoder implements the Evideo(xvideo) function, which uses a spatio-temporal convolutional network or video Transformer architecture to generate semantic representations of videos through frame-level feature extraction and temporal modeling. The design of these two encoders follows the same output specifications as the text and image encoders, ensuring that the semantic representations of all modalities have a unified mathematical structure.
[0031] In the embodiment of the present application, strict unified norms are formulated for the outputs of all encoders to ensure that the semantic representations of different modal data have complete compatibility. The outputs of all encoders are unified into 768-dimensional L2 normalized vector representations, each dimension of the vector has a value range of [-1, 1], and the vector length is strictly equal to 1. This unified vector representation ensures that the data of different modalities can be effectively compared and associated in a unified semantic space, and the semantic correlation between different modal data can be directly evaluated by cosine similarity calculation, with a similarity value range of [-1, 1], wherein 1 represents complete similarity, -1 represents complete opposition, and 0 represents irrelevance.
[0032] The fusion device integrates different modal information by performing a dynamic weight adaptive algorithm of planned perception based on the outputs of the encoders.
[0033] In the embodiment of the present application, the unified semantic vector of each entity is obtained from the source data of the entity, specifically, the single-modal semantic vector of each modal data contained in the source data of the entity after being converted by the corresponding encoder, that is, the single-modal semantic vector of each type of entity data converted by the encoder is pre-stored in the vector library.
[0034] In the embodiment of the present application, the single-modal semantic vector of a certain entity in the vector library satisfies the following conditions: u =v0 u +∑ t≠u (A ut ×v0 t ), wherein v u is the optimized vector of the u-th modality of the entity, v0 u is the initial vector of the u-th modality of the entity, which is independently encoded based on the u-th modality encoder, v0 t is the initial vector of the t-th modality of the entity, which is independently encoded based on the t-th modality encoder, A ut is the alignment weight between the u-th modality and the t-th modality of the entity, and has a value range of [0, 1], reflecting the semantic correlation strength of the two modalities under the same entity, u and t have values of 1 to p, and t≠u, and p is the total number of modalities contained in the source data of the entity.
[0035] wherein A ut =exp(sum(v0 u , v0 t ) / τ) / ∑ v=1 n (exp(sum(v0 u , v0 v ), sum(v0 u , v0 t ) is the semantic similarity between the u-th modality and the t-th modality, sum(v0 u, v0 v is the semantic similarity between the u-th modality and the v-th modality, v is an integer from 1 to p, τ is a temperature coefficient, τ>0, and the default value is 0.07, used to adjust the steepness of the weight distribution.
[0036] In a specific implementation, the query engine can identify the cross-modal data pairs with semantic relevance by calculating the vector distance of different modal data in the unified semantic space, and adjust the strength of these associations through the weight mechanism of the fusioner. The alignment algorithm adopts a batch processing mode, and processes 32 sample cross-modal data pairs each time, and accelerates the construction of the similarity matrix and the calculation of the alignment weight through GPU parallel computing. The query engine maintains a cross-modal alignment cache to store the alignment results of the last 1000 queries, which is used to accelerate the processing of similar queries. Specifically, when a new query is similar to a historical query, the query engine can directly use the alignment result in the cache to avoid repeated complex cross-modal alignment calculation, thereby improving the query processing efficiency.
[0037] Further, the multi-modal database query engine based on unified semantic representation provided by the embodiments of the present application can include:
[0038] The query analysis and semantic encoding module 1 is configured to perform semantic analysis on the query request input by the user to obtain corresponding semantic structure features, and perform unified semantic encoding on the semantic structure features to generate a corresponding encoding result.
[0039] If the semantic structure features contain single-modal data, the corresponding semantic vector of the single-modal data is directly generated as the encoding result. If the semantic structure features contain multi-modal data, the semantic vectors of each modality are generated and fused to generate a fused semantic vector as the encoding result. The query request includes single-modal queries and multi-modal queries.
[0040] In the embodiment of the present application, the semantic structure features include query intention, key entity and modal information, and specifically can include: modal combination features: record the modal types and combination relationships contained in the query, such as "text + image", "audio single modal", "text-image-audio", which are represented by modal type coding, for example, text = 1, image = 2, audio = 3, etc.; query intention features: determine the query type through an intention classification model, for example, similarity retrieval / attribute filtering / association analysis / multi-round dialogue continuation, etc., output the intention label and confidence (such as "similarity retrieval, confidence 0.92"); key entity features: extract the core entities in the query (such as "product ID = 123", "person 'Zhang San'", "scene 'conference room'"), map them to knowledge graph IDs through an entity linking tool to form an entity set; constraint condition features: analyze the explicit / implicit constraints in the query (such as "time range: recent 7 days", "result quantity <= 10", "resource priority: low delay"), and convert them into structured constraint key-value pairs.
[0041] In the embodiment of the present application, the query analysis and semantic encoding module includes a plurality of modal encoders and a fusioner; the i-th modal encoder is used for unified semantic encoding of the i-th modal to generate a corresponding encoding result, and i takes a value from 1 to n, and n is the total number of modalities contained in the query. The fusioner is used for fusing the encoding results of the modal encoders to obtain a corresponding encoding result, which satisfies the following condition: F = ∑ n i=1 β i ×E i , F is the fusion result, E i is the encoding result of the i-th modal encoder, and β i is the fusion weight of the i-th modal, wherein ∑ n i=1 β i = 1 and β i ≥ 0.
[0042] The query execution plan generation module 2 is used for generating an optimal query execution plan based on the semantic features in the encoding result, the data distribution characteristics of the vector library, the query complexity and the query engine resource status, wherein the optimal query execution plan includes a vector retrieval strategy, a modal weight allocation and a computing resource scheduling scheme.
[0043] In the embodiments of the present application, the semantic feature refers to the core attribute of the query or data at the semantic level, reflecting the meaning, association and constraint requirements of the content, and directly determining the relevance and matching accuracy of the retrieval. Its core performance includes: (1) modal composition and association relationship: single-modal semantic feature: the core meaning of a single modality such as the theme of a text (“technology” “sports”), the subject of an image (“portrait” “landscape”), the scene of an audio (“speech” “music”), etc.; multi-modal semantic feature: semantic association of cross-modal content (such as “the matching degree of the visual features of the red sports car in the text description ‘red sports car’ and the image” “the semantic consistency of the video picture and the matching caption”), the higher the association strength, the easier to guarantee the accuracy of multi-modal collaborative retrieval. (2) Semantic density and core entity: semantic density: the amount of effective semantic information contained in unit data (such as a 500-word news with more than 5 core entities is high density, and an abstract picture without a clear subject is low density); core entity: a semantic unit that plays a decisive role in data (such as the query “2023 Nobel Physics Prize winner's speech video”, “2023 Nobel Physics Prize winner” and “speech video” are core entities), the more explicit the core entity, the stronger the directionality of the retrieval. (3) Semantic strength of constraint condition: strong constraint: strict limitation on semantic range (such as “time = 2024” “place = Beijing” “emotional tendency = positive”), directly narrowing the retrieval boundary, which needs to be matched first; weak constraint: fuzzy or non-mandatory semantic requirement (such as “result quantity ≥ 10” “relevant is OK”), which can be appropriately relaxed when efficiency is prioritized. (4) Distribution characteristics of semantic vector:
[0044] Aggregation in vector space: whether vectors of the same type are aggregated in the space (such as “animal” vectors are aggregated in space A area, and “plant” vectors are aggregated in B area), the higher the aggregation, the faster the retrieval can be achieved through “area anchoring”; and the offset degree of the target library: the distance between the query vector and the global semantic center of the vector library (low offset degree means that the query content is common in the library, and a general retrieval strategy can be used; high offset degree requires targeted expansion of the retrieval range).
[0045] Semantic features determine “what to retrieve” and “how to match more relevant”, which are the core basis for selecting vector retrieval strategies (such as high-precision indexing or fast rough screening), and modal weight allocation (such as prioritizing text semantics or image features).
[0046] Data distribution characteristics refer to the statistical and physical properties of data in vector libraries in terms of storage, access, association, etc., reflecting the "storage structure, quantity ratio, access frequency, and cross-modal association status of data", which directly affects the efficiency and resource consumption of queries. Its core performance includes: (1) Storage and index structure: index type: different modal data adapted index (such as text commonly used "inverted index + vector index" mixed structure, image uses HNSW index (high-efficiency approximate retrieval), audio uses IVF index (bucket retrieval, suitable for large-scale data)); Sharding and replication: whether the data is stored by modality / time / topic sharding (such as log data sharded by "day"), the number of replicas (multiple replicas can improve concurrent access capability, but increase storage cost). (2) Quantity and proportion distribution: modality proportion: the quantity proportion of each modality data in the library (such as text accounts for 60%, image accounts for 30%, audio accounts for 10%), the modality with high proportion needs to optimize the index performance first; scale level: the magnitude of single modality data (such as 10 million text, 5 million images), large-scale data (> 10 million) needs to enable distributed retrieval strategy. (3) Access frequency distribution: hot data proportion: the proportion of data accessed frequently in the past 24 hours (such as access volume > 50% of total access volume), hot data needs to be cached to memory or SSD to improve access speed; access peak characteristics: the time regularity of data access (such as text retrieval peak in the daytime, image retrieval peak at night), used for predictive scheduling of resources. (4) Cross-modal association density co-occurrence frequency: the proportion of different modal data association pairs (such as "news text-image" "product description-product video") in total data (association density > 30% is high association), high association data can establish joint index to reduce the time-consuming of cross-modal matching; association stability: the persistence of cross-modal association (such as the life cycle of "text-image" association pair > 30 days is stable association), stable association can precompute matching results to reduce real-time computation cost.
[0047] Data distribution characteristics determine "how to efficiently retrieve" and "how to allocate resources", and are the key basis for selecting computing resource scheduling schemes (such as preferentially using hot data cache, distributed node load balancing) and optimizing retrieval path (such as high association data using joint index). Query complexity uses a quantitative scoring model (0-10 points), combined with modality combination, constraint strength, precision requirement, and sub-task characteristics for multi-dimensional judgment. The specific standards are as follows:
[0048] (1) Low complexity (1-3 points)
[0049] Applicable scenario: single modality query, and only contains 1 weak constraint (such as "result quantity > 5 items"), no special precision requirement (default similarity threshold 0.6-0.7);
[0050] Subtask characteristics: Corresponding to 1-2 linear subtasks (such as "text retrieval → basic sorting"), single task computing load < 20% of 1 CPU core.
[0051] (2) Medium complexity (4-7 points)
[0052] Applicable scenarios:
[0053] Dual-modal query + 1-2 strong constraints (such as "time range = recent 7 days + region = Beijing") + regular accuracy (threshold 0.7-0.8);
[0054] Single-modal query + ≥ 3 strong constraints + regular accuracy;
[0055] Subtask characteristics: Corresponding to 2-3 subtasks (such as "image retrieval → text attribute matching → cross-validation"), allowing partial task parallelism (such as dual-modal retrieval can be executed in parallel).
[0056] (3) High complexity (8-10 points)
[0057] Applicable scenarios:
[0058] Three or more modal queries + ≥ 2 layers of nested constraints (such as "first filter'red car' images, then match 'price < 200,000' and 'rating > 4.5 stars' text attributes") + high accuracy requirements (threshold > 0.85);
[0059] Dual-modal query + ≥ 3 strong constraints + high accuracy requirements;
[0060] Subtask characteristics: Corresponding to ≥ 3 subtasks, and containing 1 or more high-computing tasks (such as "video frame feature extraction" "cross-modal attention fusion"), forcing parallel execution strategy.
[0061] Query engine resource status is evaluated in real time through the following quantitative indicators, supporting resource adaptation of execution plans:
[0062] (1) Core load indicators (including threshold determination)
[0063] CPU utilization: ≤ 60% for low load, 60%-80% for medium load, > 80% for high load;
[0064] Memory remaining: ≥ 40% for sufficient, 20%-40% for tight, < 20% for severe shortage;
[0065] IO throughput: disk IO < 100 MB / s or network IO < 500 Mbps is determined as IO congestion;
[0066] Index service response latency: ≤ 50 ms for smooth, 50-100 ms for slight delay, > 100 ms for service congestion.
[0067] (2) Parallel computing potential assessment
[0068] Basic parallel capability: count the number of idle CPU cores, support multi-task parallelism when there are 4 or more cores, and limit parallelism when there are less than 2 cores;
[0069] Accelerated resource adaptation: if the GPU memory is greater than or equal to 4GB and the utilization rate is less than 30%, enable vector retrieval acceleration (applicable to high complexity queries);
[0070] Distributed scheduling: calculate the load difference of each node (CPU utilization standard deviation < 10% is balanced, and tasks can be evenly distributed; > 30% preferentially schedule to low-load nodes).
[0071] A multi-modal query execution module 3 is configured to execute a query operation on the vector library according to the optimal query execution plan, and obtain candidate results similar to the semantic structure of the query request.
[0072] A query result fusion module 4 is configured to perform a weighted fusion operation on the multi-path results, and generate a weighted score for each candidate result.
[0073] An output module 5 is configured to sort and output each candidate result in descending order of the weighted score.
[0074] Further, the query engine is also associated with a dynamically updated modal fusion weight list, each row of data of the modal fusion weight list including a semantic structure feature vector of a corresponding query request and a fusion weight of each modality; wherein for a user input query request, a semantic structure feature vector thereof is generated by a semantic encoding conversion module, a cosine similarity between the feature vector and each semantic structure feature vector in the modal fusion weight list is calculated, and the fusion weight of each modality corresponding to the record with a similarity greater than a set threshold is taken as the fusion weight of each modality of the current query request.
[0075] In the embodiment of the application, the set threshold can be an empirical value, for example, it can be 0.8.
[0076] In another embodiment of the application, each row of data of the modal fusion weight list further includes a weight stability index.
[0077] In the context of multimodal database query engines, the weight stability index is a parameter used to quantify the reliability and applicability of the fusion weights of each modality in a record. Its core function is to assist in determining whether the fusion weights of a record can be stably reused in a new query when the semantic structure feature vectors of the new query are similar to those of records in the list, avoiding fusion deviations caused by the accidental effectiveness of weights. When the semantic similarity between a new query and multiple records exceeds a threshold, the engine will prioritize the record weight with the higher stability index as a reference, rather than relying solely on similarity ranking. For example: Record A: similarity 0.85, stability 0.9 (based on a large number of samples, strong generalization); Record B: similarity 0.88, stability 0.6 (based on a small number of samples, weak generalization). In this case, the engine may prioritize the weight of Record A to ensure the reliability of the fusion result. In this embodiment of the invention, for a user-input query request, if there are no similar records in the modality fusion weight list, a preset initial weight allocation strategy is adopted to assign initial fusion weights to each modality data of the query request and to fuse the modality data.
[0078] In this embodiment of the invention, the initial fusion weights of each modality can be empirical values. In an illustrative embodiment, the initial fusion weights of each modality satisfy the following condition: β0 i =[c1×(1 / n)+c2×T i +c3×SP i +c4×RF i ] / ∑ n i=1 [c1×(1 / n)+c2×T j +c3×SP j +c4×RF j ], where β0 i Let T be the initial fusion weights for the i-th mode. i Let T be the mode type coefficient of the i-th mode, where T is the mode type coefficient if the i-th mode is a structured mode. i =1.1, if the i-th mode is an unstructured mode, T i =0.95, SP i SP represents the query intent adaptation coefficient for the i-th modality, where SP is the visual / text modality in a retrieval query if the i-th modality is a visual / text modality. i =1.08, if the i-th modality is a structured modality in an analytical query, SP i =1.15, other cases, SP i =1.0. RF i Let be the system resource coefficient for the i-th mode. When the resource occupancy rate of the current mode is greater than 70%, RF i =0.85, when the efficiency of the last 3 executions is greater than or equal to 0.8, RF i =1.1, other cases, RFi =1.0. c1 to c4 are preset coefficients, in one illustrative embodiment, c1 =0.4, c2 =0.25, c3 =0.2, c4 =0.15.
[0079] Further, the modal fusion weight list is obtained by the following steps:
[0080] S10, initializing an intermediate fusion weight record set, setting a weight convergence determination threshold, a minimum sample size, and a similar query group clustering threshold.
[0081] S11, for the current obtained multi-modal query request, extracting its semantic structure features and generating a feature vector through a query analysis and semantic coding module, and matching a corresponding intermediate fusion weight record table from the current intermediate fusion weight record set:
[0082] If there is a matching record table, execute S12;
[0083] If there is no matching record table, create an empty intermediate fusion weight record table containing the semantic structure feature vector for the query request, and execute S13;
[0084] S12, based on the initial fusion weight of each modal data corresponding to the corresponding intermediate fusion weight record table, fusing each modal data of the multi-modal query request; execute S14;
[0085] S13, using a preset initial weight allocation strategy, allocating initial fusion weights to each modal data of the multi-modal query request, and fusing each modal data; execute S14;
[0086] S14, based on the fusion result, executing the corresponding query and obtaining query execution data and query results;
[0087] S15, based on the query execution data and the query results, calculating the actual fusion weight of each modal data and the reward function value; the actual fusion weight of each modal data is determined based on the semantic correlation coefficient, the execution plan perception coefficient, and the resource efficiency coefficient;
[0088] S16, storing the actual fusion weight and the reward function value into the corresponding intermediate fusion weight record table, using an improved gradient descent algorithm, updating the key hyperparameters in the execution plan perception coefficient and the resource efficiency coefficient, with the goal of maximizing the cumulative reward value;
[0089] S17, convergence judgment is performed on the intermediate fusion weight record table corresponding to the multi-modal query request, if the fluctuation amplitude of the reward function value of the last continuous m is less than or equal to the set fluctuation amplitude, the average value of the actual fusion weight corresponding to the m reward functions is calculated as the final fusion weight, and the semantic structure feature vector, the final fusion weight and the convergence stability of the query request are added to the current modal fusion weight list, and S11 is executed; otherwise, the semantic structure feature vector, the actual fusion weight and the convergence stability of the query request are added to the current modal fusion weight list, S11 is executed, and the sample continues to be accumulated.
[0090] In the embodiment of the application, the semantic structure feature vector can generate a fixed dimension semantic structure feature vector V (the dimension is consistent with the vector dimension in the intermediate fusion weight record set, such as 512 dimensions) through the USR model. The generation rule is as follows:
[0091] The modal combination feature is converted into a 128-dimensional sub-vector through one-hot encoding + embedding layer;
[0092] The query intention feature is converted into a 64-dimensional sub-vector through the pre-training vector of the intention label plus confidence weighting;
[0093] The key entity feature is converted into a 256-dimensional sub-vector through the embedding vector (based on knowledge graph pre-training) of the entity ID plus weighted summation;
[0094] The constraint condition feature is converted into a 64-dimensional sub-vector through the hash encoding + rule mapping of the constraint key-value pair;
[0095] Finally, the feature vector V is obtained through splicing and L2 normalization, ensuring that the feature vectors of different queries are comparable in a unified semantic space.
[0096] In the embodiment of the application, the intermediate fusion weight record set is a structured data set storing historical queries and corresponding fusion weight configurations. Each record table contains the following core fields: feature vector, modal fusion weight list, execution statistics (average time consumption / resource consumption), update timestamp, and matching times.
[0097] In the embodiment of the application, if the cosine similarity between the feature vector corresponding to a certain intermediate fusion weight record table in the current intermediate fusion weight record set and the feature vector of the multi-modal query request is greater than a set threshold, it indicates that there is a matching record table, otherwise, it indicates that there is no matching record table.
[0098] In the embodiment of the application, the empty intermediate fusion weight record table contains the following initial information:
[0099] Feature vector: semantic structure feature vector of the current query;
[0100] Modal fusion weight list: initialized to default value (e.g., single-modal query β i = 1.0, multi-modal mixed query β i = 1 / n);
[0101] Execution statistics: null (to be filled in later after execution);
[0102] Creation timestamp: current query engine time;
[0103] Matching times: 0;
[0104] Associated query ID: unique identifier of the current query (used for subsequent execution result backtracking). In the embodiments of the present application, S14 can specifically include:
[0105] (1) Hierarchical query execution:
[0106] First stage: perform approximate nearest neighbor search (using HNSW algorithm) in the vector library based on the semantic vectors of each modality, and return TopK (default K = 50) candidate results;
[0107] Second stage: perform modality consistency check (cross-modality feature matching degree ≥ 0.5) and semantic conflict filtering (e.g., "text description 'daytime' conflicts with image feature 'night scene'") on the candidate results;
[0108] Third stage: according to the initial weight, the results that pass the check are sorted again to generate the final query result set.
[0109] (2) Execution data collection:
[0110] The collection dimensions include:
[0111] Time data: time consumed by each modality, time consumed in the fusion stage, total execution time;
[0112] Resource data: CPU occupancy rate, IO operation times, network transmission volume of each modality, and total resource consumption;
[0113] Quality data: semantic similarity score of each modality result, user perception quality score of the fusion result (generated by a pre-trained quality evaluation model).
[0114] (3) Result storage:
[0115] The query result set and execution data are stored in association, and each result contains: multi-modal content (text / image / audio segment), contribution degree of each modality (based on initial weight calculation), and semantic matching confidence (0-100 points).
[0116] Further, the actual fusion weight βc i satisfies the following conditions:
[0117] βc i =(α i ×γ i ×δ i ) / ∑ n j=1 (α j ×γ j ×δ j )。
[0118] wherein, α i is a semantic correlation coefficient of the i-th modality, determined based on cosine similarity of a query vector corresponding to a multi-modal query request and an i-th modality vector in a vector library, αi=max(Similarity(q,v ik )), wherein k is a sample index of the i-th modality in the vector library, q is the query vector, v ik is the k-th data vector of the i-th modality.
[0119] In the embodiment of the application, the calculation method of the actual fusion weight of each modality is also referred to as a three-factor weight calculation method.
[0120] γ i is an execution plan perception coefficient of the i-th modality, γ i =log(1+C i ×P i ×I i ) / log(1+C b ×P b ×I b ), wherein C i is an execution resource consumption coefficient of the i-th modality, C i =1 / (w1×IO i +w2×CPU i +w3×Net i ), IO i is a disk IO resource consumption value of the i-th modality, which is a standardized value (range [0, 1]) and is calculated by a ratio of the query engine IO performance upper limit, considering disk read / write times, seek time, data transmission volume and IO queue waiting time. CPU i is a CPU resource consumption value of the i-th modality, which is a standardized value (range [0, 1]) and is determined by a ratio of a single-core full-load computing capacity, based on weighted calculation of CPU core number, calculation time length and instruction complexity of the modality data. Net iis the network transmission resource consumption value of the ith mode, which is a standardized value (range [0, 1]) combining data transmission volume, network delay and bandwidth occupancy, and is calculated by the ratio of the value to the maximum network transmission capacity of the query engine. w1, w2 and w3 are resource weight coefficients for dynamically adjusting the influence weight of the three types of costs: w1 is increased when the disk IO load of the query engine is too high, w2 is increased when the CPU resource is tight, and w3 is increased when the network is congested. The weight values are updated in real time by the query engine resource monitoring module.
[0121] P i is the execution efficiency coefficient of the ith mode, P i = min ((PO i / TO i ) x (AC i / REC i ), 1), where PO i is the number of parallel executable operations of the ith mode, such as the shard computing operation in distributed query and the batch comparison operation in vector retrieval, which is determined based on the operation type (CPU intensive / IO intensive). TO i is the total number of operations of the ith mode, including parallel operations and serial operations, such as serial conversion steps of data preprocessing. PO i / TO i represents the proportion of parallel operations, with a value range of [0, 1], and a higher value indicates greater parallel potential. AC i is the actual available CPU core number of the ith mode, REC i is the total CPU core number pre-allocated for the ith mode of the ith mode, which is an initial resource quota based on the query plan, and AC i / REC i represents core utilization rate, with a value range of [0, 1], and a higher value indicates better matching between resource allocation and actual demand.
[0122] I i is the index utilization efficiency coefficient of the ith mode, I i = ISR i / TSR i x SL i , ISR i is the actual number of rows read by index scan of the ith mode, TSR i is the total number of rows that need to be read if full table scan is used for the ith mode, ISR i / TTSR i represents the proportion of index scan rows to total table rows, with a value range of (0, 1], and a smaller value indicates more significant index filtering effect, such as 10,000 rows scanned by index from 1,000,000 rows, with a proportion of 0.01. SL iis the selection coefficient of the i-th modality, which ranges from 0 to 1, and reflects the fitting accuracy of the index to the query condition. i = AR i × FE i , AR i is the estimated accuracy ratio of the i-th modality, which is the ratio of the estimated number of scanned rows to the actual number of scanned rows; FE i is the index filtering efficiency of the i-th modality, which is the proportion of irrelevant rows filtered out by the index to the total number of rows, and ranges from 0 to 1, where a higher value indicates a stronger ability of the index to exclude irrelevant data.
[0123] C b is the reference execution resource consumption coefficient of the i-th modality, P b is the reference execution efficiency coefficient of the i-th modality, I b is the reference index utilization efficiency coefficient of the i-th modality, C b , P b , and I b are dynamically updated based on the sliding window statistics method.
[0124] δ i is the resource efficiency coefficient of the i-th modality, δ i = exp (-λ × S i ), S i is the resource consumption intensity of the i-th modality, which ranges from 0 to 1 after standardization, where 0 represents no consumption and 1 represents resource consumption reaching the upper limit of the query engine. λ is the resource sensitivity coefficient, which is used to adjust the sensitivity of the function to resource consumption; exp () is the natural exponential function; α j is the semantic relevance coefficient of the j-th modality, γ j is the execution plan perception coefficient of the j-th modality, δ j is the resource efficiency coefficient of the j-th modality, where j ranges from 1 to n. Further, in an embodiment of the present application, the reward function satisfies the following conditions: R = k1 × Q + k2 × E + k3 × F, where Q is the result quality score, which ranges from 0 to 1, E is the global time consumption optimization coefficient, F is the global resource utilization efficiency coefficient, k1, k2, and k3 are reward weight coefficients, k1 + k2 + k3 = 1, and k1, k2, and k3 are dynamically adjusted according to the load of the query engine (e.g., k2 + k3 = 0.7 when the load is high), where in an illustrative embodiment, k1 = 0.4, k2 = k3 = 0.3. Wherein, E = 1 / (1 + (ET - ET0) / ET0), where ET is the actual execution time of the current query, and ET0 is the benchmark execution time of similar queries; F = 1 / (1 + (RC - RC0) / RC0), where RC is the comprehensive resource consumption value of the current query, and RC0 is the benchmark resource consumption value of similar queries.
[0125] In this embodiment of the invention, the result quality score can be comprehensively evaluated by a multi-dimensional quality evaluation system. This stage combines quantitative and qualitative indicators such as user interaction behavior analysis, semantic relevance score and user satisfaction feedback, and is calculated using the analytic hierarchy process.
[0126] Furthermore, S16 may specifically include:
[0127] (1) Record table update:
[0128] Add an entry to the intermediate fusion weight record table currently being queried, including: actual fusion weight, reward function value R, execution data snapshot, and update timestamp.
[0129] (2) Hyperparameter optimization objective:
[0130] With the goal of maximizing the cumulative reward value ΣR within the window, optimize the following key hyperparameters:
[0131] Weighting allocation in the perception coefficient of the execution plan;
[0132] Sensitivity coefficient in resource efficiency coefficient;
[0133] The weights k1, k2, and k3 in the reward function.
[0134] (3) Improve the execution of the gradient descent algorithm:
[0135] The Adam optimizer is used, and the learning rate is dynamically adjusted (initially 0.01, and halved if the cumulative reward value decreases for 3 consecutive times).
[0136] In each iteration, the loss function L=-ΣR is calculated, and the hyperparameters are updated through gradient backpropagation.
[0137] Add an L2 regularization term (coefficient 0.001) to prevent overfitting and ensure that the hyperparameters are within a reasonable range in physical terms (e.g., λ∈[0.5,3]).
[0138] The optimization cycle is synchronized with the sliding window (executed every 5 minutes) to ensure that hyperparameters are adapted to the recent query engine status.
[0139] In this embodiment of the invention, the fluctuation range can be set as an empirical value, for example, 0.05.
[0140] Further, to enhance the robustness of the system, the application designs a complete exception handling and degradation mechanism. When the execution plan information is missing, the system adopts the preset default gamma coefficient value, or only based on the alpha and delta factors to calculate the weight. When the real-time weight calculation is abnormal, the system quickly switches to the backup weight configuration based on historical statistics. When the weight configuration is abnormal or the performance is down, the system automatically corrects the parameters and adapts. These mechanisms ensure that the algorithm can still run normally in a heterogeneous or incomplete information environment, and maintain the stability of the query engine in various complex semantic query scenarios.
[0141] Below, the modules of the query engine provided by the embodiments of the application are described in detail.
[0142] Further, in the embodiments of the application, the query analysis and semantic coding module 1 is based on the USR model, and in particular, uses the encoder to uniformly process various input methods in the query understanding field. This module supports structured queries, programmatic interface calls, and natural language queries, and converts them into vector representations in the semantic space through the calling of corresponding coding functions.
[0143] The query analysis and semantic coding module realizes uniform processing and semantic conversion of different input methods by integrating various query understanding methods, including extended structured query language interfaces, programmatic API interfaces, and natural language query interfaces.
[0144] Among them, the extended structured query language interface maintains compatibility with traditional SQL syntax, and introduces a set of operators specifically for multi-modal semantic queries. These operators are designed directly based on the semantic space of the USR model. The semantic operators include IMAGE_SIMILAR_TO for image similarity queries, TEXT_CONTAINS for text semantic inclusion queries, AUDIO_MATCHES for audio matching queries, etc. Through the unified semantic space of the USR model, users can query related text content through images, or find similar images through text descriptions.
[0145] The programmatic API interface adopts modern API design concepts, providing flexible and powerful programmatic access capabilities for various applications to query the engine, enabling developers to seamlessly integrate multimodal query functions into existing application architectures. The main improvement of the programmatic API interface is the native support for multimodal content. Developers can submit complex query requests containing text, images, audio, and other modalities through a unified interface, and the query engine will automatically handle the semantic association and fusion between different modalities. In addition, the programmatic API interface provides endpoints with multiple complexity levels from basic query operations to advanced semantic analysis, ensuring ease of use for simple application scenarios and meeting the flexibility needs of complex application scenarios. The response design of the programmatic API interface is carefully optimized, providing not only standardized query results but also rich metadata information such as query processing time, relevance score, confidence, etc., facilitating subsequent result processing and user experience optimization at the application layer.
[0146] The natural language query interface realizes intelligent conversion from natural language expressions to structured semantic queries, significantly reducing the technical threshold for users to use multimodal queries to query the engine. The natural language query interface adopts a multi-stage natural language understanding pipeline, breaking down complex language processing tasks into a series of step-by-step refinement processing steps such as preprocessing, intent recognition, entity extraction, and query reconstruction. The language processing capabilities of the query engine not only lie in supporting multiple natural languages, but more importantly, they achieve deep semantic understanding of user query intent, accurately identifying users' real query requirements and key query parameters from unstructured natural language expressions. The cross-modal semantic alignment mechanism plays a key role in this process. Through the cross-modal alignment algorithm integrated in the fusioner, the query engine can establish semantic associations between different modalities of content, enabling accurate understanding and processing of natural language queries containing multimodal references.
[0147] Intent recognition is achieved through an intent recognition model that supports fine-grained query intent classification, enabling the differentiation of different types of query requirements such as similarity search, content retrieval, and association analysis, providing accurate semantic guidance for subsequent query processing. The query engine integrates knowledge graph-based entity processing enhancement techniques, effectively addressing common ambiguity and incomplete expression issues in natural language queries through entity recognition, disambiguation, and concept expansion. In addition, the query engine can understand and process omitted expressions, pronoun references, and context dependencies in multi-turn dialogues, maintaining dialogue state information to provide a coherent and consistent interactive experience for users. When the query engine's understanding of user query intent is uncertain, an interactive clarification mechanism is activated, actively asking questions to collaborate with users to ensure the accuracy of query understanding. This human-machine collaborative query understanding mode significantly improves the processing effect in complex query scenarios.
[0148] The semantic encoding conversion of the query content is to prepose the execution plan perception mechanism to the query understanding stage, realizing the deep fusion of semantic understanding and execution optimization. The query engine applies the corresponding encoder to all types of queries to convert the query content into a vector representation in a unified semantic space. Different modal query inputs are processed by specially designed encoders. The text query is mapped to a semantic vector by a language model encoder, the image query extracts visual semantic features by a visual encoder, the audio query is semantically encoded by an acoustic model, and the video query generates a comprehensive semantic representation by a spatiotemporal feature fusion technology.
[0149] When a user submits a composite query containing multiple modalities, the query engine not only calculates the vector representation of each modality by the corresponding encoder, but more importantly, introduces a pre-analysis mechanism of the query execution plan in the fusion process. The weight calculation in this stage is based on the preliminary estimation of the historical query statistical information or the preset value to provide the basis for the cost estimation of the subsequent query optimizer. The query engine analyzes the historical access mode of the modal data, estimates the parallel execution potential and resource consumption characteristics, and calculates the preliminary three-factor weight parameters, so that the generation process of the query vector has considered the efficiency factors of subsequent execution. This design of integrating database query optimization theory into the semantic understanding process enables the query engine to optimize the execution strategy while understanding the user's intention, lays a foundation for real-time dynamic weight adjustment in the subsequent query execution stage, and realizes the unification of semantic accuracy and execution efficiency.
[0150] The innovation value of the whole semantic encoding conversion process lies in the establishment of an integrated processing mechanism from query understanding to execution optimization, which ensures that the query content of different modalities can not only be accurately represented in a unified semantic space, but also lay a foundation for subsequent efficient execution, which is the forward-looking optimization capability that traditional query engines do not have.
[0151] The query understanding mechanism based on the USR model maps different modalities and different types of queries to a unified semantic space, providing a basis for subsequent query optimization and execution.
[0152] In the embodiment of the application, the query execution plan generation module 2 is used to include the semantic characteristics of the query into the optimization decision process, realizing the leap of query optimization technology from traditional structure perception to semantic perception. The module receives the standardized semantic vector representation from the query parsing module, generates an optimized execution plan for multi-modal data queries by analyzing the semantic characteristics of the query, the data distribution characteristics and the resource status of the query engine. This semantic perception optimization capability enables the query engine to intelligently optimize according to the semantic content of the data rather than only the structural characteristics, and provides an efficient execution strategy for multi-modal semantic queries.
[0153] Further, in the embodiment of the application, the query execution plan generation module implements semantic query optimization in cooperation with the fusioner, and the module makes optimization decisions based on the dynamic weight adaptive mechanism of the fusioner. The query engine uses the weight configuration calculated by the fusioner to guide the semantic fusion strategy, and provides execution environment feedback information for the weight calculation of the fusioner by analyzing the execution plan features of each modality data, such as execution cost, parallelism and index utilization.
[0154] The core basis of the work of the query execution plan generation module is the encoder and the execution plan-aware fusioner defined in the USR model. The query execution plan generation module first receives the query vector and the weight configuration generated by the fusioner, identifies the types and dynamic weight distribution information involved therein, and then selects the most suitable execution path based on these information. When the query contains multiple modalities, the query execution plan generation module uses the dynamic weight results calculated by the fusioner to deeply analyze the semantic correlation, execution cost features and resource consumption patterns of each modality, so as to determine the optimal modality processing order, parallel execution strategy and inter-modality cooperation mechanism, and generate the optimal execution plan. The query execution plan generation module particularly focuses on providing execution environment feedback for the fusioner, and provides accurate execution plan perception parameters for the dynamic adjustment of the weight of the fusioner by analyzing the index utilization, access pattern complexity and parallel execution potential of each modality data.
[0155] Further, in the embodiment of the application, the integration of the query execution plan generation module and the PostgreSQL query optimizer is realized through a standardized interface. The query engine designs a complete execution plan feature standardization interface specification, which defines the accurate conversion mechanism from the real database execution plan to the weight calculation parameters. The interface specification uses JSON format as the data exchange standard, and includes core data structures such as execution node structure, cost model and resource configuration.
[0156] The query engine performs execution plan analysis based on the EXPLAIN ANALYZE output of the PostgreSQL database. The key fields are extracted from the JSON format execution plan: the actual_time field is used to calculate the execution cost feature, the shared_hit_blocks and shared_read_blocks fields are used to calculate the IO resource consumption value feature, and the workers_planned and workers_launched fields are used to calculate the parallel execution efficiency feature. The index utilization efficiency feature is calculated by counting the proportion of the number of “Index Scan”, “Index Only Scan” and “Seq Scan” nodes in the execution plan, and the index selectivity effect is evaluated by combining the difference between the “PlanRows” and “Actual Rows” fields, to comprehensively generate a quantitative indicator of index utilization efficiency.
[0157] The specific execution plan parsing process achieves complete feature extraction through six consecutive processing stages. The execution plan acquisition stage uses the PostgreSQL EXPLAIN command, specifying the ANALYZE, BUFFERS, and FORMATJSON parameters, to obtain a JSON data structure containing detailed statistical information about the execution plan. This data structure includes all key performance indicators and resource consumption information during the query execution process. The node traversal stage recursively traverses each node of the execution plan tree using a depth-first search algorithm. It identifies operation nodes involving different modalities of data by analyzing the "Node Type" field of each node and establishes a parent-child relationship mapping table between nodes for subsequent feature aggregation calculations. The cost feature extraction stage extracts the actual execution time from the "ActualTotal Time" field of each node, calculates the cache hit rate by the ratio of "Shared HitBlocks" to "Shared Read Blocks," and evaluates the usage of temporary storage based on the "Temp Read Blocks" and "Temp Written Blocks" fields. The parallel feature analysis phase calculates the actual efficiency ratio of parallel execution by comparing the "WorkersPlanned" and "WorkersLaunched" fields, analyzes the "Parallel Aware" flag to determine the degree of parallelization of operations, and counts the number of worker processes actually involved in execution based on the "Worker Number" field. The index utilization evaluation phase calculates the proportion of "Index Scan," "Index OnlyScan," and "Seq Scan" nodes in the execution plan, and evaluates the estimation accuracy of the query optimizer and the index selectivity effect by comparing the differences between the "Plan Rows" and "Actual Rows" fields. The feature vector generation phase transforms the raw data extracted in the previous phases according to a predefined standardized formula to generate feature vectors that meet the input requirements of the fusion weight calculation model, ensuring the comparability of feature data for different queries and modalities.
[0158] The interface specification defines a standardized feature vector format: the ExecutionFeature structure contains three main components: cost_vector, parallel_vector, and index_vector. Each vector is normalized to ensure that execution plan features can be compared and calculated within a uniform numerical range.
[0159] Taking the execution plan parsing of a multi-modal query as an example, when the query involves an associated query of a text data table and an image data table, the execution plan generated by PostgreSQL usually contains a connection operation node for associating the two data tables, in which the text data table can be accessed by index scanning and the image data table can be accessed by sequential scanning. The query engine extracts key feature information from the execution plan, including the actual execution time of the connection node, the number of cache block hits and the number of disk read blocks, the execution time of the index scanning node and the number of parallel work processes, the execution time, the planned number of rows and the actual number of rows of the sequential scanning node, etc. Through a standardization calculation formula, the query engine converts these original feature data into standardized parameters required for three-factor weight calculation, including the resource consumption coefficient, the parallel efficiency coefficient and the index utilization efficiency coefficient of the text modality, and the corresponding coefficients of the image modality, and finally generates the execution plan perception coefficients of each modality for weight calculation.
[0160] The query engine can obtain the execution plan information from the query optimizer of PostgreSQL and convert it into a feature vector required for weight calculation. This integration mechanism has good scalability, and through an abstracted execution plan parsing interface, the query engine can adapt to the execution plan formats of other relational database query engines (such as MySQL, Oracle, etc.), realizing the semantic query optimization capability across database platforms. This integration mechanism ensures that the dynamic weight adaptive algorithm with execution plan perception can fully utilize the achievements of the underlying database query optimization, realizing the deep integration of semantic query and traditional query optimization.
[0161] The semantic-aware query analysis identifies query features from the semantic representation of the query, including the types of modalities involved, the semantic complexity of the query, etc. The intelligent index selection mechanism selects the optimal index strategy according to the semantic characteristics of the query, supporting multiple index technologies suitable for semantic vectors.
[0162] The semantic-aware execution plan generation is based on the results of the dynamic weight adaptive algorithm of the fusioner. For a query containing multiple modalities, the execution plan directly uses the three-factor weight results calculated by the fusioner to guide the formulation of the execution strategy. According to the weight factors such as semantic relevance, execution cost and resource efficiency of each modality provided by the fusioner, the query engine dynamically adjusts the processing priority and resource allocation of each modality in the execution plan.
[0163] Taking the execution plan generation process of a cross-modal query as an example, when processing a composite query containing text retrieval and image matching, the query engine first evaluates the access characteristics of each modal data table by analyzing the query engine statistics of PostgreSQL, including index configuration, data distribution characteristics, and access pattern complexity, etc. The text data table is usually configured with an index structure based on vector similarity and has good selectivity characteristics, while the image data table may lack an effective index structure due to the complex distribution characteristics of high-dimensional vector data and needs to be accessed in a sequential scanning manner. Based on the pre-analysis results of the execution plan, the query engine calculates the execution plan awareness coefficient of each modality, which reflects the relative execution efficiency of different modalities in the current database environment. The optimizer generates a targeted execution plan based on the comparative analysis results of the weight coefficients, which takes the modality operation with higher execution efficiency as the driving table, preferentially executes the high-efficiency filtering operation to reduce the data set size for subsequent processing, and then performs computationally intensive modality matching processing on the filtered candidate result set. Through this execution order arrangement, the overall query performance is effectively optimized. The specific execution plan tree structure selects appropriate join strategies according to the execution characteristics of each modality, which may use nested loop join, hash join, or sorted merge join, etc. Through the optimized plan structure arrangement, the overall execution cost is significantly reduced.
[0164] The multi-modal query execution module 3 interacts with the database storage engine according to the optimal execution plan to complete the actual query processing, breaking through the limitation of traditional query execution engines that only support structured data operations.
[0165] The multi-modal query execution module implements a dynamic weight self-adaptive mechanism during execution, which can adjust the weight configuration of the fusioner in real time according to the actual execution situation. The query engine directly uses execution feedback to optimize the semantic fusion strategy, dynamically updates the three-factor weight parameters by real-time collection of key indicators such as execution cost, resource consumption, and result quality of each modality.
[0166] The multi-modal query execution module directly operates the semantic encoding vectors defined in the USR model, and implements a dynamic weight adjustment mechanism based on execution plan awareness. When processing a single-modal query, the execution engine optimizes the execution strategy for the specific modality encoder. For multi-modal queries, the execution engine implements a dynamic weight fusion mechanism defined in the fusioner, which dynamically adjusts the weight configuration by real-time monitoring of performance indicators during execution.
[0167] At the beginning of execution, the query engine generates an initial weight configuration based on historical statistical data and preset parameters, assigns corresponding weight values to each execution path, and starts the multi-path parallel execution process.
[0168] During the execution of the query by the multi-modal query execution module, the execution information of each execution path in the optimal query execution plan is monitored in real time, and the weight and resource of each execution path are adjusted based on the real-time monitoring result; the execution information includes actual execution time, execution efficiency and query result quality score. Specifically, during the execution process, the execution state of each path can be monitored in real time through the statistical collector of PostgreSQL, and the key performance indicators are collected for dynamic adjustment and judgment.
[0169] When the actual execution time of a certain execution path deviates significantly from the estimated time, the query engine determines whether to trigger weight adjustment through a preset deviation threshold. For paths with performance deviation within the normal fluctuation range, the query engine maintains the current weight configuration; for paths with performance deviation significantly exceeding the threshold, the query engine starts the weight recalculation process. After the query engine detects that the execution efficiency of a path deviates significantly from the expected value, the weight recalculation process is triggered immediately. The weight adjustment mechanism adjusts the weight distribution according to the actual execution effect of each path. The weight of the path with relatively improved execution efficiency is increased accordingly, the weight of the path with relatively decreased execution efficiency is decreased accordingly, and the weight of the path with stable execution efficiency is maintained relatively balanced. In the embodiment of the application, an exponential smoothing algorithm can be used to adjust the weight of the execution path.
[0170] The execution engine reassigns the query engine resources according to the adjusted weight configuration, optimizes the execution priority of each path, improves the overall query performance through intelligent resource scheduling strategy, and ensures the stability of the semantic matching quality.
[0171] In a specific embodiment of the application, the weight adjustment of the execution path is triggered when any of the following conditions is met: the performance deviation detection deviates more than 10% between the actual execution time and the estimated time of a certain modality; the resource state change detection changes more than 15% in the CPU or memory utilization of the query engine; the query quality degradation detection is less than 85% of the historical average value for three consecutive times.
[0172] The semantic similarity search operation is the core operation of the multi-modal query execution module, and the query engine realizes efficient semantic vector retrieval capability. The execution engine can process similarity calculation based on semantic vectors, support cross-modal similarity search in a unified semantic space, and is the technical basis for realizing cross-modal query functions such as "image searching text" and "text searching image".
[0173] The semantic similarity search operation is directly implemented based on the encoder and the fusioner defined in the USR model. The query engine uses cosine similarity as the basic measurement method, and evaluates the semantic similarity by calculating the cosine value of the angle between the query vector and the candidate vector. For single-modal queries, the query engine directly calculates the similarity between the query vector and the candidate data vector. For multi-modal queries, the query engine first generates a unified query vector by applying the fusion mechanism defined in the fusioner implementation, and then calculates the similarity with the candidate vector. This similarity calculation based on unified semantic representation ensures the semantic consistency and accuracy of the query results.
[0174] The core innovation of the multi-modal query execution module lies in the implementation of a coordinated processing mechanism for semantic search and structured filtering. The query engine determines the optimal execution order of semantic filtering and structured filtering through a dynamic weight self-adaptive algorithm based on execution plan perception, and balances the query precision and execution efficiency through a multi-stage execution strategy. This execution capability based on unified semantic representation enables the query engine to handle complex multi-modal query requirements while maintaining excellent performance.
[0175] After the completion of the multi-modal query execution module, the query engine faces a key technical challenge: how to process heterogeneous query results from different execution paths. The root of this challenge lies in the fact that even after precise query optimization and execution, different execution paths will still produce result sets with significant quality differences, which will directly affect the accuracy of the final query results and user experience if not intelligently fused.
[0176] The query result fusion module 4 implements an intelligent result fusion mechanism based on execution feedback, feeding the performance data collected during execution into the weight optimization of the fusioner to form a complete execution-feedback-optimization closed loop.
[0177] Consider a specific query scenario: a user queries "find teaching resources about machine learning algorithms". After all the optimization steps described above, the three execution paths of the query engine return the following results:
[0178] The text retrieval path returns text resources containing "Machine Learning Algorithm Details" and "Deep Learning Tutorial", which perform well in text semantic matching and have generally high relevance scores. The image retrieval path returns visual resources containing algorithm flowcharts and mathematical formula pictures, which perform well in visual content matching but have relatively low relevance scores due to the complexity of image semantic understanding. The cross-modal association path returns multimedia resources containing video tutorials and interactive demonstrations, which combine text and visual information and have moderate relevance scores. Without result fusion, directly sorting according to the original scores of each path will result in serious quality problems. The results of the text path will occupy the top positions due to the highest scores, but these results may lack intuitive visual aids, which is not conducive to algorithm understanding. The results of the image path will be placed in the back due to lower scores, but these visual resources are of great value for understanding complex algorithms. Although cross-modal resources are comprehensive, they may be overlooked due to moderate scores.
[0179] The more critical problem lies in the differences in execution efficiency. Assuming that during the current query execution process, the text retrieval path has excellent execution efficiency due to high index hit rate, the image retrieval path has lower execution efficiency due to the need to process high-dimensional vectors, and the execution efficiency of the cross-modal association path is moderate. If these execution efficiency differences are ignored and only based on semantic score sorting, the actual reliability of the results of each path cannot be reflected.
[0180] The query result fusion module solves this problem by applying a dynamic weight mechanism. The query engine assigns a higher weight to the text path based on its high execution efficiency, a lower weight to the image path based on its low execution efficiency, and a moderate weight to the cross-modal path based on its moderate execution efficiency. After weight adjustment, the comprehensive score of the text path results is further improved, the comprehensive score of the image path results is moderately reduced, and the comprehensive score of the cross-modal path results remains stable. The final fusion sorting not only ensures the priority display of high-quality text resources but also ensures that valuable visual and multimedia resources obtain reasonable sorting positions, providing users with comprehensive and reliable query results.
[0181] The query result fusion module realizes the unified scoring of multi-path results through a weighted combination formula. For candidate results from different execution paths, the query result fusion module obtains the weighted score of the candidate result as follows: Score final =∑ z h=1 f h ×Score h , where Score final represents the weighted score of the candidate result, f hdenotes the dynamic weight coefficient of the execution path h, h takes values from 1 to z, z is the number of execution paths, Score h denotes the semantic relevance score of the candidate result in the execution path h. This kind of weighted combination ensures that the path with high execution efficiency has greater influence on the final ranking, while the importance of semantic matching quality is maintained.
[0182] The semantic-aware result ranking realizes the semantic similarity calculation between the query and the result based on the unified semantic space of the USR model. For a single-modal query, the query engine calculates the similarity between the query vector and the result vector using the corresponding encoder. For a multi-modal query, the query engine calculates the similarity between the fused query representation and the result representation using the fuser, ensuring that the ranking result accurately reflects the semantic intent of the query.
[0183] The query result fusion module solves the problem of unified integration of different modal results through a standardized fusion process. The fusion process includes four consecutive processing steps. The modality recognition step identifies the modality type of each query result, determining whether it comes from text retrieval, image retrieval, or cross-modal association, etc. different processing paths. The weight matching step associates each query result with the dynamic weight coefficient of its corresponding execution path, establishing a mapping relationship between the query result and the weight configuration. The score calculation step calculates the weighted score based on the original semantic relevance score of the query result and the corresponding path weight coefficient. This weighted score reflects both the semantic matching quality and the reliability of the execution path.
[0184] The output module 5 is used to perform unified ranking based on the weighted scores of all query results, generating a final fusion result list, ensuring that high-quality results from different modalities can obtain appropriate ranking positions according to their actual value.
[0185] Based on the same inventive concept, the embodiments of the present application provide a multi-modal database query method based on unified semantic representation, which is implemented based on the aforementioned query engine, as shown in Figure 2 The method comprises:
[0186] The query request input by the user is semantically parsed to obtain corresponding semantic structure features; the query request includes single-modal queries and multi-modal queries; the semantic structure features include query intent, key entities, and modality information.
[0187] The semantic structure features are uniformly semantically encoded to generate corresponding encoding results, wherein if the semantic structure features contain single-modal data, the corresponding single-modal semantic vector is directly generated as the encoding result; if the semantic structure features contain multi-modal data, the semantic vectors of each modality are generated and the generated semantic vectors are fused to generate a fused semantic vector as the encoding result.
[0188] Based on the semantic features in the encoding result, data distribution characteristics of the vector library, query complexity and query engine resource status, an optimal query execution plan is generated, the optimal query execution plan including vector retrieval strategy, modal weight allocation and computing resource scheduling scheme.
[0189] According to the optimal query execution plan, a query operation is performed on the vector library to obtain candidate results similar to the query request semantics.
[0190] A weighted fusion operation of multi-path results is performed to generate a weighted score of each candidate result.
[0191] Each candidate result is output in descending order of the weighted score.
[0192] In summary, the multi-modal data query scheme based on unified semantic representation provided by the embodiments of the present application has at least the following advantages:
[0193] (1) The PostgreSQL database execution plan theory is deeply integrated into the whole process of multi-modal semantic query processing, and the dynamic weight adaptive mechanism of execution plan perception realizes the unified optimization of semantic accuracy and execution efficiency. The scheme establishes a unified semantic space through the USR model, and the three-factor dynamic weight calculation model of the fusioner directly incorporates execution plan features such as execution cost, parallelism and resource utilization into the semantic fusion process, so that the query engine can start optimizing the execution strategy while understanding the semantics.
[0194] (2) A complete verification index system was established to ensure the technical effectiveness and engineering feasibility of each core innovation point through quantitative performance evaluation. The verification of the unified semantic representation of the USR model was evaluated through three key dimensions. The cross-modal semantic consistency index was evaluated by calculating the cosine similarity of different modal data with the same semantic content in the unified semantic space, requiring a similarity greater than 0.8 to verify the accuracy of semantic alignment. The semantic vector quality index was evaluated through the semantic similarity retrieval task, requiring the Top-10 retrieval accuracy rate to reach more than 85% to verify the effectiveness of semantic encoding. The cross-modal query success rate index was evaluated through cross-modal query tasks such as "searching for text by image" and "searching for images by text", requiring the query success rate to reach more than 80% to verify the cross-modal understanding ability. The validation of the execution plan-aware dynamic weight adaptive mechanism covers four core performance indicators: query execution efficiency improvement requires an average query time reduction of over 20% compared to the static weight scheme, verified through comparative experiments; weight adjustment response time requires weight update latency to be controlled within 100 milliseconds, verifying the real-time performance of dynamic adjustments; query engine resource utilization requires CPU and memory utilization to be improved by over 15% compared to traditional schemes, verifying resource optimization effects; and query result quality stability requires the result quality score fluctuation to not exceed 5% during dynamic weight adjustment, verifying the stability of the optimization process. The validation of multimodal query execution and result fusion capabilities is evaluated through three key indicators: multimodal query processing capability requires support for composite queries containing text, images, and audio modalities, with a query processing success rate of over 90%; result fusion quality is assessed through user satisfaction evaluation, requiring a user satisfaction score of 4.0 or higher (out of 5); and cross-modal association accuracy is evaluated using manually labeled cross-modal association datasets, requiring an association accuracy rate of over 75%.
[0195] (3) Theoretical analysis based on PostgreSQL database shows that the execution plan-aware dynamic weight adaptive algorithm has significant theoretical advantages over the traditional static weight scheme in multimodal query scenarios, mainly reflected in the comprehensive optimization of query execution efficiency, semantic matching quality and query engine resource utilization.
[0196] (4) Theoretical Comparative Analysis verified the optimization principle of this technical solution through the differences in algorithm mechanisms. Traditional static weighting schemes adopt fixed weight allocation strategies, which cannot be dynamically adjusted according to actual execution conditions and lack adaptability when facing different query scenarios and query engine states. The execution plan-aware dynamic weighting scheme adopts the adaptive weight adjustment mechanism proposed in this technical solution, which can dynamically adjust the weight allocation of each modality according to the real-time characteristics of the PostgreSQL execution plan, and adapt to different query scenarios and query engine environments through intelligent weight optimization strategies.
[0197] (5) The execution plan perception mechanism incorporates key features of the database execution plan into the weight calculation process, realizing the fusion of semantic understanding and execution optimization. The dynamic weight adjustment mechanism optimizes through real-time monitoring and feedback, enabling the weight configuration to reflect the current execution state and query engine environment. The intelligent weight allocation strategy considers multiple dimensions such as semantic relevance, execution efficiency, and resource utilization through a three-factor dynamic weight calculation model.
[0198] (6) The applicability analysis of the technical solution shows that the solution can effectively handle large-scale multi-modal data cross-modal query requirements, and through the dynamic weight adjustment mechanism, intelligent optimization is performed based on real-time feedback of the PostgreSQL execution plan, achieving significant improvement in query performance while maintaining semantic matching accuracy.
[0199] (7) This technical innovation realizes a fundamental technical leap from traditional database query engines to intelligent semantic query engines by establishing a complete execution perception semantic query processing mechanism. The technical solution has important practical value in intelligent retrieval systems, multimedia content management platforms, and cross-modal data analysis applications, providing core technical support and theoretical foundation for multi-modal data processing applications based on PostgreSQL databases. This solution provides an innovative technical path for solving the unified query and intelligent retrieval of large-scale heterogeneous data by deeply combining database execution plan theory and multi-modal semantic fusion technology. The embodiment of the present application also provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being configured to execute the method described in the embodiment of the present application.
[0200] The embodiment of the present application also provides a computer-readable storage medium storing computer executable instructions for executing the method described in the embodiment of the present application.
[0201] It should be understood that various forms of flow shown above can be reordered, added or deleted steps. For example, the steps described in the present application can be executed in parallel, sequentially or in different order, as long as the desired results of the technical solution disclosed in the present application can be achieved, which is not limited herein.
[0202] The above specific embodiments do not constitute a limitation on the scope of protection of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements and improvements made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A multi-modal database query engine based on a unified semantic representation, characterized in that, The query engine is associated with a vector library storing unified semantic vectors of a plurality of entities, and comprises: a query analysis and semantic coding module configured to perform semantic analysis on a query request input by a user to obtain corresponding semantic structural features, and to perform unified semantic coding on the semantic structural features to generate a corresponding coding result, wherein, if the semantic structural features contain single-modal data, a semantic vector corresponding to the single-modal data is directly generated as the coding result; if the semantic structural features contain multi-modal data, semantic vectors of each modality are generated and fused to generate a fused semantic vector as the coding result; wherein the query request comprises single-modal query and multi-modal query; and the semantic structural features comprise query intent, key entity and modality information; a query execution plan generation module configured to generate an optimal query execution plan based on semantic features in the coding result, data distribution characteristics of the vector library, query complexity and query engine resource status, the optimal query execution plan comprising vector retrieval strategy, modality weight allocation and computing resource scheduling scheme; a multi-modal query execution module configured to perform query operation on the vector library according to the optimal query execution plan to obtain candidate results similar to the semantic of the query request; a query result fusion module configured to perform weighted fusion operation on the multi-path results to generate a weighted score of each candidate result; an output module configured to output each candidate result in descending order of the weighted score.
2. The unified semantic representation based multi-modal database query engine according to claim 1, wherein, The query analysis and semantic coding module comprises a plurality of modal encoders and a fusioner; an i-th modal encoder is configured to uniformly code semantics of an i-th modality to generate a corresponding coding result, i is an integer from 1 to n, and n is a total number of modalities included in the query request; and the fusioner is configured to fuse the coding results of the modal encoders to obtain a corresponding coding result, the coding result satisfying the following condition: F = ∑ n i=1 β i ×E i , F is a fusion result, E i is the coding result of the i-th modal encoder, β i is a fusion weight of the i-th modality, wherein ∑ n i=1 β i = 1 and β i ≥ 0.
3. The unified semantic representation based multi-modal database query engine of claim 2, wherein, The query engine is also associated with a dynamically updated modality fusion weight list, each row of data in the modality fusion weight list comprising a semantic structural feature vector of a corresponding query request and fusion weights of each modality; wherein, for a query request input by a user, a semantic structural feature vector thereof is generated by a semantic coding conversion module, a cosine similarity between the feature vector and each semantic structural feature vector in the modality fusion weight list is calculated, and the fusion weights of each modality corresponding to the record with a similarity greater than a set threshold are taken as the fusion weights of each modality of the current query request.
4. The unified semantic representation based multi-modal database query engine of claim 3, wherein, The modality fusion weight list is obtained by the following steps: S10, initializing an intermediate fusion weight record set; S11, for a multi-modal query request currently obtained, extracting a semantic structural feature thereof by the query analysis and semantic coding module and generating a feature vector, and matching a corresponding intermediate fusion weight record table from the current intermediate fusion weight record set: if there is a matched record table, performing S12; if there is no matched record table, creating an empty intermediate fusion weight record table containing the semantic structural feature vector for the query request, and performing S13; S12, based on the initial fusion weights of each modality data corresponding to the corresponding intermediate fusion weight record table, fusing each modality data of the multi-modal query request; performing S14; S13, using a preset initial weight allocation strategy to allocate initial fusion weights to each modality data of the multi-modal query request, and fusing each modality data; performing S14; S14, performing corresponding query based on the fusion result, and obtaining query execution data and query result; S15, calculating actual fusion weights of each modality data and reward function values based on query execution data and query results; The actual fusion weights of each modality data are determined based on a semantic correlation coefficient, an execution plan perception coefficient and a resource efficiency coefficient; S16, storing the actual fusion weights and the reward function values into a corresponding intermediate fusion weight record table, and updating key hyperparameters in the execution plan perception coefficient and the resource efficiency coefficient by using an improved gradient descent algorithm, so as to maximize cumulative reward values; S17, performing convergence determination on the intermediate fusion weight record table corresponding to the multi-modal query request, if a fluctuation amplitude of the last continuous m reward function values is less than or equal to a set fluctuation amplitude, calculating an average value of the actual fusion weights corresponding to the m reward functions as a final fusion weight, and adding the semantic structure feature vector, the final fusion weight and convergence stability of the query request into a current modality fusion weight list, and performing S11; Otherwise, adding the semantic structure feature vector, the actual fusion weight and the convergence stability of the query request into the current modality fusion weight list, and performing S11.
5. The unified semantic representation based multi-modal database query engine according to claim 4, wherein, actual fusion weight of the i-th modality βc i satisfies the following condition: βc i =(α i ×γ i ×δ i ) / ∑ n j=1 (α) j ×γ j ×δ j ), α i γ is the semantic relevance coefficient of the i-th modality, determined based on the cosine similarity between the query vector corresponding to the multimodal query request and the i-th modality vector in the vector library; i Let γ be the execution plan perception coefficient for the i-th mode. i =log(1+C) i ×P i ×I i ) / log(1+C b ×P b ×I b ), where C i Let C be the execution resource consumption coefficient for the i-th mode. i =1 / (w1×IO i +w2×CPU i +w3×Net i ), IO i The disk I / O resource consumption value for the i-th mode, CPU i Net represents the CPU resource consumption value for the i-th mode. i P represents the network transmission resource consumption value for the i-th mode, where w1, w2, and w3 are resource weight coefficients; i Let I be the execution efficiency coefficient for the i-th mode. i C is the index utilization efficiency coefficient for the i-th mode. b P is the reference execution resource consumption coefficient for the i-th mode. b Let I be the reference execution efficiency coefficient for the i-th mode. b The reference index for the i-th mode utilizes the efficiency coefficient; δ i Let δ be the resource efficiency coefficient for the i-th mode. i =exp(-λ×S i ), S i Let α be the resource consumption intensity of the i-th mode, λ be the resource sensitivity coefficient, and exp() be the natural exponential function; j Let γ be the semantic relevance coefficient of the j-th modality. j Let δ be the execution plan perception coefficient for the j-th mode. j Let be the resource efficiency coefficient for the j-th mode, where j ranges from 1 to n.
6. The unified semantic representation based multi-modal database query engine according to claim 4, wherein, The reward function satisfies the following condition: R=k1×Q+k2×E+k3×F, wherein Q is a result quality score, E is a global time consumption optimization coefficient, F is a global resource utilization efficiency coefficient, k1, k2 and k3 are reward weight coefficients, and k1+k2+k3=1, wherein E=1 / (1+(ET-ET0) / ET0), ET is an actual execution time of the current query, ET0 is a benchmark execution time of a similar query, F=1 / (1+(RC-RC0) / RC0), RC is a comprehensive resource consumption value of the current query, and RC0 is a benchmark resource consumption value of a similar query.
7. The unified semantic representation based multi-modal database query engine according to claim 4, wherein, During execution of the multi-modal query execution module, execution information of each execution path in the optimal query execution plan is monitored in real time, and the weight and resource of each execution path are adjusted based on the real-time monitoring result; the execution information includes actual execution time, execution efficiency and query result quality score.
8. The unified semantic representation based multi-modal database query engine of claim 7, wherein, The weight of the execution path is adjusted by using an exponential smoothing algorithm.
9. The unified semantic representation based multi-modal database query engine according to claim 2, wherein, The single-modal semantic vector of a certain entity in the vector library satisfies the following condition: v u = v0 u + ∑ t≠u (A ut × v0 t ), wherein v u is the optimization vector of the u-th modality of the entity, v0 u is the initial vector of the u-th modality of the entity, generated based on the independent encoding of the u-th modality encoder, v0 t is the initial vector of the t-th modality of the entity, A ut is the alignment weight between the u-th modality and the t-th modality of the entity, u and t are valued from 1 to p, and t≠u, p is the total number of modalities contained in the source data of the entity.
10. A method for multi-modal database query based on unified semantic representation, characterized in that, The method is implemented based on the query engine of any one of claims 1 to 9, and the method comprises: performing semantic analysis on a query request input by a user to obtain corresponding semantic structure features; the query request includes a single-modal query and a multi-modal query; the semantic structure features include query intention, key entities and modality information; performing unified semantic coding on the semantic structure features to generate corresponding coding results; if the semantic structure features include single-modal data, a semantic vector corresponding to the single-modal data is directly generated as the coding result; if the semantic structure features include multi-modal data, semantic vectors of each modality are generated, and the generated semantic vectors are fused to generate a fused semantic vector as the coding result; Based on semantic features in the encoding result, data distribution characteristics of the vector library, query complexity, and query engine resource status, an optimal query execution plan is generated, which includes a vector retrieval strategy, a modality weight allocation, and a computing resource scheduling scheme; According to the optimal query execution plan, a query operation is performed on the vector library to obtain candidate results similar to the query request semantics; A weighted fusion operation of the multi-path results is performed to generate a weighted score of each candidate result; Each candidate result is output in descending order of the weighted score.
Citation Information
Patent Citations
Cross-modal search system
CN113946726A
Simulation method and system for intelligent fusion and dynamic prediction of electronic warfare information
CN120012609A