Multi-modal database query engine and method based on unified semantic representation
By using a multimodal database query engine with unified semantic representation, the problem of insufficient semantic understanding in multimodal data queries of traditional databases is solved. It realizes efficient semantic association and execution optimization across modal data, thereby improving query efficiency and accuracy.
Patent Information
- Application Number
- CN202511097961.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-08-06
AI Technical Summary
Traditional database query engines lack semantic understanding capabilities when processing multimodal data, and cannot effectively perform cross-modal semantic associations, resulting in low query efficiency. Existing technologies cannot simultaneously optimize semantic relevance and physical execution efficiency.
A multimodal database query engine based on unified semantic representation is adopted. The query parsing and semantic encoding modules parse user requests into semantic structural features and generate unified semantic vectors. The optimal query execution plan is generated by combining the data distribution characteristics and resource status of the vector library, and the query operation is executed and the results are fused.
It achieves semantic association of cross-modal data and improves execution efficiency. The query process is optimized through dynamic weight adaptive algorithm, which improves the intelligence and accuracy of multimodal data query and supports intelligent retrieval of massive heterogeneous data.
Smart Images

Figure CN120929654A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of database technology, and in particular to a multimodal database query engine and method based on unified semantic representation. Background Technology
[0002] With the deepening of digital transformation, enterprises and organizations have accumulated massive amounts of multimodal data, including text, images, audio, and video. How to effectively manage and query this heterogeneous data has become a major challenge in the current database technology field. Traditional database query engines are mainly designed for structured data and have limited support for multimodal data queries, especially in terms of semantic understanding and association. There is a technological gap between current multimodal AI research and database engineering practice. Multimodal AI models focus on semantic accuracy optimization but lack awareness of underlying data access costs and execution efficiency. Traditional database query optimizers have mature cost estimation and execution plan generation capabilities but lack semantic understanding capabilities and cannot handle cross-modal semantic associations. Existing multimodal data processing solutions cannot simultaneously optimize semantic relevance and physical execution efficiency. Specifically, heterogeneous data types lack a unified semantic representation mechanism, semantic associations cannot be optimized in conjunction with execution plans, and traditional query optimization techniques cannot perceive semantic features for intelligent decision-making. Therefore, it is necessary to establish a unified optimization mechanism that can simultaneously perceive semantic features and execution features to achieve a synergistic improvement in semantic accuracy and execution efficiency. Summary of the Invention
[0003] To address the aforementioned technical problems, the technical solution adopted by this invention is as follows: According to a first aspect of the present invention, a multimodal database query engine based on unified semantic representation is provided, the query engine being associated with a vector library, the vector library storing unified semantic vectors of multiple entities, the query engine comprising: The query parsing and semantic encoding module is used to perform semantic parsing on user-input query requests to obtain corresponding semantic structure features, and to perform unified semantic encoding on the semantic structure features to generate corresponding encoding results. Specifically, if the semantic structure features contain unimodal data, a corresponding unimodal semantic vector is directly generated as the encoding result; if the semantic structure features contain multimodal data, semantic vectors for each modality are generated and fused to generate a fused semantic vector as the encoding result. The query request includes unimodal queries and multimodal queries; the semantic structure features include query intent, key entities, and modal information.
[0004] The query execution plan generation module is used to generate an optimal query execution plan based on the semantic features in the encoding results, the data distribution characteristics of the vector library, the query complexity, and the query engine resource status. The optimal query execution plan includes a vector retrieval strategy, modality weight allocation, and computing resource scheduling scheme.
[0005] The multimodal query execution module is used to perform query operations on the vector library according to the optimal query execution plan, and obtain candidate results that are semantically similar to the query request.
[0006] The query result fusion module is used to perform weighted fusion operations on multi-path results and generate weighted scores for each candidate result; An output module is used to output candidate results in descending order of weighted scores. According to a second aspect of the present invention, a multimodal database query method based on unified semantic representation is provided. The method is implemented based on the query engine provided in the first aspect of the present invention, and includes: The user-input query request is semantically parsed to obtain the corresponding semantic structure features; the query request includes unimodal queries and multimodal queries; the semantic structure features include query intent, key entities and modal information.
[0007] The semantic structure features are uniformly semantically encoded to generate corresponding encoding results. If the semantic structure features contain single-modal data, the corresponding single-modal semantic vector is directly generated as the encoding result. If the semantic structure features contain multimodal data, the semantic vectors of each modality are generated and the generated semantic vectors are fused to generate a fused semantic vector as the encoding result.
[0008] Based on the semantic features in the encoding results, the data distribution characteristics of the vector library, the query complexity, and the query engine resource status, an optimal query execution plan is generated. The optimal query execution plan includes a vector retrieval strategy, modality weight allocation, and a computing resource scheduling scheme.
[0009] According to the optimal query execution plan, a query operation is performed on the vector library to obtain candidate results that are semantically similar to the query request.
[0010] Perform a weighted fusion operation on the multi-path results to generate a weighted score for each candidate result.
[0011] Output the candidate results in descending order of weighted scores.
[0012] The present invention has at least the following beneficial effects: The multimodal database query engine and method based on unified semantic representation provided in this invention map heterogeneous data to a unified semantic space through an encoder and an execution plan-aware fusion engine. The fusion engine employs a three-factor dynamic weight adaptive algorithm, combining database execution plan theory with multimodal semantic fusion. The semantically aware query optimization framework incorporates semantic characteristics and execution plan features into the optimization decision process based on the weight calculation results of the fusion engine. The multimodal query execution and result fusion mechanism, based on the dynamic weight configuration of the fusion engine, supports cross-modal query operations. This query engine effectively solves key technical challenges in traditional multimodal data querying, such as superficial semantic understanding, difficulty in cross-modal association, and low query efficiency, particularly achieving significant breakthroughs in intelligent weight allocation and precise execution optimization. It provides a complete technical solution for the intelligent retrieval and application of massive heterogeneous data.
[0013] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0014] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0015] Figure 1 This is a structural block diagram of a multimodal database query engine based on unified semantic representation provided in an embodiment of the present invention; Figure 2 A flowchart of a multimodal database query method based on unified semantic representation provided in an embodiment of the present invention. Detailed Implementation
[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0017] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of this invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0018] It should be noted that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the steps as sequential processes, many of these steps can be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the steps can be rearranged. A process can be terminated when its operation is complete, but it may also have additional steps not included in the figures. A process can correspond to a method, function, procedure, subroutine, subroutine, etc.
[0019] This invention aims to provide a multimodal database query engine based on unified semantic representation that integrates multimodal semantic understanding capabilities into the query engine architecture, enabling unified semantic representation and query processing of multiple modal data.
[0020] The query engine in this embodiment of the invention is associated with a vector library, which stores unified semantic vectors for multiple entities.
[0021] In this embodiment of the invention, the vector library is built based on the Unified Semantic Representation (USR) model. The USR model includes multiple encoders and fusion machines. The encoders are used to convert data of different forms into a unified semantic vector representation, enabling the query engine to process data of multiple modalities within the same semantic framework. In this embodiment, a dedicated encoder implementation is designed for each major data modality to ensure that data of each modality can be accurately mapped to a unified semantic space. This may include, for example, a semantic encoder, an image encoder, a video encoder, and an audio encoder.
[0022] The text encoder implements the Etext(xtext) function defined in the USR model, adhering to strict technical specifications to ensure the accuracy and consistency of semantic encoding. The text preprocessing stage employs standardized word segmentation and cleaning processes, including special character removal, unified encoding formatting, and sentence segmentation, ensuring consistent input text formatting. The semantic encoding stage uses a pre-trained language model based on the Transformer architecture, extracting deep semantic features from the text through a multi-layer self-attention mechanism to generate a 768-dimensional semantic vector representation. The vector normalization stage performs L2 normalization on the generated semantic vectors, ensuring a vector magnitude of 1, allowing comparison of semantic vectors from different texts on a unified unit sphere. The quality verification stage verifies encoding quality through semantic similarity testing, ensuring that semantically similar texts are close in the vector space, while texts with significant semantic differences are far apart.
[0023] Image encoders implement the Eimage(ximage) function, which is more complex than text encoders, requiring the handling of multi-level semantic abstraction of visual information. Image preprocessing includes operations such as size normalization, color space conversion, and data augmentation to unify the input image into a 224×224 pixel RGB format. The feature extraction stage uses a pre-trained visual model based on ResNet or VisionTransformer architecture to extract visual features from the image through convolutional neural networks or self-attention mechanisms. The semantic mapping stage maps visual features to a unified 768-dimensional semantic space through fully connected layers, ensuring consistency with the output dimension of the text SEncoder. The cross-modal alignment stage optimizes the semantic consistency between image vectors and corresponding text description vectors through contrastive learning methods, ensuring that images and texts describing the same content are located close to each other in the semantic space.
[0024] The audio encoder implements the Eaudio(xaudio) function, employing an audio coding model based on Wav2Vec or a similar architecture, and extracts semantic features of the audio through time-frequency analysis and deep learning methods. The video encoder implements the Evideo(xvideo) function, employing a spatiotemporal convolutional network or video Transformer architecture, and generates semantic representations of the video through frame-level feature extraction and temporal modeling. The design of these two encoders follows the same output specifications as the text and image encoders, ensuring that the semantic representations of all modalities have a unified mathematical structure.
[0025] In this embodiment of the invention, a strict unified specification is established for the output of all encoders to ensure complete compatibility of semantic representations of data from different modalities. The output of all encoders is uniformly represented as a 768-dimensional L2-normalized vector, with each dimension of the vector having a value range of [-1, 1] and a vector magnitude strictly equal to 1. This unified vector representation ensures that data from different modalities can be effectively compared and correlated in a unified semantic space. The semantic relevance between different modalities can be directly evaluated through cosine similarity calculation, with a similarity value range of [-1, 1], where 1 represents complete similarity, -1 represents complete opposites, and 0 represents irrelevance.
[0026] The fusion unit integrates information from different modalities based on the encoder output by using a plan-aware dynamic weight adaptive algorithm.
[0027] In this embodiment of the invention, the unified semantic vector of each entity is obtained through the source data of the entity. Specifically, it is the single-modal semantic vector of each modal data contained in the source data of the entity after being converted by the corresponding encoder. That is, the vector library pre-stores the single-modal semantic vectors of various entity data after being converted by the encoder.
[0028] In this embodiment of the invention, the unimodal semantic vector of a certain entity in the vector library satisfies the following condition: v u =v0 u +∑ t≠u (A) ut ×v0 t ), where v u Let v0 be the optimization vector for the u-th mode of this entity. u v0 is the initial vector of the u-th mode of this entity, independently encoded based on the u-th mode encoder. t Let A be the initial vector of the t-th mode of the entity, independently encoded based on the t-th mode encoder. ut The alignment weight between the u-th and t-th modal of the entity is [0,1], which reflects the semantic association strength between the two modalities under the same entity. The values of u and t are from 1 to p, and t≠u. p is the total number of modalities contained in the source data of the entity.
[0029] Among them, A ut =exp(sum(v0) u v0 t ) / τ) / ∑ v=1 n (exp(sum(v0)) u v0 v ), sum(v0) u v0 t Let be the semantic similarity between the u-th mode and the t-th mode, and sum(v0) uv0 v ) represents the semantic similarity between the u-th mode and the v-th mode, where v ranges from 1 to p, and τ is the temperature coefficient, where τ>0 and the default value is 0.07, used to adjust the steepness of the weight distribution.
[0030] In its implementation, the query engine identifies semantically related cross-modal data pairs by calculating the vector distance between different modalities in a unified semantic space, and adjusts the strength of these associations through the weighting mechanism of the fusion engine. The alignment algorithm employs a batch processing mode, processing 32 samples of cross-modal data pairs at a time, accelerating the construction of the similarity matrix and the calculation of alignment weights through GPU parallel computing. The query engine maintains a cross-modal alignment cache, storing the alignment results of the most recent 1000 queries to accelerate the processing of similar queries. Specifically, when a new query is similar to a historical query, the query engine can directly use the alignment results in the cache, avoiding repeated complex cross-modal alignment calculations and thus improving query processing efficiency.
[0031] Furthermore, the multimodal database query engine based on unified semantic representation provided in this embodiment of the invention may include: The query parsing and semantic encoding module 1 is used to perform semantic parsing on the query request input by the user, obtain the corresponding semantic structure features, and perform unified semantic encoding on the semantic structure features to generate the corresponding encoding results.
[0032] Wherein, if the semantic structure feature contains unimodal data, the corresponding unimodal semantic vector is directly generated as the encoding result; if the semantic structure feature contains multimodal data, the semantic vectors of each modality are generated and the generated semantic vectors are fused to generate a fused semantic vector as the encoding result; wherein, the query request includes unimodal query and multimodal query.
[0033] In this embodiment of the invention, the semantic structure features include query intent, key entities, and modal information, specifically including: Modal combination features: recording the modal types and combination relationships contained in the query, such as "text + image", "audio single modality", "text-image-audio", represented by modal type encoding, such as text=1, image=2, audio=3, etc.; Query intent features: determining the query type through an intent classification model, such as similarity retrieval / attribute filtering / association analysis / multi-turn dialogue continuation, etc., and outputting intent labels and confidence scores (such as "similarity retrieval, confidence score 0.92"); Key entity features: extracting core entities in the query (such as "product ID=123", "person 'Zhang San'", "scene 'meeting room'"), mapping them to knowledge graph IDs through entity linking tools to form an entity set; Constraint features: parsing explicit / implicit constraints in the query (such as "time range: last 7 days", "number of results ≤ 10", "resource priority: low latency"), and converting them into structured constraint key-value pairs.
[0034] In this embodiment of the invention, the query parsing and semantic encoding module includes multiple modality encoders and a fusion unit. The i-th modality encoder is used to perform unified semantic encoding on the i-th modality to generate the corresponding encoding result, where i ranges from 1 to n, and n is the total number of modalities included in the query. The fusion unit is used to fuse the encoding results of each modality encoder to obtain the corresponding encoding result, which satisfies the following condition: F = ∑ n i=1 β i ×E i F represents the fusion result, E i β is the encoding result of the i-th modal encoder. i Let be the fusion weights for the i-th mode, where ∑ n i=1 β i =1 and β i ≥0.
[0035] The query execution plan generation module 2 is used to generate an optimal query execution plan based on the semantic features in the encoding results, the data distribution characteristics of the vector library, the query complexity, and the query engine resource status. The optimal query execution plan includes a vector retrieval strategy, modality weight allocation, and computing resource scheduling scheme.
[0036] In this embodiment of the invention, semantic features refer to the core attributes of queries or data at the semantic level, reflecting the meaning, association, and constraints of the content, and directly determining the relevance and matching accuracy of the retrieval. Its core manifestations include: (1) Modal composition and association: Single-modal semantic features: such as the theme of the text ("technology", "sports"), the subject of the image ("portrait", "landscape"), the scene of the audio ("speech", "music"), etc., the core meaning of a single modality; Multimodal semantic features: semantic association across modal content (such as "the matching degree of the visual features of the text description 'red sports car' and the red sports car in the image" and "the semantic consistency between the video image and the accompanying subtitles"). The higher the association strength, the easier it is to ensure the accuracy of multimodal collaborative retrieval. (2) Semantic density and core entities: Semantic density: The amount of effective semantic information contained in a unit of data (e.g., a 500-word news article with more than 5 core entities is high density, and an abstract painting without a clear subject is low density); Core entity: The semantic unit that plays a decisive role in the data (e.g., in the query "2023 Nobel Prize in Physics laureate's speech video", "2023 Nobel Prize in Physics laureate" and "speech video" are core entities). The clearer the core entity, the stronger the retrieval orientation. (3) Semantic strength of constraints: Strong constraints: Strict limitation on the semantic scope (e.g., "time = 2024", "location = Beijing", "sentiment = positive"), directly narrowing the retrieval boundary, requiring priority matching; Weak constraints: Vague or non-essential semantic requirements (e.g., "number of results ≥ 10", "relevant is sufficient"), which can be appropriately relaxed when efficiency is the priority. (4) Distribution characteristics of semantic vectors: Clustering in vector space: Whether vectors of the same semantic class cluster in the space (e.g., "animal" class vectors cluster in region A and "plant" class vectors cluster in region B). High clustering allows for quick retrieval through "region anchoring". Offset from the target library: The distance between the query vector and the global semantic center of the vector library (low offset indicates that the query content is common in the library and a general retrieval strategy can be used; high offset requires targeted expansion of the retrieval scope).
[0037] Semantic features determine "what to retrieve" and "how to match more relevant information," and are the core basis for choosing vector retrieval strategies (such as high-precision indexing or fast coarse screening) and modality weight allocation (such as prioritizing matching text semantics or image features).
[0038] Data distribution characteristics refer to the statistical and physical attributes of data in the vector library in terms of storage, access, and association. They reflect the "storage structure, quantity ratio, access popularity and cross-modal association status of data", which directly affect the efficiency of query and resource consumption. Its core manifestations include: (1) storage and index structure: index type: indexes adapted to different modal data (such as the "inverted index + vector index" hybrid structure commonly used for text, HNSW index for images (efficient approximate retrieval), and IVF index for audio (bucket retrieval, suitable for large-scale data)); sharding and replication: whether the data is sharded and stored according to modality / time / topic (such as log data sharded by "day"), and the number of replicas (multiple replicas can improve concurrent access capabilities, but increase storage costs). (2) Quantity and proportion distribution: Modal proportion: The proportion of each modal data in the database (e.g., text accounts for 60%, images account for 30%, audio accounts for 10%). Modalities with a high proportion should be prioritized for index performance optimization; Scale level: The magnitude of single modal data (e.g., 10 million text entries, 5 million images). For ultra-large-scale data (>100 million entries), a distributed retrieval strategy should be enabled. (3) Access popularity distribution: Hot data proportion: The proportion of data that has been accessed frequently in the past 24 hours (e.g., access volume > 50% of total access volume) in each modal. Hot data should be cached in memory or SSD to improve access speed; Access peak characteristics: The time pattern of data access (e.g., peak text retrieval during the day, peak image retrieval at night), used for predictive resource scheduling. (4) Cross-modal association density co-occurrence frequency: the proportion of association pairs of different modal data (such as "news text-image" and "product description-product video") in the total data (association density > 30% is high association). High association data can be used to build a joint index to reduce the time consumption of cross-modal matching; association stability: the persistence of cross-modal association (such as "text-image" association pair with a lifespan > 30 days is stable association). Stable association can pre-calculate matching results to reduce real-time computing costs.
[0039] Data distribution characteristics determine "how to retrieve efficiently" and "how to allocate resources," which are key factors in selecting computing resource scheduling schemes (such as prioritizing hot data caching and distributed node load balancing) and optimizing retrieval paths (such as using composite indexes for highly correlated data). Query complexity is assessed using a quantitative scoring model (0-10 points), combining modality combinations, constraint strength, accuracy requirements, and subtask characteristics for multi-dimensional evaluation. Specific criteria are as follows: (1) Low complexity (1-3 points) Applicable scenarios: Single-modal queries with only one weak constraint (such as "number of results ≥ 5"), with no special precision requirements (default similarity threshold 0.6-0.7). Subtask characteristics: Corresponds to 1-2 linear subtasks (such as "text retrieval → basic sorting"), and the computational load of a single task is less than 20% of the load of 1 CPU core.
[0040] (2) Medium complexity (4-7 points) Applicable scenarios: Bimodal query + 1-2 strong constraints (such as "time range = last 7 days + region = Beijing") + normal precision (threshold 0.7-0.8). Single-modal query + ≥3 strong constraints + standard precision; Subtask characteristics: There are 2-3 subtasks (such as "image retrieval → text attribute matching → cross-validation"), and some tasks can be performed in parallel (such as bimodal retrieval can be executed in parallel).
[0041] (3) High complexity (8-10 points) Applicable scenarios: Trimodal or higher queries + ≥2 levels of nested constraints (e.g., "first filter 'red car' images, then match text attributes 'price < 200,000' and 'rating > 4.5 stars'") + high precision requirements (threshold > 0.85); Dual-modal query + ≥3 strong constraints + high precision requirements; Subtask characteristics: There are ≥3 subtasks, including more than 1 high computational task (such as "video frame feature extraction" or "cross-modal attention fusion"), which will force the parallel execution strategy to be triggered.
[0042] The query engine resource status is evaluated in real time using the following quantitative indicators to support resource adaptation of the execution plan: (1) Core load indicators (including threshold determination) CPU utilization: ≤60% is low load, 60%-80% is medium load, and >80% is high load; Memory remaining: ≥40% is sufficient, 20%-40% is strained, <20% is severely insufficient; IO throughput: Disk IO < 100MB / s or network IO < 500Mbps is considered IO congestion; Index service response latency: ≤50ms is smooth, 50-100ms is slight delay, and >100ms is service congestion.
[0043] (2) Assessment of Parallel Computing Potential Basic parallel capabilities: Count the number of idle CPU cores; support multi-task parallelism when there are 4 or more cores, and limit parallelism when there are fewer than 2 cores. Accelerate resource adaptation: If GPU memory is ≥4GB and utilization is <30%, enable vector search acceleration (suitable for high-complexity queries). Distributed scheduling: Calculate the load difference of each node (CPU utilization standard deviation <10% is balanced and tasks can be evenly distributed; >30% is prioritized for scheduling to low-load nodes).
[0044] The multimodal query execution module 3 is used to perform a query operation on the vector library according to the optimal query execution plan, and obtain candidate results that are semantically similar to the query request.
[0045] The query result fusion module 4 is used to perform a weighted fusion operation on multi-path results and generate a weighted score for each candidate result; Output module 5 is used to output the candidate results in descending order of weighted scores.
[0046] Furthermore, the query engine is also associated with a dynamically updated modality fusion weight list. Each row of the modality fusion weight list includes the semantic structure feature vector of the corresponding query request and the fusion weight of each modality. Specifically, for a query request input by the user, its semantic structure feature vector is generated by the semantic encoding conversion module. The cosine similarity between this feature vector and each semantic structure feature vector in the modality fusion weight list is calculated. The modality fusion weights corresponding to the records with similarity greater than a set threshold are used as the fusion weights of each modality of the current query request.
[0047] In this embodiment of the invention, the threshold value can be an empirical value, such as 0.8.
[0048] In another embodiment of the present invention, each row of data in the modality fusion weight list further includes a weight stability index.
[0049] In the context of multimodal database query engines, the weight stability index is a parameter used to quantify the reliability and applicability of the fusion weights of each modality in a record. Its core function is to assist in determining whether the fusion weights of a record can be stably reused in a new query when the semantic structure feature vectors of the new query are similar to those of records in the list, avoiding fusion deviations caused by the accidental effectiveness of weights. When the semantic similarity between a new query and multiple records exceeds a threshold, the engine will prioritize the record weight with the higher stability index as a reference, rather than relying solely on similarity ranking. For example: Record A: similarity 0.85, stability 0.9 (based on a large number of samples, strong generalization); Record B: similarity 0.88, stability 0.6 (based on a small number of samples, weak generalization). In this case, the engine may prioritize the weight of Record A to ensure the reliability of the fusion result. In this embodiment of the invention, for a user-input query request, if there are no similar records in the modality fusion weight list, a preset initial weight allocation strategy is adopted to assign initial fusion weights to each modality data of the query request and to fuse the modality data.
[0050] In this embodiment of the invention, the initial fusion weights of each modality can be empirical values. In an illustrative embodiment, the initial fusion weights of each modality satisfy the following condition: β0 i =[c1×(1 / n)+c2×Ti +c3×SP i +c4×RF i ] / ∑ n i=1 [c1×(1 / n)+c2×T j +c3×SP j +c4×RF j ], where β0 i Let T be the initial fusion weights for the i-th mode. i Let T be the mode type coefficient of the i-th mode, where T is the mode type coefficient if the i-th mode is a structured mode. i =1.1, if the i-th mode is an unstructured mode, T i =0.95, SP i SP represents the query intent adaptation coefficient for the i-th modality, where SP is the visual / text modality in a retrieval query if the i-th modality is a visual / text modality. i =1.08, if the i-th modality is a structured modality in an analytical query, SP i =1.15, other cases, SP i =1.0. RF i Let be the system resource coefficient for the i-th mode. When the resource occupancy rate of the current mode is greater than 70%, RF i =0.85, when the efficiency of the last 3 executions is greater than or equal to 0.8, RF i =1.1, other cases, RF i =1.0. c1 to c4 are preset coefficients. In an illustrative embodiment, c1=0.4, c2=0.25, c3=0.2, and c4=0.15.
[0051] Furthermore, the modal fusion weight list is obtained through the following steps: S10, Initialize the intermediate fusion weight record set, and set the weight convergence judgment threshold, minimum sample size, and similar query group clustering threshold.
[0052] S11, for the currently acquired multimodal query request, extract its semantic structure features and generate a feature vector through the query parsing and semantic encoding module, and match the corresponding intermediate fusion weight record table from the current intermediate fusion weight record set: If a matching record table exists, execute S12; If no matching record table exists, create an empty intermediate fusion weight record table containing semantic structure feature vectors for the query request, and execute S13; S12, based on the initial fusion weights of each modal data corresponding to the intermediate fusion weight record table, fuse the modal data of the multimodal query request; execute S14; S13, using a preset initial weight allocation strategy, assign initial fusion weights to each modal data of the multimodal query request, and fuse the data of each modality; execute S14; S14, Execute the corresponding query based on the fusion result, and obtain the query execution data and query results; S15, calculate the actual fusion weight and reward function value of each modality based on the query execution data and query results; the actual fusion weight of each modality is determined based on the semantic relevance coefficient, execution plan awareness coefficient and resource efficiency coefficient; S16, store the actual fusion weights and reward function values into the corresponding intermediate fusion weight record table, and use an improved gradient descent algorithm to update the key hyperparameters in the execution plan perception coefficient and resource efficiency coefficient with the goal of maximizing the cumulative reward value. S17. Perform a convergence determination on the intermediate fusion weight record table corresponding to the multimodal query request. If the fluctuation range of the most recent m consecutive reward function values is less than or equal to the set fluctuation range, calculate the average value of the actual fusion weights corresponding to the m reward functions as the final fusion weight, and add the semantic structure feature vector of the query request, the final fusion weight, and the convergence stability to the current modal fusion weight list, and execute S11; otherwise, add the semantic structure feature vector of the query request, the actual fusion weight, and the convergence stability to the current modal fusion weight list, execute S11, and continue to accumulate samples.
[0053] In this embodiment of the invention, the semantic structure feature vector can be generated into a fixed-dimensional semantic structure feature vector V (the dimension is consistent with the vector dimension in the intermediate fusion weight record set, such as 512 dimensions) through the USR model. The generation rules are as follows: Modal combination features are converted into 128-dimensional sub-vectors through one-hot encoding and embedding layers; The query intent features are generated into 64-dimensional sub-vectors by adding confidence weights to the pre-trained vectors of the intent tags; Key entity features are weighted and summed using the embedding vectors of entity IDs (based on knowledge graph pre-training) to generate 256-dimensional sub-vectors; The constraint features are generated into 64-dimensional sub-vectors through hash encoding of constraint key-value pairs and rule mapping; Finally, the feature vector V is obtained by concatenation and L2 normalization, ensuring that the feature vectors of different queries are comparable in a unified semantic space.
[0054] In this embodiment of the invention, the intermediate fusion weight record set is a structured dataset that stores historical queries and corresponding fusion weight configurations. Each record table contains the following core fields: feature vector, modality fusion weight list, execution statistics (average time / resource consumption), update timestamp, and number of matches.
[0055] In this embodiment of the invention, if the cosine similarity between the feature vector corresponding to a certain intermediate fusion weight record table in the current intermediate fusion weight record set and the feature vector of the multimodal query request is greater than a set threshold, it indicates that there is a matching record table; otherwise, it indicates that there is no matching record table.
[0056] In this embodiment of the invention, the empty intermediate fusion weight record table contains the following initial information: Feature vector: The semantic structure feature vector of the current query; Modality fusion weight list: initialized to default values (e.g., single-modality query β). i =1.0, multimodal mixed query β i =1 / n); Execution statistics: null (to be filled in later). Creation timestamp: Current query engine time; Match count: 0; Related Query ID: A unique identifier for the current query (used for subsequent execution result backtracking). In this embodiment of the invention, S14 may specifically include: (1) Hierarchical query execution: Phase 1: Based on the semantic vectors of each modality, perform an approximate nearest neighbor search in the vector library (using the HNSW algorithm) and return TopK (default K=50) candidate results; The second stage involves performing modal consistency checks (cross-modal feature matching degree ≥ 0.5) and semantic conflict filtering on the candidate results (e.g., "the text description 'daytime' conflicts with the image feature 'night scene'"). The third stage: Based on the initial weights, the results that pass the verification are sorted a second time to generate the final query result set.
[0057] (2) Perform data collection: The data collection dimensions include: Time data: execution time of each modality individually, execution time of the fusion phase, and total execution time; Resource data: CPU utilization, number of I / O operations, network traffic, and total resource consumption for each mode; Quality data: semantic similarity scores for each modality and user-perceived quality scores for the fusion results (generated by a pre-trained quality assessment model).
[0058] (3) Result storage: The query result set is stored in association with the execution data. Each result contains: multimodal content (text / image / audio clips), contribution of each modality (calculated based on the initial weights), and semantic matching confidence (0-100 points).
[0059] Furthermore, the actual fusion weight βc of the i-th modei The following conditions must be met: βc i =(α i ×γ i ×δ i ) / ∑ n j=1 (α) j ×γ j ×δ j ).
[0060] Where, α i Let αi be the semantic relevance coefficient of the i-th modality, determined based on the cosine similarity between the query vector corresponding to the multimodal query request and the i-th modality vector in the vector library, αi = max(Similarity(q,v)). ik ), where k is the sample index of the i-th modality in the vector library, q is the query vector, and v ik This is the k-th data vector of the i-th mode.
[0061] In this embodiment of the invention, the method for calculating the actual fusion weight of each modality is also referred to as the three-factor weight calculation method.
[0062] γ i Let γ be the execution plan perception coefficient for the i-th mode. i =log(1+C) i ×P i ×I i ) / log(1+C b ×P b ×I b ), where C i Let C be the execution resource consumption coefficient for the i-th mode. i =1 / (w1×IO i +w2×CPU i +w3×Net i ), IO i This represents the disk I / O resource consumption value for the i-th mode, using a standardized value (range [0,1]). It comprehensively considers the number of disk read / write operations, seek time, data transfer volume, and I / O queue waiting time, and is calculated as a ratio to the upper limit of the query engine's I / O performance. CPU i The CPU resource consumption value for the i-th mode is a standardized value (range [0,1]), calculated based on the number of CPU cores occupied by the data in this mode, the computation time, and the instruction complexity, and determined by the ratio to the single-core full-load computing capacity. iThe network transmission resource consumption value for the i-th mode is a standardized value (range [0,1]), calculated by combining data transmission volume, network latency, and bandwidth utilization, and then compared with the maximum network transmission capacity of the query engine. w1, w2, and w3 are resource weight coefficients used to dynamically adjust the impact weights of the three types of costs: w1 is increased when the query engine's disk I / O load is too high, w2 is increased when CPU resources are scarce, and w3 is increased when there is network congestion. The weight values are updated in real time through the query engine resource monitoring module.
[0063] P i Let P be the execution efficiency coefficient of the i-th mode. i =min((PO) i / TO i )×(AC i / REC i ),1), where PO i The number of parallelizable operations for the i-th modality, such as sharding operations in distributed queries or batch comparison operations in vector retrieval, is determined based on the operation type (CPU-intensive / IO-intensive). i The total number of operations for the i-th mode includes parallel and serial operations, such as the serial conversion step in data preprocessing. PO i / TO i This represents the percentage of parallel operations (ranging from 0 to 1). A higher value indicates greater parallel potential. i REC represents the actual number of CPU cores available in the i-th mode. i The total number of CPU cores pre-allocated for the i-th mode, where AC is the initial resource quota based on the query plan. i / REC i This represents the core utilization rate, with a value range of [0,1]. A higher value indicates a better match between resource allocation and actual needs.
[0064] I i I is the index utilization efficiency coefficient for the i-th mode. i =ISR i / TSR i ×SL i ISR i TSR represents the number of rows actually read by the i-th mode through index scanning. i For the i-th mode, if a full table scan is used, the ISR is the total number of rows that need to be read. i / TTSR i This represents the proportion of rows scanned by the index out of the total number of rows in the table. The value ranges from (0,1). A smaller value indicates a more significant index filtering effect. For example, if the index scans 10,000 rows out of 1 million rows, the proportion is 0.01. iSL is the selectivity coefficient for the i-th mode, with a value range of [0,1], reflecting the accuracy of the index's adaptation to query conditions. i =AR i ×FE i AR i Let be the estimation accuracy ratio for the i-th mode, and be the ratio of the estimated number of rows scanned to the actual number of rows scanned; FE i represents the index filtering efficiency of the i-th mode, and represents the proportion of irrelevant rows filtered out by the index to the total number of rows in the table, ranging from [0,1]. A higher value indicates that the index has a stronger ability to exclude irrelevant data.
[0065] C b P is the reference execution resource consumption coefficient for the i-th mode. b Let I be the reference execution efficiency coefficient for the i-th mode. b The reference index for the i-th mode utilizes the efficiency coefficient, C b P b and I b All are dynamically updated based on the sliding window statistical method.
[0066] δ i Let δ be the resource efficiency coefficient for the i-th mode. i =exp(-λ×S i ), S i Let be the resource consumption intensity of the i-th mode, with a standardized value ranging from [0,1], where 0 represents no consumption and 1 represents resource consumption reaching the query engine's limit. λ is the resource sensitivity coefficient, used to adjust the function's sensitivity to resource consumption; exp() is the natural exponential function; α j Let γ be the semantic relevance coefficient of the j-th mode. j Let δ be the execution plan perception coefficient for the j-th mode. j Let be the resource efficiency coefficient for the j-th mode, where j ranges from 1 to n. Further, in this embodiment of the invention, the reward function satisfies the following condition: R = k1 × Q + k2 × E + k3 × F, where Q is the result quality score, taking the value [0,1], E is the global time consumption optimization coefficient, F is the global resource utilization efficiency coefficient, k1, k2, and k3 are reward weight coefficients, k1 + k2 + k3 = 1, and k1, k2, and k3 are dynamically adjusted according to the query engine load (e.g., k2 + k3 = 0.7 under high load). In an illustrative embodiment, k1 = 0.4, k2 = k3 = 0.3. Wherein, E = 1 / (1 + (ET - ET0) / ET0), where ET is the actual execution time of the current query, and ET0 is the baseline execution time of similar queries; F = 1 / (1 + (RC - RC0) / RC0), where RC is the comprehensive resource consumption value of the current query, and RC0 is the baseline resource consumption value of similar queries.
[0067] In this embodiment of the invention, the result quality score can be comprehensively evaluated by a multi-dimensional quality evaluation system. This stage combines quantitative and qualitative indicators such as user interaction behavior analysis, semantic relevance score and user satisfaction feedback, and is calculated using the analytic hierarchy process.
[0068] Furthermore, S16 may specifically include: (1) Record table update: Add an entry to the intermediate fusion weight record table currently being queried, including: actual fusion weight, reward function value R, execution data snapshot, and update timestamp.
[0069] (2) Hyperparameter optimization objective: With the goal of maximizing the cumulative reward value ΣR within the window, optimize the following key hyperparameters: Weighting allocation in the perception coefficient of the execution plan; Sensitivity coefficient in resource efficiency coefficient; The weights k1, k2, and k3 in the reward function.
[0070] (3) Improve the execution of the gradient descent algorithm: The Adam optimizer is used, and the learning rate is dynamically adjusted (initially 0.01, and halved if the cumulative reward value decreases for 3 consecutive times). In each iteration, the loss function L=-ΣR is calculated, and the hyperparameters are updated through gradient backpropagation. Add an L2 regularization term (coefficient 0.001) to prevent overfitting and ensure that the hyperparameters are within a reasonable range in physical terms (e.g., λ∈[0.5,3]). The optimization cycle is synchronized with the sliding window (executed every 5 minutes) to ensure that hyperparameters are adapted to the recent query engine status.
[0071] In this embodiment of the invention, the fluctuation range can be set as an empirical value, for example, 0.05.
[0072] Furthermore, to enhance the robustness of the system, this invention designs a complete exception handling and degradation mechanism. When execution plan information is missing, the system uses a preset default γ coefficient value, or calculates weights based solely on α and δ factors. When real-time weight calculation is abnormal, the system quickly switches to a backup weight configuration based on historical statistics. When weight configuration is abnormal or performance degrades, the system automatically performs parameter correction and adaptive adjustment. These mechanisms ensure that the algorithm can still operate normally in heterogeneous or incomplete information environments, maintaining the stability of the query engine in various complex semantic query scenarios.
[0073] The following is a detailed description of each module of the query engine provided in the embodiments of the present invention.
[0074] Furthermore, in this embodiment of the invention, the query parsing and semantic encoding module 1 is based on the USR model, and in particular, utilizes an encoder to uniformly process various input methods in the query understanding domain. This module supports multiple input methods such as structured queries, programmatic interface calls, and natural language queries, and converts them into vector representations in the semantic space by calling the corresponding encoding functions.
[0075] The query parsing and semantic encoding module integrates multiple query understanding methods to achieve unified processing and semantic conversion of different input methods, including extended structured query language interface, programmatic API interface and natural language query interface.
[0076] The extended Structured Query Language (USR) interface, while maintaining compatibility with traditional SQL syntax, introduces a set of operators specifically designed for multimodal semantic queries. These operators are directly based on the semantic space of the USR model. Semantic operators include: IMAGE_SIMILAR_TO for image similarity queries, TEXT_CONTAINS for text semantic inclusion queries, and AUDIO_MATCHES for audio matching queries. Through the unified semantic space of the USR model, users can query related text content based on images, or find similar images based on text descriptions.
[0077] The programmatic API adopts modern API design principles, providing flexible and powerful programmatic access capabilities for various application query engines. This allows developers to seamlessly integrate multimodal query functionality into existing application architectures. The main improvement of the programmatic API lies in its native support for multimodal content. Developers can submit complex query requests containing multiple modalities such as text, images, and audio through a unified interface, and the query engine automatically handles the semantic relationships and fusion between different modalities. Furthermore, the programmatic API provides endpoints at multiple complexity levels, from basic query operations to advanced semantic analysis. This design ensures ease of use for simple application scenarios while meeting the flexibility requirements of complex applications. The programmatic API response design is carefully optimized, providing not only standardized query results but also rich metadata information, such as query processing time, relevance score, and confidence level, facilitating subsequent result processing and user experience optimization at the application layer.
[0078] The Natural Language Query Interface (NLE) enables intelligent conversion from natural language expressions to structured semantic queries, significantly lowering the technical barrier for users of the multimodal query engine. The NLE employs a multi-stage natural language understanding pipeline, breaking down complex language processing tasks into a series of progressively refined steps, including preprocessing, intent recognition, entity extraction, and query reconstruction. The query engine's language processing capabilities are not only reflected in its support for multiple natural languages, but more importantly, in its deep semantic understanding of user query intent. It can accurately identify the user's true query needs and key query parameters from unstructured natural language expressions. A cross-modal semantic alignment mechanism plays a crucial role in this process. Through the cross-modal alignment algorithm integrated into the fusion processor, the query engine can establish semantic relationships between different modalities, achieving accurate understanding and processing of natural language queries containing multimodal references.
[0079] Intent recognition is achieved through an intent recognition model, which supports fine-grained query intent classification and can distinguish different types of query needs, such as similarity search, content retrieval, and association analysis, providing precise semantic guidance for subsequent query processing. The query engine integrates knowledge graph-based entity processing enhancement technology, effectively solving common ambiguities and incomplete expressions in natural language queries through entity recognition, disambiguation, and concept expansion. Furthermore, the query engine can understand and process omitted expressions, pronoun references, and contextual dependencies in multi-turn dialogues, providing users with a consistent interactive experience by maintaining dialogue state information. When the query engine's understanding of the user's query intent is uncertain, an interactive clarification mechanism is activated, collaborating with the user through proactive questioning to ensure the accuracy of query understanding. This human-computer collaborative query understanding mode significantly improves processing performance in complex query scenarios.
[0080] The semantic encoding transformation of query content involves bringing the execution plan awareness mechanism forward to the query understanding stage, achieving a deep integration of semantic understanding and execution optimization. The query engine applies corresponding encoders to all types of queries, converting the query content into vector representations in a unified semantic space. Query inputs of different modalities are processed through specially designed encoders: text queries use a language model encoder to achieve semantic vector mapping; image queries utilize a visual encoder to extract visual semantic features; audio queries undergo semantic encoding using an acoustic model; and video queries generate a comprehensive semantic representation through spatiotemporal feature fusion technology.
[0081] When a user submits a complex query containing multiple modalities, the query engine not only calculates the vector representation of each modality separately through the corresponding encoder, but more importantly, it introduces a pre-analysis mechanism for the query execution plan during the fusion process. The weight calculation at this stage is based on historical query statistics or preset values for preliminary estimation, providing a basic cost estimation basis for the subsequent query optimizer. By analyzing historical access patterns of each modality's data, estimating parallel execution potential, and resource consumption characteristics, the query engine calculates preliminary three-factor weight parameters, ensuring that the efficiency factors of subsequent execution are considered during the generation of the query vector. This design, which integrates database query optimization theory into the semantic understanding process, allows the query engine to begin optimizing execution strategies while understanding the user's intent, laying the foundation for real-time dynamic weight adjustment in the subsequent query execution stage and achieving a balance between semantic accuracy and execution efficiency.
[0082] The innovative value of the entire semantic encoding and conversion process lies in establishing an integrated processing mechanism from query understanding to execution optimization, ensuring that query content of different modalities can not only be accurately represented in a unified semantic space, but also lay the foundation for subsequent efficient execution. This is a forward-looking optimization capability that traditional query engines do not possess.
[0083] The query understanding mechanism based on the USR model maps different modalities and types of queries to a unified semantic space, providing a foundation for subsequent query optimization and execution.
[0084] In this embodiment of the invention, the query execution plan generation module 2 incorporates the semantic characteristics of the query into the optimization decision-making process, achieving a leap from traditional structure-aware to semantic-aware query optimization technology. This module receives a standardized semantic vector representation from the query parsing module and generates an optimized execution plan for multimodal data queries by analyzing the semantic features of the query, data distribution characteristics, and query engine resource status. This semantic-aware optimization capability enables the query engine to intelligently optimize based on the semantic content of the data, rather than just its structural features, providing an efficient execution strategy for multimodal semantic queries.
[0085] Furthermore, in this embodiment of the invention, the query execution plan generation module implements semantic query optimization in collaboration with the fusion engine. This module makes optimization decisions based on the dynamic weight adaptive mechanism of the fusion engine. The query engine uses the weight configuration calculated by the fusion engine to guide the semantic fusion strategy. By analyzing the execution plan characteristics such as execution cost, parallelism, and index utilization of each modality of data, it provides execution environment feedback information for the weight calculation of the fusion engine.
[0086] The core foundation of the query execution plan generation module lies in the encoder and execution plan-aware fusion engine defined in the USR model. The module first receives the query vector and weight configuration generated by the fusion engine, identifies the types involved, and dynamic weight distributions, and then selects the most suitable execution path based on this information. When a query contains multiple modalities, the module utilizes the dynamic weight results calculated by the fusion engine to deeply analyze the semantic relevance, execution cost characteristics, and resource consumption patterns of each modality. This determines the optimal modality processing order, parallel execution strategy, and inter-modality cooperation mechanism, generating the optimal execution plan. The module places particular emphasis on providing execution environment feedback to the fusion engine. By analyzing the index utilization, access pattern complexity, and parallel execution potential of each modality's data, it provides precise execution plan-aware parameters for the fusion engine's dynamic weight adjustment.
[0087] Furthermore, in this embodiment of the invention, the integration of the query execution plan generation module with the PostgreSQL query optimizer is achieved through a standardized interface. The query engine has designed a complete standardized interface specification for execution plan features, which defines a precise conversion mechanism from real database execution plans to weight calculation parameters. The interface specification uses JSON format as the data exchange standard and includes core data structures such as execution node structure, cost model, and resource configuration.
[0088] The query engine parses the execution plan based on the EXPLAINANALYZE output of the PostgreSQL database. Key fields are extracted from the JSON-formatted execution plan: the `actual_time` field is used to calculate execution cost characteristics; the `shared_hit_blocks` and `shared_read_blocks` fields are used to calculate IO resource consumption characteristics; and the `workers_planned` and `workers_launched` fields are used to calculate parallel execution efficiency characteristics. Index utilization efficiency characteristics are calculated by statistically analyzing the proportion of "Index Scan," "Index Only Scan," and "Seq Scan" nodes in the execution plan. Simultaneously, the differences between the "PlanRows" and "Actual Rows" fields are used to evaluate index selectivity, resulting in a comprehensive quantitative indicator of index utilization efficiency.
[0089] The specific execution plan parsing process achieves complete feature extraction through six consecutive processing stages. The execution plan acquisition stage uses the PostgreSQL EXPLAIN command, specifying the ANALYZE, BUFFERS, and FORMATJSON parameters, to obtain a JSON data structure containing detailed statistical information about the execution plan. This data structure includes all key performance indicators and resource consumption information during the query execution process. The node traversal stage recursively traverses each node of the execution plan tree using a depth-first search algorithm. It identifies operation nodes involving different modalities of data by analyzing the "Node Type" field of each node and establishes a parent-child relationship mapping table between nodes for subsequent feature aggregation calculations. The cost feature extraction stage extracts the actual execution time from the "ActualTotal Time" field of each node, calculates the cache hit rate by the ratio of "Shared HitBlocks" to "Shared Read Blocks," and evaluates the usage of temporary storage based on the "Temp Read Blocks" and "Temp Written Blocks" fields. The parallel feature analysis phase calculates the actual efficiency ratio of parallel execution by comparing the "WorkersPlanned" and "WorkersLaunched" fields, analyzes the "Parallel Aware" flag to determine the degree of parallelization of operations, and counts the number of worker processes actually involved in execution based on the "Worker Number" field. The index utilization evaluation phase calculates the proportion of "Index Scan," "Index OnlyScan," and "Seq Scan" nodes in the execution plan, and evaluates the estimation accuracy of the query optimizer and the index selectivity effect by comparing the differences between the "Plan Rows" and "Actual Rows" fields. The feature vector generation phase transforms the raw data extracted in the previous phases according to a predefined standardized formula to generate feature vectors that meet the input requirements of the fusion weight calculation model, ensuring the comparability of feature data for different queries and modalities.
[0090] The interface specification defines a standardized feature vector format: the ExecutionFeature structure contains three main components: cost_vector, parallel_vector, and index_vector. Each vector is normalized to ensure that execution plan features can be compared and calculated within a uniform numerical range.
[0091] Taking the execution plan parsing of multimodal queries as an example, when a query involves a join query on both a text table and an image table, the execution plan generated by PostgreSQL typically includes join operation nodes to join the two tables. The text table may be accessed using an index scan, while the image table may be accessed using a sequential scan. The query engine extracts key feature information from the execution plan, including the actual execution time, cache block hits, and disk read blocks of the join nodes; the execution time and number of parallel worker processes of the index scan nodes; and the execution time, planned row count, and actual row count of the sequential scan nodes. Through standardized calculation formulas, the query engine converts this raw feature data into standardized parameters required for three-factor weight calculation, including resource consumption coefficients, parallel efficiency coefficients, and index utilization efficiency coefficients for the text modality, as well as corresponding coefficients for the image modality. Finally, it generates execution plan-aware coefficients for each modality for weight calculation.
[0092] The query engine can obtain execution plan information from the PostgreSQL query optimizer and convert it into feature vectors required for weight calculation. This integration mechanism has good extensibility; through an abstract execution plan parsing interface, the query engine can adapt to the execution plan formats of other relational database query engines (such as MySQL and Oracle), achieving cross-database platform semantic query optimization capabilities. This integration mechanism ensures that the execution plan-aware dynamic weight adaptive algorithm can fully utilize the results of underlying database query optimization, achieving a deep integration of semantic query and traditional query optimization.
[0093] Semantic-aware query analysis identifies query features from the semantic representation of the query, including the modality types involved and the semantic complexity of the query. The intelligent index selection mechanism chooses the optimal indexing strategy based on the semantic characteristics of the query, supporting various indexing techniques suitable for semantic vectors.
[0094] The semantically aware execution plan is generated based on the dynamic weight adaptive algorithm results of the fusion engine. For queries involving multiple modalities, the execution plan directly uses the three-factor weight results calculated by the fusion engine to guide the formulation of the execution strategy. The query engine dynamically adjusts the processing priority and resource allocation of each modality in the execution plan according to the weight factors such as semantic relevance, execution cost, and resource efficiency of each modality provided by the fusion engine.
[0095] Taking the execution plan generation process for cross-modal queries as an example, when processing composite queries involving text retrieval and image matching, the query engine first evaluates the access characteristics of each modal data table by analyzing PostgreSQL's query engine statistics, including index configuration, data distribution characteristics, and access pattern complexity. Text data tables typically have a vector similarity-based index structure and good selectivity, while image data tables, due to the complex distribution characteristics of high-dimensional vector data, may lack an effective index structure and require sequential scanning for access. Based on the pre-analysis results of the execution plan, the query engine calculates the execution plan perception coefficient for each modality, which reflects the relative execution efficiency of different modalities in the current database environment. The optimizer generates a targeted execution plan based on the comparative analysis results of the weight coefficients. This plan uses the modal operations with higher execution efficiency as the driving table, prioritizes the execution of efficient filtering operations to reduce the size of the dataset for subsequent processing, and then performs computationally intensive modality matching processing on the filtered candidate result set. This execution order arrangement effectively optimizes the overall query performance. The specific execution plan tree structure selects an appropriate connection strategy based on the execution characteristics of each modality. It may adopt different modes such as nested loop connection, hash connection or sorted merge connection. Through the optimized plan structure arrangement, the overall execution cost is significantly reduced.
[0096] The multimodal query execution module 3 interacts with the database storage engine to complete the actual query processing according to the optimal execution plan, breaking through the limitation of traditional query execution engines that only support structured data operations.
[0097] The multimodal query execution module implements a dynamic weight adaptation mechanism during execution, enabling real-time adjustment of the fusion engine's weight configuration based on actual execution conditions. The query engine directly uses execution feedback to optimize the semantic fusion strategy, dynamically updating the three-factor weight parameters by collecting key indicators such as execution cost, resource consumption, and result quality for each modality in real time.
[0098] The multimodal query execution module directly manipulates the semantic encoding vectors defined in the USR model to implement a dynamic weight adjustment mechanism based on execution plan awareness. When processing unimodal queries, the execution engine optimizes the execution strategy for the encoder of the specific modality. For multimodal queries, the execution engine implements the dynamic weight fusion mechanism defined in the fusion unit, dynamically adjusting the weight configuration by monitoring performance metrics during execution in real time.
[0099] At the start of execution, the query engine generates an initial weight configuration based on historical statistics and preset parameters, assigns corresponding weight values to each execution path, and initiates a multi-path parallel execution process.
[0100] During the execution of queries by the multimodal query execution module, the execution information of each execution path in the optimal query execution plan is monitored in real time, and the weight and resources of each execution path are adjusted based on the real-time monitoring results. The execution information includes actual execution time, execution efficiency, and query result quality score. Specifically, during execution, the execution status of each path can be monitored in real time using PostgreSQL's statistical collector, and key performance indicators are collected for dynamic adjustment and judgment.
[0101] When the actual execution time of an execution path deviates significantly from the estimated time, the query engine determines whether a weight adjustment needs to be triggered based on a preset deviation threshold. For paths with performance deviations within the normal fluctuation range, the query engine maintains their current weight configuration; for paths with performance deviations significantly exceeding the threshold, the query engine initiates a weight recalculation process. Upon detecting a path whose execution efficiency significantly deviates from expectations, the query engine immediately triggers a weight recalculation process. The weight adjustment mechanism adjusts the weight allocation according to the actual execution effect of each path. Paths with relatively improved execution efficiency will have their weights increased, paths with relatively decreased execution efficiency will have their weights decreased, and paths with stable execution efficiency will have their weights maintained at a relatively balanced level. In this embodiment of the invention, an exponential smoothing algorithm can be used to adjust the weights of the execution paths.
[0102] The execution engine reallocates query engine resources based on the adjusted weight configuration, optimizes the execution priority of each path, improves overall query performance through intelligent resource scheduling strategies, and ensures the stability of semantic matching quality.
[0103] In a specific embodiment of the present invention, the weight adjustment of the execution path is triggered when any of the following conditions are met: performance deviation detection: the actual execution time of a certain modality deviates from the estimated time by more than 10%; resource status change detection: the CPU or memory utilization of the query engine changes by more than 15%; query quality degradation detection: the quality score of the query result is lower than 85% of the historical average for three consecutive times.
[0104] Semantic similarity search is the core operation of the multimodal query execution module, and the query engine implements efficient semantic vector retrieval capabilities. The execution engine can handle similarity calculations based on semantic vectors and supports cross-modal similarity searches in a unified semantic space. This is the technical foundation for realizing cross-modal query functions such as "searching for text by image" and "searching for images by text".
[0105] The semantic similarity search operation is directly implemented based on the encoder and fusion unit defined in the USR model. The query engine uses cosine similarity as the basic metric, evaluating semantic similarity by calculating the cosine of the angle between the query vector and candidate vectors. For unimodal queries, the query engine directly calculates the similarity between the query vector and candidate data vectors. For multimodal queries, the query engine first applies the fusion mechanism defined in the fusion unit implementation to generate a unified query vector, and then calculates its similarity with the candidate vectors. This similarity calculation based on a unified semantic representation ensures the semantic consistency and accuracy of the query results.
[0106] The core innovation of the multimodal query execution module lies in its implementation of a coordinated processing mechanism for semantic search and structured filtering. The query engine determines the optimal execution order of semantic and structured filtering through an execution plan-aware dynamic weight adaptive algorithm, and achieves a balance between query accuracy and execution efficiency through a multi-stage execution strategy. This execution capability based on a unified semantic representation enables the query engine to handle complex multimodal query requirements while maintaining excellent performance.
[0107] After the multimodal query execution module is completed, the query engine faces a key technical challenge: how to handle heterogeneous query results from different execution paths. The root of this challenge lies in the fact that even after sophisticated query optimization and execution processes, different execution paths can still produce result sets with significantly different quality. Without intelligent fusion, this will directly impact the accuracy of the final query results and the user experience.
[0108] The query result fusion module 4 implements an intelligent result fusion mechanism based on execution feedback, which feeds back the performance data collected during the execution process to the weight optimization of the fusion unit, forming a complete execution-feedback-optimization closed loop.
[0109] Consider a specific query scenario: a user queries "find teaching resources about machine learning algorithms". After all the aforementioned optimization steps, the three execution paths of the query engine return the following results: The text retrieval path returned text resources such as "Detailed Explanation of Machine Learning Algorithms" and "Basic Tutorials on Deep Learning." These results performed well in text semantic matching, with relevance scores generally at a high level. The image retrieval path returned visual resources such as algorithm flowcharts and images of mathematical formulas. These results performed well in visual content matching, but due to the complexity of image semantic understanding, their relevance scores were relatively low. The cross-modal association path returned multimedia resources such as video tutorials and interactive demonstrations. These results combined text and visual information, with relevance scores at a moderate level. If the results are not fused and are directly sorted according to the raw scores of each path, serious quality issues will arise. Text path results will occupy the top positions due to their highest scores, but these results may lack intuitive visual aids, hindering algorithm understanding. Image path results will be ranked lower due to their lower scores, but these visual resources are valuable for understanding complex algorithms. Cross-modal resources, while comprehensive, may be overlooked due to their moderate scores.
[0110] A more critical issue lies in the differences in execution efficiency. Let's assume that during the current query execution process, the text retrieval path exhibits excellent efficiency due to its high index hit rate, the image retrieval path is less efficient due to the need to process high-dimensional vectors, and the cross-modal association path has moderate efficiency. Ignoring these efficiency differences and relying solely on semantic scoring for ranking will fail to reflect the actual reliability of the results for each path.
[0111] The query result fusion module addresses this issue by applying a dynamic weighting mechanism. The query engine assigns higher weights to text paths based on their high execution efficiency, lower weights to image paths based on their low execution efficiency, and medium weights to cross-modal paths based on their moderate execution efficiency. After weighting adjustments, the overall score for text path results is further improved, the overall score for image path results is moderately reduced, and the overall score for cross-modal path results remains stable. The final fusion ranking ensures that high-quality text resources are prioritized while also guaranteeing that valuable visual and multimedia resources receive appropriate ranking positions, providing users with comprehensive and reliable query results.
[0112] The query result fusion module uses a weighted combination formula to achieve a unified score for results from multiple paths. For candidate results from different execution paths, the query result fusion module obtains the weighted score of the candidate result in the following way: Score final =∑ z h=1 f h ×Score h Among them, Score final f represents the weighted score of the candidate results. hThis represents the dynamic weight coefficient of execution path h, where h ranges from 1 to z, and z is the number of execution paths. (Score) h This represents the semantic relevance score of the candidate result in execution path h. This weighted combination method ensures that paths with high execution efficiency have a greater influence on the final ranking, while maintaining the importance of semantic matching quality.
[0113] Semantic-aware result ranking is based on the unified semantic space of the USR model to calculate the semantic similarity between queries and results. For unimodal queries, the query engine uses the corresponding encoder to calculate the similarity between the query vector and the result vector. For multimodal queries, the query engine uses a fusion engine to calculate the similarity between the fused query representation and the result representation, ensuring that the ranking results accurately reflect the semantic intent of the query.
[0114] The query result fusion module addresses the unified integration of results from different modalities through a standardized fusion process. This process comprises four consecutive steps: The modality identification step identifies the modality type of each query result, determining whether it originates from a different processing path, such as text retrieval, image retrieval, or cross-modal association. The weight matching step associates each query result with the dynamic weight coefficients of its corresponding execution path, establishing a mapping between the query results and the weight configuration. The scoring calculation step calculates a weighted score based on the original semantic relevance score of the query result and the corresponding path weight coefficients. This weighted score reflects both the semantic matching quality and the reliability of the execution path.
[0115] Output module 5 is used to perform unified sorting based on the weighted scores of all query results, and generate the final fusion result list to ensure that high-quality results from different modalities can obtain appropriate sorting positions according to their actual value.
[0116] Based on the same inventive concept, embodiments of the present invention provide a multimodal database query method based on unified semantic representation. This method is implemented based on the aforementioned query engine, such as... Figure 2 As shown, the method includes: The user-input query request is semantically parsed to obtain the corresponding semantic structure features; the query request includes unimodal queries and multimodal queries; the semantic structure features include query intent, key entities and modal information.
[0117] The semantic structure features are uniformly semantically encoded to generate corresponding encoding results. If the semantic structure features contain single-modal data, the corresponding single-modal semantic vector is directly generated as the encoding result. If the semantic structure features contain multimodal data, the semantic vectors of each modality are generated and the generated semantic vectors are fused to generate a fused semantic vector as the encoding result.
[0118] Based on the semantic features in the encoding results, the data distribution characteristics of the vector library, the query complexity, and the query engine resource status, an optimal query execution plan is generated. The optimal query execution plan includes a vector retrieval strategy, modality weight allocation, and a computing resource scheduling scheme.
[0119] According to the optimal query execution plan, a query operation is performed on the vector library to obtain candidate results that are semantically similar to the query request.
[0120] Perform a weighted fusion operation on the multi-path results to generate a weighted score for each candidate result.
[0121] Output the candidate results in descending order of weighted scores.
[0122] In summary, the multimodal data query scheme based on unified semantic representation provided by the embodiments of the present invention has at least the following advantages: (1) The PostgreSQL database execution plan theory is deeply integrated into the entire process of multimodal semantic query processing. The dynamic weight adaptive mechanism that is aware of the execution plan achieves unified optimization of semantic accuracy and execution efficiency. The solution establishes a unified semantic space through the USR model. The three-factor dynamic weight calculation model of the fusion engine directly incorporates execution plan features such as execution cost, parallelism and resource utilization into the semantic fusion process, so that the query engine can start optimizing the execution strategy while understanding the semantics.
[0123] (2) A complete verification index system was established to ensure the technical effectiveness and engineering feasibility of each core innovation point through quantitative performance evaluation. The verification of the unified semantic representation of the USR model was evaluated through three key dimensions. The cross-modal semantic consistency index was evaluated by calculating the cosine similarity of different modal data with the same semantic content in the unified semantic space, requiring a similarity greater than 0.8 to verify the accuracy of semantic alignment. The semantic vector quality index was evaluated through the semantic similarity retrieval task, requiring the Top-10 retrieval accuracy rate to reach more than 85% to verify the effectiveness of semantic encoding. The cross-modal query success rate index was evaluated through cross-modal query tasks such as "searching for text by image" and "searching for images by text", requiring the query success rate to reach more than 80% to verify the cross-modal understanding ability. The validation of the execution plan-aware dynamic weight adaptive mechanism covers four core performance indicators: query execution efficiency improvement requires an average query time reduction of over 20% compared to the static weight scheme, verified through comparative experiments; weight adjustment response time requires weight update latency to be controlled within 100 milliseconds, verifying the real-time performance of dynamic adjustments; query engine resource utilization requires CPU and memory utilization to be improved by over 15% compared to traditional schemes, verifying resource optimization effects; and query result quality stability requires the result quality score fluctuation to not exceed 5% during dynamic weight adjustment, verifying the stability of the optimization process. The validation of multimodal query execution and result fusion capabilities is evaluated through three key indicators: multimodal query processing capability requires support for composite queries containing text, images, and audio modalities, with a query processing success rate of over 90%; result fusion quality is assessed through user satisfaction evaluation, requiring a user satisfaction score of 4.0 or higher (out of 5); and cross-modal association accuracy is evaluated using manually labeled cross-modal association datasets, requiring an association accuracy rate of over 75%.
[0124] (3) Theoretical analysis based on PostgreSQL database shows that the execution plan-aware dynamic weight adaptive algorithm has significant theoretical advantages over the traditional static weight scheme in multimodal query scenarios, mainly reflected in the comprehensive optimization of query execution efficiency, semantic matching quality and query engine resource utilization.
[0125] (4) Theoretical Comparative Analysis verified the optimization principle of this technical solution through the differences in algorithm mechanisms. Traditional static weighting schemes adopt fixed weight allocation strategies, which cannot be dynamically adjusted according to actual execution conditions and lack adaptability when facing different query scenarios and query engine states. The execution plan-aware dynamic weighting scheme adopts the adaptive weight adjustment mechanism proposed in this technical solution, which can dynamically adjust the weight allocation of each modality according to the real-time characteristics of the PostgreSQL execution plan, and adapt to different query scenarios and query engine environments through intelligent weight optimization strategies.
[0126] (5) The execution plan awareness mechanism incorporates key features of the database execution plan into the weight calculation process, achieving the integration of semantic understanding and execution optimization. The dynamic weight adjustment mechanism enables the weight configuration to reflect the current execution status and query engine environment through real-time monitoring and feedback optimization. The intelligent weight allocation strategy comprehensively considers multiple dimensions such as semantic relevance, execution efficiency, and resource utilization through a three-factor dynamic weight calculation model.
[0127] (6) The applicability analysis of the technical solution shows that the solution can effectively handle the cross-modal query requirements of large-scale multimodal data. Through the dynamic weight adjustment mechanism, it can perform intelligent optimization based on the real-time feedback of the PostgreSQL execution plan, and achieve a significant improvement in query performance while maintaining the accuracy of semantic matching.
[0128] (7) This technological innovation achieves a fundamental technological leap from traditional database query engines to intelligent semantic query engines by establishing a complete execution-aware semantic query processing mechanism. The technical solution has significant practical value in fields such as intelligent retrieval systems, multimedia content management platforms, and cross-modal data analysis applications, providing core technical support and theoretical foundation for multimodal data processing applications based on PostgreSQL databases. This solution, through the deep integration of database execution plan theory and multimodal semantic fusion technology, provides an innovative technical path for solving the problem of unified querying and intelligent retrieval of large-scale heterogeneous data. This embodiment of the invention also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being configured to execute the methods described in this embodiment of the invention.
[0129] This invention also provides a computer-readable storage medium storing computer-executable instructions for performing the methods described in this invention.
[0130] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this invention can be achieved, and this is not limited herein.
[0131] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A multimodal database query engine based on unified semantic representation, characterized in that, The query engine is associated with a vector library, which stores unified semantic vectors for multiple entities. The query engine includes: The query parsing and semantic encoding module is used to perform semantic parsing on user-input query requests to obtain corresponding semantic structure features, and to perform unified semantic encoding on the semantic structure features to generate corresponding encoding results. Specifically, if the semantic structure features contain unimodal data, a corresponding unimodal semantic vector is directly generated as the encoding result; if the semantic structure features contain multimodal data, semantic vectors for each modality are generated and fused to generate a fused semantic vector as the encoding result. The query request includes unimodal queries and multimodal queries; the semantic structure features include query intent, key entities, and modal information. The query execution plan generation module is used to generate an optimal query execution plan based on the semantic features in the encoding results, the data distribution characteristics of the vector library, the query complexity, and the query engine resource status. The optimal query execution plan includes a vector retrieval strategy, modality weight allocation, and computing resource scheduling scheme. The multimodal query execution module is used to perform query operations on the vector library according to the optimal query execution plan and obtain candidate results that are semantically similar to the query request. The query result fusion module is used to perform weighted fusion operations on multi-path results and generate weighted scores for each candidate result; The output module is used to sort the candidate results in descending order of their weighted scores.
2. The multimodal database query engine based on unified semantic representation according to claim 1, characterized in that, The query parsing and semantic encoding module includes multiple modality encoders and a fusion unit. The i-th modality encoder performs unified semantic encoding on the i-th modality to generate the corresponding encoding result, where i ranges from 1 to n, and n is the total number of modalities included in the query request. The fusion unit fuses the encoding results from each modality encoder to obtain the corresponding encoding result, which satisfies the following condition: F = ∑ n i=1 β i ×E i F represents the fusion result, E i β is the encoding result of the i-th modal encoder. i Let be the fusion weights for the i-th mode, where ∑ n i=1 β i =1 and β i ≥0.
3. The multimodal database query engine based on unified semantic representation according to claim 2, characterized in that, The query engine is also associated with a dynamically updated modality fusion weight list. Each row of the modality fusion weight list includes the semantic structure feature vector of the corresponding query request and the fusion weight of each modality. Specifically, for a query request input by the user, its semantic structure feature vector is generated by the semantic encoding conversion module. The cosine similarity between this feature vector and each semantic structure feature vector in the modality fusion weight list is calculated. The modality fusion weights corresponding to the records with similarity greater than a set threshold are used as the fusion weights of each modality of the current query request.
4. The multimodal database query engine based on unified semantic representation according to claim 3, characterized in that, The modality fusion weight list is obtained through the following steps: S10, Initialize the intermediate fusion weight record set; S11, for the currently acquired multimodal query request, extract its semantic structure features and generate a feature vector through the query parsing and semantic encoding module, and match the corresponding intermediate fusion weight record table from the current intermediate fusion weight record set: If a matching record table exists, execute S12; If no matching record table exists, create an empty intermediate fusion weight record table containing semantic structure feature vectors for the query request, and execute S13; S12, Based on the initial fusion weights of each modal data corresponding to the intermediate fusion weight record table, fuse the modal data of the multimodal query request; execute S14; S13, using a preset initial weight allocation strategy, assign initial fusion weights to each modal data of the multimodal query request, and fuse the modal data; execute S14; S14, Execute the corresponding query based on the fusion result, and obtain the query execution data and query results; S15, Calculate the actual fusion weights and reward function values of each modality based on the query execution data and query results; The actual fusion weights of each modality of data are determined based on the semantic relevance coefficient, the execution plan awareness coefficient, and the resource efficiency coefficient. S16, store the actual fusion weights and reward function values into the corresponding intermediate fusion weight record table, and use an improved gradient descent algorithm to update the key hyperparameters in the execution plan perception coefficient and resource efficiency coefficient with the goal of maximizing the cumulative reward value. S17, perform convergence determination on the intermediate fusion weight record table corresponding to the multimodal query request. If the fluctuation range of the most recent m consecutive reward function values is less than or equal to the set fluctuation range, calculate the average value of the actual fusion weights corresponding to the m reward functions as the final fusion weight, and add the semantic structure feature vector of the query request, the final fusion weight and the convergence stability to the current modal fusion weight list, and execute S11. Otherwise, add the semantic structure feature vector, actual fusion weights, and convergence stability of the query request to the current modality fusion weight list, and execute S11.
5. The multimodal database query engine based on unified semantic representation according to claim 4, characterized in that, The actual fusion weight βc of the i-th mode i The following conditions must be met: βc i =(α i ×γ i ×δ i ) / ∑ n j=1 (α) j ×γ j ×δ j ), α i γ is the semantic relevance coefficient of the i-th modality, determined based on the cosine similarity between the query vector corresponding to the multimodal query request and the i-th modality vector in the vector library; i Let γ be the execution plan perception coefficient for the i-th mode. i =log(1+C) i ×P i ×I i ) / log(1+C b ×P b ×I b ), where C i C is the execution resource consumption coefficient for the i-th mode. i =1 / (w1×IO i +w2×CPU i +w3×Net i ), IO i The disk I / O resource consumption value for the i-th mode, CPU i Net represents the CPU resource consumption value for the i-th mode. i P represents the network transmission resource consumption value for the i-th mode, where w1, w2, and w3 are resource weight coefficients; i Let I be the execution efficiency coefficient for the i-th mode. i C is the index utilization efficiency coefficient for the i-th mode. b P is the reference execution resource consumption coefficient for the i-th mode. b Let I be the reference execution efficiency coefficient for the i-th mode. b The reference index for the i-th mode utilizes the efficiency coefficient; δ i Let δ be the resource efficiency coefficient for the i-th mode. i =exp(-λ×S i ), S i Let α be the resource consumption intensity of the i-th mode, λ be the resource sensitivity coefficient, and exp() be the natural exponential function; j Let γ be the semantic relevance coefficient of the j-th mode. j Let δ be the execution plan perception coefficient for the j-th mode. j Let be the resource efficiency coefficient for the j-th mode, where j ranges from 1 to n.
6. The multimodal database query engine based on unified semantic representation according to claim 4, characterized in that, The reward function satisfies the following conditions: R = k1 × Q + k2 × E + k3 × F, where Q is the result quality score, E is the global time consumption optimization coefficient, F is the global resource utilization efficiency coefficient, k1, k2 and k3 are reward weight coefficients, k1 + k2 + k3 = 1, where E = 1 / (1 + (ET - ET0) / ET0), where ET is the actual execution time of the current query and ET0 is the baseline execution time of similar queries; F = 1 / (1 + (RC - RC0) / RC0), where RC is the comprehensive resource consumption value of the current query and RC0 is the baseline resource consumption value of similar queries.
7. The multimodal database query engine based on unified semantic representation according to claim 4, characterized in that, During the execution of a query by the multimodal query execution module, the execution information of each execution path in the optimal query execution plan is monitored in real time, and the weight and resources of each execution path are adjusted based on the real-time monitoring results; the execution information includes actual execution time, execution efficiency, and query result quality score.
8. The multimodal database query engine based on unified semantic representation according to claim 7, characterized in that, The weights of the execution path are adjusted using an exponential smoothing algorithm.
9. The multimodal database query engine based on unified semantic representation according to claim 2, characterized in that, The unimodal semantic vector of a certain entity in the vector library satisfies the following condition: v u =v0 u +∑ t≠u (A) ut ×v0 t ), Among them, v u Let v0 be the optimization vector for the u-th mode of this entity. u v0 is the initial vector of the u-th mode of this entity, independently encoded based on the u-th mode encoder. t Let A be the initial vector for the t-th mode of this entity. ut Let u be the alignment weight between the u-th and t-th modes of the entity, where u and t take values from 1 to p, and t ≠ u, where p is the total number of modes contained in the source data of the entity.
10. A multimodal database query method based on unified semantic representation, characterized in that, The method is implemented based on the query engine according to any one of claims 1 to 9, and the method includes: The user-input query request is semantically parsed to obtain the corresponding semantic structure features; the query request includes unimodal queries and multimodal queries; the semantic structure features include query intent, key entities, and modal information; The semantic structure features are uniformly semantically encoded to generate corresponding encoding results. If the semantic structure features contain single-modal data, the corresponding single-modal semantic vector is directly generated as the encoding result. If the semantic structure features contain multi-modal data, the semantic vectors of each modality are generated and the generated semantic vectors are fused to generate a fused semantic vector as the encoding result. Based on the semantic features in the encoding results, the data distribution characteristics of the vector library, the query complexity, and the query engine resource status, an optimal query execution plan is generated. The optimal query execution plan includes a vector retrieval strategy, modality weight allocation, and a computing resource scheduling scheme. According to the optimal query execution plan, a query operation is performed on the vector library to obtain candidate results that are semantically similar to the query request. Perform a weighted fusion operation on the multi-path results to generate a weighted score for each candidate result; Output the candidate results in descending order of weighted scores.
Citation Information
Patent Citations
Cross-modal search system
CN113946726A
Simulation method and system for intelligent fusion and dynamic prediction of electronic warfare information
CN120012609A
Cross-modal retrieval method for semantic and vector fusion in data space
CN120386902A
Cited By
Storage method for consistency verification of automobile extended-guarantee multi-source data
CN121116966A
Intelligent retrieval method and system for CAD design features based on knowledge graph
CN121255870A
Picture storage method, picture retrieval method and related devices
CN121434426A