Metadata-driven intelligent document retrieval and generation system
Through the metadata-driven intelligent document retrieval system, metadata pre-filtering and hybrid index structure are used to optimize the retrieval path, which solves the problems of low efficiency and high result noise of multi-dimensional compound queries under the distributed index architecture, and realizes efficient and accurate document retrieval.
Patent Information
- Application Number
- CN202510594134.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-05-09
AI Technical Summary
In the existing technology, the distributed index architecture lacks a metadata pre-filtering mechanism in multi-dimensional composite query scenarios, resulting in low retrieval efficiency and excessively noisy results. It is unable to effectively combine structured metadata for pre-screening, increasing computing resource consumption and data transmission delays.
A metadata-driven intelligent document retrieval and generation system is adopted, including a metadata extraction module, a dynamic index construction module, an intelligent routing module and a distributed retrieval module. Through metadata pre-filtering, a hybrid index structure, an improved consistent hashing algorithm and a time series prediction model, the retrieval path is optimized and redundant calculations are reduced.
It significantly improves the response efficiency of multi-dimensional composite queries, reduces redundant computing overhead, improves the relevance and accuracy of retrieval results, and solves the problems of low retrieval efficiency and noise accumulation under traditional distributed architecture.
Smart Images

Figure CN120104624B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, and in particular to an intelligent document retrieval and generation system based on metadata driving. BACKGROUND
[0002] In the prior art, for retrieval systems of large-scale unstructured data, a distributed index architecture is usually adopted to achieve efficient storage and query. However, when the index data is divided into multiple independently stored blocks, the retrieval process needs to traverse all data blocks to complete feature matching, resulting in a significant increase in computing resource consumption, especially in the case of multi-dimensional composite query scenarios, the system response efficiency presents exponential decline. For example, in the retrieval application of enterprise information database, the existing method cannot effectively combine the structured metadata such as application agency and time range to pre-screen the target data blocks, resulting in the system performing redundant retrieval on data blocks irrelevant to the time interval and non-associated application subject, not only increasing the data transmission delay, but also introducing a large amount of noise results. This technical defect is due to the lack of metadata pre-filtering mechanism in the existing index structure, which cannot quickly locate the highest correlation data subset through logical constraint conditions at the initial stage of retrieval, resulting in that the retrieval accuracy and efficiency are difficult to meet the actual application requirements. SUMMARY
[0003] In view of the deficiencies in the prior art, the present application provides an intelligent document retrieval and generation system based on metadata driving, which solves the problems of low retrieval efficiency and high result noise caused by the lack of metadata pre-filtering mechanism in multi-dimensional composite query under distributed index architecture.
[0004] To solve the above technical problems, the specific technical solutions of the present application are as follows:
[0005] The present application provides an intelligent document retrieval and generation system based on metadata driving, comprising:
[0006] A metadata extraction module is used to extract key fields from an input document and map them into structured key-value pairs, and output the structured key-value pairs to a streaming pipeline;
[0007] A dynamic index construction module is used to receive the structured key-value pairs transmitted by the streaming pipeline, inject the received structured key-value pairs into a preset memory buffer, classify the structured key-value pairs in the preset memory buffer by field type to obtain classification results of discrete or continuous type, construct a skip list index for discrete fields, construct a red-black tree index for continuous fields, integrate the skip list index and the red-black tree index to generate hybrid index data, and input the hybrid index data into a preset logical sharding strategy unit, and the preset logical sharding strategy unit outputs sharding strategy parameters and hybrid index data;
[0008] The intelligent routing module is configured to receive the mixed index data output by the preset logical sharding strategy unit, calculate the hash value of the primary key field in the structured key-value pair generated by the metadata extraction module through a preset hash function, perform logical sharding on the mixed index data according to the hash value of the primary key field, obtain the sharded index data, generate a mapping relationship table of the sharded data and the storage nodes through a preset improved consistent hashing algorithm, and transmit the sharded index data to the corresponding storage nodes according to the mapping relationship table.
[0009] The distributed retrieval module is configured to receive the sharded index data in response to a query request, call a time series prediction model trained based on historical query logs, generate a query probability prediction value for the sharded index data through the time series prediction model, filter a target shard set in combination with the shard strategy parameters, and retrieve the associated document fragments of the target shard set from the corresponding storage nodes according to the mapping relationship table of the sharded data and the storage nodes.
[0010] The result generation module is configured to extract the entity relationship in the associated document fragments through a preset semantic parsing engine to generate a semantic graph structure, splice the node vector of the semantic graph structure and the metadata feature vector in the structured key-value pair, input the same into a multi-head attention model, and generate structured response data.
[0011] Further, the metadata-driven intelligent document retrieval and generation system provided by the application comprises a dynamic index construction module.
[0012] The data receiving unit is configured to receive the structured key-value pair from the metadata extraction module and transmit the metadata key-value pair of the new document to the difference packaging unit through a streaming pipeline.
[0013] The difference packaging unit is configured to package the metadata key-value pair of the new document into a difference descriptor comprising an operation type and a version identifier, and inject the difference descriptor into the skip list and the red-black tree mixed storage area of the preset memory buffer of the dynamic index construction module.
[0014] The update unit is configured to periodically read the difference descriptor of the memory buffer, the persistent storage form of the mixed index data is a B+ tree index, a gap filling model is used to identify the key-value interval gap of the B+ tree index, the incremental data in the difference descriptor is classified and inserted into the key-value interval gap according to the field type, and hierarchical merging is performed between the difference descriptor and the B+ tree index to generate an updated index structure.
[0015] Further, the metadata-driven intelligent document retrieval and generation system provided by the present application, the intelligent routing module comprises: a shard decision unit, configured to receive mixed index data output by the dynamic index construction module, construct a metadata knowledge graph based on the structured key-value pairs generated by the metadata extraction module, the metadata knowledge graph comprises entity nodes and relationship edges, input the metadata knowledge graph into a preset graph neural network for dynamic embedding to generate high-dimensional vector representation of metadata entities, calculate the semantic correlation degree of metadata combinations based on the cosine similarity between the high-dimensional vector representations, generate a query heat map based on the metadata combinations and their query frequencies recorded in the historical query log, combine the semantic correlation degree of the metadata combinations with the time window access frequency distribution in the query heat map, and the physical proximity relationship in the storage node topology data to generate shard boundary parameters including virtual node weight factors and shard capacity thresholds;
[0016] a routing execution unit, configured to adjust the mapping density of the virtual node ring according to the virtual node weight factors in the shard boundary parameters, map metadata primary keys to target physical nodes through a preset improved consistent hashing algorithm, and optimize the routing path based on the physical proximity relationship in the storage node topology data;
[0017] a replica pre-distribution unit, configured to receive a hotspot probability value output by a long short-term memory prediction model trained based on the historical query log, and trigger cache shard creation of the target node according to the correlation degree weight in the shard strategy parameters when the hotspot probability value exceeds a dynamically adjusted threshold, the cache shard comprising data blocks associated with the hotspot metadata combination.
[0018] Further, the metadata-driven intelligent document retrieval and generation system provided by the present application, the shard decision unit is further configured to:
[0019] input the metadata knowledge graph into a preset graph neural network for dynamic embedding to generate high-dimensional vectors representing the semantic correlation of metadata;
[0020] perform bottom-up hierarchical clustering analysis on the high-dimensional vectors, construct a distance matrix by calculating the Euclidean distance between clusters, and merge adjacent clusters according to a dynamically adjusted clustering radius threshold to identify metadata communities with semantic correlation; wherein the metadata communities with semantic correlation are defined by the entity relationship edges in the metadata knowledge graph;
[0021] The vector space distance distribution of the metadata colony is calculated by the high-dimensional vector calculation of the semantic association metadata colony generated by the hierarchical clustering analysis, the time window access frequency distribution in the real-time query heat map is generated based on the metadata combination and the metadata combination query frequency recorded in the historical query log, the physical proximity relationship data in the storage node topology data is collected, the vector space distance distribution of the metadata colony, the time window access frequency distribution in the real-time query heat map and the physical proximity relationship data in the storage node topology data are used to construct a sharding discriminant model based on a machine learning framework, and the sharding discriminant model receives the following input features:
[0022] The vector space distance distribution of the metadata colony;
[0023] The time window access frequency distribution in the real-time query heat map;
[0024] The physical proximity relationship data in the storage node topology data;
[0025] The input features are combined through a feature cross algorithm, and based on the training data labeled by historical sharding decision records, a preset importance sorting algorithm is used to screen key factors affecting the sharding decision;
[0026] Output includes sharding capacity threshold and correlation weight, wherein:
[0027] The sharding capacity threshold is determined by the vector space distance distribution of the clustering result and the physical proximity relationship in the storage node topology data;
[0028] The correlation weight is dynamically adjusted by the time window access frequency distribution in the real-time query heat map and the feature cross result.
[0029] Further, the metadata-driven intelligent document retrieval and generation system provided by the application, the distributed retrieval module is further used for:
[0030] In the index change operation, a vector label including a logical timestamp generated by a version controller inside the distributed retrieval module is attached;
[0031] The partial order relationship of the vector labels received by different nodes is determined by a vector clock algorithm, and the operation execution order is determined;
[0032] When updating the index copy asynchronously, the conflict operation is version-merged according to the execution order, and the merged index copy is subjected to consistency check by a version merging engine inside the distributed retrieval module, so as to maintain the causal consistency of the cross-node index.
[0033] Further, the metadata-driven intelligent document retrieval and generation system provided by the application further comprises a load balancing module;
[0034] The receiving node monitors the storage pressure, network bandwidth and CPU load index data collected by the agent, and inputs the data into a preset multi-dimensional weight evaluation model in the load balancing module for quantitative calculation;
[0035] The quantitative result is input into a reinforcement learning engine based on a Q-learning algorithm to dynamically generate a sharding weight matrix;
[0036] According to the weight distribution result of the sharding weight matrix, a topology optimal routing path is calculated by a shortest path algorithm, and a routing decision instruction is sent to a query dispatcher of the intelligent routing module to adjust the shunt ratio of the query request.
[0037] Further, the metadata-driven intelligent document retrieval and generation system provided by the application receives implicit association data in the query log through a reverse feedback channel, and writes the implicit association data into a graph database of the metadata knowledge graph;
[0038] The dynamic adjustment module in the load balancing module analyzes the updated association relationship of the knowledge graph, and recalculates the semantic similarity threshold of the metadata shard;
[0039] When it is detected that the shard load deviates from the balance threshold, an elastic migration pipeline is triggered to migrate the associated data block in parallel according to the topology structure in the storage node topology data, and the data change operation log of the source node and the target node is synchronized through a double-write controller during the migration process.
[0040] Further, the metadata-driven intelligent document retrieval and generation system provided by the application, the result generation module comprises:
[0041] The preprocessing unit receives the associated document fragments returned by the distributed retrieval module, extracts entity relationships through a semantic parsing engine, and outputs a semantic graph structure with annotations;
[0042] The text generation unit receives the semantic graph structure, fuses the metadata features in the structured key-value pairs generated by the metadata extraction module in a neural network based on a multi-head attention mechanism, and generates a preliminary response text;
[0043] The post-processing unit performs prefix aggregation processing on the preliminary response text, identifies the implicit theme distribution through a preset latent Dirichlet distribution model, merges the response paragraphs of the same theme, and outputs an optimized structured response document.
[0044] Further, the metadata-driven intelligent document retrieval and generation system provided by the application, the preprocessing unit is further used for:
[0045] The preset Node2Vec algorithm is used for graph embedding calculation on entities in the metadata knowledge graph to generate high-dimensional vector representations.
[0046] The high-dimensional vector representations are input into a similarity calculator to construct a semantic adjacency matrix of the document fragments by cosine similarity measurement.
[0047] A sliding window processor traverses the semantic adjacency matrix, extracts periodic correlation patterns in the time dimension by a preset Fourier transform, and outputs a feature vector set with time sequence labels.
[0048] Further, the metadata-driven intelligent document retrieval and generation system also comprises an abnormality recovery mechanism, which comprises:
[0049] A checkpoint generation unit periodically captures the index structure state of the dynamic index construction module, generates a snapshot file comprising B+ tree node distribution and stores it in a prewrite log.
[0050] A difference synchronization controller calls a Merkle tree comparison algorithm during data migration, compares the vector space hash value differences between the source node and the target node, generates a difference data set and synchronizes it to the target node.
[0051] A lease coordinator compares the logical time stamps of the snapshot files of each node through the version vector after the network partition is restored, triggers the incremental synchronizer to execute the playback of the missing operation log, and synchronizes the index state of each node through the anti-entropy protocol.
[0052] The present application has the following advantages:
[0053] The present application filters out non-associated data shards in the index construction stage through the synergistic effect of the metadata pre-filtering mechanism and the hybrid index structure, reduces the number of shards that need to be traversed in the distributed retrieval stage, and significantly reduces the redundant calculation overhead; the intelligent routing module stores the high correlation degree shards to the physically adjacent nodes based on the semantic clustering results and the improved consistent hashing algorithm, combines the LSTM model to predict the hot query path, preferentially accesses the cache shards of the local or adjacent nodes, and suppresses the network delay caused by cross-node full scanning; the distributed retrieval module maintains the time sequence correctness of the index changes through the causal consistency protocol, dynamically integrates the metadata features and document semantics through the multi-head attention mechanism of the generation model, filters the expired data and non-associated content; the post-processing unit optimizes the response structure by using the topic model, combines the repeated semantic paragraphs and inserts standardized identifiers, and forms accurate technical document output. The above technical scheme realizes three-level synergy of metadata-driven pre-screening, path optimization and semantic reconstruction, effectively improves the response efficiency and result relevance of multi-dimensional complex queries, and solves the technical defects of low retrieval efficiency and noise accumulation under the traditional distributed architecture. BRIEF DESCRIPTION OF DRAWINGS
[0054] In order to more clearly illustrate the technical solutions of the present application, the drawings required for use in the embodiments will be briefly introduced as follows. Obviously, for those skilled in the art, other drawings can also be obtained from the drawings without creative labor.
[0055] Figure 1 The system architecture diagram of the metadata-driven intelligent document retrieval and generation system provided by the embodiments of the present application is shown. DETAILED DESCRIPTION
[0056] In order to make the purpose, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be described clearly and completely in combination with the specific embodiments of the present application and corresponding drawings. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application. The technical solutions provided by the embodiments of the present application are described in detail below in combination with the drawings. In order to better understand the purpose of the present application, the present application will be further described in detail.
[0057] Please refer to Figure 1 The present application provides a metadata-driven intelligent document retrieval and generation system, which comprises:
[0058] A metadata extraction module is used to extract key fields from an input document and map them into structured key-value pairs, and output the structured key-value pairs to a streaming pipeline.
[0059] The metadata extraction module parses the text content of the input document through a natural language processing engine, and uses a pre-trained named entity recognition model to identify the organization name, timestamp and technology classification label. During the identification process, a regular expression verification layer is called to check the format of the extracted fields, such as converting the time information in “2023 X Company A Technology Tender Book” to the ISO 8601 standard format “2023-01-01”, and filtering abnormal data that does not conform to the preset pattern. The fields that pass the verification are mapped into key-value pairs after type annotation, with the key corresponding to the metadata type identifier (such as “organization”, “time”, “technology classification”), and the value converted into a standardized expression through semantic normalization processing, such as mapping “ABC” to a unified technology classification code.
[0060] The structured key-value pair is appended with a unique document identifier and timestamp meta-information, and pushed to the streaming pipeline through an asynchronous message queue. The message queue divides data partitions according to the primary key hash value, ensuring that the same primary key metadata maintains temporal consistency during transmission. The streaming pipeline sets a back pressure control mechanism to dynamically adjust the data transmission rate to match the processing capacity of the downstream module, preventing memory buffer overflow. During transmission, the key-value pair is serialized and compressed to reduce network bandwidth occupation, and a cyclic redundancy check code field is appended to ensure data transmission integrity.
[0061] The key-value pair data output to the streaming pipeline carries a uniform format context identifier, which includes the data source node address and processing stage marker, forming an end-to-end data traceability link. The downstream module traces the metadata flow state in the system through the identifier, supporting data rollback and debugging analysis in abnormal scenarios. Failed data packets trigger an exception handling process, recording error types and original text segments to a log database for subsequent model optimization and rule iteration.
[0062] The dynamic index construction module is configured to receive the structured key-value pairs transmitted by the streaming pipeline, inject the received structured key-value pairs into a preset memory buffer, classify the structured key-value pairs in the preset memory buffer by field type to obtain discrete or continuous classification results, construct a skip list index for discrete fields and a red-black tree index for continuous fields, integrate the skip list index and the red-black tree index to generate hybrid index data, and input the hybrid index data into a preset logical sharding strategy unit, which outputs sharding strategy parameters and hybrid index data.
[0063] After receiving the structured key-value pairs transmitted by the streaming pipeline, the dynamic index construction module injects the data into a preset memory buffer through a memory allocator. The buffer is divided into a skip list partition and a red-black tree partition, and the key-value pairs are classified according to the field type classification rule: discrete fields (such as institution name, technical classification code) are allocated to the skip list partition, and continuous fields (such as timestamp, numerical sequence) are allocated to the red-black tree partition. During the classification process, a field type parser is called to match the field type identifier based on the metadata pattern definition file, generate a classification result tag and append it to the head of the key-value pair.
[0064] The skip list index builder implements a multi-level index construction on discrete fields, with 4 levels of index hierarchy, and node storage of key value hash pointers and version number information. The skip list insertion operation uses a probabilistic promotion algorithm to dynamically adjust the node hierarchy, achieving an O(log n) time complexity for accurate matching queries. The red-black tree index builder maintains global order for continuous fields, and through color marking and rotation operations, the tree structure is balanced when inserting new keys. Leaf nodes record timestamp interval boundary values, supporting efficient execution of range queries such as "2020-2023". After generating hybrid index data, the index merger integrates the pointer topology relationship of skip list and red-black tree into a unified index view, writes it to the distributed file system and adds version number identification.
[0065] After loading the hybrid index data, the logical sharding strategy unit calls the semantic clustering analysis engine to calculate the sharding capacity threshold based on the co-occurrence frequency (such as the concentration of a certain institution in a specific technical field) and the correlation strength between fields. In the clustering process, the K-means algorithm is used to divide the index key value space, and the high-density area is identified as the core area of the shard. The shard strategy parameter generator outputs a set of instructions including shard identifier, capacity threshold, and correlation weight factor to drive the downstream routing module to implement physical sharding mapping. The version manager records index version change logs, supports incremental updates and historical state rollback, forming a closed-loop index management mechanism.
[0066] The intelligent routing module is configured to receive the hybrid index data output by the preset logical sharding strategy unit, calculate the hash value of the primary key field in the structured key-value pair generated by the metadata extraction module based on a preset hash function, perform logical sharding on the hybrid index data based on the hash value of the primary key field, obtain the sharded index data, generate a mapping relationship table between the sharded data and the storage nodes by using a preset improved consistent hashing algorithm, and transmit the sharded index data to the corresponding storage nodes based on the mapping relationship table.
[0067] After receiving the hybrid index data output by the logical sharding strategy unit, the intelligent routing module extracts the primary key field (such as the combination of institution name and technology classification code) in the structured key-value pair using the primary key parser, and generates the primary key hash value by calling a double hash function. The hash function uses the MurmurHash algorithm combined with modulo operation to generate a 32-bit hash code, which is uniformly distributed in the virtual node space through the hash ring mapping mechanism, reducing the hash skew phenomenon. The shard cutter divides the hybrid index data into logical shards based on the hash value range, and dynamically determines the capacity threshold of each shard boundary based on the shard strategy parameters, generating a shard metadata table including the shard identifier and version number.
[0068] The improved consistent hashing algorithm introduces a virtual node dynamic weight factor, which is dynamically adjusted based on the real-time load status of the storage node (such as CPU utilization, disk remaining space). The virtual node manager allocates fewer virtual nodes to high-load nodes and more virtual nodes to low-load nodes according to the weight factor, thereby balancing data distribution. The hash ring reconfigurator periodically updates the mapping relationship between virtual nodes and physical nodes, generates a mapping relationship table including node address, shard identifier and version number, and stores it in the distributed configuration center.
[0069] When the shard data is transmitted, the routing executor selects the target physical node according to the mapping relationship table, and preferentially selects the topologically adjacent node to reduce network delay. The data encapsulator packages the shard index data and the shard metadata table into a transmission unit, adds a cyclic redundancy check code and an operation log sequence number, and writes it into the distributed message queue through zero-copy technology. The replica pre-distribution listener monitors the access frequency of the hot shard, and when the frequency exceeds the preset threshold, triggers the asynchronous preloading of the cached shard. After the target node receives the shard data, it establishes an in-memory cache index and updates the local routing table, synchronizes the incremental change log of the source node through the double-write controller, and maintains data consistency.
[0070] The distributed retrieval module is used to receive the indexed data after sharding in response to a query request, call a time series prediction model trained based on historical query logs, generate query probability prediction values for the indexed data after sharding through the time series prediction model, and filter a target shard set in combination with the shard strategy parameters. According to the mapping relationship table between the shard data and the storage node, the associated document fragments of the target shard set are retrieved from the corresponding storage node;
[0071] When the distributed retrieval module responds to a query request, the query parser extracts the metadata constraint conditions (such as institution name, time range and technical classification label) in the request, and generates a query feature vector. The time series prediction model loads the time series data in the historical query logs, inputs the feature vector into the trained LSTM neural network, and outputs the query probability prediction values of each shard. The probability value reflects the potential frequency of accessing the target shard in the future time window. The shard filter combines the capacity threshold and the associated weight factor in the shard strategy parameters to sort and filter the shards below the dynamically adjusted threshold, and generates a target shard identifier set.
[0072] The mapping relationship resolver parses the physical node address list corresponding to the target shard according to the mapping relationship table of the sharded data and the storage node. The retrieval coordinator sends an index scan request to the target node cluster in parallel, and the request package encapsulates the shard identifier, version number and query condition. The node locally executes a multi-level cache strategy, preferentially accesses hot shard data in memory, quickly excludes non-matching key values through a Bloom filter, and reduces disk I / O operations. For cross-node shard retrieval, the version comparator calls a causal consistency protocol, compares the logical timestamp sequences of the results returned by different nodes, merges conflicting document fragments according to the operation execution order, and generates a globally consistent intermediate result set.
[0073] The result aggregator performs deduplication and sorting operations on the intermediate result set, calculates a comprehensive relevance score based on the metadata characteristics (such as timestamp proximity and technical classification similarity) of the associated document fragments, and selects document fragments with a score higher than a preset threshold. The retrieval result is returned to the result generation module through zero-copy transmission technology, and the query log recorder is triggered to update the historical access record, forming a feedback data closed loop to optimize the training accuracy of the subsequent prediction model.
[0074] The result generation module is configured to extract entity relationships in the associated document fragments through a preset semantic parsing engine to generate a semantic graph structure, concatenate node vectors of the semantic graph structure with metadata feature vectors in the structured key-value pair, input a multi-head attention model, and generate structured response data.
[0075] After the result generation module receives the associated document fragments returned by the distributed retrieval module, the semantic parsing engine calls a dependency syntax analyzer to extract entity relationship triples in the fragments, such as parsing the entities "X Company", "ABC" and "500,000" and the relationship "tender amount" from "X Company tendered 500,000 for ABC technology in 2023". The parsed triples are aligned with the structured key-value pairs generated by the metadata extraction module, such as mapping "ABC" to a standard technical classification code, to generate a semantic graph structure with a timestamp and a source identifier.
[0076] The cross-modal fusion layer concatenates the semantic graph node vector and the metadata feature vector, and inputs an encoder-decoder model based on the multi-head attention mechanism. The encoder captures the internal context association of the document fragments through the multi-head self-attention mechanism, and the decoder introduces the metadata features as key-value pairs in the generation stage, dynamically adjusts the influence weight of the technical classification label and the time dimension feature on the generation result through the cross-attention mechanism. For example, when the timestamp feature weight is increased, the generation text preferentially reflects the time series trend analysis content.
[0077] The generation model uses a beam search algorithm to traverse the candidate response sequence, suppresses semantic repetitive fragments through an n-gram language model, and avoids missing key entities by combining a coverage penalty mechanism. The initially generated response text carries paragraph markers and entity reference tags, such as " <section1>Technical indicators <entity ref="ABC"> ...< / entity> ".
[0078] The structured response document is written back to the storage node through the version controller, updating the document association relationship edge weight in the metadata knowledge graph. The output interface converts the document into PDF or XML format conforming to the ISO standard, and adds metadata traceability information (such as "Data source: Node A - Version V2") for users to verify the relevance of the results. The anomaly detector monitors semantic conflicts or format errors during generation, triggering a rollback mechanism to re-call the retrieval module to obtain supplementary data segments and iteratively optimize the generation results.
[0079] The metadata extraction module parses the unstructured text content from the input document, uses named entity recognition algorithms to identify institution names, timestamps, and technology classification tags, and verifies the format validity of the extracted fields through regular expressions to filter noise data that does not conform to the preset pattern. The extracted fields are mapped to key-value pairs after type annotation, where the key corresponds to the metadata type identifier, and the value is converted to a standard format through semantic normalization processing, such as converting time information to a preset standard string. The key-value pair is attached with a document unique identifier and timestamp meta information, and is pushed to the streaming pipeline through an asynchronous message queue. The queue divides data partitions according to the metadata primary key hash value, ensuring the order of data with the same primary key during transmission.
[0080] After receiving the key-value pair data from the streaming pipeline, the dynamic index construction module injects the data into the memory buffer according to the field type classification rules. Discrete fields are stored in skip list index partitions, and the skip list accelerates the exact matching of field names through multiple index layers. The index node stores the hash pointer and version number of the key-value pair, supporting fast insertion and query in high-concurrency scenarios. Continuous fields are written to red-black tree index partitions, and the red-black tree maintains the global order of field values, optimizing the performance of range queries through self-balancing rotation operations. The leaf node records the start and end boundaries of the timestamp interval. After generating the mixed index data, the logical sharding strategy unit performs semantic clustering analysis on the index structure, calculates the sharding capacity threshold based on the co-occurrence frequency and association strength between fields, and generates sharding strategy instructions including shard identifiers and topology constraint parameters.
[0081] The intelligent routing module parses the sharding strategy instruction, uses a double hash function to calculate the hash of the metadata primary key, and generates a virtual node identifier. The improved consistent hashing algorithm introduces a virtual node dynamic weight factor, which adjusts the distribution density of the virtual node on the hash ring according to the real-time load state of the storage node, so that high correlation shards are mapped to physically adjacent nodes. The routing execution unit encapsulates the indexed data after sharding into data blocks, adds the shard version number and operation log sequence number, and transmits them to the target storage node through the distributed message middleware. The replica pre-distribution unit monitors the hot access pattern in the historical query log, calls the LSTM prediction model to generate the access probability distribution in the future time window, and triggers the cache replica preloading of the hot shard when the probability exceeds the dynamic threshold, reducing the cross-node retrieval delay.
[0082] The distributed retrieval module responds to the query request, parses the metadata constraint item in the query condition, and calls the timing prediction model to calculate the access probability value of the target shard. The probability value generates a candidate shard set in combination with the shard strategy parameters, and the retrieval coordinator sends index scan requests to the target nodes in parallel according to the mapping relationship table of shards and storage nodes. The node locally executes a multi-level cache strategy, preferentially accesses hot shard data in memory, filters invalid key values through a Bloom filter, and reduces disk I / O overhead. The causal consistency protocol is activated during cross-node retrieval, and the vector clock algorithm compares the logical timestamp sequence of the returned results from different nodes, merges conflicting versions according to the operation execution order, and generates a globally consistent query result set.
[0083] The result generation module performs semantic analysis on the document fragments returned by retrieval, extracts entity relationship triples in the fragments through dependency syntax analysis engine, and constructs a weighted semantic graph structure. In the encoding stage, the generation model concatenates the semantic graph node vector and the metadata feature vector, calculates the correlation weight between cross-modal features through the multi-head attention mechanism, and dynamically fuses the technical classification label and time dimension information. In the decoding stage, the beam search algorithm is used to generate candidate response sequences, and the n-gram language model is used to suppress semantic repeated fragments, and the initial structured text is output. The post-processing unit performs topic clustering analysis on the text, identifies the implicit topic distribution through the latent Dirichlet distribution model, merges the response paragraphs of the same topic, inserts chapter identifiers and cross-reference links, and generates structured response output conforming to the technical document specification.
[0084] The load balancing module collects resource utilization indicators of the storage nodes in real time, a multi-dimensional evaluation model quantifies the storage pressure, network bandwidth and CPU load of the nodes to generate a node health score matrix. A reinforcement learning engine dynamically optimizes the weight distribution strategy based on the Q-learning algorithm, a shortest path algorithm calculates the optimal routing path according to the weight matrix, and the proportion of the query request is adjusted. The elastic migration pipeline triggers the data migration task when detecting that the load of the shard deviates from the balance threshold, the dual-write controller synchronizes the data changes of the source node and the target node, and the read-write consistency in the migration process is ensured through the operation log sequence number.
[0085] The abnormality recovery mechanism maintains system reliability through checkpoint snapshots and Merkle tree comparison. The checkpoint generation unit periodically persists index structure snapshots, and the snapshot file records the key value distribution and version number information of the B+ tree node. The differential synchronization controller constructs a Merkle tree hash structure during data migration, locates the data differences between the source node and the target node, and generates an incremental patch package. After the network partition is restored, the lease coordinator compares the snapshot logical time stamps of each node through the version vector, triggers the incremental synchronization operation to make the distributed index state converge.
[0086] Specifically, the metadata-driven intelligent document retrieval and generation system provided by the application comprises a dynamic index construction module.
[0087] A data receiving unit is configured to receive the structured key-value pairs from the metadata extraction module, and transmit the metadata key-value pairs of the new document to a differential packaging unit through a streaming pipeline.
[0088] A differential packaging unit is configured to package the metadata key-value pairs of the new document into a differential descriptor comprising an operation type and a version identifier, and inject the differential descriptor into a skip list and a red-black tree mixed storage area of a preset memory buffer of the dynamic index construction module.
[0089] An updating unit is configured to periodically read the differential descriptor of the memory buffer, persist the mixed index data in the form of a B+ tree index, identify the key value interval gap of the B+ tree index by using a gap filling model, insert the incremental data in the differential descriptor into the key value interval gap according to the field type, and perform hierarchical merging with the B+ tree index to generate an updated index structure.
[0090] The data receiving unit receives the structured key-value pairs pushed by the metadata extraction module through the streaming pipeline, and uses an asynchronous transmission protocol to achieve high-throughput transmission of data. The key-value pairs are attached with a unique identifier of the document and a type marker of the operation, and the data partition is divided according to the hash value of the primary key to ensure that the data of the same primary key maintains temporal consistency during transmission. The streaming pipeline sets a data verification layer to verify the data integrity through a cyclic redundancy check code, filter abnormal data packets that fail the verification, and prevent invalid data from entering the index construction process. A back pressure control mechanism is used during transmission to dynamically adjust the data pushing rate and prevent memory buffer overflow.
[0091] The difference packaging unit analyzes the received key-value pair data, identifies the operation types including insertion, update, and deletion, and generates a difference descriptor including the operation code, field identifier, and version sequence number. The descriptor is compressed and stored in a binary encoding format, with a timestamp and data source identifier attached to the header and a checksum field attached to the tail. The difference data of discrete fields is packaged as a skip list node descriptor, recording the hierarchical pointer and key value hash value; the difference data of continuous fields is packaged as a red-black tree node descriptor, recording the key value interval boundary and balance factor. The packaged descriptor is injected into the mixed storage area of the memory buffer, and the skip list partition uses multi-level indexing to accelerate the fast positioning of discrete fields, and the red-black tree partition maintains the order of continuous fields through rotation balance.
[0092] The update unit starts a background merging thread to periodically scan the difference descriptors in the memory buffer, and uses a gap filling model to identify the key value distribution gaps of the persistent B+ tree index. The difference data of discrete fields is matched with the idle slot of the B+ tree leaf node according to the hash value, and the insertion triggers node split detection, and when the number of node key values exceeds the preset threshold, the split operation is performed to generate a new child node and update the parent node pointer. The difference data of continuous fields is compared with the existing key value range of the B+ tree through an interval merging algorithm, and non-overlapping intervals are inserted into the corresponding leaf node, and overlapping intervals are subjected to version conflict detection, the latest version data is retained and the operation log is recorded. During the merging process, a lock-free snapshot technology is used to create a read-only view of the index structure for query service access, enabling parallel execution of index update and query service.
[0093] After the hierarchical merging is completed, the update unit generates a new version of the B+ tree index and stores it persistently in the distributed file system. The version manager records the version number and change log of the index structure, and supports index recovery in abnormal states through snapshot rollback mechanism. The merging result triggers incremental update of the metadata knowledge graph, and the graph database synchronously loads the new field association relationship to provide data basis for subsequent semantic sharding decision. The index state monitor collects the distribution density and query hit rate of the B+ tree node in real time, dynamically adjusts the parameter threshold of the gap filling model, and optimizes the data processing efficiency of the next merging period.
[0094] Specifically, the metadata-driven intelligent document retrieval and generation system includes a shard decision unit that receives mixed index data output by a dynamic index construction module, constructs a metadata knowledge graph based on structured key-value pairs generated by a metadata extraction module, and inputs the metadata knowledge graph into a preset graph neural network for dynamic embedding to generate high-dimensional vector representations of metadata entities. The metadata knowledge graph includes entity nodes and relationship edges. The shard decision unit calculates the semantic correlation degree of metadata combinations based on the cosine similarity between the high-dimensional vector representations, generates a query heat map based on metadata combinations and their query frequencies recorded in historical query logs, and generates shard boundary parameters including virtual node weight factors and shard capacity thresholds based on the semantic correlation degree of metadata combinations, time window access frequency distribution in the query heat map, and physical proximity relationships in storage node topology data.
[0095] A routing execution unit adjusts the mapping density of a virtual node ring based on the virtual node weight factors in the shard boundary parameters, maps metadata primary keys to target physical nodes through a preset improved consistent hashing algorithm, and optimizes routing paths based on physical proximity relationships in storage node topology data.
[0096] A replica pre-distribution unit receives hotspot probability values output by a long short-term memory prediction model trained based on historical query logs, and triggers cache shard creation of target nodes based on correlation degree weights in shard strategy parameters when the hotspot probability values exceed a dynamically adjusted threshold. The cache shard includes data blocks associated with hotspot metadata combinations.
[0097] When the shard decision unit constructs the metadata knowledge graph, it parses the entity types and relationship attributes in the structured key-value pairs, maps the organization names to entity nodes, uses technical classification tags and timestamps as node attributes, and establishes co-occurrence relationship edges between entities. The graph neural network uses a graph convolution network layer to aggregate the features of multiple-hop neighbor nodes, generates node sequence samples through a biased random walk strategy, and uses a Skip-gram model to train and generate high-dimensional vector representations. The vector representation stores the semantic correlation strength between entities, for example, organizations in similar technical fields have a small Euclidean distance in the vector space. The semantic correlation degree calculation module inputs the vector into a cosine similarity calculator to generate a correlation matrix of metadata combinations, and the matrix dimension corresponds to the interaction strength between different metadata entities.
[0098] The query heat map construction module analyzes historical query logs, and the access frequency of metadata combinations is counted according to a time window, and a sliding window mechanism is used to capture periodic access patterns. The time window access frequency and the semantic correlation degree matrix are weighted and fused to generate a comprehensive correlation degree score. The storage node topology analyzer collects the physical connection delay and bandwidth data between nodes to construct a node proximity graph. The shard boundary parameter generator fuses the comprehensive correlation degree score and the physical proximity relationship, calculates the virtual node weight factor through a linear weighting model, dynamically adjusts the shard capacity threshold combined with the disk capacity data of the storage node, and forms a shard strategy instruction set including topological constraints.
[0099] The routing execution unit reconstructs the topology structure of the virtual node ring after parsing the shard strategy instruction. The improved consistent hashing algorithm allocates more virtual nodes to the shard with a high weight factor, and increases the distribution density of the shard on the hash ring. The metadata primary key is subjected to double hashing calculation to generate a positioning identifier, and after being mapped to the virtual node ring, the target physical node with the lowest delay is selected according to the physical proximity graph. The routing path optimizer calculates the shortest transmission path between nodes by using the Dijkstra algorithm, and generates an optimal routing path table by using a path cost function to comprehensively consider the bandwidth utilization rate and the transmission hop number. The routing instruction is packaged into a data packet including a shard identifier and a version number, and is distributed to the target node through a low-delay communication protocol.
[0100] The replica pre-distribution unit loads a long short-term memory prediction model, the model input layer receives the metadata combination sequence in the historical query log, and generates training samples through time window slicing. The gating cycle unit of the hidden layer captures the access pattern dependency relationship across time steps, and outputs the hotspot probability value of each metadata combination in the future time window. The dynamic threshold adjuster calculates the adaptive change curve of the probability threshold according to the node load state and the cache capacity. When the hotspot probability value exceeds the threshold, the replica creation instruction triggers the cache shard preloading of the target node, the data block selector extracts the data block with the highest correlation weight from the persistent storage, establishes a memory-level cache index and updates the routing meta-information table. The cache shard keeps data synchronization with the source shard through a double-write mechanism, and the operation log serial number ensures the consistency of incremental update.
[0101] Specifically, the metadata-driven intelligent document retrieval and generation system disclosed by the application is also used for:
[0102] inputting the metadata knowledge graph into a preset graph neural network for dynamic embedding to generate a high-dimensional vector representing the semantic correlation of the metadata;
[0103] performing a bottom-up hierarchical clustering analysis on the high-dimensional vectors, constructing a distance matrix by calculating the Euclidean distance between clusters, and merging adjacent clusters according to a dynamically adjusted clustering radius threshold to identify a semantic association metadata community; wherein the semantic association metadata community is defined by entity relationship edges in the metadata knowledge graph;
[0104] The high-dimensional vector generated by the hierarchical clustering analysis generates a vector space distance distribution of the metadata community, generates a time window access frequency distribution in a real-time query heat map based on the metadata combination recorded in the historical query log and the metadata combination query frequency, collects physical proximity relationship data in the storage node topology data, and constructs a machine learning framework-based shard discrimination model based on the vector space distance distribution of the metadata community, the time window access frequency distribution in the real-time query heat map, and the physical proximity relationship data in the storage node topology data. The shard discrimination model receives the following input features:
[0105] The vector space distance distribution of the metadata community;
[0106] The time window access frequency distribution in the real-time query heat map;
[0107] The physical proximity relationship data in the storage node topology data;
[0108] The input features are combined through a feature cross algorithm, and based on training data labeled by historical shard decision records, a preset importance sorting algorithm is used to screen key factors affecting shard decision;
[0109] Output shard boundary decision parameters including shard capacity threshold and correlation degree weight, wherein:
[0110] The shard capacity threshold is determined by the vector space distance distribution of the clustering result and the physical proximity relationship in the storage node topology data;
[0111] The correlation degree weight is dynamically adjusted by the time window access frequency distribution in the real-time query heat map and the feature cross result.
[0112] When the metadata knowledge graph is input into the graph neural network, the graph convolution network layer is used to aggregate the multi-hop neighbor features of the entity nodes, and the node representation is updated through the message passing mechanism. The entity relationship edge attribute participates in feature propagation as a weight parameter, generating a high-dimensional vector including high-order semantic association. A negative sampling strategy is implemented in the dynamic embedding process to randomly mask some node relationships to enhance the generalization ability of the model, so that entities with similar semantics show an aggregated distribution in the vector space. The high-dimensional vector is stored in a distributed vector database, with a timestamp label added to support temporal correlation analysis.
[0113] The hierarchical clustering analysis module performs a bottom-up aggregation operation on the high-dimensional vectors, and in the initial stage, each vector is taken as an independent cluster, and a distance matrix between clusters is constructed based on the Euclidean distance. The dynamic radius adjuster adjusts the radius of the cluster corresponding to the high-frequency access metadata according to the access frequency distribution of the real-time query heat map, so as to improve the clustering density of the hotspot data area. In the iterative merging process, the merging operation is triggered when the distance value of the adjacent clusters is less than the current threshold value, and after merging, the cluster center vector is updated to be the weighted average value of the member vectors, and the metadata community division result reflecting the semantic correlation strength is generated. The cluster identifier and the correlation strength coefficient are written into the knowledge graph, and the community attribute label of the entity node is updated.
[0114] When the shard discrimination model is constructed, the feature engineering module performs standardization processing on the input features. The vector space distance distribution feature is converted into a probability distribution curve through kernel density estimation, the time window access frequency distribution feature eliminates short-term fluctuation noise by using an exponential smoothing method, and the physical proximity relationship data is converted into a delay weight matrix between nodes. The feature cross algorithm adopts a polynomial cross strategy, generates an interaction item feature of vector distance and access frequency, and enhances the capturing ability of the model to complex correlation patterns. After data cleaning, the historical shard decision record is taken as a training set, the feature importance sorter calculates the contribution degree of each feature based on a gradient boosting decision tree model, and a key factor combination affecting the shard capacity threshold is screened out.
[0115] In the shard boundary decision parameter generation stage, the capacity threshold calculator combines the compactness index of the vector distribution in the cluster and the proportion of the remaining disk space of the storage node to dynamically adjust the maximum data carrying capacity of the shard. The correlation degree weight calculation module fuses the trend slope of the time window access frequency and the feature cross result, and generates a weight coefficient by using a sliding window weighting algorithm. The parameter optimizer triggers an online learning mechanism according to a node topology change event, feeds the deviation data of the actual routing path and the predicted path back to the shard discrimination model, and iteratively updates the weight distribution strategy. The finally generated decision parameter set is converted into a shard instruction by the strategy engine, and drives the routing module to adjust the mapping rule of the virtual node ring.
[0116] The parameter output module encapsulates the shard capacity threshold and the correlation degree weight into a strategy configuration file, and adds a version number and an effective timestamp. When the configuration file is distributed to the storage node cluster, a two-phase commit protocol is used to ensure the atomicity of the configuration update of all nodes. The version rollback manager stores the historical parameter version persistently, and automatically triggers the configuration rollback operation when detecting that the shard load is abnormal, so as to maintain the stability of the routing decision. After the strategy takes effect, the monitoring agent collects the shard query hit rate and the node load index in real time, generates a performance report, and feeds back to the shard discrimination model, to form a closed-loop optimization link.
[0117] Specifically, the metadata-driven intelligent document retrieval and generation system provided by the application, the distributed retrieval module is further used for:
[0118] In the index change operation, a vector tag including a logical timestamp generated by a version controller inside the distributed retrieval module is attached;
[0119] The operation execution order is determined by comparing the partial order relationship of the vector tags received by different nodes through the vector clock algorithm;
[0120] In the asynchronous update of the index copy, the conflicting operations are version-merged according to the execution order, and the merged index copy is subjected to consistency checking by a version-merge engine inside the distributed retrieval module, so as to maintain the causal consistency of the cross-node index.
[0121] When generating a logical timestamp, the version controller of the distributed retrieval module constructs a vector tag based on the incremental sequence of the node identifier and the local event counter. The vector tag adopts a composite structure of [node ID: event counter], records the generation node and operation serial number of the index change operation, and is attached to the data change request header through the operation log. The data packet encapsulates the operation type, field identifier and key value data, and is attached with a globally unique request identifier during transmission, forming the basic data for operation tracing.
[0122] The vector clock algorithm is activated when a node receives an index change request, parses the vector tag carried by the operation and updates the local vector clock state. The local clock maintains the latest value of each node event counter, and constructs a partial order graph by comparing the value relationship of each node counter in the vector tag. When detecting concurrent operation conflicts, the conflict detector marks the operation pairs with causal dependency, and generates an operation execution sequence according to the preset node priority strategy. For concurrent operations without direct causal relationship, the operation type priority strategy is adopted, in which the deletion operation is executed before the update operation, and the update operation is executed before the insertion operation. After processing, the node broadcasts the updated vector clock state to the related copy nodes to synchronize the latest operation sequence information.
[0123] The version-merge engine implements multi-version index management during asynchronous update, creates branch versions for operations with conflicts and attaches version identifiers. The merging process adopts an operation playback mechanism, and according to the execution order determined by the vector clock algorithm, re-applies the change operation to the index copy according to the logical timeline. For multiple update operations involving the same field, the field-level version fuser retains the last valid operation result and generates a field version history chain. After merging, the consistency checker compares the vector clock states of the index copies of different nodes, synchronizes the missing change operation logs through the anti-entropy protocol, and eliminates the state differences between nodes. The checking result triggers the index snapshot generation task, and the merged index structure is persisted and stored, and the version association relationship in the metadata knowledge graph is updated.
[0124] The conflict resolution strategy refines the logic for specific scenarios in the merging process. For field-level numerical conflicts, the latest write priority principle is adopted, and the latest valid value is determined based on the logical timestamp in the vector label. For structural conflicts, such as key value interval overlap caused by index node splitting, the interval merger implements a boundary calibration algorithm to redivide the key value range and adjust the parent node pointer. The merged index structure is verified by the balance factor detector to maintain the balance of the red-black tree and B+ tree, triggering self-balancing operations to maintain query performance. The operation log compressor periodically cleans up synchronized historical log entries, freeing up storage space and improving log playback efficiency.
[0125] The consistency maintenance mechanism runs continuously during distributed retrieval. When processing query requests, the version resolver selects index replicas that meet causal consistency based on vector clock state, excluding branch versions with incomplete merging operations. The routing coordinator prioritizes accessing the replica node with the latest logical timestamp during cross-node queries, filtering outdated data responses through version number comparison. The exception recovery module activates offline operation caching in network partition scenarios, and re-injects the cached operation logs into the merging process after partition recovery, ensuring the integrity of operations during interruptions. The monitoring agent collects real-time index merging success rate and conflict resolution time consumption indicators, which are fed back to the shard discrimination model to optimize shard strategy parameters, forming a closed-loop performance optimization link.
[0126] Specifically, the metadata-driven intelligent document retrieval and generation system described in the application further comprises a load balancing module;
[0127] The storage pressure, network bandwidth, and CPU load indicator data collected by the receiving node monitoring agent are input into the pre-set multi-dimensional weight evaluation model in the load balancing module for quantitative calculation;
[0128] The quantification results are input into the reinforcement learning engine based on the Q-learning algorithm to dynamically generate a shard weight matrix;
[0129] According to the weight allocation results of the shard weight matrix, the topologically optimal routing path is calculated through the shortest path algorithm, and the routing decision instruction is sent to the query scheduler of the intelligent routing module to adjust the shunt ratio of the query request.
[0130] The node monitoring agent of the load balancing module deploys a lightweight data collector to capture the disk usage, network interface throughput, and CPU core utilization rate indicators of the storage node in real time. The collector uses a sliding window mechanism to smooth the original data, generating a time series of indicators after eliminating transient fluctuations. The indicator sequence is transmitted to the multi-dimensional weight evaluation model through a message queue, and the model performs logarithmic normalization on the storage pressure indicators, converts the network bandwidth indicators into residual bandwidth ratios, and calculates the CPU load indicators as weighted average values of core utilization rates, generating a standardized score matrix.
[0131] After the multi-dimensional weight evaluation model outputs the node health score, the reinforcement learning engine initializes the state-action value table of the Q-learning algorithm. The state space is defined as the discretization combination of the node score matrix, and the action space corresponds to the set of shard weight adjustment strategies. In the exploration stage, the ε-greedy strategy is used to select random actions, and in the execution stage, the Q value table entries are updated according to the historical query response delay feedback. The dynamic weight generator calculates the node affinity weight and path cost factor of each shard according to the updated Q value table, generates a shard weight matrix including weight priority, and the matrix dimension matches the storage node topology structure.
[0132] When the shortest path algorithm analyzes the storage node topology graph, the node affinity weight in the shard weight matrix is mapped to the edge weight parameter. The improved Dijkstra algorithm traverses the feasible paths in the topology graph, and the path cost function integrates the transmission hop count, bandwidth utilization penalty term and node load balancing factor to filter the candidate path set that meets the minimum delay constraint. The optimal path selector selects the transmission path with the lowest comprehensive cost from the candidate path set according to the real-time network congestion state, and generates a routing decision instruction including the target node address and priority identifier.
[0133] The routing decision instruction is sent to the query scheduler through the high-priority message queue, and the instruction parser extracts the path weight parameters to dynamically adjust the query request distribution ratio in the thread pool. The shunt controller routes high-priority queries to low-load nodes according to the weight coefficient, while limiting the request reception rate of overloaded nodes. When the elastic scaler monitors the node load change, it triggers the online update task of the weight matrix, and the reinforcement learning engine recalculates the Q value table entries to form a closed-loop optimization mechanism. After the routing strategy takes effect, the performance monitor collects the query delay and node load balancing coefficient to generate an optimization report feedback to the weight evaluation model for iterative optimization of the scoring calculation rules.
[0134] The abnormal processing mechanism is activated when a node fails. After the fault detector marks the unavailable node, the path recalculation engine immediately removes the failed node from the topology graph and re-executes the shortest path algorithm to generate a backup routing scheme. The dual-active node switch migrates the shard data of the failed node to the backup node, and maintains data consistency during the migration process through the operation log double-write mechanism. After the load balancing state is restored, the incremental synchronizer compares the data differences between nodes and resends the missing query requests to the updated routing path to ensure service continuity.
[0135] Specifically, the intelligent document retrieval and generation system based on metadata driving provided by the application is configured to receive implicit association data in the query log through a reverse feedback channel, and write the implicit association data into a graph database of a metadata knowledge graph.
[0136] The dynamic adjustment module in the load balancing module parses the updated association relationship of the knowledge graph, and recalculates the semantic similarity threshold of the metadata shard.
[0137] When the shard load deviates from the balance threshold, the elastic migration pipeline is triggered to migrate the associated data blocks in parallel according to the topology structure in the storage node topology data. During the migration process, the double-write controller synchronizes the data change operation logs of the source node and the target node.
[0138] The reverse feedback channel of the load balancing module deploys a query log analysis engine to implement association rule mining on the metadata combination patterns in historical query requests. The analysis engine uses the Apriori algorithm to identify high-frequency co-occurring metadata entity pairs, converting implicit association relationships into attribute graph structure data, including entity nodes, co-occurrence frequency weights, and time decay factors. The converted graph data is written into the graph database of the metadata knowledge graph through a batch import interface, updating the weight coefficients of entity relationship edges and enhancing the knowledge graph's ability to represent dynamic query patterns.
[0139] The dynamic adjustment module triggers a semantic analysis task after the knowledge graph is updated. The graph embedding engine uses the Node2Vec algorithm to generate high-dimensional vector representations of metadata entities. The similarity threshold calculator calculates the average clustering tightness of metadata entities within the current shard based on the Euclidean distance distribution in the vector space, and combines the real-time load status of the storage nodes to generate a dynamically adjusted semantic similarity threshold through a sliding window weighting algorithm. The threshold update event triggers a shard reevaluation process, marking metadata combinations with entity similarity below the new threshold as a candidate set for migration.
[0140] The elastic migration pipeline is activated when the load monitor detects that the shard load deviates from the balance threshold. The migration controller generates a parallel migration task queue based on the storage node topology data. The data block splitter divides the shard to be migrated into multiple migration units based on physical storage location, and each unit is attached with a version identifier and source node positioning information. The double-write controller intercepts write requests during the migration process and distributes data change operation logs to both the source node and the target node, ensuring the sequential consistency of double-write operations through operation log sequence numbers. After migration is complete, the consistency checker compares the data differences between the two nodes, analyzes the unsynchronized incremental operation logs for compensation writing, and updates the shard positioning records in the routing metadata table.
[0141] In the data synchronization process, the version conflict detector identifies the data version difference between the source node and the target node, and solves the conflict by using a merging strategy based on logical timestamps. The conflict merger preferentially retains the operation records with high version numbers, and writes the data with low version numbers into a rollback log for subsequent auditing. The migration state tracker monitors the processing progress of each migration unit in real time, and when a node failure or network interruption is detected, triggers a breakpoint resume mechanism to resume data transmission from the latest successful checkpoint. After the migration task is completed, the load balancing module reevaluates the resource utilization of the node cluster, generates a new shard weight matrix and feeds it back to the reinforcement learning engine, forming a closed-loop optimization link.
[0142] Specifically, the intelligent document retrieval and generation system based on metadata driving provided by the application comprises:
[0143] The preprocessing unit receives the associated document fragments returned by the distributed retrieval module, extracts entity relationships through the semantic parsing engine, and outputs the annotated semantic graph structure;
[0144] The text generation unit receives the semantic graph structure, fuses the metadata features in the structured key-value pairs generated by the metadata extraction module in the neural network based on the multi-head attention mechanism, and generates a preliminary response text;
[0145] The post-processing unit performs prefix aggregation processing on the preliminary response text, identifies the implicit theme distribution through a preset latent Dirichlet distribution model, merges the response paragraphs with the same theme, and outputs an optimized structured response document.
[0146] After the preprocessing unit receives the associated document fragments returned by the distributed retrieval module, the semantic parsing engine calls the dependency syntax analyzer to identify the subject-predicate-object structure in the document fragments, and uses the pre-trained named entity recognition model to extract technical feature entities. The entity relationship extractor analyzes the subordinate relationship and interactive action between entities based on the combination of rule templates and deep learning models, and generates a semantic graph structure with weighted labels. The semantic graph node stores entity types and normalized identifiers, the edge attribute records relationship types and confidence scores, and additional document source identifiers and timestamp meta information, forming a traceable semantic association network.
[0147] After the text generation unit loads the semantic graph structure, the encoder layer cross-modally splices the metadata feature vector and the semantic graph node vector. The multi-head attention mechanism is activated in the decoding stage to calculate the attention weight distribution of the query vector and each semantic graph node, dynamically adjusting the influence intensity of technical classification labels and timestamp features on the generation result. The generator uses a beam search strategy to traverse the candidate response sequence space, suppresses semantic repeated fragments through an n-gram language model, and avoids missing key entities by combining a coverage penalty mechanism, and outputs a preliminary response text including technical descriptions and data traceability information.
[0148] The post-processing unit performs paragraph segmentation on the preliminary response text, the latent Dirichlet allocation model analyzes the co-occurrence patterns of terms in the paragraph, identifies the implicit topic distribution and calculates the topic similarity matrix. The prefix aggregator clusters the paragraphs of the same topic based on the similarity threshold, and selects the highest information density expression combination using the maximum boundary correlation algorithm. The structured reorganization engine performs logical sorting on the optimized paragraph set, inserts metadata-guided chapter identifiers and cross-reference hyperlinks, and generates structured output that meets the technical document specification. The version controller writes the final response document back to the storage node, synchronously updates the document association index in the metadata knowledge graph, and forms a closed-loop data processing link.
[0149] Specifically, the metadata-driven intelligent document retrieval and generation system disclosed by the present application, the preprocessing unit is also used for:
[0150] Based on the preset Node2Vec algorithm, the entity in the metadata knowledge graph is graph embedded to calculate the high-dimensional vector representation;
[0151] The high-dimensional vector representation is input into the similarity calculator to construct the semantic adjacency matrix of the document fragment by measuring the cosine similarity;
[0152] The sliding window processor traverses the semantic adjacency matrix, extracts the periodic correlation pattern in the time dimension by the preset Fourier transform, and outputs the feature vector set with time sequence markers.
[0153] When the preprocessing unit applies the Node2Vec algorithm to perform graph embedding calculation on the entities in the metadata knowledge graph, a biased random walk strategy is used to generate node sequence samples. The return parameter and the weight of the in-out parameter are adjusted during the walk process to control the tendency of the walk path to the breadth-first or depth-first strategy, and to capture the structural proximity and semantic similarity between entities. The generated node sequence is input into the Skip-gram model for training, and the discriminability of the high-dimensional vector representation is optimized by the negative sampling technique, so that entities with similar technical fields are distributed in a clustered manner in the vector space. The vector representation is attached with entity type identifiers and timestamp markers and stored in a distributed vector database for calling by downstream modules.
[0154] The similarity calculator performs normalization processing on the high-dimensional vector to eliminate the influence of vector length difference on similarity measurement. The cosine similarity calculation module constructs a fully connected semantic adjacency matrix, and the matrix elements represent the association strength of the corresponding entity pair in the semantic space. Sparse processing is performed during matrix construction to filter low correlation data by a preset threshold, generating a compressed adjacency matrix structure. The matrix is divided into blocks according to the time dimension and attached with time window identifiers, and stored as a time sequence block matrix to provide structured input for periodic pattern analysis.
[0155] The sliding window processor traverses the time series block matrix along the time axis direction, sets configurable window span and sliding step parameters. The matrix data in the window is converted into frequency domain signal through fast Fourier transform, and the periodic correlation mode corresponding to the frequency domain energy peak value is extracted. The mode recognizer maps the frequency domain features back to the time domain, and generates a time sequence label with a starting time, a period length and an intensity value in combination with the original time stamp. The feature vector set generation module performs dimension reduction processing on the extracted mode, retains the key feature dimensions by using principal component analysis algorithm, and outputs a feature vector set carrying time context information for the generation model to perform cross-modal fusion.
[0156] The association relationship between the time sequence label and the feature vector is realized through a time stamp alignment engine. The engine analyzes the time window identifier of the feature vector, and matches it with the entity time attribute in the metadata knowledge graph. The aligned data is injected into the time sequence index structure of the graph database, supporting fast retrieval based on time range. The anomaly detector monitors the mutation event of the periodic mode, and when the mode intensity deviates from the historical baseline, triggers a dynamic update mechanism of the metadata knowledge graph, enhancing the adaptability of the system to the time sequence correlation change.
[0157] Specifically, the metadata-driven intelligent document retrieval and generation system disclosed by the application further comprises an abnormality recovery mechanism, which comprises:
[0158] A checkpoint generation unit periodically captures the index structure state of the dynamic index construction module, generates a snapshot file including B+ tree node distribution and stores it in a prewrite log;
[0159] A difference synchronization controller calls a Merkle tree comparison algorithm during data migration, compares the vector space hash value difference between the source node and the target node, generates a difference data set and synchronizes it to the target node;
[0160] A lease coordinator compares the logical time stamps of the snapshot files of each node through a version vector after network partition recovery, triggers an incremental synchronizer to execute playback of the missing operation log, and synchronizes the index states of each node through an anti-entropy protocol.
[0161] The checkpoint generation unit deploys a periodic task scheduler to trigger the index structure state capture process at regular intervals. The index traverser scans the B+ tree node distribution in the dynamic index construction module, records the key value range of the leaf node and the pointer topology relationship of the non-leaf node. The snapshot serializer converts the node metadata and key value mapping relationship into binary format, and writes it into the independent storage partition of the prewrite log after adding the logical time stamp and version number. The log manager implements a generation storage strategy, retains complete snapshots and incremental operation logs of the last three checkpoint periods, and forms recovery baseline data based on the time window.
[0162] The difference synchronization controller generates a hash tree structure of the source node and the target node when the data migration task starts. The leaf node stores the vector space hash code of the key value data, and the non-leaf node is generated by the hash value of its child node. The tree comparator compares the hash values layer by layer from the root node to locate the sub-tree branch with differences. The difference data extractor generates a minimum difference data set according to the inconsistent sub-tree path, encapsulates the difference key value as an incremental patch package, and sends it to the target node through the compression transmission protocol. The data verifier of the target node reconstructs the local Merkle tree after receiving the patch package, and updates the local index version after completing the final consistency check.
[0163] The lease coordinator starts the state coordination protocol in the network partition recovery stage, and requests the latest checkpoint snapshot metadata from each node. The version vector parser extracts the logical timestamp sequence in the snapshot, and constructs a version vector table including the maximum event counter of each node. The conflict detector identifies the node pair with version divergence through the vector clock algorithm, and preferentially selects the node with the latest logical timestamp as the synchronization reference source. The incremental synchronizer pulls the operation log in the missing time interval from the reference source node, and applies the log to the target node index structure in operation order through the transaction replay mechanism. The anti-entropy protocol is activated after synchronization is completed, and the key value distribution state of each node index copy is compared, and a compensation synchronization request is initiated for the residual difference data until all node version vectors reach the convergence state.
[0164] During the abnormal recovery process, the breakpoint continuation controller monitors the data transmission interruption event, and records the offset of the successfully transmitted data block. When the network recovers, the difference data package is continued to be transmitted from the breakpoint position, avoiding repeated transmission to improve efficiency. The version rollback manager rolls back to the latest stable version according to the snapshot file in the pre-write log when detecting the abnormality of the index structure after synchronization, while retaining the abnormal version log for subsequent analysis. The monitoring agent collects the recovery success rate and synchronization time consumption indicators in real time, and feeds back to the load balancing module to optimize the sharding migration strategy, forming a closed loop mechanism of abnormal handling and system optimization.
[0165] The technical features of the present application are explained as follows:
[0166] Metadata extraction module: a functional unit that automatically identifies and extracts key fields (such as institution name, timestamp, and technology classification label) from unstructured documents. Through natural language processing technology, the document content is parsed, the original text is converted into a standardized key-value pair format (for example, "application date: January 1, 2023" is mapped to {"date":"2023-01-01"}), and regular expression verification is performed on the field format to filter invalid data.
[0167] Structured Key-Value Pairs: Standardized data units organized in key-value format, where the key represents a metadata type (e.g., institution, time, technology category), and the value corresponds to the specific content. For example, {"institution": "XYZ Corporation", "tech_field": "ABC"}, which is used for subsequent index construction and fast matching during retrieval.
[0168] Streaming Pipeline: An asynchronous data transmission channel based on message queues (e.g., Kafka), used to transmit metadata key-value pairs in order to downstream modules. By hashing partitioning based on primary keys, it ensures that metadata from the same institution or technology field maintains temporal consistency during transmission, with additional timestamp markers to support incremental data processing.
[0169] Skip List Index: A hierarchical linked list structure designed for discrete fields (e.g., institution names), enabling fast insertion and precise querying with a time complexity of O(log n) by establishing multiple levels of indexing (e.g., layering by institution name initials). For example, an institution name index supports quick targeting of data blocks through initials.
[0170] Red-Black Tree Index: A self-balancing binary search tree used to maintain the order of continuous fields (e.g., timestamps). By coloring nodes and rotating operations, it maintains tree height balance, supporting efficient range queries (e.g., "documents from 2020 to 2023"), with a time complexity of O(log n).
[0171] Logical Sharding Strategy Unit: A decision module that divides data storage ranges based on metadata semantic associations. By analyzing field co-occurrence frequencies (e.g., the concentration of documents from a certain institution in a specific technology field) and query heat maps, it generates sharding capacity thresholds and association weight parameters to guide data sharding storage.
[0172] Improved Consistent Hashing Algorithm: Introduces a virtual node dynamic scaling mechanism based on the traditional hash ring. According to the sharding association weight, it adjusts the distribution density of virtual nodes on the hash ring, so that high-association-degree shards (e.g., documents from the same technology field) are mapped to physically adjacent storage nodes, reducing cross-node query frequency.
[0173] LSTM Prediction Model: Long Short-Term Memory neural network used to analyze time series patterns in historical query logs (e.g., a surge in retrieval volume for a certain technology field in a specific month), predict future hot metadata combinations and their access probabilities, and guide the preloading of cache shards.
[0174] Causal consistency protocol: a mechanism for maintaining the order of operations in a distributed system. By recording the timing of operations with vector clocks (such as [node ID: event counter]), comparing the clock states of different nodes to determine the execution priority of conflicting operations (such as deleting operations taking precedence over updating operations), and ensuring the causal order of data changes.
[0175] Multi-head attention mechanism: a feature fusion technique in neural networks used for semantic correlation analysis of generated models. By calculating the attention weights of document fragment semantics and metadata features (such as institutional attributes and time dimensions) in parallel, dynamically adjusting the influence of different features on the generated results, and suppressing irrelevant content output.
[0176] Merkle tree comparison algorithm: a data consistency verification method based on hash trees. By comparing the hash values of the source node and the target node (such as the root node hash difference indicating underlying data inconsistency), locating the smallest data set that needs to be synchronized, and generating an incremental patch package for differential synchronization.
[0177] Elastic migration pipeline: dynamically adjust the transmission channel of data distribution. When detecting unbalanced load of shards, parallelly migrate associated data blocks (such as all bidding information of a certain institution) according to the storage node topology, attach version identifiers to the migration units to avoid data loss, and maintain read-write consistency through a dual-write controller.
[0178] Latent Dirichlet allocation model: a topic modeling algorithm used to identify the implicit topic distribution in response text. By analyzing the co-occurrence patterns of words in paragraphs, calculating the topic similarity matrix, merging paragraphs with the same topic, and optimizing the output structure.
[0179] In specific implementation, the invention targets the multi-dimensional composite query scenario of enterprise information databases, and realizes efficient retrieval and low-noise output through multi-module collaboration. The metadata extraction module deploys a natural language processing engine, uses a pre-trained named entity recognition model to parse input documents, identifies institutional names, timestamps, and technical classification labels, verifies field formats through regular expressions, and converts "2023 X Company A Technology Bidding Document" into structured key-value pairs {"Institution": "X Company", "Time": "2023-01-01", "Technical Classification": "ABC"}. The streaming pipeline is implemented based on the Kafka message queue, and data is transmitted according to the primary key hash partitioning, with transmission delay controlled within milliseconds, ensuring the order and integrity of time-series data.
[0180] The dynamic index construction module divides the skip list and the red-black tree partition in the memory buffer, builds the skip list index for the discrete field such as the institution name, sets the node level to 4 levels, and realizes the accurate matching with the time complexity of O(log n); the continuous field such as the timestamp builds the red-black tree index, maintains the ordered key value interval to support the range query of “2020-2023 years”. The logical sharding strategy unit implements hierarchical clustering analysis on the index data, and when it is detected that the document concentration of a certain institution in a specific technical field exceeds 80%, an associated shard is generated and marked as a high-priority storage area. After the mixed index data is calculated by the sharding strategy, the intelligent routing module triggers the sharding decision process, the graph neural network generates a high-dimensional vector for the “X company-ABC” entity pair in the metadata knowledge graph, combines the monthly average access volume of 2000 times of the combination in the historical query, calculates the sharding capacity threshold and the associated weight, and drives the improved consistent hashing algorithm to route the shard to the node cluster adjacent to the physical topology.
[0181] The distributed retrieval module responds to the query of “X company 2020-2023 ABC field bidding information”, the LSTM prediction model predicts the access probability value of the metadata combination to be 0.92 according to the historical log, the retrieval coordinator preferentially accesses the cache shard of the local node, and filters the documents of non-target time interval through the Bloom filter. The causal consistency protocol compares the version vectors when cross-node retrieval, and returns 50 globally consistent associated document fragments after merging the conflict operations. The semantic analysis engine of the result generation module extracts the technical parameters and contract amount entities in the fragments, and constructs a weighted semantic graph; the generation model fuses the institution attribute features through the multi-head attention mechanism, and outputs the structured response document including the technical index comparison table and the time trend chart, the post-processing unit merges the repeated bidding process description paragraphs, and finally generates the output file conforming to the ISO document specification. When the load balancing module monitors that the CPU utilization of the node exceeds 75%, it triggers elastic migration to parallelly migrate 30% of the shard data to the low-load node, and the migration process maintains the response time of the query service below 500 ms through the double-write controller. The abnormal recovery mechanism quickly synchronizes the missing 5 operation logs based on the Merkle tree difference comparison after the network interruption for 5 minutes, so that the cluster restores data consistency within 2 minutes.
[0182] The application optimizes the index construction phase by a metadata-driven pre-filtering mechanism and a hybrid index structure to screen out non-associated data shards. A metadata extraction module parses the organization name, timestamp and technical classification label from the input document, maps them to structured key-value pairs and injects them into the streaming pipeline; a dynamic index construction module constructs a skip list and a red-black tree hybrid index based on the field type, the skip list index realizes fast and accurate matching of discrete fields, and the red-black tree maintains the order of continuous fields to support range queries. When the index is generated, the logical sharding strategy unit pre-screens high correlation shards based on the semantic correlation degree of the metadata combination, excludes non-target time interval and non-associated organization data shards at the initial stage of retrieval through the pre-filtering mechanism, reduces the number of shards to be traversed in the distributed retrieval stage, and reduces the redundant calculation overhead.
[0183] The intelligent routing module implements hierarchical clustering analysis based on the metadata vector generated by the graph neural network, generates shard boundary parameters combining the query heat map and the storage node topology characteristics. The improved consistent hashing algorithm introduces a virtual node dynamic weight factor, directs the high correlation shards to the physically adjacent nodes, combines the LSTM prediction model to predict the hot query mode, and preloads the cache shards in the target node. This mechanism predicts the query path based on semantic correlation, preferentially accesses the associated shard data of the local node or adjacent node, avoids the network transmission delay and invalid calculation caused by cross-node full shard scanning, and suppresses the access frequency of noise shards.
[0184] The distributed retrieval module and the result generation module cooperate to realize semantic-level noise filtering. The retrieval stage calls the causal consistency protocol to maintain the time sequence correctness of the cross-node index, the vector clock algorithm merges the conflict operation version, and excludes the interference of expired data; the generation stage fuses the metadata features and document semantics through the multi-head attention mechanism, and suppresses the output of non-associated content through cross-modal feature weighting. The post-processing unit adopts the latent Dirichlet distribution model to identify the implicit theme distribution of the response text, merges the repeated semantic paragraphs and inserts structured identifiers to generate the optimized technical document. The above three-level cooperative mechanism effectively improves the accuracy and response efficiency of multi-dimensional complex queries through index pre-screening, path optimization and semantic reconstruction.
Claims
1. Metadata-driven intelligent document retrieval and generation system, characterized by: include: A metadata extraction module is used to extract key fields from input documents and map them into structured key-value pairs, and output the structured key-value pairs to a streaming pipeline; A dynamic index construction module is used to inject structured key-value pairs into a preset memory buffer, classify the structured key-value pairs in the preset memory buffer by field type and generate mixed index data, and input the mixed index data into a preset logical sharding strategy unit, which outputs sharding strategy parameters and mixed index data; An intelligent routing module is configured to receive mixed index data, calculate the hash value of the primary key field in the structured key-value pair generated by the metadata extraction module using a preset hash function, logically shard the mixed index data according to the hash value of the primary key field to obtain sharded index data, generate a mapping relationship table between the sharded data and the storage nodes using a preset improved consistent hash algorithm, and transmit the sharded index data to the corresponding storage nodes according to the mapping relationship table; A distributed retrieval module is configured to receive sharded index data in response to a query request, invoke a time series prediction model trained based on historical query logs, generate a query probability prediction value for the sharded index data through the time series prediction model, filter the target shard set based on the sharding strategy parameters, and retrieve the associated document fragments of the target shard set from the corresponding storage node based on a mapping relationship table between shard data and storage nodes; The result generation module is used to extract the entity relationship in the associated document fragment through a preset semantic parsing engine to generate a semantic graph structure, splice the node vector of the semantic graph structure with the metadata feature vector in the structured key-value pair, and input it into a multi-head attention model to generate structured response data.
2. The metadata-driven intelligent document retrieval and generation system according to claim 1, characterized in that: The dynamic index building module includes: A data receiving unit is configured to receive structured key-value pairs from the metadata extraction module and transmit the metadata key-value pairs of the newly added documents to the difference encapsulation unit through a streaming pipeline; A difference encapsulation unit is used to encapsulate the metadata key-value pair of the newly added document into a difference descriptor including an operation type and a version identifier, and inject the difference descriptor into a skip list and red-black tree mixed storage area of a preset memory buffer of a dynamic index construction module; An update unit is used to periodically read the difference descriptors in the memory buffer. The persistent storage form of the mixed index data is a B+ tree index. A gap filling model is used to identify the key value interval gaps of the B+ tree index. The incremental data in the difference descriptor is inserted into the key value interval gaps according to the field type classification, and hierarchically merged with the B+ tree index to generate an updated index structure.
3. The metadata-driven intelligent document retrieval and generation system according to claim 2, characterized in that: The intelligent routing module includes: a sharding decision unit, which is used to receive the mixed index data output by the dynamic index construction module, construct a metadata knowledge graph based on the structured key-value pairs generated by the metadata extraction module, the metadata knowledge graph including entity nodes and relationship edges, input the metadata knowledge graph into a preset graph neural network for dynamic embedding to generate a high-dimensional vector representation of the metadata entity, calculate the semantic relevance of the metadata combination based on the cosine similarity between the high-dimensional vector representations, generate a query heat map based on the metadata combination and its query frequency recorded in the historical query log, and combine the semantic relevance of the metadata combination with the time window access frequency distribution in the query heat map and the physical proximity relationship in the storage node topology data to generate shard boundary parameters including a virtual node weight factor and a shard capacity threshold; a routing execution unit, configured to adjust the mapping density of the virtual node ring according to the virtual node weight factor in the shard boundary parameter, map the metadata primary key hash value to the target physical node through a preset improved consistent hashing algorithm, and optimize the routing path based on the physical proximity relationship in the storage node topology data; The replica pre-distribution unit is used to receive the hotspot probability value output by the long-short-term memory prediction model trained based on historical query logs. When the hotspot probability value exceeds the dynamically adjusted threshold, the creation of a cache shard of the target node is triggered according to the association weight in the sharding strategy parameter. The cache shard includes data blocks associated with the hotspot metadata combination.
4. The metadata-driven intelligent document retrieval and generation system according to claim 3, characterized in that: The shard decision unit is also used to: The metadata knowledge graph is input into the preset graph neural network for dynamic embedding to generate a high-dimensional vector representing the semantic relevance of the metadata; Performing a bottom-up hierarchical clustering analysis on the high-dimensional vectors, constructing a distance matrix by calculating the Euclidean distance between clusters, and merging adjacent clusters according to a dynamically adjusted cluster radius threshold to identify semantically related metadata communities; wherein the semantically related metadata communities are defined by entity relationship edges in the metadata knowledge graph; The high-dimensional vectors of semantically associated metadata communities generated by the hierarchical clustering analysis are used to calculate the vector space distance distribution of the metadata communities. Based on the metadata combinations and metadata combination query frequencies recorded in the historical query logs, a time window access frequency distribution in the real-time query heat map is generated. Physical proximity relationship data in the storage node topology data is collected. Based on the vector space distance distribution of the metadata communities, the time window access frequency distribution in the real-time query heat map, and the physical proximity relationship data in the storage node topology data, a sharding discriminant model based on a machine learning framework is constructed. The sharding discriminant model receives the following input features: a vector space distance distribution of the metadata community; The time window access frequency distribution in the real-time query heat map; Physical proximity relationship data in the storage node topology data; The input features are combined through a feature crossover algorithm, and based on the training data annotated with historical sharding decision records, a preset importance ranking algorithm is used to screen the key factors that affect the sharding decision; The output includes the shard boundary decision parameters of the shard capacity threshold and the relevance weight, where: The shard capacity threshold is determined by the vector space distance distribution of the clustering results and the physical proximity relationship in the storage node topology data; The association weight is dynamically adjusted by the time window access frequency distribution and feature cross-results in the real-time query heat map.
5. The metadata-driven intelligent document retrieval and generation system according to claim 4 is characterized in that: The distributed retrieval module is further used to: Attach a vector tag including a logical timestamp generated by the version controller inside the distributed retrieval module to the index change operation; The vector clock algorithm is used to compare the partial order of vector labels received by different nodes to determine the order in which operations are executed. When the index replica is updated asynchronously, the conflicting operations are version-merged according to the execution order, and the consistency of the merged index replica is checked by the version merging engine in the distributed retrieval module.
6. The metadata-driven intelligent document retrieval and generation system according to claim 5, characterized in that: Also includes: Load balancing module; Receive the storage pressure, network bandwidth and CPU load indicator data collected by the node monitoring agent, and input it into the multi-dimensional weight evaluation model preset in the load balancing module for quantitative calculation; Input the quantization results into the reinforcement learning engine based on the Q-learning algorithm to dynamically generate the sharding weight matrix; According to the weight distribution result of the shard weight matrix, the topological optimal routing path is calculated by the shortest path algorithm, and the routing decision instruction is sent to the query scheduler of the intelligent routing module.
7. The metadata-driven intelligent document retrieval and generation system according to claim 6, characterized in that: The load balancing module is configured to receive implicit correlation data in the query log through a reverse feedback channel, and write the implicit correlation data into the graph database of the metadata knowledge graph; The dynamic adjustment module in the load balancing module analyzes the updated associations in the knowledge graph and recalculates the semantic similarity thresholds of the metadata shards. When it is detected that the shard load deviates from the balance threshold, the elastic migration pipeline is triggered to migrate the associated data blocks in parallel according to the topological structure in the storage node topology data. During the migration process, the data change operation logs of the source node and the target node are synchronized through the dual-write controller.
8. The metadata-driven intelligent document retrieval and generation system according to claim 7, characterized in that: The result generation module includes: The preprocessing unit receives the associated document fragments returned by the distributed retrieval module, extracts entity relationships through the semantic parsing engine, and outputs an annotated semantic graph structure; a text generation unit that receives the semantic graph structure, fuses metadata features in the structured key-value pairs in a neural network based on a multi-head attention mechanism, and generates a preliminary response text; The post-processing unit performs prefix aggregation processing on the preliminary response text, identifies the implicit topic distribution through a preset latent Dirichlet distribution model, merges the response paragraphs of the same topic, and outputs an optimized structured response document.
9. The metadata-driven intelligent document retrieval and generation system according to claim 8, characterized in that: The pre-processing unit is further configured to: Based on the preset Node2Vec algorithm, the entities in the metadata knowledge graph are embedded in the graph to generate high-dimensional vector representations; Inputting the high-dimensional vector representation into a similarity calculator and constructing a semantic adjacency matrix of the document fragment using a cosine similarity metric; The sliding window processor traverses the semantic adjacency matrix, extracts the periodic correlation pattern in the time dimension through a preset Fourier transform, and outputs a set of feature vectors with time series labels.
10. The metadata-driven intelligent document retrieval and generation system according to claim 9, characterized in that: It also includes an abnormal recovery mechanism, which includes: The checkpoint generation unit periodically captures the index structure status of the dynamic index building module, generates a snapshot file including the B+ tree node distribution, and stores it in the pre-write log; The difference synchronization controller calls the Merkle tree comparison algorithm during data migration to compare the vector space hash value differences between the source node and the target node, generate a difference data set, and synchronize it to the target node; After the network partition is recovered, the lease coordinator compares the logical timestamps of the snapshot files of each node through the version vector, triggers the incremental synchronizer to replay the missing operation log, and synchronizes the index status of each node through the anti-entropy protocol.
Citation Information
Patent Citations
Hybrid indexing method based on big-data model metadata
CN107273443A
Fragmented database routing method, system and device and storage medium
CN111367884A
Cited By
Metadata-based data object detection and classification in cloud computing environments for data security posture management
US12694107B1