Intelligent data storage and knowledge base integrated management system enabling AI large model
The intelligent data storage and knowledge base integration management system, powered by AI big data models, solves the problems of high data storage costs and cumbersome and error-prone operations in existing technologies. It achieves efficient multi-source heterogeneous data storage and knowledge base management, improving resource utilization and query speed.
Patent Information
- Application Number
- CN202511131539.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-11-21
AI Technical Summary
Existing data storage and knowledge base integrated management systems lack the powerful ability to utilize large AI models, resulting in high data storage costs, cumbersome and error-prone operations, and low efficiency, making it difficult to achieve deep integration and intelligent upgrade of data storage and knowledge base management.
The intelligent data storage and knowledge base integration management system, powered by AI big data models, includes a data hierarchical storage module, a hybrid index construction module, an integrated knowledge base establishment module, and an integrated knowledge base indexing module. It calculates data value scores through multi-dimensional evaluation indicators, establishes a dynamic storage hierarchical mechanism, constructs a dynamic knowledge graph, and introduces an incremental update mechanism to achieve multi-modal indexing and self-diagnosis.
It achieves efficient multi-source heterogeneous data storage, reduces overall storage costs, improves resource utilization and knowledge coverage, enhances query speed and accuracy, and forms a self-healing closed-loop mechanism.
Smart Images

Figure CN120994844A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, specifically to an intelligent data storage and knowledge base integrated management system empowered by AI large-scale models. Background Technology
[0002] Currently, with the rapid development of artificial intelligence technology, especially the explosive application of large-scale AI models such as large language models, the scale of data faced by enterprises, institutions, and organizations is growing exponentially. Data formats are becoming increasingly complex, encompassing structured data (such as relational databases), semi-structured data (such as logs, JSON / XML), and massive amounts of unstructured data (such as text, images, audio, and video). Simultaneously, to support intelligent decision-making, automated processes, and knowledge-driven applications, building efficient and integrated knowledge base systems has become crucial. These knowledge bases not only need to store static factual information but also need to integrate tacit knowledge from structured and unstructured data sources from diverse data sources. While existing data storage systems and knowledge base technologies are relatively mature in their respective fields, they have not yet fully and effectively utilized the powerful capabilities of large-scale AI models to achieve deep integration of data storage and knowledge base management, intelligent upgrades, and the proactive construction of dynamic knowledge.
[0003] Traditional data storage and knowledge base integration management methods generally collect data from various data sources (databases, logs, documents, etc.) through ETL (Extract, Transform, Load). After preprocessing operations such as cleaning, deduplication, and format conversion, different storage technologies are selected according to data types and access requirements. Then, a rule engine is used to map and populate the knowledge base, providing functions such as indexing, partitioning, backup and recovery.
[0004] Traditional data storage and knowledge base integration management methods are mainly based on static data design, lacking the ability to intelligently compress, layer, or archive data using the information inherent in the data content itself. Data storage costs are high, and reliance on manual or predefined processes leads to cumbersome and error-prone operations, resulting in complex and inefficient management. Summary of the Invention
[0005] To address the problems in related technologies, this invention provides an AI-powered intelligent data storage and knowledge base integration management system to overcome the aforementioned technical issues in existing related technologies.
[0006] To solve the aforementioned technical problem, the present invention is achieved through the following technical solution:
[0007] This invention provides an intelligent data storage and knowledge base integration management system for AI large-scale models, specifically including: a data hierarchical storage module, a hybrid index construction module, an integrated knowledge base establishment module, and an integrated knowledge base index module;
[0008] The data hierarchical storage module is used to acquire multi-source heterogeneous data and standardize it, calculate data value scores based on multi-dimensional evaluation indicators, map them to corresponding storage locations, and establish a dynamic storage hierarchical mechanism.
[0009] The hybrid index construction module is used to perform indexing processing on different types of hierarchical data in the dynamic storage hierarchical mechanism using a three-layer hybrid index structure, and adjust the storage location to obtain a multi-source heterogeneous data storage layer.
[0010] The integrated knowledge base establishment module is used to extract multimodal features from the multi-source heterogeneous data storage layer, construct a dynamic knowledge graph, and introduce an incremental update mechanism to update it in real time to obtain the integrated knowledge base.
[0011] The integrated knowledge base indexing module is used to input natural language and perform semantic parsing, obtain multimodal index results by using hybrid indexing in the integrated knowledge base, and establish a self-diagnosis mechanism for the integrated knowledge base.
[0012] Preferably, the acquisition and standardization of multi-source heterogeneous data includes:
[0013] Acquire multi-source heterogeneous data, identify the types of multi-source heterogeneous data, and form an initial multi-source heterogeneous data set; set data detection rules, mark discarded data, and fill in missing multi-source heterogeneous data to generate a preprocessed multi-source heterogeneous data set, and then standardize it to obtain a standardized multi-source heterogeneous data set.
[0014] Preferably, the establishment of the dynamic storage hierarchy mechanism includes:
[0015] In the standardized multi-source heterogeneous data set, data value scores are calculated based on multi-dimensional evaluation indicators;
[0016] By setting a grading threshold, the data value score is divided into hot data layer, warm data layer and cold data layer according to the grading threshold. Multi-source heterogeneous data is stored in the corresponding storage location to establish a dynamic storage grading mechanism.
[0017] Preferably, the establishment of the dynamic storage hierarchy mechanism includes:
[0018] In the dynamic storage hierarchical mechanism, a new binary array is set for the hot data layer;
[0019] For the warm data layer and cold data layer in the dynamic storage grading mechanism, an adaptive block partitioning strategy is adopted to partition the data into blocks for different data types, resulting in warm data layer data blocks and cold data layer data blocks.
[0020] Based on the new binary array, a multi-path index tree is built to obtain a new hot data layer;
[0021] A query range is set in the data block of the temperature data layer, and a value range distribution histogram is established to obtain a new temperature data layer;
[0022] The cold data layer data blocks are used to create a cold data bitmap, which is then compressed to obtain a new cold data layer.
[0023] Preferably, the multi-source heterogeneous data storage layer includes:
[0024] Set a time window, detect the number of accesses to the new hot data layer, the new warm data layer, and the new cold data layer within the time window, determine whether the upgrade and downgrade conditions are triggered, complete the storage location adjustment, and obtain a multi-source heterogeneous data storage layer.
[0025] Preferably, the construction of the dynamic knowledge graph includes:
[0026] Extract the entity types of multimodal data features from the multi-source heterogeneous data storage layer, label the relationships between entity types to obtain relationship labels, and set relationship labels between entity types to obtain initial triples;
[0027] In the multi-source heterogeneous data storage layer, select fusion objects, find duplicate entity types in the initial triples, and merge the relation tags corresponding to the duplicate entity types according to the fusion objects to construct a dynamic knowledge graph.
[0028] Preferably, the obtained integrated knowledge base includes:
[0029] An incremental update mechanism is set up and trigger conditions are set sequentially. Entity types and relationship tags are added to the dynamic knowledge graph for updating, resulting in an integrated knowledge base.
[0030] Preferably, obtaining multimodal index results using hybrid indexing in the integrated knowledge base includes:
[0031] Input natural language and perform multi-intent joint recognition, set single intent query, compound intent query and hidden logic query, and obtain semantic parsing results;
[0032] The semantic parsing results are input into the integrated knowledge base for retrieval, and the corresponding index results in natural language are obtained, which are denoted as multimodal index results.
[0033] Preferably, the self-diagnostic mechanism for establishing an integrated knowledge base includes:
[0034] The integrated knowledge base is updated in real time, and its knowledge integrity and logical consistency are checked. Coverage thresholds and percentage thresholds are set, and the integrated knowledge base is compared with the coverage thresholds and percentage thresholds to establish a self-diagnosis mechanism for the integrated knowledge base.
[0035] The present invention has the following beneficial effects:
[0036] 1. This invention acquires heterogeneous data from multiple sources and processes it in a standardized manner, unifying the processing of relational data, documents, streaming data, images, etc. It can break down data silos, perform joint analysis, and automatically classify and store data according to thresholds. The dynamic classification mechanism enables value-driven and precise storage. High-frequency data is protected by high-speed media, while low-frequency data is automatically downgraded to low-cost storage, resulting in a decrease in overall storage costs and an increase in resource utilization.
[0037] 2. This invention employs a three-layer hybrid index structure to index different types of layered data, achieving dynamic matching between data value and access efficiency. This avoids performance waste caused by mixing hot and cold data. It then determines whether to trigger upgrade / downgrade conditions and adjusts the storage location, considering sub-millisecond response for hot data, intelligent association retrieval for warm data, and low-cost availability for cold data. Simultaneously, it automatically migrates data through access frequency thresholds, and the dynamic upgrade / downgrade mechanism achieves intelligent resource allocation.
[0038] 3. This invention extracts multimodal features, integrates entity relationships to generate triples, constructs a dynamic knowledge graph, and integrates multi-source features such as text and image descriptions. It solves the limitations of traditional knowledge graphs in multi-source heterogeneous data fusion, improves knowledge coverage, and introduces an incremental update mechanism for real-time updates, reducing tedious and error-prone operations and avoiding high-frequency invalid updates.
[0039] 4. This invention improves semantic matching accuracy by inputting natural language and performing semantic parsing, adopts multi-intent joint recognition, greatly reduces the response of complex multimodal queries through data hierarchical retrieval, improves the speed of composite queries by hybrid indexes, and dynamically monitors coverage and conflict rates, automatically triggers incremental updates, and forms a self-healing closed-loop mechanism. Attached Figure Description
[0040] To more clearly illustrate the technical solutions of the embodiments of the invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the invention. For those skilled in the art, the drawings can be obtained from these drawings without creative effort.
[0041] Figure 1 A flowchart of an AI-powered intelligent data storage and knowledge base integrated management system for this invention;
[0042] Figure 2 This invention provides a flowchart illustrating the intelligent data storage and knowledge base integration management method empowered by AI large models. Detailed Implementation
[0043] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0044] Traditional data storage and knowledge base integration management methods are mainly based on static data design, lacking the ability to intelligently compress, layer, or archive data using the information inherent in the data content itself. Data storage costs are high, and reliance on manual or predefined processes leads to cumbersome and error-prone operations, resulting in complex and inefficient management.
[0045] To solve the above technical problems, such as Figure 1 As shown, this embodiment of the invention provides an AI-powered intelligent data storage and knowledge base integration management system, specifically including: a data layered storage module, a hybrid index construction module, an integrated knowledge base establishment module, and an integrated knowledge base index module. The data layered storage module is used to acquire multi-source heterogeneous data and standardize it, calculate data value scores based on multi-dimensional evaluation indicators, map them to corresponding storage locations, and establish a dynamic storage hierarchical mechanism. The hybrid index construction module is used to index different types of layered data in the dynamic storage hierarchical mechanism using a three-layer hybrid index structure, adjust storage locations, and obtain a multi-source heterogeneous data storage layer. The integrated knowledge base establishment module is used to extract multi-modal features from the multi-source heterogeneous data storage layer, construct a dynamic knowledge graph, and introduce an incremental update mechanism for real-time updates, resulting in an integrated knowledge base. The integrated knowledge base index module is used to input natural language and perform semantic parsing, obtain multi-modal index results in the integrated knowledge base using a hybrid index, and establish an integrated knowledge base self-diagnosis mechanism.
[0046] In a specific embodiment, during the intelligent upgrade project of a leading e-commerce platform, the platform had an average daily order volume of over 3 million, an inventory of over 30 million products (covering electronics, apparel, fresh produce, etc.), and over 20 data source types (including real-time transaction streams, multilingual reviews, product images / videos, supply chain logs, etc.), providing a solid data foundation for the embodiments of this invention.
[0047] In the specific implementation process of the above embodiments, firstly, multi-source heterogeneous data is acquired and standardized. Data value scores are calculated based on multi-dimensional evaluation indicators, and grading thresholds are set to map the data value scores to corresponding storage locations, establishing a dynamic storage grading mechanism. This method uniformly processes relational data, documents, streaming data, images, etc., breaking down data silos. Through joint analysis, it automatically grades storage according to thresholds. The dynamic grading mechanism achieves value-driven precise storage. High-frequency data uses high-speed media to ensure performance, while low-frequency data is automatically downgraded to low-cost storage, resulting in a decrease in overall storage costs and an increase in resource utilization. Secondly, a three-layer hybrid index structure is used to index different types of layered data, and then it is determined whether upgrade / downgrade conditions are triggered, adjusting the storage location to obtain a multi-source heterogeneous data storage layer. This method considers sub-millisecond response for hot data, intelligent association retrieval for warm data, and low-cost availability for cold data, using a three-layer hybrid index structure to achieve dynamic matching between data value and access efficiency. To avoid performance waste caused by mixing hot and cold data, this method automatically migrates data based on access frequency thresholds and achieves intelligent resource allocation through a dynamic upgrade and downgrade mechanism. Multimodal feature extraction is then performed, entity relationships are fused to generate triples, a dynamic knowledge graph is constructed, and an incremental update mechanism is introduced for real-time updates, resulting in an integrated knowledge base. This method integrates multi-source features such as text and image descriptions, overcoming the limitations of traditional knowledge graphs in multi-source heterogeneous data fusion, improving knowledge coverage, and the timed update mechanism reduces tedious and error-prone operations, avoiding high-frequency invalid updates. Finally, natural language is input and semantic parsing is performed. A hybrid index is used in the integrated knowledge base to obtain multimodal index results, and a self-diagnostic mechanism for the integrated knowledge base is established. This method uses multi-intent joint recognition to improve semantic matching accuracy, significantly reduces the response time of complex multimodal queries through data hierarchical retrieval, improves the speed of composite queries through hybrid indexes, and dynamically monitors coverage and conflict rates, automatically triggering incremental updates to form a self-healing closed-loop mechanism.
[0048] Furthermore, to better illustrate the technical solutions of the embodiments of the present invention, such as... Figure 2 As shown, this paper describes in detail the intelligent data storage and knowledge base integration management system empowered by AI large-scale models, combining the methods of intelligent data storage and knowledge base integration management with AI large-scale models. Specifically, it includes the following:
[0049] S1. Acquire multi-source heterogeneous data and standardize the processing. Calculate the data value score based on multi-dimensional evaluation indicators. Set a grading threshold to map the data value score to the corresponding storage location and establish a dynamic storage grading mechanism.
[0050] S1 includes the following steps:
[0051] S11. Acquire multi-source heterogeneous data, identify the types of multi-source heterogeneous data, and obtain multi-source heterogeneous data such as relational databases, unstructured documents, and real-time streaming data. Use text descriptions to describe images, tables, etc., to form an initial multi-source heterogeneous data set. Set data detection rules. When there are α consecutive missing multi-source heterogeneous data in the initial multi-source heterogeneous data set, mark it as discarded data. Use interpolation to fill the missing multi-source heterogeneous data in the initial multi-source heterogeneous data set. Otherwise, the initial multi-source heterogeneous data set meets the data detection rules, and a preprocessed multi-source heterogeneous data set is generated.
[0052] The timestamps in the preprocessed multi-source heterogeneous data set are standardized, and conversion rules are set according to different multi-source heterogeneous data types to convert currency, percentage, unit system, etc. in real time. After data normalization, garbled characters are cleaned, tabs, consecutive spaces and other redundant data are deleted, and consecutive duplicate data are merged to complete the standardization process and output a standardized multi-source heterogeneous data set.
[0053] S12. In the standardized multi-source heterogeneous data set, the number of data accesses within a historical unit time period is counted, the time interval from the time the multi-source heterogeneous data was generated to the current time is calculated and recorded as the data timeliness, the storage cost of the multi-source heterogeneous data is calculated, the value of the multi-source heterogeneous data is quantified, and the data value score A = ω1·log(a+1) + ω2·b - ω3·c is calculated based on multi-dimensional evaluation indicators, where ω1, ω2 and ω3 represent the score weights, a represents the number of data accesses, b represents the data timeliness, and c represents the storage cost.
[0054] S13. Set the grading thresholds as a first threshold, a second threshold, and a third threshold. When the data value score is greater than or equal to the first threshold, the corresponding multi-source heterogeneous data is recorded as high-frequency access data and added to the hot data layer. When the data value score is less than the first threshold but greater than or equal to the second threshold, the corresponding multi-source heterogeneous data is recorded as medium-frequency access data and added to the warm data layer. When the data value score is less than the second threshold but greater than or equal to the third threshold, the corresponding multi-source heterogeneous data is recorded as low-frequency access data and added to the cold data layer. Otherwise, the corresponding multi-source heterogeneous data is added to the archive layer and Blu-ray archived. Set the storage locations to include the hot data layer, the warm data layer, and the cold data layer, and assign different storage media types to different types of tiered data in the storage locations in sequence. Then, for new multi-source heterogeneous data, calculate the data value score and store it in the corresponding storage location to establish a dynamic storage grading mechanism.
[0055] In this embodiment, multi-source heterogeneous data is acquired and standardized. Data value scores are calculated based on multi-dimensional evaluation indicators. A grading threshold is set to map the data value scores to corresponding storage locations, establishing a dynamic storage grading mechanism. This method uniformly processes relational data, documents, streaming data, images, etc., breaking down data silos. Through joint analysis, it automatically grades storage according to thresholds. The dynamic grading mechanism achieves value-driven, precise storage. High-frequency data uses high-speed media to ensure performance, while low-frequency data is automatically downgraded to low-cost storage, resulting in a decrease in overall storage costs and an increase in resource utilization. Specifically, for example, an e-commerce platform uses this system to manage product data, including: relational data: product prices in MySQL. The system utilizes various data sources, including grids, inventory, unstructured documents (product description text, user reviews), and real-time streaming data (user click behavior logs). It standardizes timestamps, converts currencies, and cleans redundancy by removing consecutive spaces and garbled characters from reviews. Image standardization converts product images into text descriptions, outputting a collection of 5000 standardized multi-source heterogeneous data entries. A value score is calculated for each product data record, with 1,200 visits per day (historical 24-hour clicks), a timeliness of 8 hours, and a storage cost of 0.5 yuan / GB / day. An expert voting method is used to weigh the data value and obtain initial weights. Specifically, following the Delphi method, 30 experts score the data, and the average is taken. Then, the Q-lea reinforcement learning framework is used. In the initial state, the Q-table, spatial actions, and reward function are set. Periods T0 (initial state) and T1 are selected to simulate the state changes and weight adjustments over two time periods. Period T0 (initial state): Hot data access frequency level 1 (medium), storage cost level 1 (medium), data timeliness level 1 (medium), initial weights ω1 = 0.4, ω2 = 0.3, and ω3 = 0.3 respectively. System performance: latency 50ms, storage cost 1000 yuan, hot data hit rate 80%. Period T1: Data access frequency level 2 (high), storage cost level 2 (high), data timeliness level 2 (high). Based on the current state and the Q-table (initially 0), spatial actions are selected. Actions, such as using ε-greedy (ε = 0.3), randomly select action 0: increase ω1, decrease ω2, resulting in adjusted weights: ω1 = 0.45, ω2 = 0.25, and ω3 = 0.3. After implementing the new weights, the system runs for one cycle (30 minutes), with a latency of 45ms (decreasing), storage cost of 1100 yuan (increasing, more data is stored in the hot layer), and a data hit rate of 85%. Calculate the reward function value and update the initial Q-table. Repeat the above process, and the Q-table gradually accumulates experience. Based on the state, select the optimal spatial action to adjust the weights, setting the weights to 0.6, 0.3, and 0.1 respectively, resulting in a data value score A = 0.6 × 7.09 + 2.4 - 0.05 = 6.604; For example, set the threshold for hot data layer to ≥8, warm data layer to [5, 8), cold data layer to [3, 5), and archive layer to <3, use Blu-ray storage, and store them in the corresponding storage locations to establish a dynamic storage hierarchy mechanism;
[0056] S2. For different types of layered data in the dynamic storage grading mechanism, a three-layer hybrid index structure is used to index the different types of layered data respectively, and then it is determined whether the upgrade or downgrade condition is triggered, and the storage location is adjusted to obtain a multi-source heterogeneous data storage layer.
[0057] S2 includes the following steps:
[0058] S21. In the dynamic storage hierarchical mechanism, the hot data layer in the hot data layer is recorded as hot data layer data; a binary array with an initial value of 0 and several hash functions are set, the hot data layer data corresponds to the binary array, the hash functions are used to calculate the position of all hot data layer data, and the initial value of 0 at the corresponding position is modified to 1, and a new binary array is obtained.
[0059] In the dynamic storage tiering mechanism, the warm data layer and cold data layer are respectively denoted as warm data layer data and cold data layer data. The data types of the warm data layer data and cold data layer data are identified, and an adaptive block partitioning strategy is adopted to perform block processing for different data types to obtain warm data layer data blocks and cold data layer data blocks.
[0060] S22. A three-layer hybrid index structure is used to index different types of layered data to obtain a new hot data layer, a new warm data layer, and a new cold data layer. The specific steps are as follows:
[0061] S221. Calculate the location of the hot data layer data again using the hash function, and query the new binary array. If the query result is 0, the query request returns 0, indicating the hot data layer data does not exist; otherwise, the query request returns 1, indicating the hot data layer data exists. Divide the hot data layer data into several data pages, each containing leaf nodes. Build an empty tree, starting from the root node of the empty tree, and search for leaf nodes from bottom to top. Store the hot data layer data in the leaf nodes, and split the leaf nodes until the root node, maintaining the tree's balance to obtain a multi-path index tree. Use the multi-path index tree as an index to obtain a new hot data layer.
[0062] S222. For the warm data layer data blocks, obtain the number of queries for the warm data layer data blocks, calculate the ratio of simultaneous queries of two warm data layer data blocks to queries of a single warm data layer data block, and obtain the correlation coefficient; set a correlation threshold, when the correlation coefficient is greater than the correlation threshold, merge and store the corresponding warm data layer data blocks, otherwise store the corresponding warm data layer data blocks separately, and merge them sequentially to obtain several warm data layer data groups; find the maximum and minimum values in the warm data layer data groups to obtain the query interval, and establish a value range distribution histogram based on the distribution frequency of the warm data layer data groups in each query region; use the value range distribution histogram as an index to obtain a new warm data layer;
[0063] S223. For the cold data layer data blocks, each cold data layer data block is treated as a row, and the number of cold data layer data blocks is the number of columns, to obtain a cold data bitmap; traverse the columns of the cold data bitmap to obtain unique cold data, create an array with a length equal to the number of rows, and initialize the array to 0 to obtain an initial array; traverse the rows of the cold data bitmap, and when a row of the cold data bitmap contains cold data equal to the unique cold data, change the corresponding position in the initial array from 0 to 1 until the cold data bitmap is completely traversed, then identify consecutive 0s and 1s for compression to obtain a compressed bitmap; the cold data layer uses the compressed bitmap as an index to obtain a new cold data layer;
[0064] S23. Set a time window and detect the number of accesses to the new hot data layer, the new warm data layer, and the new cold data layer within the time window to obtain the number of accesses per unit time; set a first access number threshold and a second access number threshold. When the number of accesses per unit time is less than the first access number threshold, a downgrade condition is triggered; when the number of accesses per unit time is greater than the second access number threshold, an upgrade condition is triggered; otherwise, it remains unchanged. Specifically, when the downgrade condition is triggered, hot data in the new hot data layer is transferred to the new warm data layer, and warm data in the new warm data layer is transferred to the new cold data layer; when the downgrade condition is triggered, cold data in the new cold data layer is transferred to the new warm data layer, and warm data in the new warm data layer is transferred to the new hot data layer.
[0065] After completing the storage location adjustment, a multi-source heterogeneous data storage layer is obtained by combining a new hot data layer, a new warm data layer, and a new cold data layer.
[0066] In this embodiment, a three-layer hybrid index structure is used to index different types of layered data, and then the storage location is adjusted according to whether the upgrade / downgrade condition is triggered, resulting in a multi-source heterogeneous data storage layer. This method considers sub-millisecond response of hot data, intelligent association retrieval of warm data, and low-cost availability of cold data. The three-layer hybrid index structure realizes dynamic matching between data value and access efficiency, avoiding performance waste caused by hot and cold data mixing. At the same time, data is automatically migrated through access frequency thresholds, and the dynamic upgrade / downgrade mechanism realizes intelligent resource allocation. Specifically, for example, hot data layer index initialization: for 100 popular product IDs in the hot data layer, a binary array (length 10000) is created, using 3 hash functions. Calculate hash positions: at positions 583, 2194, and 7621, change the values from 0 to 1 to construct a Bloom filter; Warm / cold data chunking: for example, warm data layer user browsing records are chunked by user ID (500 users per chunk), and cold data layer historical orders are chunked by order date (30 days of data per chunk); Hot data layer index (multi-path index tree): when querying a product, the Bloom filter returns 1 (exists), dividing the product data into data pages (20 SKUs per page), constructing a B+ tree index, reducing query latency from 2ms to 0.3ms; Warm data layer index (value range distribution histogram): analyze user behavior blocks, block A (users 1-500): query mobile phone... Block A (users 501-1000): Headphones accounted for 52% of searches. When the correlation threshold is greater than 30%, the data blocks are considered to be merged. The correlation coefficient is calculated: the proportion of simultaneous searches for mobile phones and headphones reaches 45%. At this time, Block A and Block B are merged into "3C Product Group", and a value range distribution histogram is established. Cold data layer index (bitmap compression): For example, column 1 may have three values: 'active', 'inactive', 'pending', a total of 5 rows, row 1 = 'active', row 2 = 'inactive', row 3 = 'active', row 4 = 'pending', row 5 = 'inactive', resulting in a matrix [1, 0, 1]. [0, 0; 0, 1, 0, 0, 1; 0, 0, 0, 1, 0], identify consecutive 0s or 1s for compression; dynamic upgrade / downgrade trigger: set to check once every 15 minutes (time window 15min), if the number of accesses in the cold data layer is 182 times / 15min, which is greater than the second threshold of 150 times → upgrade is triggered, data is extracted from the cold data layer, the Bloom filter is reconstructed, and inserted into the B+ tree of the hot data layer; if the number of accesses in the hot data layer is 3 times / 15min, which is less than the first threshold of 50 times → downgrade is triggered, the node is deleted from the B+ tree, and merged into the "3C product group" of the warm data layer according to the user behavior correlation; the peak concurrent processing capacity is 12,000 upgrade / downgrade operations, of which the average query latency of the hot data layer is 0.The average query latency for the warm data layer is 3ms, the average query latency for the cold data layer is 8ms, and the average query latency for the hot data layer is 120ms. Compared to before the upgrade / downgrade was triggered, the data storage cost was reduced to half of the original solution, while ensuring 99.9% availability for business continuity. The first and second thresholds were obtained by acquiring historical multi-source heterogeneous data storage layers and analyzing the data transfer conditions between the hot data layer, the warm data layer, and the cold data layer.
[0067] S3. The multi-source heterogeneous data storage layer performs multimodal feature extraction, integrates entity relationships to generate triples, constructs a dynamic knowledge graph, and introduces an incremental update mechanism to update in real time, thereby obtaining an integrated knowledge base.
[0068] S3 includes the following steps:
[0069] S31. For the multi-source heterogeneous data in the multi-source heterogeneous data storage layer, the text is segmented into paragraphs and sentences, images, tables and other text descriptions are collected, and then word segmentation is performed according to sentences to identify multimodal data features. The association anchor points of multimodal data features are established, the entity types of multimodal data features are extracted, and the relationships between entity types are labeled to obtain relationship labels.
[0070] For the entity types and relation labels, an entity relation extraction method is established based on data time sequence, behavioral logic, relation reasoning, etc., and relation labels are set between entity types; the entity types are used as nodes and relation labels are used as edges, and the edges are used to connect the nodes to obtain the initial triples;
[0071] S32. Knowledge fusion is performed on the initial triples to construct a dynamic knowledge graph, and an incremental update mechanism is introduced to update it in real time to obtain an integrated knowledge base. The specific steps are as follows:
[0072] S321. Obtain the data credibility, data timeliness, and semantic contradictions of the multi-source heterogeneous data in the multi-source heterogeneous data storage layer; sort the data credibility according to priority, select the multi-source heterogeneous data corresponding to the highest priority, and find the corresponding entity features as the fusion object; select the latest multi-source heterogeneous data as the fusion object according to the data timeliness; add the semantic contradictions to the manual review queue.
[0073] Find duplicate entity types in the initial triples, and merge the relation tags corresponding to the duplicate entity types according to the fusion object to construct a dynamic knowledge graph;
[0074] S322. Set the incremental update mechanism to include real-time update, timed update, and event-driven update, and set the trigger conditions in sequence to obtain new multi-source heterogeneous data, convert it into entity type and relation label, determine the trigger conditions corresponding to the new multi-source heterogeneous data, add the corresponding entity type and relation label to the dynamic knowledge graph for updating, and obtain the integrated knowledge base.
[0075] In this embodiment, multimodal feature extraction is performed, entity relationships are fused to generate triples, a dynamic knowledge graph is constructed, and an incremental update mechanism is introduced for real-time updates to obtain an integrated knowledge base. This method integrates multi-source features such as text and image descriptions, solving the limitations of traditional knowledge graphs in multi-source heterogeneous data fusion, improving knowledge coverage, and the timed update mechanism reduces tedious and error-prone operations and avoids high-frequency invalid updates. Specifically, for example, the hot data layer is the real-time user click stream ("User A clicks on the phone"), the warm data layer is the product description document ("The phone is equipped with a Snapdragon chip"), and the cold data layer is the historical user reviews ("Battery life is insufficient"). The entity type of the above are all products, and the relationship tags are equipped with, featured, and exist. The defects are linked to the chip, screen, and battery, forming an initial triplet: (phone, equipped with, Snapdragon chip), (phone, equipped with, 6.1-inch OLED screen), (phone, defective, insufficient battery life); credibility conflict: official documentation (high credibility): "Battery life 28 hours", user comments (low credibility): "Actual battery life less than 10 hours"; timeliness conflict: 2023 data: "Equipped with Snapdragon 888 chip", 2025 data: "Upgraded to Snapdragon 8G3 chip"; official documentation and 2025 data are selected as the fusion objects to construct a dynamic knowledge graph; regular updates: analysis of 12,000 comments from the previous day to discover new relationships (phone, compatible accessories, charger);
[0076] S4. Input natural language and perform semantic parsing. Use hybrid indexing in the integrated knowledge base to obtain multimodal index results and establish a self-diagnosis mechanism for the integrated knowledge base.
[0077] S4 includes the following steps:
[0078] S41. Input natural language and perform multi-intent joint recognition. When the natural language matches a single intent query, directly extract entity type and relation label. When the natural language matches a compound intent query, perform word segmentation on the natural language and then extract entity type and relation label. When the natural language matches a hidden logic query, perform logical negation on the natural language and then extract entity type and relation label to obtain semantic parsing results.
[0079] S42. Input the semantic parsing result into the integrated knowledge base for retrieval, identify the data value score of the semantic parsing result, find the corresponding storage location in the multi-source heterogeneous data storage layer, and trace back to the data source corresponding to the semantic parsing result according to the three-layer hybrid index structure to obtain the index result corresponding to the natural language, which is recorded as the multimodal index result.
[0080] The integrated knowledge base is updated in real time. The knowledge integrity and logical consistency of the integrated knowledge base are checked. The coverage rate of entity types in the integrated knowledge base and the proportion of conflicting triples in the integrated knowledge base are calculated. Coverage threshold and proportion threshold are set. When the coverage rate of entity types in the integrated knowledge base is less than the coverage threshold, or the proportion of conflicting triples in the integrated knowledge base is greater than the proportion threshold, an alarm is triggered to complete the self-diagnosis of the integrated knowledge base and establish a self-diagnosis mechanism for the integrated knowledge base.
[0081] In this embodiment, natural language is input and semantic parsing is performed. A hybrid index is used in the integrated knowledge base to obtain multimodal index results, and a self-diagnostic mechanism for the integrated knowledge base is established. This method uses multi-intent joint recognition to improve semantic matching accuracy, significantly reduces the response time of complex multimodal queries through data hierarchical retrieval, improves the speed of composite queries through hybrid indexing, and dynamically monitors coverage and conflict rates, automatically triggering incremental updates to form a self-healing closed-loop mechanism. Specifically, for example, a single intent: "What is the screen size of the iPhone 15?", outputs (iPhone 15, equipped with, 6.1-inch screen); a composite intent: "Compare the chips and prices of the iPhone 15 and Huawei Mate 60", outputs [(iPhone 15, equipped with, A15 chip), (Mate 60, price, 5499 yuan)]; hidden logic: "Apple phones should not be priced below 5000 yuan", outputs (Apple phone, price, ≥5000 yuan). For example, a user asks "Display negative review images about iPhone 15 battery life issues". Semantic parsing results: (iPhone 15, defective, battery issue) + associated media type image; Hot data layer: Bloom filter confirms no image data → skip; Warm data layer: Value range histogram locates the "battery issue" comment block (hit rate 83%); Cold data layer: Bitmap compression quickly excludes non-image data; Multimodal result generation; Image 1: "Standby screenshot uploaded by user A (battery depleted in 1 hour); Entity coverage = existing entities / total number of entities to be covered (coverage threshold 95%); Conflict 3" Tuple percentage = number of contradictory data / total number of triples (percentage threshold 0.05%). All thresholds are set by the minimum standard of the standard integrated knowledge base. The detection showed that the coverage of the new "VisionPro Headset" was only 82%, less than 95%. The reason was: lack of "compatible device" relationship. Conflict was detected: 15% of users reported "overheating", while the official statement was "normal temperature". The conflict percentage was 0.7% (>0.5%) → triggering an alarm. After actual testing by the quality inspection department, it was marked as "defect of specific batch". The self-diagnosis of the integrated knowledge base was completed.
[0082] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic point described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristic points described may be combined in any suitable manner in one or more embodiments or examples.
[0083] The preferred embodiments of the invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention.
Claims
1. An AI-powered intelligent data storage and knowledge base integration management system, characterized in that: include: The data tiered storage module is used to acquire multi-source heterogeneous data and standardize it, calculate data value scores based on multi-dimensional evaluation indicators, map them to corresponding storage locations, and establish a dynamic storage tiering mechanism. The hybrid index building module is used to perform indexing on different types of hierarchical data in the dynamic storage hierarchical mechanism using a three-layer hybrid index structure, and adjust the storage location to obtain a multi-source heterogeneous data storage layer. The integrated knowledge base building module is used to extract multimodal features from the multi-source heterogeneous data storage layer, construct a dynamic knowledge graph, and introduce an incremental update mechanism to update it in real time to obtain the integrated knowledge base; An integrated knowledge base indexing module is used to input natural language and perform semantic parsing. It uses a hybrid index to obtain multimodal index results in the integrated knowledge base and establishes a self-diagnostic mechanism for the integrated knowledge base.
2. The AI-powered intelligent data storage and knowledge base integration management system according to claim 1, characterized in that, The acquisition and standardization of multi-source heterogeneous data includes: Acquire multi-source heterogeneous data, identify the types of multi-source heterogeneous data, and form an initial multi-source heterogeneous data set; set data detection rules, mark discarded data, and fill in missing multi-source heterogeneous data to generate a preprocessed multi-source heterogeneous data set, and then standardize it to obtain a standardized multi-source heterogeneous data set.
3. The AI-powered intelligent data storage and knowledge base integration management system according to claim 2, characterized in that, The establishment of the dynamic storage hierarchy mechanism includes: In the standardized multi-source heterogeneous data set, data value scores are calculated based on multi-dimensional evaluation indicators; By setting a grading threshold, the data value score is divided into hot data layer, warm data layer and cold data layer according to the grading threshold. Multi-source heterogeneous data is stored in the corresponding storage location to establish a dynamic storage grading mechanism.
4. The AI-powered intelligent data storage and knowledge base integration management system according to claim 3, characterized in that, The establishment of the dynamic storage hierarchy mechanism includes: In the dynamic storage hierarchical mechanism, a new binary array is set for the hot data layer; For the warm data layer and cold data layer in the dynamic storage grading mechanism, an adaptive block partitioning strategy is adopted to partition the data into blocks for different data types, resulting in warm data layer data blocks and cold data layer data blocks. Based on the new binary array, a multi-path index tree is built to obtain a new hot data layer; A query range is set in the data block of the temperature data layer, and a value range distribution histogram is established to obtain a new temperature data layer; The cold data layer data blocks are used to create a cold data bitmap, which is then compressed to obtain a new cold data layer.
5. The AI-powered intelligent data storage and knowledge base integration management system according to claim 4, characterized in that, The obtained multi-source heterogeneous data storage layer includes: Set a time window, detect the number of accesses to the new hot data layer, the new warm data layer, and the new cold data layer within the time window, determine whether the upgrade and downgrade conditions are triggered, complete the storage location adjustment, and obtain a multi-source heterogeneous data storage layer.
6. The AI-powered intelligent data storage and knowledge base integration management system according to claim 5, characterized in that, The construction of the dynamic knowledge graph includes: Extract the entity types of multimodal data features from the multi-source heterogeneous data storage layer, label the relationships between entity types to obtain relationship labels, and set relationship labels between entity types to obtain initial triples; In the multi-source heterogeneous data storage layer, select fusion objects, find duplicate entity types in the initial triples, and merge the relation tags corresponding to the duplicate entity types according to the fusion objects to construct a dynamic knowledge graph.
7. The AI-powered intelligent data storage and knowledge base integration management system according to claim 6, characterized in that, The integrated knowledge base includes: An incremental update mechanism is set up and trigger conditions are set sequentially. Entity types and relationship tags are added to the dynamic knowledge graph for updating, resulting in an integrated knowledge base.
8. The AI-powered intelligent data storage and knowledge base integration management system according to claim 7, characterized in that, The method of obtaining multimodal index results using hybrid indexing in the integrated knowledge base includes: Input natural language and perform multi-intent joint recognition, set single intent query, compound intent query and hidden logic query, and obtain semantic parsing results; The semantic parsing results are input into the integrated knowledge base for retrieval, and the corresponding index results in natural language are obtained, which are denoted as multimodal index results.
9. The AI-powered intelligent data storage and knowledge base integration management system according to claim 8, characterized in that, The self-diagnostic mechanism for establishing an integrated knowledge base includes: The integrated knowledge base is updated in real time, and its knowledge integrity and logical consistency are checked. Coverage thresholds and percentage thresholds are set, and the integrated knowledge base is compared with the coverage thresholds and percentage thresholds to establish a self-diagnosis mechanism for the integrated knowledge base.
10. An AI-powered intelligent data storage and knowledge base integration management method, characterized in that: Specifically, it includes: S1. Acquire multi-source heterogeneous data and standardize the processing. Calculate the data value score based on multi-dimensional evaluation indicators. Set a grading threshold to map the data value score to the corresponding storage location and establish a dynamic storage grading mechanism. S2. For different types of layered data in the dynamic storage grading mechanism, a three-layer hybrid index structure is used to index the different types of layered data respectively, and then it is determined whether the upgrade or downgrade condition is triggered, and the storage location is adjusted to obtain a multi-source heterogeneous data storage layer. S3. The multi-source heterogeneous data storage layer performs multimodal feature extraction, integrates entity relationships to generate triples, constructs a dynamic knowledge graph, and introduces an incremental update mechanism to update in real time, thereby obtaining an integrated knowledge base. S4. Input natural language and perform semantic parsing. Use hybrid indexing in the integrated knowledge base to obtain multimodal index results and establish a self-diagnosis mechanism for the integrated knowledge base.
Citation Information
Cited By
Network additional storage intelligent data management method and system
CN122045162A