Data Lake Storage Optimization Methods and Platforms for Enterprise Business Integration

By acquiring enterprise business integration data streams, performing joint explicit and implicit information modeling and storage optimization, and generating a storage optimization memory, the problem of low data storage efficiency for enterprise business integration is solved, achieving more efficient data lake storage and management.

CN120973835BActive Publication Date: 2026-03-06江苏鑫埭信息科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511470373.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-15
Publication Date
2026-03-06
Estimated Expiration
2045-10-15

AI Technical Summary

Technical Problem

Existing technologies for enterprise business integration data storage are characterized by low efficiency, high redundancy, low storage utilization, and insufficient ability to process and access real-time data.

Method used

By acquiring enterprise business integration data streams, extracting multi-source heterogeneous data sets, performing joint modeling of explicit and implicit information, generating a storage optimization memory, and performing prior storage identification and verification based on the memory to optimize data lake storage.

Benefits of technology

It improves the storage efficiency of the data lake, solves the problem of low data storage efficiency for enterprise business integration, and achieves more efficient data management and access capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120973835B_ABST
    Figure CN120973835B_ABST
Patent Text Reader

Abstract

This invention discloses a data lake storage optimization method and platform for enterprise business integration, relating to the field of data processing technology. The method includes: acquiring the business integration data stream of the target enterprise; extracting a multi-source heterogeneous data set from the business integration data stream; performing explicit and implicit information joint modeling to obtain a set of explicit and implicit information joint association results; performing iterative memory storage optimization analysis, and performing data lake storage and generating a storage optimization memory for the multi-source heterogeneous data set of the business according to the target storage optimization strategy; when receiving real-time business data, performing prior storage identification based on the storage optimization memory to obtain prior storage results, and verifying the prior storage results; if the verification is successful, storing the real-time business data based on the prior storage results. This solves the technical problem of low data storage efficiency in enterprise business integration in existing technologies, achieving the technical effect of improving data lake storage efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, specifically to a data lake storage optimization method and platform for enterprise business integration. Background Technology

[0002] In the process of enterprise digital transformation, data from business systems often exhibits characteristics of multi-source and heterogeneity, with data generated by different departments and application platforms differing in structure, format, and semantics. As businesses rapidly expand, enterprises need to centrally manage and analyze this scattered data during business integration to support strategic decision-making and real-time operations. However, existing data storage methods often lack a unified optimization mechanism, resulting in a high proportion of redundant data in data lakes, low storage utilization, and insufficient processing and access capabilities for real-time data. This not only affects the efficient use of storage resources but also restricts the collaborative analysis and intelligent application of enterprise business data. Summary of the Invention

[0003] This application provides a data lake storage optimization method and platform for enterprise business integration, which solves the technical problem of low data storage efficiency for enterprise business integration in the prior art.

[0004] A first aspect of this application provides a data lake storage optimization method for enterprise business integration, the method comprising:

[0005] The system acquires the business integration data stream of the target enterprise and extracts a multi-source heterogeneous data set from the data stream. It then traverses the multi-source heterogeneous data set to perform explicit and implicit information joint modeling, obtaining a set of explicit and implicit information joint association results, which includes a set of explicit and implicit business multi-source heterogeneous information and a set of implicit business multi-source heterogeneous information. Based on the set of explicit and implicit information joint association results, iterative memory storage optimization analysis is performed, and data lake storage and a storage optimization memory are generated for the multi-source heterogeneous data set according to the target storage optimization strategy. When receiving real-time business data, prior storage identification is performed based on the storage optimization memory to obtain prior storage results, and the prior storage results are verified. If the verification is successful, the real-time business data is stored based on the prior storage results.

[0006] A second aspect of this application provides a data lake storage optimization platform for enterprise business integration, the platform comprising:

[0007] Data Acquisition Module: Acquires the business integration data stream of the target enterprise and extracts a multi-source heterogeneous data set from the business integration data stream; Information Modeling Module: Traverses the multi-source heterogeneous data set to perform explicit and implicit information joint modeling and obtains a set of explicit and implicit information joint association results, wherein the set of explicit and implicit information joint association results includes a set of business multi-source heterogeneous explicit information and a set of business multi-source heterogeneous implicit information; Storage Optimization Analysis Module: Performs iterative memory storage optimization analysis based on the set of explicit and implicit information joint association results, and performs data lake storage and generates a storage optimization memory bank for the multi-source heterogeneous data set according to the target storage optimization strategy; Data Storage Module: When receiving real-time business data, performs prior storage identification based on the storage optimization memory bank, obtains prior storage results, and verifies the prior storage results. If the verification is successful, the real-time business data is stored based on the prior storage results.

[0008] One or more technical solutions provided in this application have at least the following technical effects or advantages:

[0009] First, the business integration data flow of the target enterprise is acquired, and a multi-source heterogeneous data set is extracted from it. Next, the multi-source heterogeneous data set is traversed to perform explicit and implicit information joint modeling, obtaining a set of explicit and implicit information joint association results. This set includes both explicit and implicit information sets from the multi-source heterogeneous data set. Then, iterative memory storage optimization analysis is performed based on this set, and data lake storage and a storage optimization memory are generated for the multi-source heterogeneous data set according to the target storage optimization strategy. Finally, when receiving real-time business data, prior storage identification is performed based on the storage optimization memory to obtain prior storage results. These results are then verified; if successful, the real-time business data is stored based on the prior storage results. This approach solves the technical problem of low data storage efficiency in enterprise business integration data storage in existing technologies, achieving the technical effect of improving data lake storage efficiency. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 This is a schematic diagram of a data lake storage optimization method for enterprise business integration provided in an embodiment of this application;

[0012] Figure 2This is a schematic diagram of the data lake storage optimization platform structure for enterprise business integration provided in an embodiment of this application.

[0013] Figure labeling: Data acquisition module 11, information modeling module 12, storage optimization and analysis module 13, data storage module 14. Detailed Implementation

[0014] This application solves the technical problem of low data storage efficiency for enterprise business integration by providing a data lake storage optimization method and platform for enterprise business integration.

[0015] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0016] It should be noted that the terms "comprising" and "having" are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or modules that are not explicitly listed or that are inherent to these processes, methods, products, or devices.

[0017] Example 1, as Figure 1 As shown, this application provides a data lake storage optimization method for enterprise business integration, wherein the method includes:

[0018] Obtain the business integration data stream of the target enterprise, and extract a set of multi-source heterogeneous business data from the business integration data stream.

[0019] Establish data interface connections in the target company's business information system, financial management system, supply chain system, customer relationship management system, and other business application systems. Acquire business integration data streams in real time or periodically through standardized data acquisition modules. The business integration data streams include structured data (such as tables and logs), semi-structured data (such as XML and JSON files), and unstructured data (such as text and audio / video files), forming a raw data set covering different business modules of the enterprise.

[0020] During the data acquisition process, the business integration data stream undergoes format recognition and data preprocessing to identify the data's source system, data type, timestamp, and key identifier fields. Based on the format recognition results, a data parser is used to uniformly extract different types of data, performing structured data table parsing, semi-structured file tag parsing, and unstructured data feature extraction respectively. Through field mapping and data standardization rules, the extracted data is converted into a unified data description format and preliminarily classified according to data source, business function module, and data type. Finally, a multi-source heterogeneous data set of the target enterprise is generated as the input basis for subsequent explicit and implicit information joint modeling.

[0021] The explicit and implicit information joint modeling is performed by traversing the multi-source heterogeneous data set of the business to obtain the joint association result set of explicit and implicit information, which includes the multi-source heterogeneous explicit information set and the multi-source heterogeneous implicit information set of the business.

[0022] The system iterates through the multi-source heterogeneous data set of the business, loading corresponding parsing rules based on data type and source system. In explicit information modeling, according to preset explicit content indicators (such as field name, value range, timestamp, business tags, etc.), it extracts the directly visible attributes of each data instance, forming a multi-source heterogeneous explicit information set of the business. In implicit information modeling, for the same batch of multi-source heterogeneous business data, based on preset implicit access indicators (such as access frequency, access time period, query mode), and combining the sensitivity identification module and business value prediction module, it analyzes the usage characteristics and sensitivity of each data object. The system analyzes the business value level to generate a multi-source heterogeneous implicit information set. The explicit and implicit information sets are then input into the joint modeling module. In this module, a mapping operation is first performed to associate the explicit and implicit information within the same data object dimension, forming a one-to-one or one-to-many correspondence. Next, based on context labels, business labels, and access patterns, cross-object context association calculations are performed to identify potential connections between different data objects, forming a joint association vector. Finally, the explicit information set, implicit information set, and joint association vector are integrated into a set of explicit and implicit information joint association results.

[0023] Furthermore, the explicit and implicit information are jointly modeled by traversing the multi-source heterogeneous data set of the business, and a joint association result set of explicit and implicit information is obtained. This joint association result set includes a multi-source heterogeneous explicit information set and a multi-source heterogeneous implicit information set of the business, including:

[0024] Based on preset explicit content indicators, the raw data of the multi-source heterogeneous data set is extracted to obtain a multi-source heterogeneous raw data set; based on the business scenario information of the target enterprise, business tags are assigned to the multi-source heterogeneous data set to obtain a multi-source heterogeneous data business tag set; based on a preset association window, the multi-source heterogeneous data set is traversed to perform context association tag recognition to obtain a multi-source heterogeneous context association tag set; the multi-source heterogeneous raw data set, the multi-source heterogeneous data business tag set, and the multi-source heterogeneous context association tag set are mapped and associated to obtain a multi-source heterogeneous explicit information set; the multi-source heterogeneous data set is parsed for implicit information to obtain a multi-source heterogeneous implicit information set; the multi-source heterogeneous explicit information set and the multi-source heterogeneous implicit information set are jointly associated to obtain a joint association result set of explicit information.

[0025] First, raw data is extracted from the multi-source heterogeneous data set according to preset explicit content indicators. These indicators include raw fields, indicator items, and data text content, forming the multi-source heterogeneous raw data set. Second, business tag assignment is performed on the multi-source heterogeneous data set based on the target enterprise's business scenario information. Data from different sources is mapped to the enterprise's existing business processes, application modules, or business roles, generating a multi-source heterogeneous data business tag set. Then, contextual traversal is performed on the multi-source heterogeneous data set based on preset association windows. These preset association windows consist of contextual data with potential relevance within a specific time period or logical business scope. This traversal yields... A set of heterogeneous business context-related tags is generated. Based on this, the original heterogeneous business data set, the business tag set, and the context-related tag set are mapped and integrated to obtain a set of explicit heterogeneous business information. Further, implicit information parsing is performed on the heterogeneous business data set, including dimensions such as access frequency, peak periods, typical query patterns, sensitivity, and data value, ultimately forming a set of implicit heterogeneous business information. Finally, the explicit and implicit information sets are jointly associated to establish one-to-one or one-to-many relationships between explicit and implicit information, generating a set of joint association results.

[0026] Furthermore, implicit information parsing is performed on the aforementioned multi-source heterogeneous data set of business data to obtain a multi-source heterogeneous implicit information set of business data, including:

[0027] Based on preset implicit access indicators, implicit information is extracted from the multi-source heterogeneous data set of the business to obtain a set of implicit access indicators for multi-source heterogeneous data of the business; sensitivity identification is performed on the multi-source heterogeneous data set of the business to obtain a set of sensitivity of multi-source heterogeneous data of the business; the business value predictor is invoked to classify the value of the multi-source heterogeneous data set of the business to obtain a set of value levels of multi-source heterogeneous data of the business; the set of implicit access indicators, the set of sensitivity, and the set of value levels of multi-source heterogeneous data of the business are integrated to obtain a set of implicit information of multi-source heterogeneous data of the business.

[0028] Furthermore, the implicit access metrics include data access frequency, peak periods, and typical query patterns.

[0029] Specifically, implicit information is extracted from the multi-source heterogeneous data set of the business according to preset implicit access indicators. These implicit access indicators reflect the usage characteristics of the data in actual business scenarios, including data access frequency, peak access periods, and typical query patterns. By statistically analyzing and modeling the access status of various types of data across different time and usage dimensions, a set of implicit access indicators for multi-source heterogeneous business data is formed. Next, the multi-source heterogeneous data set is iterated through one by one, and a sensitivity identification module is invoked to analyze the privacy and security attributes of the data. Based on whether the data contains sensitive fields, sensitive text, or privacy attributes, a corresponding sensitivity level is generated, thus forming the multi-source heterogeneous business data... The system first obtains a sensitivity set; then, it calls a pre-trained business value predictor to classify the value of the multi-source heterogeneous data set. The business value predictor is based on a neural network model and is trained by combining historical business data and expert-annotated data. It can score the value of data according to its business relevance, usage frequency, and contribution to decision-making, thereby obtaining a set of value levels for the multi-source heterogeneous data. Finally, it integrates and correlates the set of implicit access indicators, sensitivity sets, and value levels of the multi-source heterogeneous data to form a set of implicit information, which serves as the implicit information input for the joint modeling of explicit and implicit information.

[0030] Based on the set of results of the joint association of explicit and implicit information, iterative memory storage optimization analysis is performed, and data lake storage and storage optimization memory are generated for the set of multi-source heterogeneous data of the business according to the target storage optimization strategy.

[0031] Furthermore, based on the set of results from the joint association of explicit and implicit information, iterative memory storage optimization analysis is performed, and according to the target storage optimization strategy, data lake storage and storage optimization memory are generated for the set of multi-source heterogeneous business data, including:

[0032] Based on data access frequency and business tags, a first round of storage optimization is performed on the multi-source heterogeneous explicit information set and the multi-source heterogeneous implicit information set of the business to obtain a hot-cold tiered storage strategy, which is then stored in the first memory unit. Based on data sensitivity and raw data, a second round of storage optimization is performed on the multi-source heterogeneous explicit information set and the multi-source heterogeneous implicit information set of the business, combined with the first memory unit, to obtain a compression and encryption strategy, which is then added to the first memory unit to obtain a second memory unit. Based on historical query patterns and business context-related tags... In conjunction with the second memory unit, a third round of storage optimization is performed on the multi-source heterogeneous explicit information set and the multi-source heterogeneous implicit information set of the business to obtain a partitioning and indexing strategy. The partitioning and indexing strategy is then added to the second memory unit to obtain a third memory unit. Based on the data value and the third memory unit, a fourth round of storage optimization is performed on the multi-source heterogeneous explicit information set and the multi-source heterogeneous implicit information set of the business to obtain a target storage optimization strategy. The target storage optimization strategy and the result set of the joint association of the explicit and implicit information are stored in an initially empty database to generate a storage optimization memory library.

[0033] Based on data access frequency and business tags, a first round of storage optimization is performed on the multi-source heterogeneous explicit information set and the multi-source heterogeneous implicit information set of the business. High-frequency access data and low-frequency access data are managed in layers, and a hot / cold tiered storage strategy is constructed and written into the first memory unit. Next, based on data sensitivity and original data characteristics, a second round of storage optimization is performed on the explicit and implicit information in conjunction with the first memory unit. Sensitive data undergoes hierarchical encryption processing, while non-sensitive but large-volume data is compressed for storage, thus obtaining a compression and encryption strategy. This strategy is then superimposed on the first memory unit to form the second memory unit. Subsequently, based on historical query patterns and business... Context-related tags are used to perform a third round of storage optimization on explicit and implicit information in conjunction with the second memory unit. By modeling and optimizing typical query paths, the data is divided into several partitions and an efficient retrieval index is established. Partition and index strategies are generated and written into the second memory unit to obtain the third memory unit. Finally, based on the data value and the third memory unit, a fourth round of storage optimization is performed on explicit and implicit information. The allocation of storage resources is dynamically adjusted according to the contribution of data to business decisions and operations to form the final target storage optimization strategy. This strategy is then stored together with the result set of the joint association of explicit and implicit information in the initially empty database to generate the storage optimization memory.

[0034] Furthermore, based on data sensitivity and raw data, a second round of storage optimization is performed on the multi-source heterogeneous explicit information set and the multi-source heterogeneous implicit information set of the business, combined with the first memory unit, to obtain a compression and encryption strategy. This compression and encryption strategy is then added to the first memory unit to obtain a second memory unit, comprising:

[0035] Using the original data as an index, the storage area of ​​the business multi-source heterogeneous explicit information set is extracted from the cold and hot tiered storage strategy of the first memory unit to obtain the set of business multi-source heterogeneous explicit information split storage area groups; using data sensitivity as an index, the split data sensitivity of each business multi-source heterogeneous implicit information in the business multi-source heterogeneous implicit information set is extracted, and combined with the set of business multi-source heterogeneous explicit information split storage area groups, the corresponding data is compressed and encrypted to obtain the compression and encryption strategy.

[0036] Specifically, using the original data as an index, the storage area for the set of business multi-source heterogeneous explicit information is extracted from the existing hot and cold tiered storage strategy in the first memory unit. The extracted storage area is then split to obtain a set of split storage area groups for business multi-source heterogeneous explicit information. Next, using data sensitivity as an index, the set of business multi-source heterogeneous implicit information is traversed one by one to extract the split data sensitivity level corresponding to each implicit information object, forming a sensitivity labeling result. Then, the sensitivity labeling result is mapped to the set of split storage area groups for business multi-source heterogeneous explicit information. For highly sensitive data objects, a multi-layer encryption algorithm is used for security processing, while compression algorithms are combined to compress and store non-sensitive but large-volume data. Finally, the compression and encryption strategies generated in this process are recorded and written into the first memory unit to form the updated second memory unit.

[0037] When receiving real-time service data, prior storage identification is performed based on the storage optimization memory to obtain prior storage results, and the prior storage results are verified. If the verification is successful, the real-time service data is stored based on the prior storage results.

[0038] During the real-time business data access phase, the input data stream is parsed and preprocessed to extract the corresponding original fields, business tags, and access features. The real-time business data is mapped to the set of explicit and implicit information joint association results in the storage optimization memory. An approximate matching algorithm is called to compare the real-time explicit and implicit information to determine the corresponding historical optimization results. Based on this, the target storage optimization strategy is retrieved to generate prior storage results. Next, the prior storage results are verified. The verification process includes calculating the simulated storage write latency and simulated query response time and comparing them with preset performance thresholds. When all simulated indicators meet the threshold conditions, the verification is deemed successful, and the real-time business data is stored according to the corresponding cold and hot tiering strategy, compression and encryption strategy, partitioning and indexing strategy, and value-oriented strategy in the prior storage results.

[0039] Furthermore, when receiving real-time service data, prior storage identification is performed based on the storage optimization memory to obtain prior storage results, including:

[0040] The real-time business data is subjected to joint explicit and implicit information modeling to obtain explicit and implicit real-time business information; the explicit and implicit real-time business information is approximated and matched with the set of joint association results of explicit and implicit information in the storage optimization memory to determine the matching joint association results of explicit and implicit information; based on the matching joint association results of explicit and implicit information, the target storage optimization strategy is retrieved in the storage optimization memory to obtain the prior storage results.

[0041] First, explicit and implicit information joint modeling is performed on real-time business data. The input real-time data is parsed to extract its original fields, business tags, and contextual features, forming explicit real-time business information. Simultaneously, implicit feature analysis is performed on the real-time business data by combining dynamic indicators such as real-time access frequency, access time period, and query pattern, forming implicit real-time business information. Then, the explicit and implicit real-time business information are approximated and matched with the set of joint association results of explicit and implicit information already stored in the storage optimization memory. The approximation matching process includes calculating similarity indices for both explicit and implicit dimensions, and selecting the optimal matching item from the multi-dimensional similarity results to determine the corresponding joint association result of explicit and implicit information. Finally, based on the matched joint association result of explicit and implicit information, the corresponding target storage optimization strategy is retrieved from the storage optimization memory. The target storage optimization strategy includes multi-dimensional storage rules such as hot and cold tiering, compression and encryption, partitioning and indexing, and value orientation. Prior storage results are generated based on the retrieval results to guide the final storage operation of real-time business data.

[0042] Furthermore, from both explicit and implicit dimensions, approximate calculations are performed on the set of explicit real-time business information, implicit real-time business information, and the joint association result set of explicit and implicit information to obtain a set of explicit real-time approximation degree and a set of implicit real-time approximation degree. The set of explicit real-time approximation degree and the set of implicit real-time approximation degree are then weighted to obtain a set of implicit real-time approximation matching degree. The joint association result of explicit and implicit information corresponding to the maximum value in the set of implicit real-time approximation matching degree is taken as the joint association result of explicit and implicit matching information.

[0043] Specifically, the explicit dimension approximation calculation includes calculating the similarity between the field names, numerical features, business tags, and context tags of real-time business data and their corresponding items in historical explicit information, resulting in a real-time explicit approximation set. The implicit dimension approximation calculation includes comparing the implicit features of real-time business data, such as access frequency, peak periods, typical query patterns, sensitivity levels, and value, with historical implicit information, resulting in a real-time implicit approximation set. After obtaining these two sets, a weighted calculation is performed on the real-time explicit and implicit approximation sets. The weight values ​​are preset or dynamically adjusted according to the target enterprise's business scenario to obtain a comprehensive evaluation index, forming a real-time approximate matching approximation set. Finally, the joint association result of explicit and implicit information corresponding to the maximum value in the real-time approximate matching approximation set is selected as the joint association result of explicit and implicit information matching, and the target storage optimization strategy is retrieved from the storage optimization memory based on this.

[0044] Furthermore, this includes:

[0045] The real-time business data is simulated and stored according to the prior storage results to obtain the simulated write latency and simulated query response time; it is determined whether the simulated write latency and simulated query response time both meet the preset requirements, and if so, the verification is successful.

[0046] After real-time business data is accessed, simulated storage operations are performed on the real-time business data based on the cold / hot tiering strategy, compression and encryption strategy, partitioning and indexing strategy included in the prior storage results. The latency time during the simulated write process is recorded to obtain the simulated write latency duration, and a simulated query response duration is generated during the simulated query process. Subsequently, the simulated write latency duration and the simulated query response duration are compared with preset performance thresholds, which are set by the business processing requirements of the target enterprise and include the maximum tolerable write latency and the maximum tolerable query response duration. When both the simulated write latency duration and the simulated query response duration meet the preset requirements, the prior storage result is deemed to have passed verification, and the real-time business data can be written into the data lake according to the optimization strategy corresponding to the prior storage result. If any indicator fails to meet the standard, the verification is deemed to have failed, triggering a rollback mechanism to re-call the default storage process or adjust the storage optimization strategy to ensure the stability and availability of data storage.

[0047] In summary, the embodiments of this application have at least the following technical effects:

[0048] First, the business integration data flow of the target enterprise is acquired, and a multi-source heterogeneous data set is extracted from it. Next, the multi-source heterogeneous data set is traversed to perform explicit and implicit information joint modeling, obtaining a set of explicit and implicit information joint association results. This set includes both explicit and implicit information sets from the multi-source heterogeneous data set. Then, iterative memory storage optimization analysis is performed based on this set, and data lake storage and a storage optimization memory are generated for the multi-source heterogeneous data set according to the target storage optimization strategy. Finally, when receiving real-time business data, prior storage identification is performed based on the storage optimization memory to obtain prior storage results. These results are then verified; if successful, the real-time business data is stored based on the prior storage results. This approach solves the technical problem of low data storage efficiency in enterprise business integration data storage in existing technologies, achieving the technical effect of improving data lake storage efficiency.

[0049] Example 2, based on the same inventive concept as the data lake storage optimization method for enterprise business integration in the foregoing examples, such as... Figure 2 As shown, this application provides a data lake storage optimization platform for enterprise business integration, wherein the platform includes:

[0050] Data acquisition module 11: Acquires the business integration data stream of the target enterprise and extracts a multi-source heterogeneous data set from the business integration data stream; Information modeling module 12: Traverses the multi-source heterogeneous data set of the business to perform explicit and implicit information joint modeling and obtains a set of explicit and implicit information joint association results, wherein the set of explicit and implicit information joint association results includes a set of business multi-source heterogeneous explicit information and a set of business multi-source heterogeneous implicit information; Storage optimization analysis module 13: Performs iterative memory storage optimization analysis based on the set of explicit and implicit information joint association results, and performs data lake storage and generates a storage optimization memory bank for the multi-source heterogeneous data set of the business according to the target storage optimization strategy; Data storage module 14: When receiving real-time business data, performs prior storage identification based on the storage optimization memory bank, obtains prior storage results, and verifies the prior storage results. If the verification is successful, the real-time business data is stored based on the prior storage results.

[0051] Furthermore, the information modeling module 12 is used to perform the following methods:

[0052] Based on preset explicit content indicators, the raw data of the multi-source heterogeneous data set is extracted to obtain a multi-source heterogeneous raw data set; based on the business scenario information of the target enterprise, business tags are assigned to the multi-source heterogeneous data set to obtain a multi-source heterogeneous data business tag set; based on a preset association window, the multi-source heterogeneous data set is traversed to perform context association tag recognition to obtain a multi-source heterogeneous context association tag set; the multi-source heterogeneous raw data set, the multi-source heterogeneous data business tag set, and the multi-source heterogeneous context association tag set are mapped and associated to obtain a multi-source heterogeneous explicit information set; the multi-source heterogeneous data set is parsed for implicit information to obtain a multi-source heterogeneous implicit information set; the multi-source heterogeneous explicit information set and the multi-source heterogeneous implicit information set are jointly associated to obtain a joint association result set of explicit information.

[0053] Furthermore, the information modeling module 12 is used to perform the following methods:

[0054] Based on preset implicit access indicators, implicit information is extracted from the multi-source heterogeneous data set of the business to obtain a set of implicit access indicators for multi-source heterogeneous data of the business; sensitivity identification is performed on the multi-source heterogeneous data set of the business to obtain a set of sensitivity of multi-source heterogeneous data of the business; the business value predictor is invoked to classify the value of the multi-source heterogeneous data set of the business to obtain a set of value levels of multi-source heterogeneous data of the business; the set of implicit access indicators, the set of sensitivity, and the set of value levels of multi-source heterogeneous data of the business are integrated to obtain a set of implicit information of multi-source heterogeneous data of the business.

[0055] Furthermore, the information modeling module 12 is used to perform the following methods:

[0056] The implicit access metrics include data access frequency, peak periods, and typical query patterns.

[0057] Furthermore, the storage optimization analysis module 13 is used to perform the following methods:

[0058] Based on data access frequency and business tags, a first round of storage optimization is performed on the multi-source heterogeneous explicit information set and the multi-source heterogeneous implicit information set of the business to obtain a hot-cold tiered storage strategy, which is then stored in the first memory unit. Based on data sensitivity and raw data, a second round of storage optimization is performed on the multi-source heterogeneous explicit information set and the multi-source heterogeneous implicit information set of the business, combined with the first memory unit, to obtain a compression and encryption strategy, which is then added to the first memory unit to obtain a second memory unit. Based on historical query patterns and business context-related tags... In conjunction with the second memory unit, a third round of storage optimization is performed on the multi-source heterogeneous explicit information set and the multi-source heterogeneous implicit information set of the business to obtain a partitioning and indexing strategy. The partitioning and indexing strategy is then added to the second memory unit to obtain a third memory unit. Based on the data value and the third memory unit, a fourth round of storage optimization is performed on the multi-source heterogeneous explicit information set and the multi-source heterogeneous implicit information set of the business to obtain a target storage optimization strategy. The target storage optimization strategy and the result set of the joint association of the explicit and implicit information are stored in an initially empty database to generate a storage optimization memory library.

[0059] Furthermore, the storage optimization analysis module 13 is used to perform the following methods:

[0060] Using the original data as an index, the storage area of ​​the business multi-source heterogeneous explicit information set is extracted from the cold and hot tiered storage strategy of the first memory unit to obtain the set of business multi-source heterogeneous explicit information split storage area groups; using data sensitivity as an index, the split data sensitivity of each business multi-source heterogeneous implicit information in the business multi-source heterogeneous implicit information set is extracted, and combined with the set of business multi-source heterogeneous explicit information split storage area groups, the corresponding data is compressed and encrypted to obtain the compression and encryption strategy.

[0061] Furthermore, the data storage module 14 is used to perform the following method:

[0062] The real-time business data is subjected to joint explicit and implicit information modeling to obtain explicit and implicit real-time business information; the explicit and implicit real-time business information is approximated and matched with the set of joint association results of explicit and implicit information in the storage optimization memory to determine the matching joint association results of explicit and implicit information; based on the matching joint association results of explicit and implicit information, the target storage optimization strategy is retrieved in the storage optimization memory to obtain the prior storage results.

[0063] Furthermore, the data storage module 14 is used to perform the following method:

[0064] From both explicit and implicit dimensions, approximate calculations are performed on the set of explicit real-time business information, implicit real-time business information, and the joint association result set of explicit and implicit information to obtain a set of explicit real-time approximation degree and a set of implicit real-time approximation degree. The set of explicit real-time approximation degree and the set of implicit real-time approximation degree are then weighted to obtain a set of implicit real-time approximation matching degree. The joint association result of explicit and implicit information corresponding to the maximum value in the set of implicit real-time approximation matching degree is taken as the joint association result of explicit and implicit matching information.

[0065] Furthermore, the data storage module 14 is used to perform the following method:

[0066] The real-time business data is simulated and stored according to the prior storage results to obtain the simulated write latency and simulated query response time; it is determined whether the simulated write latency and simulated query response time both meet the preset requirements, and if so, the verification is successful.

[0067] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. The processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.

[0068] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

[0069] This specification and accompanying drawings are merely illustrative examples of this application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Therefore, if such modifications and modifications fall within the scope of this application and its equivalents, this application intends to include such modifications and modifications.

Claims

1. A method for data lake storage optimization for enterprise business integration, characterized in that, The method comprises: acquiring a business integration data stream of a target enterprise, and extracting a business multi-source heterogeneous data set from the business integration data stream; iteratively performing memory storage optimization analysis based on the explicit and implicit information joint correlation result set, and performing data lake storage on the business multi-source heterogeneous data set according to a target storage optimization strategy and generating a storage optimization memory library; when real-time business data is received, performing prior storage identification based on the storage optimization memory library, obtaining a prior storage result, verifying the prior storage result, and if the verification is passed, storing the real-time business data based on the prior storage result; performing iterative memory storage optimization analysis based on the explicit and implicit information joint correlation result set, and performing data lake storage on the business multi-source heterogeneous data set according to a target storage optimization strategy and generating a storage optimization memory library, comprising: based on data access frequency and business labels, performing first-round storage optimization on the business multi-source heterogeneous explicit information set and the business multi-source heterogeneous implicit information set, obtaining a cold and hot layered storage strategy, and storing the cold and hot layered storage strategy to a first memory unit; based on data sensitivity and raw data, combining the first memory unit to perform second-round storage optimization on the business multi-source heterogeneous explicit information set and the business multi-source heterogeneous implicit information set, obtaining a compression and encryption strategy, and adding the compression and encryption strategy to the first memory unit to obtain a second memory unit; based on historical query mode and business context association labels, combining the second memory unit to perform third-round storage optimization on the business multi-source heterogeneous explicit information set and the business multi-source heterogeneous implicit information set, obtaining a partitioning and indexing strategy, and adding the partitioning and indexing strategy to the second memory unit to obtain a third memory unit; based on data value degree and the third memory unit, performing fourth-round storage optimization on the business multi-source heterogeneous explicit information set and the business multi-source heterogeneous implicit information set, obtaining a target storage optimization strategy, and storing the target storage optimization strategy and the explicit and implicit information joint correlation result set in an initially empty database to generate a storage optimization memory library; based on data sensitivity and raw data, combining the first memory unit to perform second-round storage optimization on the business multi-source heterogeneous explicit information set and the business multi-source heterogeneous implicit information set, obtaining a compression and encryption strategy, and adding the compression and encryption strategy to the first memory unit to obtain a second memory unit, comprising: taking the raw data as an index, extracting a storage area of the business multi-source heterogeneous explicit information set from the cold and hot layered storage strategy of the first memory unit, and obtaining a business multi-source heterogeneous explicit information split storage area group set; ​ The data sensitivity is taken as an index to extract the split data sensitivity of each business multi-source heterogeneous implicit information in the business multi-source heterogeneous implicit information set, and the corresponding data is compressed and encrypted in combination with the business multi-source heterogeneous explicit information split storage area group set, to obtain a compression and encryption strategy.

2. The data lake storage optimization method for enterprise business integration of claim 1, wherein, The business multi-source heterogeneous data set is traversed to perform joint modeling of explicit and implicit information, to obtain a joint association result set of explicit and implicit information, wherein the joint association result set of explicit and implicit information includes a business multi-source heterogeneous explicit information set and a business multi-source heterogeneous implicit information set, and includes: According to a preset explicit content index, the business multi-source heterogeneous data set is subjected to original data extraction to obtain a business multi-source heterogeneous original data set; According to the business scene information of the target enterprise, the business multi-source heterogeneous data set is subjected to business label allocation to obtain a business multi-source heterogeneous data business label set; Based on a preset association window, the business multi-source heterogeneous data set is traversed to perform context association label identification to obtain a business multi-source heterogeneous context association label set; The business multi-source heterogeneous original data set, the business multi-source heterogeneous data business label set, and the business multi-source heterogeneous context association label set are mapped and associated to obtain a business multi-source heterogeneous explicit information set; The business multi-source heterogeneous data set is subjected to implicit information analysis to obtain a business multi-source heterogeneous implicit information set; The business multi-source heterogeneous explicit information set and the business multi-source heterogeneous implicit information set are jointly associated to obtain the joint association result set of explicit information.

3. The data lake storage optimization method for enterprise business integration of claim 2, wherein, The business multi-source heterogeneous data set is subjected to implicit information analysis to obtain a business multi-source heterogeneous implicit information set, including: According to a preset implicit access index, the business multi-source heterogeneous data set is subjected to implicit information extraction to obtain a business multi-source heterogeneous data implicit access index set; The business multi-source heterogeneous data set is traversed to perform sensitivity identification to obtain a business multi-source heterogeneous data sensitivity set; A business value predictor is called to perform value classification on the business multi-source heterogeneous data set to obtain a business multi-source heterogeneous data value level set; The business multi-source heterogeneous data implicit access index set, the business multi-source heterogeneous data sensitivity set, and the business multi-source heterogeneous data value level set are integrated to obtain the business multi-source heterogeneous implicit information set.

4. The data lake storage optimization method for enterprise business integration of claim 3, wherein, The implicit access index includes data access frequency, peak period, and typical query mode.

5. The data lake storage optimization method for enterprise business integration of claim 1, wherein, When receiving real-time business data, prior storage identification is performed based on the storage optimization memory bank to obtain a prior storage result, including: The real-time business data is subjected to joint modeling of explicit and implicit information to obtain real-time business explicit information and real-time business implicit information; The real-time business explicit information and the real-time business implicit information are approximately matched with the joint association result set of explicit and implicit information in the storage optimization memory bank to determine a matching joint association result of explicit and implicit information; Based on the matching joint association result of explicit and implicit information, a target storage optimization strategy in the storage optimization memory bank is searched to obtain the prior storage result.

6. The data lake storage optimization method for enterprise business integration of claim 5, wherein, Approximate calculation is performed on the real-time business explicit information, real-time business implicit information and the combined association result set of the explicit and implicit information to obtain a real-time explicit approximate degree set and a real-time implicit approximate degree set; Weighted calculation is performed on the real-time explicit approximate degree set and the real-time implicit approximate degree set to obtain a real-time approximate matching approximate degree set; The combined association result corresponding to the maximum value in the real-time approximate matching approximate degree set is taken as the matching combined association result of the explicit and implicit information.

7. The data lake storage optimization method for enterprise business integration of claim 1, wherein, It comprises: The real-time business data is simulated and stored according to the prior storage result to obtain a simulated write delay duration and a simulated query response duration; It is judged whether the simulated write delay duration and the simulated query response duration meet the preset requirements, and if so, the verification is passed.

8. A data lake storage optimization platform for enterprise business integration, characterized by, The platform for implementing the enterprise business integration-oriented data lake storage optimization method of any one of claims 1-7 comprises: A data acquisition module: acquiring the business integration data stream of a target enterprise and extracting a business multi-source heterogeneous data set from the business integration data stream; An information modeling module: traversing the business multi-source heterogeneous data set to perform explicit and implicit information joint modeling and obtaining a combined association result set of explicit and implicit information, wherein the combined association result set of explicit and implicit information comprises a business multi-source heterogeneous explicit information set and a business multi-source heterogeneous implicit information set; A storage optimization analysis module: performing iterative memory storage optimization analysis based on the combined association result set of explicit and implicit information and performing data lake storage and generating a storage optimization memory bank for the business multi-source heterogeneous data set according to a target storage optimization strategy; A data storage module: when receiving real-time business data, performing prior storage identification based on the storage optimization memory bank to obtain a prior storage result, verifying the prior storage result, and if the verification is passed, storing the real-time business data based on the prior storage result; The storage optimization analysis module is further used to perform: Based on the data access frequency and the business label, the business multi-source heterogeneous explicit information set and the business multi-source heterogeneous implicit information set are subjected to first round storage optimization, a cold and hot layered storage strategy is obtained, and the cold and hot layered storage strategy is stored to a first memory unit; based on the data sensitivity and the original data, the business multi-source heterogeneous explicit information set and the business multi-source heterogeneous implicit information set are subjected to second round storage optimization in combination with the first memory unit, a compression and encryption strategy is obtained, the compression and encryption strategy is added to the first memory unit to obtain a second memory unit; based on the historical query mode and the business context association label, the business multi-source heterogeneous explicit information set and the business multi-source heterogeneous implicit information set are subjected to third round storage optimization in combination with the second memory unit, a partition and index strategy is obtained, the partition and index strategy is added to the second memory unit to obtain a third memory unit; based on the data value degree and the third memory unit, the business multi-source heterogeneous explicit information set and the business multi-source heterogeneous implicit information set are subjected to fourth round storage optimization, a target storage optimization strategy is obtained, and the target storage optimization strategy and the explicit and implicit information joint association result set are stored to an initially empty database to generate a storage optimization memory bank; The storage optimization analysis module is also used to perform: With the original data as an index, a storage area of the business multi-source heterogeneous explicit information set is extracted from the cold and hot layered storage strategy of the first memory unit to obtain a business multi-source heterogeneous explicit information split storage area group set; with the data sensitivity as an index, a split data sensitivity of each business multi-source heterogeneous implicit information in the business multi-source heterogeneous implicit information set is extracted, and compression and encryption are performed on the corresponding data in combination with the business multi-source heterogeneous explicit information split storage area group set to obtain a compression and encryption strategy.

Citation Information

Patent Citations

  • Query optimization method and device based on data lake and storage medium

    CN117667998A

  • Management service platform system based on big data

    CN120723969A