Data acceleration processing method and system for optimizing cross-domain data weaving computing efficiency

By adopting distributed data acquisition architecture, data compression, data acceleration and intelligent scheduling technologies in cross-domain data weaving, the problems of low computing efficiency and difficult real-time real-time in traditional technologies are solved, and efficient and real-time cross-domain data processing and analysis are achieved.

CN119046019BActive Publication Date: 2025-05-23INSPUR SOFTWARE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411534312.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-31
Publication Date
2025-05-23
Estimated Expiration
2044-10-31

AI Technical Summary

Technical Problem

Traditional cross-domain data weaving solutions rely on centralized computing architectures, resulting in low data processing efficiency and difficult to meet the real-time processing requirements of large-scale data sets, especially in terms of resource heterogeneity, data transmission delay, computing resource consumption, and computing task scheduling complexity.

Method used

Adopting a distributed data acquisition architecture, data compressor, data acceleration engine and intelligent scheduling model, the scheduling strategy of computing tasks is dynamically adjusted to improve computing efficiency by optimizing TaskManager parameters, configuring data acquisition strategies, adopting Gzip data compression algorithms, utilizing data cache views and adaptive materialization acceleration, and an intelligent scheduling model based on machine learning.

Benefits of technology

It significantly improves the computing efficiency and data transmission speed in the cross-domain data weaving process, realizes real-time and accuracy of data processing, optimizes the allocation and utilization of computing resources, and enhances the flexibility and scalability of data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119046019B_ABST
    Figure CN119046019B_ABST
Patent Text Reader

Abstract

The present invention discloses a data acceleration processing method and system for optimizing the computing efficiency of cross-domain data weaving, relates to the technical field of big data processing and analysis, and adopts a scheme to solve the problem of low computing efficiency of current cross-domain data weaving: construct a distributed data acquisition architecture in the data acquisition link to achieve timed or triggered acquisition of raw data by optimizing parameters and configuring data acquisition strategies of each node; build a data compressor in the data preprocessing link to compress the collected raw data and output it to the data weaving engine; construct a data acceleration engine in the data collaborative computing link to optimize the storage and access path of data in the data weaving engine; construct an intelligent scheduling model in the data resource scheduling link to dynamically adjust the scheduling strategy of computing tasks according to the priority, resource availability and network bandwidth of computing tasks, and generate a scheduling scheme. The present invention can improve the computing efficiency and transmission speed of the cross-domain data weaving process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of big data processing and analysis, and specifically to a data acceleration processing method and system for optimizing cross-domain data weaving computing efficiency. Background Art

[0002] With the rapid development of information technology, the amount of data accumulated by enterprises has increased exponentially, and the data types have also expanded from traditional structured data to semi-structured and unstructured data, forming a complex and diverse data structure system. This surge in data volume and diversification of structure not only brings unprecedented data value mining potential to enterprises, but also brings challenges such as data silos, data redundancy, and uneven data quality.

[0003] At the business level, enterprises have an increasingly urgent need for cross-domain data integration and sharing. Data from different departments, different business lines, and even different geographical locations need to be efficiently integrated to support more accurate business decisions, more efficient operations management, and more personalized customer service. However, traditional data integration methods, such as ETL data processes, are unable to cope with the complexity of cross-domain data and it is difficult to ensure the real-time, consistency, and security of data.

[0004] At the same time, the rise of technologies such as cloud computing, big data, and artificial intelligence has provided strong technical support for the birth of cross-domain data weaving technology. Cloud computing provides flexible and scalable computing and storage resources, making cross-domain data processing and analysis possible; big data technology provides efficient data processing and analysis tools to help companies extract valuable information from massive data; artificial intelligence technology, especially machine learning and deep learning technology, can automatically process and analyze data, improving the efficiency and accuracy of data processing. Against this background, cross-domain data weaving technology came into being.

[0005] Cross-domain data weaving is a data management strategy that aims to integrate, transform and consolidate data resources from different fields (such as different data centers, cloud environments or geographical locations) to form a unified, consistent and high-quality data view. This data view is essential for supporting complex data analysis, decision making, and data mining applications. This cross-domain data weaving breaks the limitations of traditional data centralized storage, not only solves the problem of data silos, improves data availability and ease of use, but also ensures data quality and security through the implementation of data governance and security policies.

[0006] However, traditional cross-domain data weaving solutions often rely on centralized computing architectures, which have low data processing efficiency and are difficult to cope with the real-time processing requirements of large-scale data sets. The core issue is the cross-domain data collaborative computing efficiency problem. Since cross-domain data usually has the characteristics of large data volume, diverse data formats, and wide data distribution, traditional centralized data processing methods often cannot meet the requirements of real-time and high efficiency. Specifically, the computing efficiency problems of cross-domain data weaving are mainly reflected in the following aspects:

[0007] 1. Resource heterogeneity. In cross-domain data processing, the resources in each data center are often heterogeneous, that is, the computing resources, storage resources, and network resources of different data centers are different. This resource heterogeneity poses a huge challenge to the computing efficiency of cross-domain data weaving.

[0008] 2. Data transmission delay. Cross-domain data weaving involves data transmission and synchronization between multiple data centers. Due to the limitations of network bandwidth and latency, data transmission and synchronization often become key factors affecting computing performance.

[0009] 3. Computing resource consumption. Cross-domain data weaving involves a large number of data processing tasks, including data cleaning, conversion, and integration. These tasks consume a lot of computing resources, such as CPU, memory, and disk space.

[0010] 4. Computational task scheduling complexity. In cross-domain data weaving, computing tasks are usually distributed to multiple data centers for parallel processing. However, due to the heterogeneity of resources and differences in network conditions in each data center, the scheduling of computing tasks becomes very complicated.

[0011] Therefore, the computational efficiency of data weaving is one of the important challenges facing the current data management and integration field. Based on the above existing problems, and with the explosive growth of data volume and the increase in real-time data processing needs, there is an urgent need for an efficient and scalable data acceleration processing system and method to solve the computational efficiency problem of cross-domain data weaving. Summary of the invention

[0012] Traditional data weaving focuses more on the realization of data processing flow, but does not pay attention to the problem of low parallel computing performance in the scenario of large-scale parallel processing of large amounts of data. Since data weaving needs to connect, integrate and manage data distributed in different systems and databases, when the amount of data is large or the data type is complex, the computing performance may be affected, resulting in slower data processing speed, and even resource bottlenecks may occur, affecting the stability and response speed of the overall system. Based on this, the present invention provides a data acceleration processing method and system for optimizing the computing efficiency of cross-domain data weaving.

[0013] In the first aspect, the present invention provides a data acceleration processing method for optimizing the computing efficiency of cross-domain data weaving. The technical solution adopted to solve the above technical problems is as follows:

[0014] A data acceleration processing method for optimizing cross-domain data weaving computing efficiency comprises the following steps:

[0015] S1. In the data collection phase, a distributed data collection architecture is constructed. The distributed data collection architecture adopts a distributed FLink cluster. By optimizing TaskManager parameters and configuring data collection strategies for each node, the timed or triggered collection of raw data is realized.

[0016] S2. In the data preprocessing phase, a data compressor is built. The data compressor uses Gzip as the data compression algorithm to compress the original data collected by the Flink cluster and output the compressed data to the Trino engine.

[0017] S3. In the data collaborative computing phase, a data acceleration engine is constructed. The data acceleration engine uses two methods, data cache view and adaptive materialized acceleration, to optimize the storage and access path of data in the Trino engine and improve the efficiency of data collaborative computing.

[0018] S4. In the data resource scheduling link, an intelligent scheduling model is constructed. The intelligent scheduling model dynamically adjusts the scheduling strategy of the computing task according to the priority of the computing task, resource availability and network bandwidth, generates a scheduling plan, and then converts the generated scheduling plan into specific execution instructions. The instructions are transmitted to the corresponding computing resource pool for collaborative computing processing, so as to achieve reasonable allocation and efficient utilization of computing resources.

[0019] Optionally, step S1 is performed, wherein the distributed data collection architecture optimizes TaskManager parameters and configures data collection strategies of each node to achieve timed or triggered collection of raw data. This process specifically includes:

[0020] First, calculate the number of nodes for the Flink cluster based on the amount of data per second, the size of each piece of data, the number of deduplicated keys, and the state size of each key.

[0021] Secondly, configure the number of TaskManager tasks, decompose the overall data collection task according to the number of TaskManagers, and decompose it into multiple subtasks executed in parallel. The decomposed subtasks will be assigned to different collection nodes, and each node is responsible for completing a part of the collection task. The number of TaskManagers is calculated as follows: num_of_tm = ceil(parallelism / slot), where parallelism indicates the degree of parallelism and slot indicates the number of slots of each TaskManager.

[0022] Finally, configure the data collection strategy for each node to perform timed or triggered collection of the original data.

[0023] Optionally, execute step S2, where the data compressor uses Gzip as the data compression algorithm to compress the original data collected by the Flink cluster. After the data compression is completed, the data size before and after compression, decompression speed and data quality are compared as evaluation indicators to evaluate the compression effect; according to the evaluation result, the compression parameters of the data compression algorithm Gzip are adjusted, and the original data is compressed again. When the evaluation result meets the expectation, the compressed data is output to the Trino engine.

[0024] Optionally, step S3 is performed, wherein the data acceleration engine uses two methods, data cache view and adaptive materialized acceleration, to optimize the storage and access path of data in the Trino engine and improve the efficiency of data collaborative computing, wherein:

[0025] The data cache view method uses Hive as the data cache library, and temporarily stores the data with usage frequency exceeding the set threshold and across tables in the data cache library in the form of data view;

[0026] The adaptive materialization acceleration method materializes data whose usage frequency exceeds a set threshold to form a materialized view.

[0027] Optionally, execute step S4, where the intelligent scheduling model uses an intelligent algorithm of machine learning or deep learning to optimize the scheduling plan. The intelligent algorithm analyzes historical data, predicts future resource requirements, and evaluates the effectiveness of the generated scheduling plan, and outputs a scheduling plan that meets expectations.

[0028] In a second aspect, the present invention provides a data acceleration processing system for optimizing the computing efficiency of cross-domain data weaving. The technical solution adopted to solve the above technical problems is as follows:

[0029] A data acceleration processing system for optimizing cross-domain data weaving computing efficiency, comprising:

[0030] The distributed data collection architecture used in the data collection link is used for the FLink cluster with a distributed architecture. By optimizing the TaskManager parameters and configuring the data collection strategy of each node, the timed or triggered collection of the original data is realized;

[0031] The data compressor used in the data preprocessing phase uses Gzip as the data compression algorithm to compress the original data collected by the Flink cluster and output the compressed data to the Trino engine;

[0032] The data acceleration engine used in data collaborative computing uses data cache view and adaptive materialized acceleration to optimize the storage and access path of data in the Trino engine and improve the efficiency of data collaborative computing.

[0033] The intelligent scheduling model applied in the data resource scheduling link is used to dynamically adjust the scheduling strategy of computing tasks according to the priority of computing tasks, resource availability and network bandwidth, generate scheduling plans, and then convert the generated scheduling plans into specific execution instructions. The instructions are transmitted to the corresponding computing resource pool for collaborative computing processing, so as to achieve reasonable allocation and efficient utilization of computing resources.

[0034] Optionally, the distributed data collection architecture involved specifically includes:

[0035] The node calculation module is used to calculate the number of nodes of the Flink cluster scale based on the amount of data per second, the size of each data item, the number of deduplicated keys, and the state size of each key;

[0036] The configuration decomposition module is used to configure the number of TaskManager tasks. The overall data collection task is decomposed according to the number of TaskManagers, and it is decomposed into multiple subtasks executed in parallel. The decomposed subtasks will be assigned to different collection nodes. Each node is responsible for completing a part of the collection task. The number of TaskManagers is calculated as follows: num_of_tm = ceil(parallelism / slot), where parallelism indicates the degree of parallelism and slot indicates the number of slots of each TaskManager.

[0037] The strategy configuration module is used to configure the data collection strategy of each node and perform timed or triggered collection of raw data.

[0038] Optionally, the data compressor involved uses Gzip as the data compression algorithm to compress the original data collected by the Flink cluster. After the data compression is completed, the data size before and after compression, decompression speed and data quality are compared as evaluation indicators to evaluate the compression effect; according to the evaluation result, the compression parameters of the data compression algorithm Gzip are adjusted, and the original data is compressed again. When the evaluation result meets the expectation, the compressed data is output to the Trino engine.

[0039] Optionally, the data acceleration engine involved uses two methods, data cache view and adaptive materialized acceleration, to optimize the storage and access path of data in the Trino engine and improve the efficiency of data collaborative computing, including:

[0040] The data cache view method uses Hive as the data cache library, and temporarily stores the data with usage frequency exceeding the set threshold and across tables in the data cache library in the form of data view;

[0041] The adaptive materialization acceleration method materializes data whose usage frequency exceeds a set threshold to form a materialized view.

[0042] Optionally, the intelligent scheduling model involved uses intelligent algorithms of machine learning or deep learning to optimize the scheduling plan. The intelligent algorithm analyzes historical data, predicts future resource requirements, evaluates the effectiveness of the generated scheduling plan, and outputs a scheduling plan that meets expectations.

[0043] The data acceleration processing method and system of the present invention for optimizing cross-domain data weaving computing efficiency have the following beneficial effects compared with the prior art:

[0044] 1. The present invention integrates distributed computing, data compression, data acceleration and intelligent data scheduling technologies to effectively improve the computing efficiency and data transmission speed in the cross-domain data weaving process; by optimizing the data processing process, it reduces data transmission delay and computing resource consumption, and achieves a dual improvement in the real-time and accuracy of cross-domain data weaving;

[0045] 2. The present invention can speed up data processing and realize efficient data flow from collection to analysis; optimize the allocation and utilization of computing resources to avoid resource waste and overload; enhance the flexibility and scalability of data processing to meet the growing demand for data processing, aiming to provide enterprises and organizations with an efficient, secure and easy-to-use cross-domain data weaving solution to help them with digital transformation and intelligent upgrading;

[0046] 3. The present invention brings significant performance improvement to data weaving calculations by optimizing data processing procedures, enhancing computing capabilities and improving data transmission efficiency. It can greatly shorten the time for data preparation and analysis, making the processing of large-scale data sets faster and more efficient. At the same time, it also enhances the capabilities of parallel computing and distributed processing, further improving the execution speed of complex computing tasks. These optimizations enable data weaving calculations to gain insight into data value more quickly and provide timely and accurate support for decision-making, thereby gaining an advantage in a highly competitive market environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Attached Figure 1 is a flow chart for implementing the method described in the embodiment of the present invention;

[0048] Attached Figure 2 It is an implementation architecture diagram of the method or system described in the embodiment of the present invention. DETAILED DESCRIPTION

[0049] In order to make the technical solution, the technical problem solved and the technical effect of the present invention more clearly understood, the technical solution of the present invention is clearly and completely described below in conjunction with specific embodiments.

[0050] Embodiment 1: Combining with the attached Figure 1 This embodiment proposes a data acceleration processing method for optimizing cross-domain data weaving computing efficiency, including the following steps:

[0051] S1. In the data collection phase, a distributed data collection architecture (DDC) is constructed. The distributed data collection architecture (DDC) adopts a distributed FLink cluster. By optimizing TaskManager parameters and configuring the data collection strategy of each node, the timed or triggered collection of raw data is realized. This process specifically includes:

[0052] First, calculate the number of nodes for the Flink cluster based on the amount of data per second, the size of each piece of data, the number of deduplicated keys, and the state size of each key.

[0053] Secondly, configure the number of TaskManager tasks, decompose the overall data collection task according to the number of TaskManagers, and decompose it into multiple subtasks executed in parallel. The decomposed subtasks will be assigned to different collection nodes, and each node is responsible for completing a part of the collection task. The number of TaskManagers is calculated as follows: num_of_tm = ceil(parallelism / slot), where parallelism indicates the degree of parallelism and slot indicates the number of slots of each TaskManager.

[0054] Finally, configure the data collection strategy for each node to perform timed or triggered collection of the original data.

[0055] S2. In the data preprocessing stage, a data compressor (DCR) is built. The data compressor (DCR) uses Gzip as the data compression algorithm to compress the original data collected by the Flink cluster and output the compressed data to the Trino engine.

[0056] It should be added that after the data compression is completed, the data size before and after compression, decompression speed and data quality are compared as evaluation indicators to evaluate the compression effect. According to the evaluation results, the compression parameters of the data compression algorithm Gzip are adjusted, and the original data is compressed again. When the evaluation results meet expectations, the compressed data is output to the Trino engine.

[0057] S3. In the data collaborative computing phase, a data acceleration engine (DAE) is constructed. The data acceleration engine (DAE) uses two methods, data cache view (D-dcv) and adaptive materialized acceleration (D-pca), to optimize the storage and access path of data in the Trino engine and improve the efficiency of data collaborative computing.

[0058] The data cache view (D-dcv) method uses Hive as the data cache library, and temporarily stores data that exceeds the set threshold and crosses tables in the data cache library in the form of data views, facilitating fast access during data collaborative computing.

[0059] The adaptive materialization acceleration (D-pca) method materializes data whose usage frequency exceeds a set threshold to form a materialized view; when the same query appears again, the result can be obtained directly from the materialized view without repeatedly accessing the original data, thereby significantly improving the query response speed.

[0060] S4. In the data resource scheduling link, an intelligent scheduling model (ISM) is constructed. The intelligent scheduling model (ISM) dynamically adjusts the scheduling strategy of computing tasks according to the priority, resource availability and network bandwidth of computing tasks, generates a scheduling plan, and then converts the generated scheduling plan into specific execution instructions, and transmits them to the corresponding computing resource pool for collaborative computing processing, so as to achieve reasonable allocation and efficient utilization of computing resources.

[0061] It should be added that the intelligent scheduling model (ISM) uses intelligent algorithms of machine learning or deep learning to optimize scheduling plans. The intelligent algorithms analyze historical data, predict future resource requirements, evaluate the effectiveness of the generated scheduling plans, and output scheduling plans that meet expectations.

[0062] Embodiment 2: Combining with Figure 2 This embodiment proposes a data acceleration processing system for optimizing cross-domain data weaving computing performance, which includes:

[0063] The distributed data collection architecture (DDC) used in the data collection link is used for the FLink cluster with a distributed architecture. By optimizing the TaskManager parameters and configuring the data collection strategy of each node, the timed or triggered collection of raw data can be realized;

[0064] The data compressor (DCR) used in the data preprocessing phase uses Gzip as the data compression algorithm to compress the original data collected by the Flink cluster and output the compressed data to the Trino engine;

[0065] The data acceleration engine (DAE) used in the data collaborative computing link is used to optimize the storage and access path of data in the Trino engine and improve the efficiency of data collaborative computing by using two methods: data cache view (D-dcv) and adaptive materialization acceleration (D-pca). Among them: the data cache view (D-dcv) method uses Hive as the data cache library, and temporarily stores the data with a usage frequency exceeding the set threshold and across tables in the data cache library in the form of data views, which is convenient for fast access during data collaborative computing; the adaptive materialization acceleration (D-pca) method materializes the data with a usage frequency exceeding the set threshold to form a materialized view; when the same query appears again, the result can be obtained directly from the materialized view without repeatedly accessing the original data, thereby significantly improving the query response speed;

[0066] The intelligent scheduling model (ISM) applied in the data resource scheduling link is used to dynamically adjust the scheduling strategy of computing tasks according to the priority, resource availability and network bandwidth of computing tasks, generate scheduling plans, and then convert the generated scheduling plans into specific execution instructions. The instructions are transmitted to the corresponding computing resource pool for collaborative computing processing, so as to achieve reasonable allocation and efficient utilization of computing resources.

[0067] In this embodiment, the distributed data collection architecture (DDC) involved specifically includes:

[0068] The node calculation module is used to calculate the number of nodes of the Flink cluster scale based on the amount of data per second, the size of each data item, the number of deduplicated keys, and the state size of each key;

[0069] The configuration decomposition module is used to configure the number of TaskManager tasks. The overall data collection task is decomposed according to the number of TaskManagers, and it is decomposed into multiple subtasks executed in parallel. The decomposed subtasks will be assigned to different collection nodes. Each node is responsible for completing a part of the collection task. The number of TaskManagers is calculated as follows: num_of_tm = ceil(parallelism / slot), where parallelism indicates the degree of parallelism and slot indicates the number of slots of each TaskManager.

[0070] The strategy configuration module is used to configure the data collection strategy of each node and perform timed or triggered collection of raw data.

[0071] In this embodiment, the data compressor (DCR) involved uses Gzip as the data compression algorithm to compress the original data collected by the Flink cluster. After the data compression is completed, the data size before and after compression, decompression speed and data quality are compared as evaluation indicators to evaluate the compression effect; according to the evaluation result, the compression parameters of the data compression algorithm Gzip are adjusted, and the original data is compressed again. When the evaluation result meets the expectation, the compressed data is output to the Trino engine.

[0072] In this embodiment, the intelligent scheduling model (ISM) involved uses intelligent algorithms of machine learning or deep learning to optimize the scheduling plan. The intelligent algorithm analyzes historical data, predicts future resource requirements, evaluates the effectiveness of the generated scheduling plan, and outputs a scheduling plan that meets expectations.

[0073] In summary, the data acceleration processing method and system of the present invention for optimizing the computing efficiency of cross-domain data weaving effectively improves the computing efficiency and data transmission speed in the cross-domain data weaving process by integrating distributed computing, data compression, data acceleration and intelligent data scheduling technology, and solves the technical defects of traditional data weaving that cause the data processing speed to slow down when the data volume is large or the data type is complex, and even cause resource bottlenecks, affecting the stability and response speed of the overall system.

[0074] The above specific examples are used to explain the principles and implementation methods of the present invention in detail. These examples are only used to help understand the core technical content of the present invention. Based on the above specific embodiments of the present invention, any improvements and modifications made by technicians in this technical field without departing from the principles of the present invention should fall within the scope of patent protection of the present invention.

Claims

1. A data acceleration processing method for optimizing cross-domain data weaving computing performance, characterized in that: The steps include: S1. In the data collection phase, a distributed data collection architecture is constructed. The distributed data collection architecture adopts a distributed FLink cluster. By optimizing TaskManager parameters and configuring the data collection strategy of each node, the timed collection or triggered collection of raw data is realized. This process specifically includes: first, the number of nodes of the Flink cluster scale is calculated by the amount of data per second, the size of each data, the number of deduplicated keys, and the state size of each key; second, the number of TaskManager tasks is configured, and the overall data collection task is decomposed according to the number of TaskManagers, and it is decomposed into multiple subtasks executed in parallel. The decomposed subtasks will be assigned to different collection nodes, and each node is responsible for completing a part of the collection task; the number of TaskManagers is calculated as follows: num_of_tm = ceil(parallelism / slot), parallelism represents the degree of parallelism, and slot represents the number of slots of each TaskManager; finally, the data collection strategy of each node is configured to perform timed collection or triggered collection of raw data; S2. In the data preprocessing phase, a data compressor is built. The data compressor uses Gzip as the data compression algorithm to compress the original data collected by the Flink cluster and output the compressed data to the Trino engine. S3. In the data collaborative computing phase, a data acceleration engine is constructed. The data acceleration engine uses two methods, data cache view and adaptive materialized acceleration, to optimize the storage and access path of data in the Trino engine and improve the efficiency of data collaborative computing. S4. In the data resource scheduling link, an intelligent scheduling model is constructed. The intelligent scheduling model dynamically adjusts the scheduling strategy of the computing task according to the priority of the computing task, resource availability and network bandwidth, generates a scheduling plan, and then converts the generated scheduling plan into specific execution instructions. The instructions are transmitted to the corresponding computing resource pool for collaborative computing processing, so as to achieve reasonable allocation and efficient utilization of computing resources.

2. The data acceleration processing method for optimizing cross-domain data weaving computing efficiency according to claim 1 is characterized in that: Execute step S2, the data compressor uses Gzip as the data compression algorithm to compress the original data collected by the Flink cluster. After the data compression is completed, the data size before and after compression, the decompression speed and the data quality are compared as evaluation indicators to evaluate the compression effect; according to the evaluation result, the compression parameters of the data compression algorithm Gzip are adjusted, and the original data is compressed again. When the evaluation result meets the expectation, the compressed data is output to the Trino engine.

3. The data acceleration processing method for optimizing cross-domain data weaving computing efficiency according to claim 1 is characterized in that: Execute step S3, the data acceleration engine uses two methods, data cache view and adaptive materialized acceleration, to optimize the storage and access path of data in the Trino engine and improve the efficiency of data collaborative computing, wherein: The data cache view method uses Hive as the data cache library, and temporarily stores the data with usage frequency exceeding the set threshold and across tables in the data cache library in the form of data view; The adaptive materialization acceleration method materializes data whose usage frequency exceeds a set threshold to form a materialized view.

4. The data acceleration processing method for optimizing cross-domain data weaving computing efficiency according to claim 1 is characterized in that: Execute step S4, the intelligent scheduling model uses machine learning or deep learning intelligent algorithms to optimize the scheduling plan. The intelligent algorithm analyzes historical data, predicts future resource requirements, and evaluates the effectiveness of the generated scheduling plan, and outputs a scheduling plan that meets expectations.

5. A data acceleration processing system for optimizing cross-domain data weaving computing performance, characterized in that: It includes: The distributed data collection architecture applied in the data collection link is used for the FLink cluster adopting the distributed architecture, and realizes the timed collection or triggered collection of the original data by optimizing the TaskManager parameters and configuring the data collection strategy of each node; wherein, the distributed data collection architecture specifically includes: a node calculation module, which is used to calculate the number of nodes of the Flink cluster scale by the amount of data per second, the size of each data, the number of deduplicated keys and the state size of each key; a configuration decomposition module, which is used to configure the number of TaskManager tasks, decompose the overall data collection task according to the number of TaskManagers, and decompose it into multiple subtasks executed in parallel. The decomposed subtasks will be assigned to different collection nodes, and each node is responsible for completing a part of the collection task, wherein the number of TaskManagers is calculated as follows: num_of_tm = ceil(parallelism / slot), parallelism represents the degree of parallelism, and slot represents the number of slots of each TaskManager; a strategy configuration module, which is used to configure the data collection strategy of each node, and perform timed collection or triggered collection of the original data; The data compressor used in the data preprocessing phase uses Gzip as the data compression algorithm to compress the original data collected by the Flink cluster and output the compressed data to the Trino engine; The data acceleration engine used in data collaborative computing uses data cache view and adaptive materialized acceleration to optimize the storage and access path of data in the Trino engine and improve the efficiency of data collaborative computing. The intelligent scheduling model applied in the data resource scheduling link is used to dynamically adjust the scheduling strategy of computing tasks according to the priority of computing tasks, resource availability and network bandwidth, generate scheduling plans, and then convert the generated scheduling plans into specific execution instructions. The instructions are transmitted to the corresponding computing resource pool for collaborative computing processing, so as to achieve reasonable allocation and efficient utilization of computing resources.

6. The data acceleration processing system for optimizing cross-domain data weaving computing efficiency according to claim 5, characterized in that: The data compressor uses Gzip as the data compression algorithm to compress the original data collected by the Flink cluster. After the data compression is completed, the data size before and after compression, the decompression speed and the data quality are compared as evaluation indicators to evaluate the compression effect; according to the evaluation result, the compression parameters of the data compression algorithm Gzip are adjusted, and the original data is compressed again. When the evaluation result meets the expectation, the compressed data is output to the Trino engine.

7. The data acceleration processing system for optimizing cross-domain data weaving computing efficiency according to claim 5, characterized in that: The data acceleration engine uses two methods, data cache view and adaptive materialized acceleration, to optimize the storage and access path of data in the Trino engine and improve the efficiency of data collaborative computing, including: The data cache view method uses Hive as the data cache library, and temporarily stores the data with usage frequency exceeding the set threshold and across tables in the data cache library in the form of data view; The adaptive materialization acceleration method materializes data whose usage frequency exceeds a set threshold to form a materialized view.

8. The data acceleration processing system for optimizing cross-domain data weaving computing efficiency according to claim 5, characterized in that: The intelligent scheduling model uses machine learning or deep learning intelligent algorithms to optimize scheduling plans. The intelligent algorithms analyze historical data, predict future resource requirements, evaluate the effects of generated scheduling plans, and output scheduling plans that meet expectations.

Citation Information

Patent Citations

  • Intelligent analysis platform design method based on big data framework

    CN113347170A

  • Method for solving joint retrieval of cross-domain heterogeneous data

    CN113886457A