Intelligent data splitting and optimizing method and system based on multiple dimensions

By adopting intelligent data splitting and optimization methods in multi-dimensional analysis, the problems of data bloating, slow computing speed and data tilt are solved, efficient data processing and analysis are achieved, and the needs of real-time business analysis are met.

CN120045613APending Publication Date: 2025-05-27BEIJING BAIJU YIXING TECH CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510114857.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing technology faces the problems of data inflation, slow computing speed and data skew in multi-dimensional analysis, and it is difficult to effectively deal with the real-time analysis needs of large-scale data.

Method used

Using multi-dimensional intelligent data splitting and optimization methods, intelligent data splitting and optimization are achieved by connecting to data sources, using ETL tools for data cleaning, task splitting strategies, dynamic scheduling and resource optimization, combined with the distributed computing power and query optimizer of the ODPS platform, intelligent data splitting and optimization are achieved.

Benefits of technology

It significantly improves computing efficiency, effectively solves the problems of data bloating and data tilt, improves the speed of data processing and analysis, and meets the needs of modern enterprises for real-time business analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045613A_ABST
    Figure CN120045613A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent data splitting and optimizing method and system based on multiple dimensions. The invention relates to the technical field of data splitting. And dividing the number and size of the sub-tasks based on dimensions and metrics required to be split, and dynamically scheduling the dependency relationship and the execution sequence among the sub-tasks. Based on the execution efficiency of the dimension combination and the data skew condition, dynamically adjusting the splitting strategy of the dimension combination, for example, increasing the number of subtasks of hot dimension combinations, optimizing the processing logic of the dimension combination with serious data skew, and the like; through a task splitting and dynamic scheduling mechanism, a large task is decomposed into a plurality of small tasks to be executed in parallel, distributed computing resources are fully utilized, and data processing and analysis time is remarkably shortened. By optimizing the data storage structure and utilizing the ETL tool to clean and compress the data, the occupied physical storage space is reduced, the storage cost is reduced, and meanwhile, the data access speed is increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data splitting, specifically to data splitting for multi-dimensional data sources using HiveCube, and particularly to a method and system for intelligent data splitting and optimization based on multiple dimensions. Background Art

[0002] With the continuous evolution of data warehouse technology, multi-dimensional analysis technology has become an important research direction. This technology allows users to perform operations such as slicing, dicing, rotating, and drilling down on data from multiple dimensions, thereby deeply revealing the internal relationships and laws of the data. However, the implementation of multi-dimensional analysis often involves the construction of complex data models and query statements, which poses high requirements on users' skills and is difficult to effectively meet the real-time analysis needs of large-scale data.

[0003] To solve the above problems, HiveCube emerged as an innovative data warehouse solution, which cleverly combines Hive and Cube technologies. HiveCube makes full use of the advantages of Hive in big data processing and can support the storage and query of large-scale data. At the same time, by introducing Cube technology, HiveCube realizes pre-computation of data and dimension combination, thereby significantly improving the aggregation and query speed of data and enhancing the overall efficiency and performance of data processing.

[0004] In addition, HiveCube also provides self-service data analysis functions. Users can build custom analysis views through simple drag-and-drop operations without writing complex query statements. This feature greatly reduces the threshold of data analysis and enables more users to participate in data analysis. Nevertheless, HiveCube still faces the following three main technical challenges in practical applications:

[0005] (1) Data inflation problem: As the number of dimensions increases, the data inflation problem becomes more and more serious. A large number of dimension combinations lead to data redundancy and waste of storage space, and at the same time increase the complexity of queries and computational burden. For example, in a data set containing 10 dimensions, if each dimension has 10 possible values, the number of possible combinations will reach 10^10, which is a huge number far beyond the processing capacity of traditional data warehouses.

[0006] (2) Slow calculation speed: Due to the complexity of data inflation and dimension combination, existing technologies often consume a large amount of computing resources and time when performing multi-dimensional analysis. For example, for a real-time data analysis task containing a large number of dimensions, a traditional data warehouse may take hours or even days to complete the calculation, which cannot meet the needs of modern enterprises for real-time business analysis.

[0007] (3) Data skew problem: In a distributed system, data skew is a common problem. When the amount of data on some nodes is much larger than that on other nodes, these nodes may become performance bottlenecks, leading to a decline in the performance of the entire system. This problem is particularly prominent in multidimensional analysis because different dimension combinations may cause extremely uneven distribution of data among nodes.

[0008] To this end, the present invention proposes a data intelligent splitting and optimization method and system based on multiple dimensions. Summary of the Invention

[0009] In view of this, the present invention hopes to provide a data intelligent splitting and optimization method and system based on multiple dimensions to solve or alleviate the technical problems existing in the prior art, that is, how to solve the data expansion problem and data skew temperature, and improve the calculation speed, and at least provide a beneficial choice for this; the technical solution of the present invention is implemented as follows:

[0010] In the first aspect, a data intelligent splitting and optimization method based on multiple dimensions:

[0011] (1) Overview:

[0012] Aiming at the data expansion, data skew and calculation speed problems in multidimensional analysis, the present invention first docks with the data source to ensure the comprehensiveness and accuracy of the data. Secondly, data storage and management are carried out on the ODPS platform side, and ETL tools are used for data cleaning to ensure the cleanliness of the data. Then, through the task splitting strategy, sub-tasks are reasonably divided, and the execution order and dependency relationship of the sub-tasks are dynamically scheduled to improve the calculation efficiency. At the same time, according to the execution efficiency of the dimension combination and the data skew situation, the splitting strategy is dynamically adjusted to further optimize the computing resources. Finally, the ODPS scheduling system is used to manage the task execution, output the processed data, and dynamically adjust the data storage structure according to the system performance monitoring results to achieve continuous optimization.

[0013] (2) Technical solution:

[0014] Dock with the data source, obtain its internal database, external API and log files, and after determining the dimensions and metrics to be split, perform the following operation steps.

[0015] 2.1 Step S1, data engineering:

[0016] Based on the data tables or partitions created on the ODPS platform side, read the original data and intermediate results stored therein;

[0017] 2.1.1 Step S100, data cleaning:

[0018] Use the ETL tool provided by the ODPS platform to preprocess the read raw data, including data cleaning, that is, removing duplicate, incorrect, and irrelevant data to ensure data accuracy and consistency. The ETL tool will perform a series of data transformation and verification operations to meet the requirements of subsequent analysis.

[0019] 2.1.2 Step S101, Data storage:

[0020] Store the cleaned data in a specified location on the ODPS platform for subsequent data processing and analysis steps to use.

[0021] 2.2 Step S2, Task splitting:

[0022] Divide the number and size of subtasks, and dynamically schedule the dependencies and execution order between subtasks.

[0023] 2.2.1 Step S200, Subtask division:

[0024] Based on the dimensions and metrics to be split, divide the number and size of subtasks for parallel processing to improve execution efficiency.

[0025] 2.2.2 Step S201, Adjustment of dimension combination splitting strategy:

[0026] Increase the number of subtasks for popular dimension combinations to disperse the computing pressure, aiming to optimize the execution efficiency and resource utilization of tasks.

[0027] 2.3 Step S3, Task scheduling:

[0028] Generate an optimal query plan, use the scheduling system on the ODPS platform to manage the execution, aggregation, and integration of subtasks, and output data;

[0029] 2.3.1 Step S300, Query plan generation:

[0030] Use the query optimizer on the ODPS platform to automatically generate an optimal query plan according to the input query statement and data statistical information, including parsing, optimizing, and rewriting the query statement.

[0031] 2.3.2 Step S301, Subtask management and scheduling:

[0032] Use the DAG task scheduling technology of the scheduling system (MaxCompute) on the ODPS platform to manage the execution, dependencies, resource allocation, and priorities of subtasks, ensuring that subtasks can be executed in the optimal order and manner to achieve efficient resource utilization and task execution.

[0033] 2.3.3 Step S302, Data aggregation and integration:

[0034] Utilize DAG task scheduling technology to merge, sort, and deduplicate data, summarize and integrate the data output by each subtask to ensure data integrity and consistency, and generate the final result dataset.

[0035] 2.3.4 Step S303, output result data:

[0036] Interact with the ODPS platform side, output the summarized and integrated data to the specified storage location or data table to ensure that the data can be correctly stored and accessed.

[0037] 2.4 Step S4, data output:

[0038] According to the data processing speed and resource utilization information of each fixed time period, dynamically adjust the data storage structure. The strategy for dynamically adjusting the data storage structure is as follows: for the case of joining a large table with a small table, use map join instead of reduce join to avoid data skew caused by shuffle operations; vice versa.

[0039] (III) Mechanism for solving technical problems:

[0040] 3.1 Mechanism for solving the data bloat problem:

[0041] Data bloat refers to the situation where the size of the physical data file is significantly higher than the actual amount of data stored. In the ODPS (OpenData Processing Service) environment, the data bloat problem is usually related to the data storage mechanism and data operation methods.

[0042] The present invention cleans the original data through an ETL tool, removes invalid and redundant data, and reduces unnecessary data storage.

[0043] 3.2 Principle of the mechanism for solving the data skew problem:

[0044] Data skew means that in distributed computing, a large number of identical keys are distributed to the same node, resulting in the amount of data processed by this node being much larger than that of other nodes, thus slowing down the overall computing speed.

[0045] The present invention analyzes the original query task, splits it according to dimensions and metrics, and decomposes the large task into multiple small tasks for parallel execution. According to the execution results of historical data and log information, dynamically adjust the splitting strategy of dimension combinations, such as increasing the number of subtasks for popular dimension combinations and optimizing the processing logic for dimension combinations with serious data skew. During the task scheduling process, distribute tasks to different computing nodes through a load balancing algorithm to avoid a single node undertaking too many tasks.

[0046] 3.3 Mechanism Principle for Improving Computational Speed:

[0047] The present invention utilizes the distributed computing ability of ODPS to split large tasks into multiple small tasks for parallel execution, thereby improving the overall computational speed. The query optimizer of ODPS will automatically generate an optimal query plan to reduce unnecessary computations and data transmissions.

[0048] In a second aspect, a multi-dimensional data intelligent splitting and optimization system:

[0049] This splitting and optimization system is used to execute the splitting and optimization method as described above, including:

[0050] (1) An application end for user-system interaction: responsible for receiving user query requests, presenting processing results, and providing corresponding interaction functions. Key components include:

[0051] (1.1) A user interface: allows users to input query conditions, view processing results, etc.

[0052] (1.2) Interaction tools: such as API interfaces, which facilitate the integration of the application with the system to achieve automated queries and data acquisition.

[0053] (2) A server end for receiving requests from the application end, invoking services on the ODPS platform end, executing data processing and analysis tasks, and returning the results to the application end; key components include:

[0054] (2.1) A request receiving module for receiving query requests from the application end;

[0055] (2.2) A task scheduling module for scheduling resources on the ODPS platform end: according to the query request, schedule resources on the ODPS platform end to execute corresponding data processing and analysis tasks. This includes task splitting, dynamically scheduling dependencies and execution orders between subtasks, etc.

[0056] (2.3) A result processing module for collecting processing results returned by the ODPS platform end: perform necessary processing and integration to generate a final result set and return it to the application end.

[0057] (3) The ODPS platform end for supporting large-scale distributed computing; its key components include:

[0058] (3.1) A data storage module for storing original data and intermediate results: create data tables or partitions on the ODPS platform end, support efficient data compression and free space management to address the data expansion problem.

[0059] (3.2) ETL (Extract, Transform, Load) tools for cleaning and transforming raw data: Remove duplicate, incorrect, and irrelevant data to ensure data accuracy and consistency.

[0060] (3.3) Query optimizers for automatically analyzing query statements and generating optimal execution plans: Reduce unnecessary data scans and calculations to improve query efficiency.

[0061] (3.4) Distributed computing engines for supporting the MapReduce distributed computing framework: Capable of processing large-scale data sets and achieving efficient data processing and analysis.

[0062] (3.5) Scheduling systems for managing the execution, summarization, and integration of subtasks: The DAG (Directed Acyclic Graph) task scheduling system of MaxCompute ensures that all subtasks can be successfully completed and the final data is output.

[0063] Compared with the prior art, the beneficial effects of the present invention are:

[0064] I. Significantly improve computing efficiency: The present invention decomposes large tasks into multiple small tasks for parallel execution through task splitting and dynamic scheduling mechanisms, making full use of distributed computing resources and significantly shortening the time for data processing and analysis.

[0065] II. Effectively solve the problem of data expansion: The present invention reduces the occupation of physical storage space and storage costs, while improving data access speed, by optimizing the data storage structure and using ETL tools for data cleaning and compression.

[0066] III. Alleviate the problem of data skew: The present invention effectively alleviates the data skew phenomenon and enables more balanced utilization of computing resources through dynamic adjustment of the splitting strategy of dimension combinations, application of load balancing algorithms, and addition of data preprocessing steps.

[0067] IV. The present invention avoids resource waste by reasonably allocating computing resources, and at the same time reduces disk I / O operations using a caching mechanism, thereby reducing operating costs. The improvement in data processing and analysis speed enables users to obtain query results faster, improving work efficiency and decision-making speed. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0069] Figure 1 Schematic diagram of the method flow of the present invention;

[0070] Figure 2 Schematic diagram of the system composition of the present invention. Detailed implementation manners

[0071] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe the detailed implementation manners of the present invention with reference to the accompanying drawings. Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below;

[0072] It should be noted that the embodiments in this specification are described in a progressive manner, and the key points of each embodiment are the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0073] Explanation of related terms:

[0074] (1) Multidimensional analysis: A data analysis method that reveals the associations and trends between data by representing the observed items and variables in a dataset as multiple dimensions

[0075] (2) Self-service data analysis: Also known as self-service business intelligence (self-service BI), it is a set of approved and supported processes, architectures, and tools that enable business users to complete their data management and analysis tasks without relying on IT and data analysis teams.

[0076] (3) HiveCube: A way provided by Hive to quickly generate multi-dimensional aggregated data, mainly used to solve the problem of frequently aggregating and summarizing data in data analysis and data warehouses. HiveCube is based on fact and dimension tables and can query and analyze data from multiple perspectives and levels.

[0077] (4) Query optimizer on the ODPS platform: An automated component that automatically analyzes query statements and generates an optimal execution plan to reduce unnecessary data scans and calculations and improve query efficiency.

[0078] (5) ETL tool on the ODPS platform: An automated component used for data extraction, transformation, and loading, which helps clean and prepare raw data for subsequent analysis and processing.

[0079] (6) Joining a large table with a small table: In database operations, a larger data table is joined with a smaller data table to obtain the associated information between the two tables.

[0080] (7) Map join: An optimized join operation that is usually completed in the map phase, used to reduce data transfer and the overhead of the shuffle phase, especially suitable for joining a large table with a small table.

[0081] (8) Reduce join: The standard join operation that needs to be completed in the reduce phase, involving data shuffle and sorting to find matching records for joining.

[0082] (9) Shuffle: In distributed computing, the shuffle phase is responsible for redistributing the data output by the map phase to different reduce nodes for subsequent calculation or aggregation operations.

[0083] (10) ODPS (Open Data Processing Service) platform side: A massive data processing platform independently developed by Alibaba, providing data storage, query, calculation, and analysis services.

[0084] Example 1: As Figure 1 shown, this example discloses a multi-dimensional data intelligent splitting and optimization method, which can be used for the classification task of multi-dimensional data sources of online car-hailing. The execution method includes the following steps S1 to S4.

[0085] In this example, regarding step S1, the data engineering stage includes:

[0086] Specifically, step S100, data cleaning: Based on the data table or partition specifically created for online car-hailing data on the ODPS platform side, read the original trip data and intermediate processing results stored therein.

[0087] Using the ETL tool provided by the ODPS platform, preprocess this raw data. The preprocessing operations include cleaning the data, that is, removing duplicate trip records, correcting incorrect geographical location information, and eliminating data fields unrelated to online car-hailing operations to ensure the accuracy and consistency of the data. The ETL tool will perform a series of data conversion and verification operations to meet the requirements of subsequent multi-dimensional analysis.

[0088] Specifically, in step S101, data storage: Store the cleaned online car-hailing data at a specified location on the ODPS platform side, namely a dedicated online car-hailing data warehouse, so that subsequent data processing and analysis steps can efficiently use this data.

[0089] In this embodiment, regarding step S2, the task splitting stage, includes:

[0090] Specifically, in step S200, subtask division: According to the requirements of online car-hailing data analysis, based on the dimensions (time, location, passenger type) and metrics (number of trips, trip duration) to be split, divide the number and size of subtasks.

[0091] Specifically, in step S201, adjustment of the dimension combination splitting strategy: For popular dimension combinations (peak hours, popular locations) of online car-hailing data, increase the corresponding number of subtasks to disperse the computing pressure and improve the processing speed. This strategy aims to optimize the execution efficiency of tasks and resource utilization.

[0092] In this embodiment, regarding step S3, the task scheduling stage, includes:

[0093] Specifically, in step S300, query plan generation: Use the query optimizer on the ODPS platform side to automatically generate an optimal query plan according to the input query statement (query the number of trips in a specific time period) and the statistical information of online car-hailing data. This includes parsing, optimizing, and rewriting the query statement to ensure the efficient execution of the query.

[0094] Specifically, in step S301, subtask management and scheduling: Utilize the scheduling system on the ODPS platform side (DAG task scheduling technology of MaxCompute) to manage the execution, dependency relationships, resource allocation, and priorities of online car-hailing data analysis subtasks. Ensure that subtasks can be executed in the optimal order and manner to achieve efficient resource utilization and task execution.

[0095] Specifically, in step S302, data aggregation and integration: After the subtasks are executed, use DAG task scheduling technology to perform operations such as merging, sorting, and deduplication on the data, and aggregate and integrate the data output by each subtask. This can ensure the integrity and consistency of the data to generate the final multi-dimensional analysis result dataset of online car-hailing.

[0096] Specifically, in step S303, output result data: Interact with the ODPS platform side to output the aggregated and integrated multi-dimensional analysis result data of online car-hailing to the specified storage location or data table. Ensure that this data can be correctly stored and accessed for subsequent report generation and business applications.

[0097] In this embodiment, regarding step S4, the data output and optimization phase: According to the data processing speed and resource utilization information of online car-hailing data in each fixed time period (such as every hour), the data storage structure is dynamically adjusted. For the join operation of large tables (such as the full-scale trip data table) and small tables (such as the trip data table for a specific time period or location), map join is used instead of reduce join to avoid data skew problems caused by shuffle operations. Conversely, if the small table becomes large enough, reduce join can also be considered to optimize performance. This strategy aims to ensure the efficiency and stability of data processing.

[0098] Embodiment 2: As Figure 2 shown, based on Embodiment 1, this embodiment will further provide a multi-dimensional data intelligent splitting and optimization system; this splitting and optimization system is used to execute the splitting and optimization method as described above, including:

[0099] (1) An application side for users to interact with the system: responsible for receiving users' query requests, displaying processing results, and providing corresponding interaction functions. The key components include:

[0100] (1.1) A user interface: allowing users to input query conditions, view processing results, etc.

[0101] (1.2) Interaction tools: such as API interfaces, which facilitate users to integrate the application with the system to achieve automated queries and data acquisition.

[0102] Among them, the execution program (javascript) of this application side is as follows:

[0103]

[0104]

[0105] Among them, the application side collects the task parameters input by users through an HTML form. The fetch function of JavaScript is used to send a POST request to the server side, and the request contains the task parameters input by users. After the server side processes the request, the application side receives and processes the returned results.

[0106] (2) A server side for receiving requests from the application side, calling services on the ODPS platform side, executing data processing and analysis tasks, and returning the results to the application side; the key components include:

[0107] (2.1) A request receiving module for receiving query requests from the application side;

[0108] (2.2) Task Scheduling Module for Scheduling Resources on the ODPS Platform Side: According to the query request, schedule the resources on the ODPS platform side to execute corresponding data processing and analysis tasks. This includes task splitting, dynamically scheduling the dependencies and execution order between subtasks, etc.

[0109] (2.3) Result Processing Module for Collecting the Processing Results Returned by the ODPS Platform Side: Perform necessary processing and integration to generate the final result set and return it to the application side.

[0110] Among them, the execution program (javascript) of the server side is as follows:

[0111]

[0112]

[0113]

[0114] In the above program, the server side uses the Spring Boot framework to define REST APIs. When the application side sends a POST request to / api / submitTask, the server side receives the task parameters in the request body. The server side calls the submitTask function in the service layer to process the task. This function constructs an SQL query and submits it to the ODPS platform side. The server side returns the task ID or error message to the application side.

[0115] (3) ODPS Platform Side for Supporting Large-Scale Distributed Computing; Its key components include:

[0116] (3.1) Data Storage Module for Storing Raw Data and Intermediate Results: Create data tables or partitions on the ODPS platform side, support efficient data compression and free space management to solve the data expansion problem.

[0117] (3.2) ETL (Extract, Transform, Load) Tool for Cleaning and Transforming Raw Data: Remove duplicate, incorrect, and irrelevant data to ensure data accuracy and consistency.

[0118] (3.3) Query Optimizer for Automatically Analyzing Query Statements and Generating the Optimal Execution Plan: To reduce unnecessary data scans and calculations and improve query efficiency.

[0119] (3.4) Distributed Computing Engine for Supporting the MapReduce Distributed Computing Framework: Capable of processing large-scale data sets and achieving efficient data processing and analysis.

[0120] (3.5) Scheduling system for managing the execution, summarization, and integration of subtasks: It is the DAG (Directed Acyclic Graph) task scheduling system of MaxCompute, ensuring that all subtasks can be successfully completed and the final data is output.

[0121] (3.6) Principle of the ODPS platform side:

[0122] Solution to the data expansion problem: Through technical means such as data storage optimization, data compression, and free space management, reduce the occupation of physical storage space and lower the storage cost.

[0123] Mitigation of the data skew problem: Through the task splitting and dynamic scheduling mechanism, decompose large tasks into multiple small tasks for parallel execution, and dynamically adjust the splitting strategy of dimension combinations according to the execution results of historical data and log information to mitigate the data skew problem.

[0124] Improvement of the computing speed: Utilize the distributed computing engine and query optimizer on the ODPS platform side to achieve efficient data processing and analysis, significantly improving the computing speed.

[0125] Among them, the execution program (javascript) of the ODPS platform side is as follows:

[0126]

[0127]

[0128] In the above program, the client code of the ODPS platform side is responsible for interacting with the server side and receiving SQL queries or data processing tasks. The client uses the ODPSSDK to submit tasks to the ODPS platform side and execute SQL queries or data processing. After the task execution is completed, the ODPS platform side generates a unique ID for each task, and the client returns this ID to the server side.

[0129] All of the above embodiments only express the implementation manners of the related practical applications of the present invention, and their descriptions are relatively specific and detailed, but they should not be construed as limitations on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the appended claims.

[0130] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0131] At the same time, those skilled in the art can understand that all or part of the processes of implementing the methods of all the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium provided in the present application and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

Claims

1. A multi-dimensional data intelligent splitting and optimization method, characterized in that: The following steps are included: S1, based on the data table or partition created on the ODPS platform, read the original data and intermediate results stored therein; S2, divides the number and size of subtasks, and dynamically schedules the dependencies and execution order between subtasks; S3 generates the optimal query plan, uses the scheduling system on the ODPS platform to manage the execution, aggregation, and integration of subtasks, and outputs data; S4, dynamically adjust the data storage structure according to the data processing speed and resource utilization information in each fixed time period.

2. The splitting and optimization method according to claim 1, characterized in that: In S1, the ETL tool provided by the ODPS platform is used to pre-process the read raw data; This includes data cleaning to remove duplicate, erroneous and irrelevant data.

3. The splitting and optimization method according to claim 1, characterized in that: In S2, the number and size of subtasks are divided based on the dimensions and metrics that need to be split; the number of subtasks for popular dimension combinations is increased to disperse the computing pressure.

4. The splitting and optimization method according to claim 1, characterized in that: In the S3, the query optimizer on the ODPS platform is used to automatically generate the optimal query plan based on the input query statement and data statistics, including parsing, optimizing and rewriting the query statement.

5. The splitting and optimization method according to claim 4, characterized in that: The DAG task scheduling technology of the scheduling system on the ODPS platform is used to manage the execution, dependencies, resource allocation, and priorities of subtasks.

6. The splitting and optimization method according to claim 4, characterized in that: In the DAG task scheduling technology, data is merged, sorted and deduplicated, and the data output by each subtask is summarized and integrated.

7. The splitting and optimization method according to claim 4, characterized in that: In the S3, the ODPS platform interacts with the data and outputs the aggregated and integrated data to the specified storage location or data table.

8. The splitting and optimization method according to claim 1, characterized in that: In S4, the strategy for dynamically adjusting the data storage structure is: when a large table is joined to a small table, map join is used instead of reduce join to avoid data skew caused by shuffle operation; and vice versa.

9. A multi-dimensional data intelligent splitting and optimization system, characterized by: The system is used to execute the splitting and optimization method according to any one of claims 1 to 8, comprising: Application end for users to interact with the system; The server receives requests from the application side, calls the services on the ODPS platform side, performs data processing and analysis tasks, and returns the results to the server side of the application side. ODPS platform used to support large-scale distributed computing.

10. The splitting and optimizing system according to claim 9, characterized in that: The ODPS platform includes: A data storage module for storing raw data and intermediate results; ETL tools for cleaning and transforming raw data; A query optimizer that automatically analyzes query statements and generates the optimal execution plan; Distributed computing engine used to support MapReduce distributed computing framework; A scheduling system for managing the execution, aggregation, and integration of subtasks.

Citation Information

Cited By

  • Climate event intelligent retrieval and question-answering method based on multi-source data fusion

    CN120277196A

  • Parallel computing performance optimization method and device, equipment and medium

    CN120560866A

  • Parallel computing performance optimization method and device, equipment and medium

    CN120560866B

  • Data checking integrated circuit chip, data checking method and device and server machine frame

    CN122364313A