Method for accelerating report generation based on large-scale parallel processing
By adopting large-scale parallel processing technology and multi-threaded dynamic scheduling modules in massive data environments, massive report calculation formulas are compressed into fewer complex calculation formulas, and using distributed parallel processing technology and batch writing method, the problem of slow report generation in traditional databases under massive data is solved, achieving fast and efficient report generation.
Patent Information
- Application Number
- CN202510237271.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-02
- Publication Date
- 2025-05-30
AI Technical Summary
In the case of massive data, traditional relational databases generate reports due to architectural limitations, resulting in slow report generation, which affects the timeliness and accuracy of business analysis.
Using a large-scale parallel processing (MPP) method, the report request analysis module is used to classify and group a large number of simple calculation formulas into fewer complex calculation formulas, and the multi-threaded dynamic scheduling module makes full use of server resources, combined with distributed parallel processing technology and batch writes of calculation results in batches to improve report generation efficiency.
Through this method, multiple units, multiple months and multiple reports can be quickly generated in massive data, solving the problem of slow report generation caused by insufficient resource utilization or overload, and significantly improving the efficiency of report generation.
Smart Images

Figure CN120066735A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of report generation and big data, and particularly relates to a method for accelerating report generation based on massive parallel processing. Background Art
[0002] With the rapid development of the business of units, the amount of data has shown explosive growth. At the same time, according to business requirements, each unit and each department usually needs to generate reports periodically. Commonly, all reports of all units are generated centrally within the monthly end, quarterly, semi-annual, first three quarters, annual, etc. cycle to verify business data, and it is necessary to ensure that the report generation task is completed in a short time to meet the timeliness of verifying business data. With the rapid development of the business of units, the generated report data has accumulated to a certain scale, and at the same time, it presents the complex characteristics of a large number of generated reports (nearly 400 reports) and a large amount of calculation for each report (there are tens of thousands of calculation formulas for each report), which poses an unprecedented challenge to the data processing and report generation efficiency that combines CPU-intensive and IO-intensive. When facing such characteristic data processing, traditional relational database management systems (RDBMS) are often unable to cope due to architectural limitations (such as single-node processing, limited concurrency ability, etc.). The symmetric multi-processing (SMP) architecture is not easy to expand, and it cannot meet the computing performance indicators in terms of CPU calculation and IO throughput, resulting in slow report generation and affecting the timeliness and accuracy of business analysis. How to quickly generate reports that meet multiple units (departments), multiple cycles, and multiple reports in massive data is the main research content of the present invention. Summary of the Invention
[0003] (1) Technical Problems to be Solved
[0004] The technical problem to be solved by the present invention is how to provide a method for accelerating report generation based on massive parallel processing to solve the problem of slow report generation in the case of existing massive data.
[0005] (2) Technical Solutions
[0006] To solve the above technical problems, the present invention proposes a method for accelerating report generation based on massive parallel processing, and the method includes the following steps:
[0007] Step 1: A report request parsing module obtains report generation request parameter information and calculation formulas of report indicators. The report request parsing module compresses a simple calculation formula of a report into a complex calculation formula through classification and grouping, and then writes it into a calculation task table, and notifies a multi-thread dynamic scheduling module that there is a new task to be processed;
[0008] Step 2: The multi-thread dynamic scheduling module obtains the MPP server resources to dynamically create N threads, retrieves N computing tasks from the computing task table, assigns one thread to each computing task, submits the tasks to the MPP server concurrently in multiple threads, updates the task status after each computing task is completed, and sends the execution result to the computing data batch processing module;
[0009] Step 3: The MPP server distributes the computing tasks to each node to execute the report generation tasks in parallel. The MPP server gives full play to the independent data storage and independent computing capabilities of each execution node, and completes the computing tasks quickly in multiple paths in parallel. After each computing task is completed, the computing result is returned to the multi-thread dynamic scheduling module;
[0010] Step 4: The computing data batch processing module temporarily writes the execution result of each task into the server memory. When the calculation results of the units and months selected for a report are all calculated, the report data in the memory is written into the database in batches at one time, improving the system write performance by reducing the disk I / O of the database server.
[0011] (III) Beneficial effects
[0012] The present invention proposes a method for accelerating report generation based on massive parallel processing. The present invention proposes a method for improving report generation based on MPP, which mainly solves the deficiencies of insufficient resource utilization or resource overload and slow report generation when generating multiple units, multiple months, and multiple reports in massive data. The achieved improvement effect is: by classifying, sorting, and compressing a large number of simple report calculation formulas into fewer complex calculation formulas through business, making full use of server resources through multi-thread dynamic scheduling, using distributed parallel processing technology, and the method of writing calculation results in batches at one time to improve the efficiency of generating multiple units, multiple months, and multiple reports.
[0013] The innovation of the present invention lies in adopting a method of classifying, sorting, and compressing a large number of report calculation formulas into fewer calculation formulas through business, making full use of server resources through multi-thread dynamic scheduling, parallel processing technology, and writing calculation results in batches at one time to quickly generate reports in massive data. Brief description of the drawings
[0014] Figure 1 It is the flowchart of the method of the present invention. Detailed implementation manners
[0015] To make the objectives, contents, and advantages of the present invention clearer, the following further describes in detail the specific implementation manners of the present invention with reference to the drawings and embodiments.
[0016] Aiming at the problem of slow report generation in the case of existing massive data, the present invention proposes a method for accelerating report generation based on massively parallel processing (MPP). This method improves the efficiency of report generation by classifying and sorting massive simple report calculation formulas into fewer complex calculation formulas through business classification, using a multi-thread dynamic scheduling strategy to make full use of server resources, massive parallel processing, and batch writing of calculation results at one time.
[0017] The present invention provides a method for accelerating report generation based on massively parallel processing (MPP), including: Step 1, the report request parsing module obtains report generation request parameter information and the calculation formulas of report indicators. To improve the efficiency of report generation, the report request parsing module classifies and sorts tens of thousands of simple calculation formulas on a report into dozens of complex calculation formulas through business classification, then writes them into the calculation task table, and notifies the multi-thread dynamic scheduling module that there are new tasks to be processed. Step 2, the multi-thread dynamic scheduling module obtains MPP server resources to dynamically create N threads, retrieves N calculation tasks from the calculation task table, assigns one thread to each calculation task, and concurrently submits the calculation tasks to the MPP by multiple threads. After each calculation task is completed, it updates the task status and sends the execution result to the calculation data batch processing module, and then dynamically creates M threads according to the MPP server resources to retrieve M pending calculation tasks and submit them to the MPP. Step 3, the MPP executes the report generation task in parallel. The MPP gives full play to the independent data storage and independent computing capabilities of each execution node, and quickly completes the calculation tasks in multiple paths in parallel. After each calculation task is completed, it returns the calculation result to the multi-thread dynamic scheduling module; Step 4, the calculation data batch processing module temporarily writes the execution result of each task into the memory. When the units and months selected for a report are all calculated, it is batch-written into the database report at one time to improve the write performance of the system by reducing the disk I / O of the database server.
[0018] The present invention realizes a method for quickly generating reports through a report request parsing module, a multi-thread dynamic scheduling module, massive parallel processing, and a calculation data batch processing module, and improves the efficiency of report generation by classifying and sorting multiple calculation formulas into fewer complex calculation formulas through business classification, making full use of server resources by multi-thread dynamic scheduling, massive parallel processing, and batch writing of calculation results at one time.
[0019] The present invention provides a method for accelerating report generation based on massively parallel processing, including the following steps:
[0020] Step 1: The report request parsing module obtains the report generation request parameter information and the calculation formulas of report metrics. To improve the report generation efficiency, the report request parsing module classifies and groups tens of thousands of simple calculation formulas on a report into dozens of complex calculation formulas, then writes them into the calculation task table, and notifies the multi-thread dynamic scheduling module that there are new tasks to be processed.
[0021] Step 2: The multi-thread dynamic scheduling module obtains the MPP server resources to dynamically create N threads, retrieves N calculation tasks from the calculation task table, assigns one thread to each calculation task, and submits the tasks to the MPP server concurrently in multiple threads. After each calculation task is completed, it updates the task status and sends the execution result to the calculation data batch processing module;
[0022] Dynamically create M threads according to the MPP server resources to retrieve M calculation tasks to be executed and submit them to the MPP server.
[0023] Step 3: The MPP server distributes the calculation tasks to each node to execute the report generation task in parallel. The MPP server gives full play to the independent data storage and independent calculation capabilities of each execution node, and completes the calculation tasks quickly in multiple paths in parallel. After each calculation task is completed, it returns the calculation result to the multi-thread dynamic scheduling module;
[0024] Step 4: The calculation data batch processing module temporarily writes the execution result of each task into the server memory. When the calculation results of the units and months selected for a report are all calculated, it writes the report data in the memory into the database in batches at one time, improving the system write performance by reducing the disk I / O of the database server.
[0025] In an embodiment of a method for accelerating report generation based on massive parallel processing (MPP) of the present invention, the report request parsing module obtains the unit information, month information, and report information of the report generation request, and reads the calculation formulas of the report metrics according to the report information. To improve the report generation efficiency, the report request parsing module classifies and organizes the calculation formulas of the report metrics according to the business (for prosecution business, classify and merge according to the case settlement reason, acceptance situation, and case settlement situation) and groups them according to the generating unit, compressing an average of R (50,000) simple calculation formulas for a report into an average of S (50) complex calculation formulas. The calculation amount for the user to generate a report is compressed from the number of units * number of months * number of reports * 50,000 calculation formulas to the number of months * number of reports * 50 calculation formulas. Then it creates a total task for the generation request, creates a sub-task for the generation request of each report, then writes them into the calculation task table, and notifies the multi-thread dynamic scheduling module that there are new tasks to be processed.
[0026] In an embodiment of a method for accelerating report generation based on massively parallel processing (MPP) of the present invention, the multi-thread dynamic scheduling module uses a Java program to respectively read the / proc / cpuinfo file information of each cluster node to obtain the total number of CPU cores and the CPU usage rate, reads the / proc / meminfo file information of each cluster node to obtain the memory usage rate, and dynamically calculates the number of allocatable threads N according to the current server resource usage situation.
[0027] N = Round(Min(CPU total cores * (1 - Max(CPU usage rate)) * 80%, CPU total cores * (1 - Max(memory usage rate)) * 80%))
[0028] Then, obtain the number of computing tasks T to be executed, and determine whether the number of computing tasks T is greater than N. If it is greater, create N threads, take N computing tasks from the computing task table, and allocate each computing task to a thread, and submit the report generation tasks concurrently by multiple threads to the MPP server. If T is not greater than N, then determine whether T is greater than 0. If it is greater, dynamically create T threads, take T computing tasks from the computing task table, and allocate each computing task to a thread, and submit the report generation tasks concurrently by multiple threads to the MPP server; if T is not greater than 0, it means that all the computing tasks to be executed have been completed.
[0029] After each computing task is completed, update the task execution status and send the execution result to the computing data batch processing module, and then dynamically create M threads according to the above method to take M pending computing tasks and submit them to the MPP.
[0030] In an embodiment of a method for accelerating report generation based on massively parallel processing (MPP) of the present invention, the MPP server executes the report generation tasks in parallel. The MPP server distributes the computing tasks to each execution node, gives full play to the independent data storage and independent computing capabilities of each execution node, and can dynamically expand horizontally and vertically to enhance the computing capabilities of the execution nodes, and complete the computing tasks quickly in multiple paths in parallel. After each execution node completes the computing, it summarizes the computing results. After each computing task is completed, the computing result is returned to the multi-thread dynamic scheduling module.
[0031] In an embodiment of a method for accelerating report generation based on massive parallel processing (MPP) of the present invention, the calculation data batch processing module first determines whether all calculation tasks of a report have been completed. If not, it temporarily writes the execution result of each task into the memory. To improve the report writing and query efficiency, each row of the report structure stores the index data of a certain unit for 12 months of a certain year for a certain index. When all the calculation results of a report index for the selected unit and month are calculated, all the index calculation results are read from the memory and written into the database at one time, so as to improve the system writing performance by reducing the disk I / O of the database server.
[0032] Table 1 Report Structure Table
[0033]
[0034] Example 1:
[0035] A method for accelerating report generation based on massive parallel processing (MPP) includes the following steps:
[0036] Step 1, the report request parsing module obtains the unit information, month information, report information, and calculation formula of the report index in the report generation request. To improve the report generation efficiency, the report request parsing module classifies and sorts the calculation formulas of the report index according to the business and groups them according to the generating unit, so as to compress tens of thousands of simple calculation formulas on a report into dozens of complex calculation formulas, and then writes them into the calculation task table, and notifies the multi-thread dynamic scheduling module that there are new tasks to be processed;
[0037] Step 2, the multi-thread dynamic scheduling module obtains the total number of CPU cores, CPU usage rate, and memory usage rate of the MPP cluster, dynamically creates N threads, fetches N calculation tasks from the calculation task table, assigns one thread to each calculation task, submits the multi-threads concurrently to the MPP. After each calculation task is completed, it updates the task status and sends the execution result data to the calculation data batch processing module, and then dynamically creates M threads according to the MPP server resources, fetches M pending calculation tasks, and submits them to the MPP;
[0038] Step 3, the MPP distributes and executes the report generation task in parallel. The MPP gives full play to the independent data storage and independent calculation capabilities of each execution node, and completes the calculation task quickly in multiple paths in parallel. After each calculation task is completed, it returns the calculation result to the multi-thread dynamic scheduling module.
[0039] Step 4, the calculation data batch processing module temporarily writes the execution result of each task into the server memory. When all the calculations of a report for the selected unit and month are completed, it is written into the database at one time, so as to improve the system writing performance by reducing the disk I / O of the database server.
[0040] Further, in the first step, the report request parsing module obtains the unit information, month information, report information, and the calculation formula of the report indicators in the report generation request. To improve the report generation efficiency, the report request parsing module classifies and organizes the calculation formulas of the report indicators according to the business and groups them by the generating unit, realizing the compression of an average of 50,000 simple calculation formulas for one report to an average of 50 complex calculation formulas. The calculation volume of the user to generate a report is compressed from the number of units * the number of months * the number of reports * 50,000 to the number of months * the number of reports * 50. A total task is created for the report generation request, a sub-task is created for each table, and then it is written into the calculation task table, and the multi-thread dynamic scheduling module is notified that there are new tasks to be processed.
[0041] Further, in the second step, the multi-thread dynamic scheduling module uses a Java program to respectively read the / proc / cpuinfo file information of each cluster node to obtain the total number of CPU cores and the CPU usage rate, reads the / proc / meminfo file information of each cluster node to obtain the memory usage rate, dynamically creates N threads according to the current server computing resource situation, fetches N calculation tasks from the calculation task table, assigns each calculation task to a thread, and submits the multi-threads concurrently to the MPP. After each calculation task is completed, the task execution status is updated and the execution result is sent to the calculation data batch processing module, and then M threads are dynamically created according to the above method to fetch M pending calculation tasks and submit them to the MPP.
[0042] Further, in the third step, the MPP distributes and executes the report generation task in parallel. The MPP distributes the calculation requests to each execution node, gives full play to the independent data storage and independent computing capabilities of each execution node, and can dynamically expand horizontally and vertically to improve the computing capabilities of the execution nodes, and quickly completes the calculation tasks in multiple paths in parallel. After each calculation task is completed, the calculation result is returned to the multi-thread dynamic scheduling module.
[0043] Further, in the fourth step, the calculation data batch processing module temporarily writes the execution result of each task into the server memory. To improve the report writing and query performance, each row of the report structure is designed as the index data of a certain unit for 12 months in a certain year. When the calculation of a report for the selected unit and month is completed, the in-memory report data is written into the database at one time, improving the system writing performance by reducing the disk I / O of the database server.
[0044] The present invention proposes a method for improving report generation based on MPP, which mainly solves the deficiencies of insufficient resource utilization or resource overload and slow report generation when generating reports for multiple units, multiple months, and multiple reports in massive data. The achieved improvement effect is as follows: by classifying and organizing the massive simple report calculation formulas into fewer complex calculation formulas, making full use of server resources through multi-thread dynamic scheduling, using distributed parallel processing technology, and writing the calculation results in batches at one time, the efficiency of generating reports for multiple units, multiple months, and multiple reports is improved.
[0045] The innovation of the present invention lies in a method of quickly generating reports in massive data by classifying and organizing the massive report calculation formulas into fewer calculation formulas, making full use of server resources through multi-thread dynamic scheduling, using parallel processing technology, and writing the calculation results in batches at one time.
[0046] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art of the present technology, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. A method for accelerating report generation based on large-scale parallel processing, characterized in that: The method comprises the following steps: Step 1: The report request parsing module obtains the report generation request parameter information and the calculation formula of the report index. The report request parsing module compresses a simple calculation formula of a report into a complex calculation formula by classification and grouping, and then writes it into the calculation task table, and notifies the multi-thread dynamic scheduling module that there is a new task to be processed; Step 2: The multi-threaded dynamic scheduling module obtains the MPP server resources to dynamically create N threads, takes N computing tasks from the computing task table, assigns one thread to each computing task, and submits tasks to the MPP server concurrently through multiple threads. After each computing task is completed, the task status is updated and the execution result is sent to the computing data batch processing module. Step 3: The MPP server distributes the computing tasks to each node and executes the report generation task in parallel. The MPP server fully utilizes the independent data storage and independent computing capabilities of each execution node, and completes the computing tasks in parallel and quickly. After each computing task is completed, the computing results are returned to the multi-threaded dynamic scheduling module. Step 4: The calculation data batch processing module temporarily writes the execution results of each task into the server memory. When the calculation results of the units and months selected in a report are completed, the report data in the memory is written into the database in batches at once, thereby improving the system write performance by reducing the database server disk IO.
2. The method for accelerating report generation based on large-scale parallel processing according to claim 1, characterized in that: In the step 1, the report request parsing module obtains the unit information, month information and report information of the report generation request, and reads the calculation formula of the report indicator according to the report information.
3. The method for accelerating report generation based on large-scale parallel processing according to claim 2, characterized in that: In the step one, the report request parsing module classifies and organizes the calculation formulas of the report indicators according to the business and groups them by the generation units to compress R simple calculation formulas of a report into S complex calculation formulas. The calculation amount of the user-generated report is compressed from the number of units * the number of months * the number of reports * R calculation formulas to the number of months * the number of reports * S calculation formulas. Then, a total task is created for the generation request, a subtask is created for each report generation request, which is then written into the calculation task table, and the multi-threaded dynamic scheduling module is notified that there is a new task to be processed.
4. The method for accelerating report generation based on large-scale parallel processing according to claim 3, characterized in that: Classification and organization based on business include: for procuratorial business, classification and merging according to the reasons for case closure, acceptance circumstances and closure circumstances.
5. The method for accelerating report generation based on large-scale parallel processing according to claim 3, characterized in that: In step 2, the multi-threaded dynamic scheduling module uses a Java program to read the / proc / cpuinfo file information of each cluster node to obtain the total number of CPU cores and CPU usage, read the / proc / meminfo file information of each cluster node to obtain memory usage, and dynamically calculate the number of allocable threads N based on the current server resource usage.
6. The method for accelerating report generation based on large-scale parallel processing according to claim 5, characterized in that: N=Min(total number of CPU cores*(1-Max(CPU usage))*80%, total number of CPU cores*(1-Max(memory usage))*80%), round up.
7. The method for accelerating report generation based on large-scale parallel processing according to claim 5, characterized in that: In the step 2, the number of computing tasks to be executed T is obtained, and it is determined whether the number of computing tasks T is greater than N. If it is greater, N threads are created, N computing tasks are taken from the computing task table, each computing task is assigned to a thread, and multiple threads concurrently submit report generation tasks to the MPP server; If T is not greater than N, determine whether T is greater than 0. If it is greater than 0, dynamically create T threads, take T computing tasks from the computing task table, assign each computing task to a thread, and submit report generation tasks to the MPP server concurrently with multiple threads; if T is not greater than 0, it means that all computing tasks to be executed have been completed.
8. The method for accelerating report generation based on large-scale parallel processing according to claim 7, characterized in that: The step 2 also includes: after each computing task is executed, updating the task execution status and sending the execution result to the computing data batch processing module, and then dynamically creating M threads to take M computing tasks to be executed and submitting them to the MPP.
9. The method for accelerating report generation based on large-scale parallel processing according to claim 7, characterized in that: In the step three, the MPP server executes the report generation task in parallel. The MPP server distributes the computing task to each execution node, fully utilizing the independent data storage and independent computing capabilities of each execution node, and dynamically expands horizontally and vertically to improve the computing capabilities of the execution nodes. The computing tasks are completed quickly in parallel in multiple ways. After each execution node completes the calculation, the calculation results are summarized; after each computing task is completed, the calculation results are returned to the multi-threaded dynamic scheduling module.
10. The method for accelerating report generation based on large-scale parallel processing according to claim 9, characterized in that: In step 4, the data batch processing module first determines whether all calculation tasks of a report have been completed. If not, the execution results of each task are temporarily written into the memory. To improve the efficiency of report writing and query, each row of the report structure stores the indicator data of an indicator of a certain unit for 12 months of a certain year. When the calculation of a report indicator of the selected unit and month is completed, all indicator calculation results are read from the memory and written into the database at once.