High-performance large-data-volume cross-library data analysis method

Through the asynchronous operation of the timing tasks and statistical analysis information database, the problems of low efficiency and high cost in cross-border data analysis are solved, efficient and flexible data integration and resource optimization are achieved, and cross-border query needs of large data volumes are adapted.

CN120492452APending Publication Date: 2025-08-15INSPUR SOFTWARE TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510543232.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing cross-base data analysis technology has bottlenecks in query performance, high risk of data leakage, complex technology stack, high learning and maintenance costs, and the cost is uncontrollable under the situation of large data volume, and the compatibility problems are serious.

Method used

The timing tasks and auxiliary statistical analysis information database are used to realize asynchronous operations of cross-border data analysis. Through timing tasks, data is extracted from multiple databases and integrated into the analysis database, data is queried and analyzed according to business needs, and resource usage is optimized by executing timing tasks staggered.

Benefits of technology

It realizes automated data integration, reduces manual intervention, improves efficiency, ensures data consistency, rationally allocates resources, reduces costs, and supports flexible changes in business needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492452A_ABST
    Figure CN120492452A_ABST
Patent Text Reader

Abstract

The invention provides a high-performance large-data-volume cross-database data analysis method, and belongs to the technical field of big data analys.The method comprises the steps that timed tasks are designed, and data needed by services in multiple databases are extracted, integrated and stored in an analysis database; according to business requirements, form data in the statistical analysis information base is inquired, counted, analyzed and processed; the efficiency can be improved and the cost can be reduced by adjusting the execution time of the timed task and performing peak shifting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of big data analysis technology, and in particular to a high-performance cross-database data analysis method for large amounts of data. Background Art

[0002] Cross-database data analysis of large volumes of data refers to the process of integrating, processing, and analyzing large datasets stored in different databases or data sources. Its core goal is to extract valuable information from multiple heterogeneous data sources to support decision-making, business optimization, and insight discovery.

[0003] The core of big data analytics lies in data integration and analysis. However, due to the acceleration of digital transformation, the era of globalization and interconnectedness, the complexity of social issues, and changes in economic structures, the diversity of data storage methods often leads to the formation of data silos, especially when data is dispersed across different platforms and tools.

[0004] With the development of digital businesses and the growth of data volumes, database environments are becoming increasingly complex. Most organizations and enterprises require different types of databases to meet business needs, and the proportion of organizations using two or even more databases is gradually increasing. This surge in the number of databases has brought various impacts, including an increase in the skills required for database maintenance and operation, as well as an increase in data migration and security issues.

[0005] Existing cross-database query technologies primarily include native cross-database query, global query language, and middleware cache query. While these technologies each have their own advantages and can meet user requirements for efficiency, flexibility, real-time performance, data governance, and compliance, they also come with drawbacks such as query performance bottlenecks, increased data leakage and privacy risks, complex technology stacks, high learning and maintenance costs, unmanageable costs when data volumes surge, and compatibility issues caused by technology lock-in.

[0006] To adapt to a multi-database environment, developers and managers must acquire new skills to manage and monitor databases. To address the challenges and demands of data dispersion, scale growth, diversity, business needs, compliance requirements, technological development, real-time requirements, cost optimization, social issues, changing consumer behavior, competitive pressures, research and education needs, and building social trust, there is an urgent need for high-performance, large-scale, cross-database data analysis methods. Summary of the Invention

[0007] To address the above technical issues, the present invention provides a high-performance, large-volume cross-database data analysis method. This method allows for asynchronous, cross-database querying, displaying, statistics, and analysis of large amounts of data. Based on defined statistical objects and permission scopes, it supports efficient statistical analysis of statistical data under different query conditions according to different dimensions.

[0008] The technical solution of the present invention is:

[0009] The present invention provides a high-performance cross-database data analysis method for large amounts of data. Through scheduled tasks and an auxiliary statistical analysis information library, cross-database and data analysis become asynchronous operations, thereby improving efficiency and reducing costs.

[0010] By designing scheduled tasks, the data required for business operations from several databases are extracted and integrated and stored in the analysis database. Based on business needs, the form data in the statistical analysis information database is queried, counted, analyzed and processed. The scheduled task execution time can be adjusted to stagger execution.

[0011] Specific technologies include:

[0012] Step 1. Design and establish databases based on business needs, such as auxiliary statistical analysis information database A (including statistical table a1 and statistical record table a2), database B (including form b that needs to query, display or calculate statistical data) that requires cross-database data statistical analysis, and database C (including form c that needs to query, display or calculate statistical data), as shown in the following table:

[0013]

[0014]

[0015] Step 2: Design a scheduled task to extract and integrate the business data required by multiple databases through internal interfaces or public interfaces, store them in the statistical analysis database, and record the statistics. For example, obtain the data of storage table b in storage database B through the internal interface, and obtain the data of storage table c in storage database C through the public interface. According to business needs, form new data at the implementation layer through comparison, traversal, associative containers and other methods, and store the cross-database integrated data in the statistical table a1 of the statistical analysis information database A. Store the records of this cross-database statistics in the statistical record table a1 of the statistical analysis information database A.

[0016] Step 3: Query, count, analyze, and process the form data in the statistical analysis information database according to business needs. That is, query and count the data in the statistical table a1 in the statistical analysis information database A according to business needs, and directly display the data or perform further processing before displaying the data.

[0017] Step 4: Adjust the execution time of scheduled tasks and execute them in staggered periods. If the system has multiple cross-database business requirements with large data and processing volume and occupies a large number of databases, you can write them into scheduled tasks and comprehensively manage the execution time to execute them in staggered periods.

[0018] The beneficial effects of the present invention are

[0019] It achieves automated cross-database integration of data, reduces manual intervention, and improves efficiency; regularly updates scheduled tasks to avoid data delays and ensure data consistency; scheduled tasks are run in staggered periods, resources are allocated reasonably, and resource optimization is achieved; the execution frequency of scheduled tasks is set according to business needs to achieve flexible scheduling; modular management of scheduled tasks can quickly manage and maintain scheduled tasks as business needs change, and has strong scalability; by executing scheduled tasks on demand and releasing resources after the task is completed, continuous occupation of computing resources and resource waste are avoided, thereby reducing costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It is a schematic diagram of the workflow of the present invention. DETAILED DESCRIPTION

[0021] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0022] Take the exam plan statistics by region function in the exam system as an example:

[0023] 1. Case data system and database structure:

[0024] According to business needs, basic information such as the organizational structure of the examination unit and administrative division information is stored in the basic information database, and examination management record information such as examination plan information and candidate information is stored in the examination management information database. At the same time, a statistical analysis information database is established to assist the high-performance large-scale cross-database data analysis method invented in this invention. The specific description and sample form are as follows:

[0025]

[0026] 2. Scheduled task design:

[0027] First, obtain the time window record data of the statistical analysis form personnel_plan_statistics from the time window record form analyze_table_time_window of the statistical analysis table in the statistical analysis information library dcap_analyse. The update time of this data is the most recent update time of the examination number statistics table personnel_plan_statistics, which is also the start time of the new update data query.

[0028] Secondly, through the internal interface, query the organizational structure table bs_org in the basic information database dcap_basic to obtain the organization information (including ancestor node ID, service level code, superior ID, organization name, province, city, district (county) three-level administrative division code, etc.).

[0029] Finally, query and integrate the test attendance statistics. To prevent performance degradation caused by obtaining a large amount of data at once, the query start time to the current time is divided into days. If the end time does not reach the current time, the following actions are performed in a loop:

[0030] (1) Using the split time range as a condition, associate the examinee information subtable exam_plan_examinee, the exam plan table exam_plan_info, and the cheating information table exam_plan_examinee_cheat in the exam management information database dcap_exam to query the number of examinees counted by exam plan;

[0031] (2) Set the organization information in the list, separate it by ancestral level, set the organization information of the ancestral level, and summarize it into the data list; use the statistical type attribute to distinguish the data of the ancestral level from the sub-level data summarized into the ancestral level;

[0032] (3) Insert the integrated data into the statistical form in batches. To avoid errors caused by inserting too much data at one time, the data needs to be divided into batches and inserted multiple times.

[0033] (4) Update the time window record data of the statistical analysis form personnel_plan_statistics in the record form analyze_table_time_window to complete a cycle.

[0034] 3. Statistical analysis business design:

[0035] First, based on the login account and input parameters, the user's organizational structure information is obtained, thereby obtaining the administrative division level of the statistical data. Based on the division level, the administrative division code to be queried and the administrative division name to be echoed are obtained, and these are processed into an associative container with the division code as the key and a list of administrative division code strings.

[0036] Next, based on the query time range passed in the parameter, calculate the time range required for statistical queries on year-on-year and month-on-month performance, and other personalized services. Query the statistical analysis form, personnel_plan_statistics, to obtain data on the number of examinees within different time ranges, grouped by region, personnel type, industry type, operation project, exam type, and special worker review type.

[0037] Finally, further data processing is performed to meet personalized business needs, and the data is integrated into the return value required by the business.

[0038] 4. Scheduled task control

[0039] A special task management function can be developed to comprehensively manage various scheduled tasks, stagger the execution time of scheduled tasks that occupy more database and running memory resources, and improve the overall operating efficiency of the system.

[0040] The above description is only a preferred embodiment of the present invention and is only used to illustrate the technical solution of the present invention, and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention are included in the scope of protection of the present invention.

Claims

1. A high-performance, large-volume cross-database data analysis method, characterized in that: By designing scheduled tasks, the data required for business operations from several databases are extracted and integrated and stored in the analysis database. Based on business needs, the form data in the statistical analysis information database is queried, counted, analyzed and processed. The scheduled task execution time can be adjusted to stagger execution.

2. The method according to claim 1, characterized in that Design and build database according to business needs.

3. The method according to claim 2, characterized in that The database includes: an auxiliary statistical analysis information database A, which includes statistical table a1 and statistical record table a2; a database B that requires cross-database statistical analysis, which includes form b that needs to query, display or count data; and a database C that includes form c that needs to query, display or count data.

4. The method according to claim 3, characterized in that The data required for business in the database is extracted and integrated through internal interfaces or public interfaces, stored in the statistical analysis database, and the statistics are recorded.

5. The method according to claim 4, characterized in that include: The data of storage table b in storage database B is obtained through the internal interface, and the data of storage table c in storage database C is obtained through the public interface.

6. The method according to claim 5, characterized in that According to business needs, new data is formed at the implementation layer through comparison, traversal, and association of containers, and the cross-database integrated data is stored in the statistical table a1 of the statistical analysis information base A, and the records of this cross-database statistics are stored in the statistical record table a1 of the statistical analysis information base A.

7. The method according to claim 6, characterized in that Based on business needs, query and analyze the data in statistical table a1 in information base A, and directly display the data or perform further processing before displaying the data.

8. The method according to claim 7, characterized in that If there are several cross-database business requirements in the system that have large data and processing volume and occupy a large amount of database, write them into scheduled tasks, and comprehensively manage the execution time to stagger the execution.

Citation Information

Patent Citations

  • Cross-database data management method, device and equipment and computer storage medium

    CN116594990A

  • Production management and control system and method based on business intelligence, computer and medium

    CN117252534A

  • Method, system and equipment for collecting data between database tables and medium

    CN118093585A

  • Risk monitor parameter automatic configuration method

    CN118170448A

  • Data multi-source joint retrieval method and device

    CN119046346A