Data mining analysis visualization method and system

By collecting and converting data from different data sources, and using a timing synchronization and priority scheduling mechanism, the data is integrated into the data warehouse, solving the problem of data silos between information systems and achieving efficient and visual data integration and analysis.

CN120030081APending Publication Date: 2025-05-23SHENYANG AIRCRAFT DESIGN INST AVIATION IND CORP OF CHINA
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510078790.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively integrate data from different information systems, resulting in data silos, unable to directly reuse data from other systems, and on the premise of meeting confidentiality requirements, the challenge of efficiently integrating data is greater.

Method used

By collecting data from different data sources, converting it into a data warehouse-adapted format, and using a timing synchronization mechanism, data is synchronized to the data warehouse. Use the priority and time slice rotation algorithm to synchronize data, and introduce a data verification mechanism to ensure the effectiveness of data synchronization. Finally, the data is visualized.

Benefits of technology

It realizes effective integration of data between systems, solves the data island problem, meets the speciality and quality requirements of data, and provides intuitive data analysis results through visual means.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030081A_ABST
    Figure CN120030081A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data mining, and particularly relates to a data mining analysis visualization method and system. The method comprises the steps of 1, collecting data of different data sources, and converting the collected data into a format adaptive to a data warehouse; 2, determining data timing synchronization; step 3, synchronizing the collected data to a data warehouse at regular time; 4, verifying a data synchronization result to ensure the validity of data synchronization; and step 5, carrying out visualization on the data. According to the data mining analysis visualization method, the problem of data isolation and the problem of data isomerism between systems can be solved, a foundation is laid for visualization work, and the requirements for data particularity and data quality can be met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of data mining technology, and in particular, relates to a data mining analysis visualization method and system. Background Art

[0002] In some scenarios, based on information security considerations, the network environment must meet physical isolation and network isolation. In the early stage of building an information system, the considerations were not comprehensive. They only focused on how to ensure that the system functions were fully realized, and did not consider the possibility of subsequent data flow and data transmission between various systems. This led to the existence of data islands between various systems and the inability to directly reuse data from other systems. How to effectively integrate the data of various systems and how to most efficiently integrate effective and accurate data for use by various units in the shortest time while meeting confidentiality requirements are issues that this unit urgently needs to solve.

[0003] Currently, there are several main types of data mining methods: descriptive statistics methods, exploratory data analysis methods, correlation analysis methods, machine learning methods, database methods, etc.

[0004] 1) Descriptive statistical methods

[0005] Data statistical analysis methods are mainly used to describe and summarize data. Commonly used methods include mean, median, mode, standard deviation, variance, range, percentile, etc., which can provide an intuitive understanding of the overall characteristics of the data set.

[0006] 2) Exploratory Data Analysis Methods

[0007] Exploratory data analysis is a method used to discover patterns and relationships in data. By drawing charts such as histograms, scatter plots, box plots, etc. to show the distribution, correlation, and outliers of data, it can help researchers find hidden patterns in the data and provide clues for further analysis.

[0008] 3) Correlation analysis method

[0009] Correlation analysis is used to measure the correlation between two or more variables. Commonly used correlation analysis methods include Pearson correlation coefficient and Spearman correlation number, which can help researchers determine the linear or nonlinear relationship between variables and evaluate their strength and direction.

[0010] 4) Machine Learning Methods

[0011] Machine learning is a branch of artificial intelligence that uses statistical and computer science methods to enable computers to learn from data and improve their performance. It can be mainly divided into supervised learning, unsupervised learning, semi-supervised learning, and reinforcement learning:

[0012] a) Supervised Learning

[0013] In this type of learning, the algorithm learns from a labeled training set, with the goal of finding a mapping function from input to output. Common algorithms include support vector machines, decision trees, etc.

[0014] b) Unsupervised Learning

[0015] The training data set for unsupervised learning does not have pre-labeled results. The algorithm discovers some structure or pattern from these data, such as clustering, dimensionality reduction, anomaly detection, etc. Common algorithms include K-means clustering, principal component analysis, etc.

[0016] c) Semi-supervised learning

[0017] Semi-supervised learning is between supervised and unsupervised learning. Only part of the data is labeled. Semi-supervised learning uses this part of the labeled data to improve the performance of unlabeled data.

[0018] d) Reinforcement Learning

[0019] Reinforcement learning is a special type of machine learning method that learns how to make the best decisions to maximize accumulated rewards through interaction with the environment. Reinforcement learning is different from traditional supervised learning and unsupervised learning because it does not involve a fixed training data set, but relies on a trial-and-error method to learn and change.

[0020] 5) Database method

[0021] Databases usually provide data preprocessing, data indexing and querying for data mining. Data preprocessing includes data cleaning, data integration, data transformation and other processes. In the data cleaning process, it is necessary to process errors, missing and duplicate values ​​in the data to ensure the accuracy and completeness of the data. Data integration is to integrate data from different data sources into one data set to facilitate data mining and analysis. Data transformation is to convert the original data, such as mapping the actual value into a code value for use in subsequent steps.

[0022] Data indexing is to improve data access speed and query efficiency by establishing appropriate index structures, and speed up query operations in data mining. Common index structures include B-tree, B+ tree, hash index, etc.

[0023] It can be seen that the results of descriptive statistics methods are too simple; exploratory data analysis methods and related analysis methods can only provide simple clues for subsequent work; machine learning methods have very high requirements on data quality, and the existing system data is insufficient in dimension and poor in quality, which makes it difficult to support machine learning work; database methods are intuitive and efficient, but how to process data, how to integrate data into visualization requirements, and how to store data in various systems have become issues that need to be considered.

[0024] Therefore, it is desired to have a technical solution to overcome or at least alleviate at least one of the above-mentioned defects of the prior art. Summary of the invention

[0025] The purpose of this application is to provide a data mining analysis visualization method and system to solve at least one problem existing in the prior art.

[0026] The technical solution of this application is:

[0027] The first aspect of the present application provides a data mining analysis visualization method, comprising:

[0028] Step 1: Collect data from different data sources and convert the collected data into a format that is compatible with the data warehouse;

[0029] Step 2: Determine the timing synchronization of data;

[0030] Step 3: Synchronize the collected data to the data warehouse regularly;

[0031] Step 4: Verify the data synchronization results to ensure the validity of data synchronization;

[0032] Step 5: Visualize the data.

[0033] In at least one embodiment of the present application, in step one, the data source includes: an Oracle database, a MySQL database, and a DAMO database.

[0034] In at least one embodiment of the present application, in step 2, determining data timing synchronization includes:

[0035] S21, configuring corresponding process control blocks for data from different data sources, and calculating the scheduling order according to the process information of the process control blocks, and placing all the process control blocks in the ready state into the ready queue according to the scheduling order;

[0036] S22, taking out one of the process control blocks from the head of the ready queue, allocating a time slice to it, and setting the state to execution, and the process control block runs on the CPU until the time slice is used up or the process is actively abandoned;

[0037] S23, if the process control block has completed execution within the time slice, remove it from the ready queue and release the resources it occupies; if the time slice is used up but the process control block has not completed execution, set its state to ready and determine the priority change parameter of the process control block;

[0038] S24, recalculating the scheduling order according to the priority change parameter of the process control block, placing all the process control blocks in the ready state into the ready queue according to the new scheduling order, and returning to step S22;

[0039] S25. Repeat steps S22 to S24 until all the process control blocks are executed or the termination condition is met.

[0040] In at least one embodiment of the present application, the process information includes: process ID, total process running time, process used CPU time, process priority, and process start time.

[0041] In at least one embodiment of the present application, in step S21, the scheduling order is:

[0042] Order(Pi)=Sec(Pi)*{(Pri(Pi)-Cpu(Pi)) / Tot(Pi)}

[0043] Among them, Order(Pi) is the scheduling order, Pri(Pi) is the process priority, Cpu(Pi) is the CPU time used by the process, Tot(Pi) is the total running time of the process, Start(Pi) is the start time of the process, and Sec(Pi) is the confidentiality parameter.

[0044] In at least one embodiment of the present application, in step S24, the scheduling order is:

[0045] NewOrder(Pi)=NewPri(Pi)*Order(Pi)

[0046] Among them, NewPri(Pi) is the priority change parameter.

[0047] In at least one embodiment of the present application, the data warehouse includes:

[0048] The basic data layer is used to store collected data and save or clean historical data according to data business requirements;

[0049] The public dimension model layer includes a detailed data layer and a summary data layer, wherein the detailed data layer is used to process the data from the basic data layer to generate detailed fact data and dimension table data, and the summary data layer is used to process the detailed fact data and the dimension table data to generate public indicator summary data;

[0050] The application data layer is used to store the personalized statistical indicator data of the data products generated by processing the basic data layer and the common dimension model layer;

[0051] The task scheduling time in the data warehouse is End(Pi), End(Pi)=Tot(Pi)+Start(Pi).

[0052] In at least one embodiment of the present application, in step 4, verifying the synchronization result of the data includes:

[0053] Process verification: Verify whether all processes have been completed. If there is an abnormally completed process, call it again and report an error to remind relevant personnel;

[0054] Data warehouse verification: Verify whether the number of newly added data items matches. If not, send reminders to relevant personnel.

[0055] In at least one embodiment of the present application, in step five, visualizing the data includes:

[0056] Visualize the data in reports;

[0057] Visualize the data in charts.

[0058] A second aspect of the present application provides a data mining analysis visualization system, comprising:

[0059] The data collection module is used to collect data from different data sources and convert the collected data into a format that is compatible with the data warehouse;

[0060] A data timing synchronization module, used to determine data timing synchronization;

[0061] The data warehouse construction module is used to synchronize the collected data to the data warehouse at regular intervals;

[0062] The data verification mechanism module is used to verify the data synchronization results to ensure the validity of data synchronization;

[0063] Data visualization module, used to visualize data.

[0064] The invention has at least the following beneficial technical effects:

[0065] The data mining and analysis visualization method of the present application can solve the data isolation problem and data heterogeneity problem between systems, lay a good foundation for visualization work, and meet the needs of data specificity and data quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] Figure 1 is a flow chart of a data mining analysis visualization method according to an embodiment of the present application;

[0067] Figure 2 is a data timing synchronization flow chart of an implementation method of the present application;

[0068] Figure 3 is a process control block structure diagram of an implementation method of the present application;

[0069] Figure 4 is a data structure diagram of a data warehouse according to one embodiment of the present application;

[0070] Figure 5 This is a diagram of the overall architecture of a data warehouse in one implementation of the present application;

[0071] Figure 6 This is a schematic diagram of a data warehouse task scheduling process in one implementation mode of the present application;

[0072] Figure 7 It is a schematic diagram of a visualization report of an implementation method of the present application;

[0073] Figure 8 It is a schematic diagram of a data mining analysis visualization system according to one embodiment of the present application. DETAILED DESCRIPTION

[0074] In order to make the purpose, technical scheme and advantages of the implementation of this application clearer, the technical scheme in the embodiment of this application will be described in more detail below in conjunction with the drawings in the embodiment of this application. In the drawings, the same or similar reference numerals throughout represent the same or similar elements or elements with the same or similar functions. The described embodiments are part of the embodiments of this application, not all of them. The embodiments described below with reference to the drawings are exemplary and are intended to be used to explain this application, and should not be construed as limitations on this application. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application. The embodiments of this application are described in detail below in conjunction with the drawings.

[0075] In the description of the present application, it should be understood that the terms "center", "longitudinal", "lateral", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as limiting the scope of protection of the present application.

[0076] The following is combined with Figures 1 to 8 This application is described in further detail.

[0077] The first aspect of the present application provides a data mining analysis visualization method, such as Figure 1 As shown, the following steps are included:

[0078] Step 1: Collect data from different data sources and convert the collected data into a format that is compatible with the data warehouse;

[0079] Step 2: Determine the timing synchronization of data;

[0080] Step 3: Synchronize the collected data to the data warehouse regularly;

[0081] Step 4: Verify the data synchronization results to ensure the validity of data synchronization;

[0082] Step 5: Visualize the data.

[0083] The data mining and analysis visualization method of this application is not planned in its initial stage of information construction, which leads to problems such as different types of information system databases, chaotic data storage in the database, and poor data quality. Therefore, in step 1, it is necessary to collect data from different types of databases from different data sources, such as Oracle database, MySQL database, DAMO database, etc., to ensure that data can be collected from multiple different types of databases and converted into a format suitable for the data warehouse.

[0084] In this application, the data mining and analysis visualization method has too many sources of data. In order to ensure that the synchronization of multi-source data does not conflict, the synchronization time needs to be reasonably sorted. Different data sources have different importance and priorities. Therefore, when performing scheduled data synchronization, it is necessary to consider placing the data with high priority at the front of the synchronization queue to ensure the success rate of its synchronization completion.

[0085] In the preferred implementation of the present application, by introducing a priority parameter, the data timing synchronization is sorted according to the time slice round-robin algorithm, combining the fairness of the time slice round-robin method and the efficiency of the priority scheduling algorithm. A priority is set for each process in the algorithm, and processes with high priority can obtain more CPU time. In order to prevent processes with high priority from occupying the CPU for a long time, each process is also limited by the time slice. When the time slice is used up, if the process is not completed, it will be suspended and wait for the next scheduling. If the priority of the processes is the same, they are scheduled in a time slice round-robin manner.

[0086] Specifically, in this embodiment, Figure 2 As shown, in step 2, determining data timing synchronization includes:

[0087] S21, configuring corresponding process control blocks for data from different data sources, and calculating the scheduling order according to the process information of the process control blocks, and placing all process control blocks in the ready state into the ready queue according to the scheduling order;

[0088] S22, take a process control block from the head of the ready queue, allocate a time slice for it, and set the state to execute. The process control block runs on the CPU until the time slice is used up or the process is actively abandoned (such as manual control abandonment or process completion);

[0089] S23. If the process control block has completed execution within the time slice, it is removed from the ready queue and the resources occupied by it are released; if the time slice is used up but the process control block has not yet completed execution, its state is set to ready and the priority change parameter of the process control block is determined;

[0090] S24, recalculate the scheduling order according to the priority change parameter of the process control block, put all the process control blocks in the ready state into the ready queue according to the new scheduling order, and return to step S22;

[0091] S25. Repeat steps S22 to S24 until all process control blocks are executed or the termination condition (such as manual closing) is met.

[0092] When the data is synchronized, in S21, the process control block is initialized. Figure 3 As shown, the process information of the process control block includes several fields, including process ID, total process running time, process used CPU time, process priority, and process start time. The fields are expandable. For process Pi, the process priority is Pri(Pi), the used CPU time is Cpu(Pi), the total running time is Tot(Pi), and the start time is Start(Pi). Due to the confidentiality requirements of the data, the confidentiality level parameter Sec(Pi) is introduced, and the parameters are set according to the confidentiality level of the system. In this embodiment, the scheduling order can be calculated based on the following formula:

[0093] Order(Pi)=Sec(Pi)*{(Pri(Pi)-Cpu(Pi)) / Tot(Pi)}

[0094] After calculating the scheduling order, all process control blocks in the ready state are placed in the ready queue according to the scheduling order.

[0095] In step S24, the calculation formula of the scheduling order is:

[0096] NewOrder(Pi)=NewPri(Pi)*Order(Pi)

[0097] Among them, NewPri(Pi) is the priority change parameter.

[0098] Each time before taking a process out of the ready queue for execution, the processes in the queue need to be re-sorted according to their priority to ensure that processes with higher priorities are scheduled first.

[0099] In the data mining and analysis visualization method of this application, in step 3, the design concept of the data warehouse follows the dimensional modeling idea, such as Figure 4 As shown, the data is divided into a basic data layer (STAGE), a common dimensional model layer (CDM) and an application data layer (ADS), wherein the CDM layer includes a detailed data layer (DWD) and a summary data layer (DWS).

[0100] like Figure 5 As shown, in this embodiment, the basic data layer is used to store the collected data and save or clean the historical data according to the data business requirements. The main function of the basic data layer is to synchronize data and store the data from the information system directly in the data warehouse. In the public dimension model layer, the detailed data layer is used to process the data from the basic data layer to generate detailed fact data and dimension table data, and the summary data layer is used to process the detailed fact data and dimension table data to generate public indicator summary data. The application data layer is used to store the statistical indicator data personalized for the data products generated by the basic data layer and the public dimension model layer. Figure 6 As shown, in the data warehouse, the process completion time is introduced as a parameter, and the calculation formula is as follows: End(Pi)=Tot(Pi)+Start(Pi). The task scheduling time in the data warehouse is adjusted to after End(Pi) to ensure that the data update is completed in the fastest time.

[0101] In the data mining and analysis visualization method of this application, in step 4, the synchronization results of the data need to be verified to ensure the validity of data synchronization. The verification work is divided into two parts: the first part is the process verification, which verifies whether all processes have ended. If there is an abnormal end of the process, it will be re-called and an error will be reported to remind relevant personnel; the second part is the data warehouse verification. After the storage process of the data warehouse is executed, the verification execution process is automatically started to detect whether the number of newly added data items matches. If not, a reminder will be sent to relevant personnel.

[0102] The data mining and analysis visualization method of this application, at the end, in step 5, can provide two modes of report visualization and chart visualization according to user needs. The report mode provides users with a detailed list of various types of information, and is attached with a drill-down form. After clicking the interaction, you can drill down to the bottom layer, and calculate statistical information such as the average, median, and percentile, which is convenient for users to view intuitively, such as Figure 7As shown. You can also construct charts, such as line charts to represent trends, bar charts, scatter plots, histograms to represent distributions, pie charts to represent proportions, etc. You can also represent information such as correlations and outliers, which is convenient for users to carry out data mining and analysis.

[0103] Based on the above data mining analysis visualization method, the second aspect of the present application provides a data mining analysis visualization system, such as Figure 8 As shown, including:

[0104] The data collection module is used to collect data from different data sources and convert the collected data into a format that is compatible with the data warehouse;

[0105] A data timing synchronization module, used to determine data timing synchronization;

[0106] The data warehouse construction module is used to synchronize the collected data to the data warehouse at regular intervals;

[0107] The data verification mechanism module is used to verify the data synchronization results to ensure the validity of data synchronization;

[0108] Data visualization module, used to visualize data.

[0109] The data mining and analysis visualization system of the present application has a data collection module that meets the adaptation requirements of different types of databases within the enterprise. The data visualization module is tightly coupled with the data warehouse, and provides users with useful page data after secondary processing at the front end, providing good interactive response.

[0110] The data mining analysis visualization method and system of the present application have the following beneficial effects:

[0111] 1. Provide a complete data mining and analysis visualization method within the enterprise through existing tools to obtain and visualize data from multi-source heterogeneous databases;

[0112] 2. Introduced the priority and time slice rotation algorithm, added a feedback verification mechanism, dynamically adjusted the priority, set up a multi-level feedback queue, and adaptively adjusted the time slice size to ensure that the task is carried out in the best way;

[0113] 3. Build a data warehouse in accordance with the confidentiality requirements within the enterprise. Link the task scheduling time within the data warehouse with the data synchronization completion time to speed up data updates;

[0114] 4. Set up a data verification mechanism, check the results of daily updates, and discover synchronization failures as quickly as possible;

[0115] 5. Convenient visualization, using drag and drop to quickly generate reports required by users, providing users with charts of overall characteristics, trends, correlations, etc. of data sets.

[0116] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.

Claims

1. A data mining analysis visualization method, characterized in that: include: Step 1: Collect data from different data sources and convert the collected data into a format that is compatible with the data warehouse; Step 2: Determine the timing synchronization of data; Step 3: Synchronize the collected data to the data warehouse regularly; Step 4: Verify the data synchronization results to ensure the validity of data synchronization; Step 5: Visualize the data.

2. The data mining analysis visualization method according to claim 1, characterized in that: In step 1, the data sources include: Oracle database, MySQL database, and DAMO database.

3. The data mining analysis visualization method according to claim 1, characterized in that: In step 2, determining the timing synchronization of data includes: S21, configuring corresponding process control blocks for data from different data sources, and calculating the scheduling order according to the process information of the process control blocks, and placing all the process control blocks in the ready state into the ready queue according to the scheduling order; S22, taking out one of the process control blocks from the head of the ready queue, allocating a time slice to it, and setting the state to execution, and the process control block runs on the CPU until the time slice is used up or the process is actively abandoned; S23, if the process control block has completed execution within the time slice, remove it from the ready queue and release the resources it occupies; if the time slice is used up but the process control block has not completed execution, set its state to ready and determine the priority change parameter of the process control block; S24, recalculating the scheduling order according to the priority change parameter of the process control block, placing all the process control blocks in the ready state into the ready queue according to the new scheduling order, and returning to step S22; S25. Repeat steps S22 to S24 until all the process control blocks are executed or the termination condition is met.

4. The data mining analysis visualization method according to claim 3, characterized in that: The process information includes: process ID, total process running time, process used CPU time, process priority, and process start time.

5. The data mining analysis visualization method according to claim 4, characterized in that: In step S21, the scheduling order is: Order(Pi)=Sec(Pi)*{(Pri(Pi)-Cpu(Pi)) / Tot(Pi)} Among them, Order(Pi) is the scheduling order, Pri(Pi) is the process priority, Cpu(Pi) is the CPU time used by the process, Tot(Pi) is the total running time of the process, Start(Pi) is the start time of the process, and Sec(Pi) is the confidentiality parameter.

6. The data mining analysis visualization method according to claim 5, characterized in that: In step S24, the scheduling order is: NewOrder(Pi)=NewPri(Pi)*Order(Pi) Among them, NewPri(Pi) is the priority change parameter.

7. The data mining analysis visualization method according to claim 6, characterized in that: The data warehouse includes: The basic data layer is used to store collected data and save or clean historical data according to data business requirements; The public dimension model layer includes a detailed data layer and a summary data layer, wherein the detailed data layer is used to process the data from the basic data layer to generate detailed fact data and dimension table data, and the summary data layer is used to process the detailed fact data and the dimension table data to generate public indicator summary data; The application data layer is used to store the personalized statistical indicator data of the data products generated by processing the basic data layer and the common dimension model layer; The task scheduling time in the data warehouse is End(Pi), End(Pi)=Tot(Pi)+Start(Pi).

8. The data mining analysis visualization method according to claim 7, characterized in that: In step 4, the data synchronization result is verified, including: Process verification: Verify whether all processes have been completed. If there is an abnormally completed process, call it again and report an error to remind relevant personnel; Data warehouse verification: Verify whether the number of newly added data items matches. If not, send reminders to relevant personnel.

9. The data mining analysis visualization method according to claim 8, characterized in that: In step 5, the data is visualized, including: Visualize the data in reports; Visualize the data in charts.

10. A data mining analysis visualization system, characterized in that: include: The data collection module is used to collect data from different data sources and convert the collected data into a format that is compatible with the data warehouse; A data timing synchronization module, used to determine data timing synchronization; The data warehouse construction module is used to synchronize the collected data to the data warehouse at regular intervals; The data verification mechanism module is used to verify the data synchronization results to ensure the validity of data synchronization; Data visualization module, used to visualize data.

Citation Information

Cited By

  • Index monitoring and early warning method based on big data

    CN120950836A