Data analysis method, data analysis system, storage medium and program product
By dynamically selecting a single-process or multi-process processing mode and adjusting the processing strategy according to the data scale, the problem of performance degradation when a single-machine single-process processing big data is solved, and the performance and resource utilization of data analysis are improved.
Patent Information
- Application Number
- CN202510600204.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-12
AI Technical Summary
With the continuous expansion of data scale, the data analysis and processing methods of single-machine single-process data have led to serious decline in system performance, frequent memory exchanges, reduced computing speed, and low resource utilization.
By calculating the resource occupancy ratio, dynamically select a single-process or multi-process processing mode. For small data sets, a single-process mode is used, and it is directly converted into an R language data frame for processing; for large data sets, multiple R processes are created based on the number of CPU cores, a TCP/IP connection pool is established, and the data is sharded and calculated in parallel.
It realizes dynamic adjustment of processing mode under different data scales, balances resource utilization and processing efficiency, and improves the performance and resource utilization of data analysis.
Smart Images

Figure CN120104358A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data analysis, and in particular to a data analysis method, a data analysis system, a storage medium, and a program product. Background Art
[0002] With the advent of the big data era, data analysis and processing are increasingly widely used. Rapid and efficient analysis and processing of massive data can obtain valuable information and decision support, especially in the fields of finance, medicine, industry, etc., where the real-time and accuracy requirements for data analysis and processing are getting higher and higher.
[0003] At present, data analysis and processing methods are usually carried out in a single-machine single-process manner. After receiving a data analysis and processing request, the system directly loads all data into the memory and then calls the statistical analysis function for calculation. This method is relatively simple to implement and can meet basic analysis and processing requirements when the data volume is small, and the development and maintenance costs are low.
[0004] However, as the scale of data continues to expand, the single-machine single-process processing method will cause a serious decline in system performance. When the amount of data to be analyzed and processed exceeds the available memory of the system, frequent memory swapping will greatly reduce the computing speed, resulting in low utilization of system resources and poor processing efficiency. Summary of the invention
[0005] The present application provides a data analysis method, a data analysis system, a storage medium, and a program product for improving the performance and resource utilization of data analysis.
[0006] In a first aspect, the present application provides a data analysis method, which is applied to a data analysis system, and the method includes: obtaining the current number of CPU cores and a task analysis request input by a user, the task analysis request including a data set to be analyzed, dividing the physical memory size occupied by the data set to be analyzed by the current available physical memory size to obtain a resource occupancy ratio; if the resource occupancy ratio is less than a preset ratio threshold, converting the data set to be analyzed into an R language data frame, calling an R function to process the R language data frame, and obtaining a first analysis result; if the resource occupancy ratio is greater than or equal to the preset ratio threshold, determining the number of sub-processes based on the current number of CPU cores, creating R processes equal to the number of sub-processes, and establishing a TCP / IP connection pool for the R processes; calculating the maximum allowable size of a single data slice according to the current available physical memory size, and dividing the data set to be analyzed into multiple data blocks according to the maximum allowable size; sending the data blocks to the corresponding R processes through the TCP / IP connection pool for parallel calculation, obtaining multiple partial calculation results, merging the multiple partial calculation results to obtain a second analysis result; and visually displaying the first analysis result or the second analysis result.
[0007] By adopting the above technical solution, the data analysis system calculates the resource occupancy ratio and dynamically selects the single-process or multi-process processing mode based on the preset ratio threshold, effectively balancing resource utilization and processing efficiency. When the data set to be analyzed is small, the single-process mode is adopted, and the data set to be analyzed is directly converted into an R language data frame for processing, avoiding the additional overhead caused by multiple processes; when the data set to be analyzed is large, multiple R processes are created based on the current number of CPU cores and a TCP / IP connection pool is established. Through data sharding and parallel computing, the multi-core computing power is fully utilized. This adaptive processing mode not only ensures the processing efficiency of small data sets, but also effectively responds to the computing needs of large data sets, and overall improves the performance and resource utilization of data analysis.
[0008] In combination with some embodiments of the first aspect, in some embodiments, after the step of obtaining the current number of CPU cores and the task analysis request input by the user, the method also includes: obtaining the inherent delay of calling the R function based on rpy2 and the time overhead of memory data sharing, and calculating the total delay of a single process call; obtaining the network delay for data transmission based on Rserve and the time overhead of actually processing data in the R process, and calculating the total delay of multiple process calls; comparing the total delay of a single process call and the total delay of multiple process calls; if the total delay of a single process call is less than the total delay of multiple process calls, adopting the single process mode; if the total delay of a single process call is greater than or equal to the total delay of multiple process calls, adopting the multi-process mode.
[0009] By adopting the above technical solutions, the data analysis system calculates the inherent delay of calling R functions based on rpy2, the time overhead of memory data sharing, the network delay of data transmission based on Rserve, and the time overhead of actual data processing in the R process. By comparing the actual delay overhead of the single-process and multi-process modes, the data analysis system can accurately evaluate the performance of the current task under different processing modes, and thus select the mode with the smallest total delay. This adaptive selection mechanism based on actual delay measurement avoids the performance loss that may be caused by the fixed use of a certain mode, and can select the optimal processing strategy according to the specific task characteristics and system status, effectively improving data processing efficiency.
[0010] In combination with some embodiments of the first aspect, in some embodiments, the maximum allowable size of a single data shard is calculated based on the currently available physical memory size, specifically including: inputting the preset shard size, the empirical coefficient and the currently available physical memory size into the shard size calculation formula to obtain the maximum allowable size of a single data shard; the shard size calculation formula is: B=min(Bmax, k×M); wherein B is used to represent the maximum allowable size of a single data shard, Bmax is used to represent the preset shard size, k is used to represent the empirical coefficient, and M is used to represent the currently available physical memory size.
[0011] By adopting the above technical solution, the data analysis system determines the maximum allowable size of a single data shard through the shard size calculation formula B=min(Bmax, k×M). The shard size calculation formula comprehensively considers three key factors: the preset shard size, the current available physical memory size, and the empirical coefficient, balancing security and efficiency. On the one hand, it avoids the risk of system crash caused by too large shards, and on the other hand, it can be flexibly adjusted according to the system resource status, making full use of available memory, and improving the robustness and efficiency of data processing.
[0012] In combination with some embodiments of the first aspect, in some embodiments, after the step of sending data blocks to corresponding R processes for parallel computing through a TCP / IP connection pool, the method also includes: calculating the network transmission time of a single data block; based on the total size of the data set to be analyzed and the network transmission time, calculating the total data transmission delay; when the current number of CPU cores is greater than a preset value, dividing the total data transmission delay by the current number of CPU cores to obtain the average delay of the R process.
[0013] By adopting the above technical solution, the data analysis system calculates the network transmission time of a single data block, calculates the total data transmission delay based on the total size of the data set to be analyzed, and calculates the average delay of the R process when the current number of CPU cores is large. This multi-level performance monitoring mechanism can accurately evaluate the data transmission efficiency and identify potential performance bottlenecks. Especially in multi-core systems, the actual effect of parallel processing can be evaluated by calculating the average delay, thereby determining whether the multi-core advantage is fully utilized.
[0014] In combination with some embodiments of the first aspect, in some embodiments, when the current number of CPU cores is greater than a preset value, after the step of dividing the total data transmission delay by the current number of CPU cores to obtain the average delay of the R process, the method also includes: obtaining the processing success rate, memory utilization rate and calculation time of the R process to complete data analysis for each data block; dividing the R process into a high-performance process group and a low-performance process group according to the processing success rate, memory utilization rate and calculation time; reclaiming the memory space of the R process in the low-performance process group, initializing the operating environment of the R process in the low-performance process group, and reloading the function library of the R process in the low-performance process group.
[0015] By adopting the above technical solutions, the data analysis system monitors the processing success rate, memory usage, and calculation time of the R process, so as to comprehensively evaluate the performance of each R process and divide each R process into two groups: high-performance and low-performance. The data analysis system performs special optimization processing on the low-performance process group, including memory recycling, environment initialization, and function library reloading. This differentiated processing strategy based on performance grouping not only ensures the continuous and stable operation of high-performance processes, but also can timely discover and optimize low-performance processes, avoid individual processes dragging down the overall performance, and improve the overall operation efficiency and stability.
[0016] In combination with some embodiments of the first aspect, in some embodiments, after the step of obtaining the current number of CPU cores and the task analysis request input by the user, the method also includes: obtaining the variable type of the data set to be analyzed; determining whether the data set to be analyzed includes categorical variables; if categorical variables are included, automatically selecting the chi-square test path, performing the chi-square independence test, and obtaining a first verification result; if categorical variables are not included, automatically selecting the variance analysis path, performing one-way variance analysis, and obtaining a second verification result; based on the first verification result or the second verification result, calculating the effect value, and generating a statistical test report.
[0017] By adopting the above technical solution, the data analysis system automatically identifies the variable type of the data set to be analyzed, and selects the corresponding statistical test method based on whether it includes categorical variables: if categorical variables are included, the data analysis system automatically performs a chi-square independence test; if categorical variables are not included, a one-way analysis of variance is performed. This automated analysis path selection not only reduces the user's operating difficulty, but also ensures the scientificity and accuracy of the statistical analysis method. At the same time, the data analysis system calculates the effect value and generates a complete statistical test report, providing users with comprehensive data analysis results, greatly improving the professionalism and usability of data analysis.
[0018] In combination with some embodiments of the first aspect, in some embodiments, after the step of sending data blocks to corresponding R processes for parallel computing through the TCP / IP connection pool, the method also includes: periodically detecting the running status of the R process; when the R process exits abnormally, obtaining the data processing progress of the R process, and saving the completed data analysis results to the intermediate result cache area, and recording the error type of the abnormal exit.
[0019] By adopting the above technical solution, the data analysis system periodically detects the running status of the R process, so that process anomalies can be discovered in time. When the R process exits abnormally, the data analysis system automatically obtains the data processing progress, saves the completed data analysis results to the intermediate result buffer, and records the error type. This meticulous exception handling mechanism not only ensures that the completed work results will not be lost, but also provides an important reference for subsequent fault analysis through error type records, greatly improving fault tolerance and reliability. Even in abnormal situations, it can protect data and analysis results to the maximum extent, ensuring the continuity and recoverability of data analysis tasks.
[0020] In a second aspect, an embodiment of the present application provides a data analysis system, the data analysis system comprising: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code comprises computer instructions, and the one or more processors call the computer instructions to enable the data analysis system to execute the method described in the first aspect and any possible implementation method of the first aspect.
[0021] In a third aspect, an embodiment of the present application provides a computer program product comprising instructions, which, when the computer program product is run on a data analysis system, enables the data analysis system to execute the method described in the first aspect and any possible implementation of the first aspect.
[0022] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, comprising instructions. When the instructions are executed on a data analysis system, the data analysis system executes the method described in the first aspect and any possible implementation manner of the first aspect.
[0023] It is understandable that the data analysis system provided in the second aspect, the computer program product provided in the third aspect, and the computer storage medium provided in the fourth aspect are all used to execute the method provided in the embodiment of the present application. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method, which will not be repeated here.
[0024] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: 1. By adopting the above technical solution, the data analysis system calculates the resource occupancy ratio and dynamically selects the single-process or multi-process processing mode based on the preset ratio threshold, effectively balancing resource utilization and processing efficiency. When the data set to be analyzed is small, the single-process mode is adopted, and the data set to be analyzed is directly converted into an R language data frame for processing, avoiding the additional overhead caused by multiple processes; when the data set to be analyzed is large, multiple R processes are created based on the current number of CPU cores and a TCP / IP connection pool is established. Through data sharding and parallel computing, the multi-core computing power is fully utilized. This adaptive processing mode not only ensures the processing efficiency of small data sets, but also effectively responds to the computing needs of large data sets, and overall improves the performance of data analysis and resource utilization.
[0025] 2. By adopting the above technical solution, the data analysis system determines the maximum allowable size of a single data shard through the shard size calculation formula B=min(Bmax, k×M). The shard size calculation formula comprehensively considers three key factors: the preset shard size, the current available physical memory size, and the empirical coefficient, balancing security and efficiency. On the one hand, it avoids the risk of system crashes caused by overly large shards, and on the other hand, it can be flexibly adjusted according to the system resource status, making full use of available memory, and improving the robustness and efficiency of data processing.
[0026] 3. By adopting the above technical solutions, the data analysis system monitors the processing success rate, memory usage, calculation time and other multi-dimensional indicators of the R process, so as to comprehensively evaluate the performance of each R process and divide each R process into two groups: high performance and low performance. The data analysis system performs special optimization processing on the low-performance process group, including memory recycling, environment initialization and function library reloading. This differentiated processing strategy based on performance grouping not only ensures the continuous and stable operation of high-performance processes, but also can timely discover and optimize low-performance processes, avoid individual processes dragging down the overall performance, and improve the overall operation efficiency and stability. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 It is a flow chart of the data analysis method in the embodiment of the present application; Figure 2 is another flow chart of the data analysis method in the embodiment of the present application; Figure 3 It is a schematic diagram of a physical device structure of a data analysis system in an embodiment of the present application. DETAILED DESCRIPTION
[0028] The terms used in the following embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to be used as limitations to the present application. As used in the specification of the present application, the singular expressions "one", "a kind of", "above", "the" and "this" are intended to also include plural expressions, unless there is a clear indication to the contrary in the context. It should also be understood that the term "and / or" used in the present application refers to any or all possible combinations comprising one or more of the listed items.
[0029] In the following, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as suggesting or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features, and in the description of the embodiments of the present application, unless otherwise specified, "plurality" means two or more.
[0030] The following is a description of the process of the method provided by this implementation. Figure 1 , which is a flow chart of the data analysis method in the embodiment of the present application.
[0031] S101, obtaining the current number of CPU cores and a task analysis request input by a user, wherein the task analysis request includes a data set to be analyzed, and dividing the physical memory size occupied by the data set to be analyzed by the current available physical memory size to obtain a resource occupancy ratio; Among them, the current number of CPU cores is used to indicate the number of processor cores in the computer that can execute instructions in parallel, including the number of physical cores and logical cores; task analysis request refers to the task application submitted by the user that requires data analysis and processing, including the data set to be processed and the data analysis and processing requirements; the data set to be analyzed refers to the original data set that needs to be analyzed and processed; the physical memory size refers to the random access memory (RAM) capacity actually installed in the computer; the current available physical memory size refers to the free memory capacity that can be allocated to the system at this moment; the resource utilization ratio refers to the ratio between the memory required for the data set to be analyzed and the available memory of the system.
[0032] When the data analysis system receives a task analysis request submitted by a user, it first needs to evaluate the system resource status to determine the subsequent processing strategy. Specifically, first, the data analysis system obtains the current number of CPU cores through the operating system interface, and at the same time receives and parses the task analysis request input by the user, and extracts the data set to be analyzed. Then, the data analysis system calculates the memory space required when the data set to be analyzed is loaded into the memory, and obtains the free memory capacity available in the system. Finally, the data analysis system divides the physical memory size occupied by the data set to be analyzed by the current available physical memory size to obtain a resource utilization rate indicator, which is used to determine whether optimization measures such as data sharding are needed.
[0033] S102, if the resource occupancy ratio is less than a preset ratio threshold, converting the data set to be analyzed into an R language data frame, calling an R function to process the R language data frame, and obtaining a first analysis result; Among them, the preset ratio threshold is a pre-set critical value of resource usage, which is used to determine whether the data sharding mechanism needs to be enabled; the R language data frame refers to the data structure used to store two-dimensional tabular data in the R language, and supports mixed storage of different types of data; the R function is used to represent various types of data processing and statistical analysis functions predefined in the R language environment; the first analysis result refers to the data analysis result obtained by direct processing in single-process mode.
[0034] When the data analysis system determines that the scale of the data set to be analyzed is relatively small and will not cause great pressure on system resources, a simple and direct single-process processing mode is adopted. Specifically, first, the data analysis system compares the calculated resource occupancy ratio with the preset ratio threshold. If the resource occupancy ratio is less than the preset ratio threshold, it indicates that the current system resources are sufficient. Next, the data analysis system converts the original data set to be analyzed into a data frame format dedicated to the R language. This process includes data type conversion, missing value processing, etc. Then, the data analysis system calls the corresponding R function to process the R language data frame according to the data analysis and processing requirements, which may include descriptive statistics, hypothesis testing, regression analysis and other operations. Finally, the data analysis system marks the obtained data analysis results as the first analysis results.
[0035] S103, if the resource occupancy ratio is greater than or equal to the preset ratio threshold, determining the number of child processes based on the current number of CPU cores, creating R processes equal to the number of child processes, and establishing a TCP / IP connection pool for the R processes; Among them, the number of subprocesses refers to the number of independent R language processes that the data analysis system will create, which is usually determined based on the current number of CPU cores; R process refers to an independently running R language runtime environment instance that can execute R language code and functions; TCP / IP connection pool refers to a resource pool used to manage multiple network connections, supporting data transmission between different processes.
[0036] When the data analysis system finds that the data set to be analyzed is large and may cause great pressure on system resources, it is necessary to enable a parallel processing mechanism. Specifically, first, the data analysis system determines whether the resource occupancy ratio reaches or exceeds the preset ratio threshold. If so, a multi-process processing strategy needs to be adopted. Then, the data analysis system calculates the optimal number of sub-processes based on the current number of CPU cores, and usually sets the number of sub-processes to a multiple or approximate number of the current number of CPU cores. Next, the data analysis system creates a corresponding number of R processes, each of which is an independent computing unit. Finally, the data analysis system establishes a TCP / IP connection pool to manage network communications between these R processes, including establishing connections, maintaining connection status, allocating connection resources, etc., in preparation for subsequent data distribution and result collection.
[0037] S104, calculating the maximum allowable size of a single data slice according to the current available physical memory size, and dividing the data set to be analyzed into multiple data blocks according to the maximum allowable size; Among them, a single data shard refers to a smaller data block obtained by splitting a large data set to be analyzed; the maximum allowed size refers to the maximum memory space that a single data shard can occupy; and a data block refers to a data subset obtained after dividing according to the maximum allowed size.
[0038] When the data analysis system needs to process a large-scale data set to be analyzed in parallel, data sharding must be performed first. Specifically, first, the data analysis system obtains the current available physical memory size, combines the preset shard size and the empirical coefficient, and calculates the maximum allowable size of a single data shard through the shard size calculation formula. The maximum allowable size must ensure that the sharded data block can be effectively processed by a single R process while avoiding memory pressure. Then, the data analysis system divides the original data set to be analyzed into multiple data blocks of similar size based on the calculated maximum allowable size of a single data shard. The division process will take into account the integrity of the data to avoid dispersing related data into different blocks. The data analysis system will also assign a unique identifier to each data block for subsequent data distribution and result association.
[0039] Optionally, in general, based on the current available physical memory size, the maximum allowable size of a single data shard can be calculated in the following way, which is not limited here: input the preset shard size, the empirical coefficient and the current available physical memory size into the shard size calculation formula to obtain the maximum allowable size of a single data shard; the shard size calculation formula is: B=min(Bmax, k×M); wherein B is used to represent the maximum allowable size of a single data shard, Bmax is used to represent the preset shard size, k is used to represent the empirical coefficient, and M is used to represent the current available physical memory size.
[0040] The following is a specific example to illustrate the calculation process of this data sharding: Assume the following parameters: Current available physical memory size (M) = 16GB = 16384MB; Default shard size (Bmax) = 4GB = 4096MB; Empirical coefficient (k) = 0.25; Use the shard size calculation formula B=min(Bmax, k×M) to calculate: B==4096MB; Therefore, the maximum allowed size of a single data shard is 4096MB (4GB).
[0041] Assume that there is a 32GB data set to be analyzed and processed, which can be divided into: Data block 1: 0-4GB; Data block 2: 4-8 GB; Data block 3: 8-12GB; Block 4: 12-16 GB; Block 5: 16-20 GB; Block 6: 20-24 GB Block 7: 24-28 GB Data block 8: 28-32GB.
[0042] S105, sending the data blocks to the corresponding R processes through the TCP / IP connection pool for parallel computing, obtaining multiple partial computing results, and merging the multiple partial computing results to obtain a second analysis result; Among them, TCP / IP connection pool refers to a resource pool used to manage multiple network connections, supporting reliable data transmission between different processes; parallel computing refers to the process in which multiple processing units execute computing tasks simultaneously; some calculation results are used to represent the intermediate results obtained after each R process completes the processing of its own data block; the second analysis result refers to the final analysis result obtained by merging multiple processes after parallel processing.
[0043] When the data sharding is completed, the data analysis system starts to perform parallel data processing. Specifically, first, the data analysis system allocates a dedicated network connection to each R process through the TCP / IP connection pool to ensure the reliability of data transmission. Then, the data analysis system sends each data block to different R processes through the corresponding network connection. The sending process includes steps such as data serialization, network transmission, and data reception confirmation. After receiving the data, each R process independently executes the same analysis and processing logic. During the processing, the data analysis system monitors the execution status and progress of each R process. When all R processes are completed, the data analysis system collects the partial calculation results generated by each R process, and combines these partial calculation results into a complete second analysis result according to a predetermined merging strategy. The merging process may include data deduplication, sorting, aggregation and other operations.
[0044] S106: Visually display the first analysis result or the second analysis result.
[0045] Among them, the first analysis result refers to the data analysis result obtained by direct processing through a single-process mode; the second analysis result refers to the data analysis result obtained after parallel processing of multiple processes; and visual display refers to converting the data analysis results into intuitive display forms such as charts and graphs.
[0046] When the data analysis system completes the data analysis processing, it needs to present the data analysis results to the user in an intuitive way. Specifically, first, the data analysis system selects the corresponding data analysis results (first analysis results or second analysis results) according to the processing mode (single-process mode or multi-process mode). Then, the data analysis system determines the characteristics of the data analysis results, including data type, data distribution, relationship between data, etc., and automatically selects the most suitable visualization type. The visualization type is used to represent different data display methods, such as line charts, bar charts, scatter plots, etc. Next, the data analysis system converts the data analysis results into the selected visualization form, and performs data normalization, coordinate axis setting, legend generation and other operations during the conversion process. Finally, the data analysis system generates an interactive visualization interface that supports users to perform operations such as data filtering, detail viewing, and view adjustment to ensure that users can easily understand and explore the data analysis results.
[0047] By adopting the above technical solution, the data analysis system calculates the resource occupancy ratio and dynamically selects the single-process or multi-process processing mode based on the preset ratio threshold, effectively balancing resource utilization and processing efficiency. When the data set to be analyzed is small, the single-process mode is adopted, and the data set to be analyzed is directly converted into an R language data frame for processing, avoiding the additional overhead caused by multiple processes; when the data set to be analyzed is large, multiple R processes are created based on the current number of CPU cores and a TCP / IP connection pool is established. Through data sharding and parallel computing, the multi-core computing power is fully utilized. This adaptive processing mode not only ensures the processing efficiency of small data sets, but also effectively responds to the computing needs of large data sets, and overall improves the performance and resource utilization of data analysis.
[0048] The following is a more detailed description of the process of the method provided by this implementation. Figure 2 , is another flow chart of the data analysis method in the embodiment of the present application.
[0049] S201. Obtain the current number of CPU cores and a task analysis request input by a user, where the task analysis request includes a data set to be analyzed.
[0050] For details, please refer to step S101, which will not be described in detail here.
[0051] S202: Obtain the variable type of the data set to be analyzed, and determine whether the data set to be analyzed includes a categorical variable.
[0052] Among them, the variable type is used to represent the data attributes of each field in the data set to be analyzed, including continuous variables, discrete variables, categorical variables, etc.; categorical variables refer to variables with a finite number of categories, such as gender, education level, etc.; continuous variables refer to variables that can take any value, such as height, weight, etc.; discrete variables refer to variables that can only take specific values, such as age, number, etc.
[0053] Before the data analysis system starts statistical analysis, it is necessary to determine the appropriate analysis method. Specifically, first, the data analysis system reads the metadata information of the data set to be analyzed and identifies the data type and value characteristics of each variable. Then, the data analysis system performs detailed type judgment on each variable, including checking whether there are predefined category labels, analyzing the discrete degree of the numerical values, evaluating the continuity of the values, etc. For variables that may be ambiguous, the data analysis system will further analyze their data distribution characteristics and statistical characteristics to accurately determine their types. Finally, the data analysis system determines whether the data set to be analyzed contains categorical variables. This judgment result will determine the statistical test method used subsequently.
[0054] S203. If a categorical variable is included, a chi-square test path is automatically selected, and a chi-square independence test is performed to obtain a first test result.
[0055] Among them, the chi-square test path refers to the processing flow of using the chi-square statistical method for data analysis; the chi-square independence test refers to the statistical method used to test whether there is a correlation between two categorical variables; the first verification result refers to the statistical analysis result obtained by the chi-square test, including the chi-square value, degrees of freedom, P value, etc.
[0056] When the data analysis system detects that the data set to be analyzed contains categorical variables, a statistical method suitable for categorical data analysis is used. Specifically, first, the data analysis system establishes a contingency table of categorical variables and counts the observed frequencies of each category combination. Then, the data analysis system calculates the expected frequency of each cell, which is used to represent the theoretical expected observation value, and compares the expected frequency with the actual frequency (the actual frequency refers to the actual observed data frequency). Next, the data analysis system calculates the chi-square statistic, which includes calculating the difference between the observed value and the expected value, standardization, and summation. The data analysis system also calculates the degrees of freedom and finds the corresponding critical value. Finally, the data analysis system determines the P value based on the calculation results, and organizes the chi-square value, degrees of freedom, P value and other information into the first verification result.
[0057] S204: If no categorical variables are included, the variance analysis path is automatically selected, and one-way variance analysis is performed to obtain a second verification result.
[0058] Among them, the variance analysis path refers to the processing flow of using the variance analysis method for data analysis; one-way variance analysis refers to the statistical method of studying the impact of one factor on the observed variable; the second verification result refers to the statistical result obtained through variance analysis, including F value, between-group variance, and within-group variance.
[0059] When the data analysis system confirms that the data set to be analyzed does not contain categorical variables, a statistical method suitable for continuous data analysis is used. Specifically, first, the data analysis system calculates the descriptive statistics of each group, including the mean, standard deviation, etc. Then, the data analysis system calculates the sum of squares between groups and the sum of squares within groups, which involves the calculation of the deviation of the data from the overall mean and the group mean. Next, the data analysis system calculates the degrees of freedom between groups and the degrees of freedom within groups, and obtains the mean square value based on this. Then, the data analysis system calculates the F statistic and determines the corresponding critical value based on the degrees of freedom. Finally, the data analysis system determines the P value and organizes the statistics such as the F value, degrees of freedom, and P value into the second verification result.
[0060] S205. Calculate the effect value based on the first verification result or the second verification result, and generate a statistical test report.
[0061] Among them, the effect value refers to the actual impact of the statistical test results, which is used to measure the size of the effect; the statistical test report refers to the technical document containing the complete statistical analysis results.
[0062] When the data analysis system completes the statistical test, it is necessary to explain the test results in depth and form a report. Specifically, first, the data analysis system selects the corresponding effect value calculation formula according to the selected test method (chi-square test or analysis of variance). Then, the data analysis system calculates the size of the effect value, and this process may involve the standardization of multiple statistics. Next, the data analysis system determines the degree of effect based on the size of the effect value and gives corresponding explanations. The data analysis system also integrates all statistical analysis results, including descriptive statistics, test statistics, effect values and other information. Finally, the data analysis system generates a standardized statistical test report, which includes complete content such as data overview, analysis method description, test results, effect analysis, and conclusion interpretation.
[0063] S206. Obtain the inherent delay of calling the R function based on rpy2 and the time overhead of memory data sharing, and calculate the total delay of a single process call; obtain the network delay of data transmission based on Rserve and the time overhead of actually processing data in the R process, and calculate the total delay of multi-process calls.
[0064] Among them, rpy2 is used to represent the interface tool for Python to call R language; inherent delay refers to the inevitable system overhead when calling R functions; the time overhead of memory data sharing refers to the time consumed in transferring and converting data in memory; the total delay of single-process call refers to the total time consumed in processing data using single-process mode; Rserve is used to represent the network service interface of R language; network delay refers to the time loss of data during network transmission; the actual time overhead of processing data refers to the actual time consumed by the R process to perform data analysis; the total delay of multi-process calls refers to the total time consumed in processing data using multi-process mode.
[0065] Before the data analysis system selects a processing mode, it is necessary to evaluate the time efficiency of different processing methods. Specifically, first, the data analysis system measures the basic delay of calling the R function through the rpy2 interface through a test program, including the time consumption of function call initialization, parameter passing, result return, etc., and obtains the inherent delay of calling the R function based on rpy2. Then, the data analysis system measures the time overhead of data sharing and conversion in memory, adds these two delays, and obtains the total delay of a single process call. Next, the data analysis system measures the network transmission delay based on Rserve, including the time consumption of connection establishment, data serialization, network transmission, data reception, etc., and obtains the network delay of data transmission based on Rserve. Finally, the data analysis system measures the time overhead of the R process actually processing data, adds these two delays, and obtains the total delay of multi-process calls.
[0066] S207: Compare the total delay of a single process call and the total delay of multiple process calls.
[0067] After the data analysis system obtains the latency data of the two processing modes, it needs to make a performance comparison. Specifically, first, the data analysis system ensures that the latency data of the two processing modes are comparable data measured under the same conditions. Then, the data analysis system compares the total latency of a single process call and the total latency of multiple process calls. This comparison process will take into account measurement errors and system fluctuations, and may require multiple measurements to take the average value.
[0068] S208: If the total delay of a single process call is less than the total delay of multiple process calls, the single process mode is adopted.
[0069] Among them, the single-process mode refers to the operating mode of using a single process to complete data processing.
[0070] When the single-process mode shows better time efficiency, the data analysis system needs to configure the corresponding processing environment. Specifically, first, the data analysis system confirms that it is running in single-process mode, at which time the created multi-process resources will be closed or recycled. Then, the data analysis system initializes the single-process operating environment, including configuring memory allocation, loading necessary function libraries, etc. Next, the data analysis system establishes the rpy2 call interface to ensure normal communication between the Python and R language environments. Finally, the data analysis system prepares the data processing flow, including the specific implementation plans for steps such as data loading, conversion, and analysis.
[0071] S209: If the total delay of a single process call is greater than or equal to the total delay of a multi-process call, the multi-process mode is adopted.
[0072] Among them, the multi-process mode refers to the operating mode of using multiple parallel processes to complete data processing.
[0073] When the multi-process mode shows better time efficiency, the data analysis system needs to configure a parallel processing environment. Specifically, first, the data analysis system confirms that it is running in multi-process mode and starts the process creation and resource allocation mechanism. Then, the data analysis system determines the optimal number of processes and resource allocation scheme based on the current number of CPU cores and memory conditions. Next, the data analysis system initializes the operating environment of each process, including configuring the Rserve service, establishing a network connection, etc. Finally, the data analysis system prepares the parallel processing flow, including specific implementation schemes for steps such as data sharding, task allocation, and result merging, and establishes a coordination mechanism between processes.
[0074] S210. If the resource occupancy ratio is greater than or equal to the preset ratio threshold, the number of child processes is determined based on the current number of CPU cores, R processes equal to the number of child processes are created, and a TCP / IP connection pool is established for the R processes; based on the current available physical memory size, the maximum allowable size of a single data slice is calculated, and the data set to be analyzed is divided into multiple data blocks according to the maximum allowable size.
[0075] For details, please refer to steps S103 and S104, which will not be described in detail here.
[0076] S211. Send the data blocks to the corresponding R processes through the TCP / IP connection pool for parallel computing.
[0077] For details, please refer to step S105, which will not be described in detail here.
[0078] S212: Calculate the network transmission time of a single data block.
[0079] The network transmission time refers to the time required for a data block to be transmitted in the network.
[0080] Specifically, first, the data analysis system selects a representative data block as a test sample. Then, the data analysis system records the timestamp when the data block transmission starts. This process includes data serialization (data serialization refers to the conversion of data into a transmittable format) and the time when the network connection is established. Next, the data analysis system monitors the entire process of data transmission, recording each time point of data packet sending, transmission and reception confirmation. Finally, the data analysis system calculates the total time from the start of sending to the completion of receiving. This total time includes the time consumed by all links such as data serialization, network transmission, and data deserialization.
[0081] S213: Calculate the total data transmission delay based on the total size of the data set to be analyzed and the network transmission time.
[0082] Among them, the total size of the data set to be analyzed refers to the capacity of the complete data set that needs to be processed; the network transmission time refers to the transmission time of a single data block; and the total data transmission delay refers to the total time required to transmit the complete data set to be analyzed.
[0083] Specifically, first, the data analysis system calculates the number of transmission batches required based on the total size of the data set to be analyzed and the size of a single data block. Then, the data analysis system considers the parallel transmission capability of the network and determines the number of data streams that can be transmitted simultaneously. Next, the data analysis system multiplies the network transmission time of a single data block by the number of transmission batches, and considers the time savings brought by parallel transmission. Finally, the data analysis system comprehensively considers factors such as network load and transmission competition, and calculates the total delay time for transmitting the complete data set to be analyzed, that is, the total data transmission delay.
[0084] S214. When the current number of CPU cores is greater than a preset value, the total data transmission delay is divided by the current number of CPU cores to obtain the average delay of the R process.
[0085] Among them, the current number of CPU cores refers to the number of processor cores available in the current system; the preset value refers to the predefined CPU core number threshold; and the average delay refers to the average transmission delay when each R process processes data.
[0086] Specifically, first, the data analysis system compares the current number of CPU cores with the preset value to confirm whether it is necessary to calculate the average latency of the R process. Then, the data analysis system considers the actual availability and performance characteristics of the CPU cores, and may need to exclude system reserved cores or performance-limited cores. Next, the data analysis system divides the previously calculated total data transmission delay by the effective number of CPU cores to obtain the average latency of each R process. This value reflects the expected data transmission time cost of each process when multiple processes are processed in parallel.
[0087] S215, obtaining the processing success rate, memory usage rate, and calculation time required to complete data analysis for each data block of the R process.
[0088] The processing success rate is used to indicate the proportion of R processes that successfully complete data analysis tasks; the memory usage rate refers to the extent to which the R process occupies system memory resources; and the calculation duration is used to indicate the actual time required for the R process to process a single data block.
[0089] Specifically, first, the data analysis system establishes a performance monitoring mechanism to track the running status of each R process in real time, and records the number of successes and failures of each R process in processing data tasks to calculate the processing success rate. Then, the data analysis system monitors the memory usage of each R process through the system interface, and records the peak and average memory usage. At the same time, the data analysis system will also accurately record the start and end time of each R process processing a data block to calculate the actual computing time.
[0090] S216. According to the processing success rate, memory usage rate, and calculation time, the R process is divided into a high-performance process group and a low-performance process group.
[0091] Among them, the high-performance process group is used to represent a set of R processes with good running status; the low-performance process group is used to represent a set of R processes with poor performance that needs to be optimized.
[0092] Specifically, first, the data analysis system determines the weight of each performance indicator to establish a comprehensive scoring mechanism. Then, the data analysis system standardizes the processing success rate, memory usage, and calculation time of each R process and calculates the weighted score. Next, the data analysis system sets the threshold standard for performance grouping, which may need to consider system load and task requirements. Finally, the data analysis system divides the process into the corresponding performance group based on the scoring results. This grouping process is dynamic and will be updated as the process performance changes.
[0093] S217. Reclaim the memory space of the R process in the low-performance process group, initialize the operating environment of the R process in the low-performance process group, and reload the function library of the R process in the low-performance process group.
[0094] Among them, memory space is used to represent the system memory resources occupied by the process; the operating environment is used to represent the configuration and parameters required for the process to run; the function library is used to represent the set of functional modules of the R language; environment initialization refers to resetting the running status of the process; resource recovery refers to releasing the system resources occupied by the process.
[0095] Specifically, first, the data analysis system forcibly reclaims the memory resources occupied by each process in the low-performance process group through the operating system interface, including cleaning up the memory cache and temporary data. Then, the data analysis system reinitializes the operating environment of these processes, including resetting environment variables, cleaning up session status, and restoring default configurations. Finally, the data analysis system reloads the necessary R language function library to ensure that the latest version of the library file is loaded, and verifies the integrity and availability of the library file.
[0096] S218. Periodically check the running status of the R process; when the R process exits abnormally, obtain the data processing progress of the R process, save the completed data analysis results to the intermediate result buffer, and record the error type of the abnormal exit.
[0097] Among them, abnormal exit refers to the situation where the R process terminates abnormally; data processing progress refers to the proportion of completed analysis tasks; intermediate result buffer refers to the storage space for temporarily storing analysis results; error type refers to the specific reason classification that causes the R process to exit.
[0098] When the data analysis system detects the running status of the R process at preset time intervals and detects a process abnormality, emergency processing is required. Specifically, first, the data analysis system captures the signal of the abnormal exit of the R process and quickly saves the current processing status. Then, the data analysis system obtains the amount of data processed by the R process and the current processing stage through the progress recording mechanism, and quickly transfers the completed analysis results to a dedicated cache area. This process needs to ensure the integrity and consistency of the data. At the same time, the data analysis system will record the specific causes of the abnormality in detail, including error codes, error messages, system status and other information.
[0099] The data analysis system in the embodiment of the present invention is described below from the perspective of hardware processing. Figure 3 , which is a schematic diagram of a physical device structure of a data analysis system in an embodiment of the present application.
[0100] It should be noted that Figure 3 The structure of the data analysis system shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.
[0101] like Figure 3 As shown, the data analysis system includes a CPU 301, which can perform various appropriate actions and processes according to the program stored in the read-only memory ROM 302 or the program loaded from the storage part 308 to the random access memory RAM 303, such as executing the method described in the above embodiment. Various programs and data required for system operation are also stored in the RAM 303. The CPU 301, ROM 302 and RAM 303 are connected to each other via a bus 304. An I / O interface 305 is also connected to the bus 304.
[0102] The following components are connected to the I / O interface 305: an input section 306 including an audio input device, a button switch, etc.; an output section 307 including a liquid crystal display (LCD) and an audio output device, an indicator light, etc.; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to the I / O interface 305 as needed. A removable medium 311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 310 as needed so that a computer program read therefrom is installed into the storage section 308 as needed.
[0103] In particular, according to an embodiment of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product, which includes a computer program carried on a computer readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 309, and / or installed from the removable medium 311. When the computer program is executed by the CPU 301, various functions defined in the present invention are executed.
[0104] It should be noted that specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, apparatus, or device.
[0105] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present invention. Each box in the flowchart or block diagram may represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of a code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the box may also occur in an order different from that marked in the accompanying drawings.
[0106] Specifically, the data analysis system of this embodiment includes a processor and a memory. The memory stores a computer program. When the computer program is executed by the processor, the data analysis method provided in the above embodiment is implemented.
[0107] As another aspect, the present invention further provides a computer-readable storage medium, which may be included in the data analysis system described in the above embodiment; or may exist independently without being assembled into the data analysis system. The above storage medium carries one or more computer programs, and when the above one or more computer programs are executed by a processor of the data analysis system, the data analysis system implements the data analysis method provided in the above embodiment.
[0108] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
[0109] As used in the above embodiments, the term "when..." may be interpreted to mean "if..." or "after..." or "in response to determining..." or "in response to detecting...", depending on the context. Similarly, the phrases "upon determining..." or "if (the stated condition or event) is detected" may be interpreted to mean "if determining..." or "in response to determining..." or "upon detecting (the stated condition or event)" or "in response to detecting (the stated condition or event)", depending on the context.
[0110] Those skilled in the art can understand that to implement all or part of the processes in the above-mentioned embodiments, the processes can be completed by computer programs to instruct related hardware, and the programs can be stored in computer-readable storage media. When the programs are executed, they can include the processes of the above-mentioned method embodiments. The aforementioned storage media include: ROM or random access memory RAM, magnetic disk or optical disk and other media that can store program codes.
Claims
1. A data analysis method, characterized in that: Applied to a data analysis system, the method comprises: Obtain the current number of CPU cores and a task analysis request input by a user, wherein the task analysis request includes a data set to be analyzed, and divide the physical memory size occupied by the data set to be analyzed by the current available physical memory size to obtain a resource occupancy ratio; If the resource occupancy ratio is less than a preset ratio threshold, converting the data set to be analyzed into an R language data frame, calling an R function to process the R language data frame, and obtaining a first analysis result; If the resource occupancy ratio is greater than or equal to a preset ratio threshold, the number of sub-processes is determined based on the current number of CPU cores, R processes equal to the number of sub-processes are created, and a TCP / IP connection pool is established for the R processes; Calculate the maximum allowable size of a single data slice according to the currently available physical memory size, and divide the data set to be analyzed into multiple data blocks according to the maximum allowable size; Sending the data blocks to corresponding R processes respectively through the TCP / IP connection pool for parallel computing to obtain multiple partial computing results, and merging the multiple partial computing results to obtain a second analysis result; Visually display the first analysis result or the second analysis result.
2. The method according to claim 1, characterized in that After the step of obtaining the current number of CPU cores and the task analysis request input by the user, the method further includes: Obtain the inherent delay of calling R functions based on rpy2 and the time overhead of memory data sharing, and calculate the total delay of single process calls; Obtain the network delay of data transmission based on Rserve and the time overhead of actual data processing in the R process, and calculate the total delay of multi-process calls; Comparing the total delay of the single process call and the total delay of the multi-process call; If the total delay of the single-process call is less than the total delay of the multi-process call, the single-process mode is adopted; If the total delay of the single-process call is greater than or equal to the total delay of the multi-process call, the multi-process mode is adopted.
3. The method according to claim 1, characterized in that The calculating the maximum allowable size of a single data slice according to the currently available physical memory size specifically includes: Inputting the preset shard size, the empirical coefficient and the current available physical memory size into a shard size calculation formula to obtain the maximum allowable size of the single data shard; The shard size calculation formula is: B=min(Bmax, k×M); Among them, B is used to represent the maximum allowed size of the single data slice, Bmax is used to represent the preset slice size, k is used to represent the empirical coefficient, and M is used to represent the current available physical memory size.
4. The method according to claim 1, characterized in that: After the step of sending the data blocks to the corresponding R processes for parallel computing through the TCP / IP connection pool, the method further includes: Calculate the network transmission time of a single data block; Calculating a total data transmission delay based on the total size of the data set to be analyzed and the network transmission duration; When the current number of CPU cores is greater than a preset value, the total data transmission delay is divided by the current number of CPU cores to obtain the average delay of the R process.
5. The method according to claim 4, characterized in that When the current number of CPU cores is greater than a preset value, after the step of dividing the total data transmission delay by the current number of CPU cores to obtain the average delay of the R process, the method further includes: Get the processing success rate, memory usage, and computing time of each data block to complete data analysis of the R process; According to the processing success rate, the memory usage rate and the calculation duration, the R process is divided into a high-performance process group and a low-performance process group; The memory space of the R process in the low-performance process group is recovered, the operating environment of the R process in the low-performance process group is initialized, and the function library of the R process in the low-performance process group is reloaded.
6. The method according to claim 1, characterized in that After the step of obtaining the current number of CPU cores and the task analysis request input by the user, the method further includes: Obtaining the variable type of the data set to be analyzed; Determining whether the data set to be analyzed includes categorical variables; If categorical variables are included, the chi-square test path is automatically selected, and the chi-square independence test is performed to obtain the first test result; If categorical variables are not included, the variance analysis path is automatically selected, and a one-way variance analysis is performed to obtain the second verification result; Based on the first verification result or the second verification result, the effect value is calculated and a statistical test report is generated.
7. The method according to claim 1, characterized in that After the step of sending the data blocks to the corresponding R processes for parallel computing through the TCP / IP connection pool, the method further includes: Periodically check the running status of the R process; When the R process exits abnormally, the data processing progress of the R process is obtained, and the completed data analysis results are saved to the intermediate result buffer area, and the error type of the abnormal exit is recorded.
8. A data analysis system, characterized in that: The data analysis system comprises: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code comprises computer instructions, and the one or more processors call the computer instructions so that the data analysis system executes the method as described in any one of claims 1-7.
9. A computer-readable storage medium comprising instructions, characterized in that: When the instructions are executed on a data analysis system, the data analysis system is caused to execute the method according to any one of claims 1 to 7.
10. A computer program product, characterized in that When the computer program product is run on a data analysis system, the data analysis system is caused to perform the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Resource scheduling method and device, computer equipment and readable storage medium
CN118708344A
Job execution method, apparatus and device in distributed scene, and program product
CN119668692A
Data processing method and apparatus, device, and storage medium
US20220342929A1