Data analysis software architecture design method

Through a hierarchical architecture and adaptive data transmission mechanism, combined with rpy2 and Rserve interfaces, cross-language call and interface response problems are solved, modular design is realized, the user experience and scalability of statistical analysis tools is improved, and the friendly and professional needs of non-programmer users are met.

CN120371266AActive Publication Date: 2025-07-25BEIJING FENGRUIKELIN MEDICAL TECH CO LTD

Patent Information

Application Number
CN202510410608.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-25
Estimated Expiration
2045-04-02

AI Technical Summary

Technical Problem

Existing statistical analysis tools cannot adaptively optimize performance in cross-language calls, resulting in unnecessary increase in communication overhead in small data sets, excessive memory usage or call failure in large data sets; the problem of parallel interface response and computing has not been effectively solved, affecting the user experience; the module is insufficient in scalability and maintenance, making it difficult to meet the friendly and professional needs of non-programmer users.

Method used

The hierarchical architecture design is adopted, including the model layer, view layer and controller layer, combined with the adaptive selection strategy of rpy2 and Rserve interfaces, multi-process parallel analysis and task scheduling are introduced, and the modular design idea is adopted to disassemble the analysis process into loosely coupled functional modules, providing adaptive data transmission and dynamic loading mechanisms.

Benefits of technology

It realizes efficient and stable cross-language calls under different data scales, avoids interface lag, provides flexible module expansion capabilities, ensures user-friendliness and professionalism, and improves system scalability and computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371266A_ABST
    Figure CN120371266A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data analysis, and particularly relates to a data analysis software architecture design method which comprises the following steps: S1, building a statistical analysis system which comprises a model layer, a view layer and a controller layer, and separating data processing, business logic and a user interface into independent layers, according to the method, delay and memory occupation of cross-language communication are remarkably reduced, calculation resources are utilized to the maximum extent while calculation accuracy is guaranteed, limitation of a fixed communication mode is broken through, heterogeneous language collaborative analysis can run efficiently in various scenes, interface operation and background calculation are carried out at the same time, and the method is high in practicability. The problem of interface lagging caused by long-time operation is thoroughly solved, a layered and modularized statistical analysis process is constructed, a complex data analysis task is disassembled into reusable and replaceable functional modules, a user can freely combine the analysis process or add new modules according to requirements, and it is ensured that the system has the long-term evolution capacity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data analysis, and in particular to a method for designing a data analysis software architecture. Background Art

[0002] The main types of statistical analysis tools widely used in the current clinical research and data analysis fields are as follows: First, commercial statistical software such as SPSS and SAS, which provide a complete GUI interface, are convenient but expensive and have limited scalability; second, open-source statistical programming languages such as R and Python, which are powerful but require professional programming skills, and there are certain thresholds for clinical research personnel to directly use. In addition, with the growth of big data and complex analysis requirements, it is often difficult for a single language or tool to balance performance and usability. For example, the R language has a rich statistical library, but it is easily limited by memory and single-thread performance when dealing with large-scale data; the data analysis libraries in the Python ecosystem are widely used, but the implementation of some professional statistical methods is insufficient. There are some existing solutions that attempt to combine the advantages of multiple technologies, such as calling R language functions in Python (for example, through the Rserve service or the rpy2 interface) to obtain stronger statistical computing capabilities. However, such cross-language calls usually fixedly adopt a certain communication method and cannot be adaptively optimized according to the data scale and computing environment, resulting in unnecessary increased call overhead for small data sets and memory or communication becoming a bottleneck for large data sets. In addition, most traditional statistical analysis software uses single-process sequential execution, and the interface often freezes when performing time-consuming analysis, resulting in a poor user experience. Generally speaking, the existing technologies are difficult to achieve a balance between user-friendliness and high-performance flexible analysis, and there is a lack of a new type of statistical analysis system that is both user-friendly for non-programmer users, can handle large data, and is convenient for expansion. 1. Performance bottleneck of cross-language calls: Existing solutions usually fixedly adopt embedded calls or network service calls when interacting between Python and R, and cannot balance the performance requirements of different-scale data. For small-scale data, enabling network services brings unnecessary communication overhead; for ultra-large-scale data, relying only on embedded calls will cause excessive memory occupation or call failure. The present invention needs to solve how to adaptively select the optimal cross-language communication method according to the data characteristics to ensure the call efficiency and stability in different situations. 2. Interface response and computational parallelism issues: Most traditional statistical software is single-threaded, and the interface is blocked when performing complex analysis, affecting user operations. Another issue that the present invention focuses on is how to ensure the smooth response of the user interface while fully utilizing multi-core to improve the speed of analysis and calculation. This requires a multi-process or multi-thread task scheduling mechanism to avoid resource contention between the GUI and the calculation, and ensure that long-running analysis tasks do not cause the application to become unresponsive. 3. Insufficient module expansion and maintainability: Existing statistical analysis tools have difficulties in function expansion. Many software writes the analysis process rigidly in a single program, lacking good modular design, resulting in the need to significantly modify the code when adding new analysis methods or improving algorithms, with high risks and difficult to maintain. In addition, there is no unified interface for different analysis steps, and the coupling degree between modules is high, restricting the reuse and combination of functions.What the present invention aims to solve is how to design a modular and extensible architecture, break down the data analysis process into loosely coupled functional units by stages, which is convenient for function plugging and unplugging while maintaining overall collaborative work. 4. Balance between user-friendliness and professionalism: For non-computer professional users such as clinical researchers, using statistical tools in a pure programming manner is not user-friendly, while graphical interface tools often have limited functions. The technical problem to be solved is to provide a friendly and intuitive user interface without reducing the professionalism and flexibility of statistical analysis. That is to say, the system should allow users to conveniently complete advanced statistical analysis configurations through the interface, and at the same time ensure that rigorous statistical methods are adopted for underlying calculations to output credible results. Summary of the Invention

[0003] The purpose of the present invention is to propose a data analysis software architecture design method to solve the shortcomings in the background technology.

[0004] To achieve the above purpose, the present invention adopts the following technical solutions:

[0005] A data analysis software architecture design method includes the following steps:

[0006] S1. Build a statistical analysis system. The statistical analysis system includes a model layer, a view layer, and a controller layer, separating data processing, business logic, and user interface into independent layers with clear responsibilities for each layer and collaborative work;

[0007] S2. Cross-language interaction and adaptive data transmission mechanism: The model layer communicates with the R language environment through an interface. The system integrates two R calling methods at the same time: the embedded rpy2 interface call and the server-based Rserve interface call, and adopts an adaptive selection strategy to optimize the interface;

[0008] S3. Multi-process parallel analysis and task scheduling: Introduce a multi-process analysis execution mechanism. The system runs data analysis tasks in independent child processes in the architecture. The main process is mainly responsible for interface interaction and task control. Use the Python multiprocessing library to establish a task scheduling manager. When a user initiates an analysis on the interface, the controller creates a new analysis process and passes the corresponding data and parameters to this process to execute the calculation code of the model layer;

[0009] S4. Modular function design and extension: Adopt the modular design concept, divide the statistical analysis process into multiple loosely coupled modules according to different stages and functions, providing good scalability and customization, including multi-level analysis module division, independent module communication, module registration, and dynamic loading.

[0010] Preferably, the model layer is responsible for data storage, analysis, and processing. It is the core computing unit of the system, covering functions such as data reading and cleaning, field extraction, automatic variable type identification, statistical description calculation, hypothesis testing, and model construction. The model layer is implemented using the object-oriented and factory method patterns. According to different data analysis requirements, corresponding data processing objects are dynamically created. For different types of data sets or statistical analysis methods, the model layer can select appropriate algorithm module instances through the factory pattern. The model layer also encapsulates the call logic for R language statistical functions and performs specific statistical calculations through a custom R function interface. After obtaining the analysis results, the model layer generates a structured result object and returns the results to the controller layer.

[0011] Preferably, the view layer provides a graphical user interface for direct interaction with users. This layer is responsible for presenting the analysis results generated by the model layer in a friendly manner and receiving user input operations. It uses the PySide6 framework to build a cross-platform GUI. The interface includes components such as a data table display area, a statistical chart drawing area, and a log output area. To adapt to diverse data formats and display forms, the view layer introduces the adapter pattern: by writing adapters, different types of underlying data are converted into forms that can be directly displayed by interface controls. The view interface design focuses on intuitiveness and usability, providing drop-down menus and forms for users to select statistical methods and set parameters, and triggering analysis runs through buttons. The result display is presented intuitively in the form of charts and tables, facilitating users to understand the analysis output.

[0012] Preferably, the controller layer is the scheduling center of the system, responsible for responding to user operations and coordinating the model and the view. The controller receives user requests from the view layer, calls the corresponding analysis function modules in the model layer to perform calculations accordingly, and then distributes the results returned by the model layer to the view layer for display. To ensure the order and efficiency of the entire system process, the controller is designed as a singleton pattern, that is, there is only one controller instance in the entire application for unified management. This can avoid resource waste caused by repeated instantiation and ensure consistent coordination of calls to the model layer and the view layer. The controller also assumes the management role of module loading and extension: through the decorator pattern and the plugin mechanism, the controller can dynamically load different analysis function modules or attach new features to existing functions at runtime. When a user needs to add a new statistical analysis method, its implementation can be encapsulated into a module and registered with the controller, and the controller can then identify and call the new module without modifying the core code.

[0013] Preferably, the rpy2 interface is called as follows: rpy2 allows direct calling of R functions and objects in a Python process, which is equivalent to embedding an R interpreter in memory. For medium and small-scale data sets and relatively fast statistical calculations, using rpy2 can avoid starting additional processes or performing network communication, with low overhead and small call latency. Data is directly passed to R through memory sharing, and the results are also synchronously returned to Python.

[0014] Preferably, the Rserve interface is called as follows: Rserve provides a service interface through the TCP / IP protocol in an independent R process. When the data set is large or the multi-threaded / multi-process computing advantages of R need to be utilized, the system will switch to the Rserve mode. Through Rserve, Python sends data to the R service process for processing, and the results are returned through the network.

[0015] Preferably, the adaptive selection strategy is as follows: By monitoring the size of the data set to be analyzed and the resource utilization of the current system, it dynamically switches between the two calling methods: When the data volume is small and the system memory is sufficient, the memory direct connection method of rpy2 is preferentially used to obtain the fastest response; When the data volume is huge or the memory is tight, Rserve is used instead to share the memory pressure. At the same time, a data sharding transmission mechanism is introduced in the Rserve mode: For ultra-large data, instead of transmitting the whole data at once, the data is automatically split into several blocks for batch transmission and processing, reducing the amount of data transmitted at one time. This batch processing mechanism significantly reduces network transmission latency and peak memory occupancy. A dynamic R environment management strategy is also designed: The controller will dynamically decide whether to start multiple Rserve worker processes to parallel process different data blocks or multiple tasks according to the current number of CPU cores and task requirements, ensuring that computing resources are fully utilized without waste. Through the above mechanisms, the collaborative call between Python and R achieves adaptive optimization and can obtain near-optimal performance in data analysis tasks of various scales.

[0016] Preferably, the interface optimization is as follows: To further reduce the cross-language call overhead, parameter passing and result return are optimized at the interface layer of the interaction between Python and R. On the one hand, efficient binary data serialization and shared memory technology are used to reduce the number of data copies; On the other hand, a unified data result format is agreed upon to simplify result parsing. Under this optimization, each time Python calls an R function, it carries the necessary parameters and as little data redundancy as possible, and the R return results are also compressed and packaged into a form that is easy for Python to process. This reduces the additional time consumption of cross-language communication, so that even if the R function is called frequently, the overall performance loss is very small.

[0017] Preferably, the multi-level analysis module is divided as follows: According to the general process of statistical analysis, the system divides the functions into three main levels: data preprocessing, data analysis, and result visualization. Each level is further divided into several independent function modules. At the data preprocessing level, there are "missing value processing", "outlier detection", and "variable type conversion" modules; at the data analysis level, there are "descriptive statistics", "hypothesis testing", "regression analysis", and "principal component analysis" modules; at the visualization level, there are "distribution histogram drawing" and "correlation matrix drawing" modules. Each module performs its own duties and can be developed and tested independently.

[0018] Preferably, the independent communication of the modules is as follows: To reduce the coupling degree between modules, each function module is connected through a uniformly defined data interface. The data transmitted between modules uses a standard format, and the input and output of the modules are agreed upon: the data output by the previous module is directly used as the input of the next module. Due to the use of the standard data format, the modules do not depend on the specific implementation of each other and only need to follow the interface contract. This means that developers can replace the internal implementation of a certain module without affecting other parts. There can be "outlier processing" modules with multiple different algorithms. As long as their input and output formats are the same, they can be registered and used in the system. The modules are connected through the controller scheduling. After the user configures the analysis process on the interface, the controller calls the corresponding modules in sequence and transmits data in turn to complete the entire pipeline; The module registration and dynamic loading are as follows: Design a module management mechanism based on the factory pattern and multi-level dictionary mapping. The system internally maintains a module registry. The first layer distinguishes the level to which the module belongs, and the second layer indexes the specific module implementation object by the module name or function key. When a new module needs to be added, the developer registers the module class or function in this dictionary mapping. When the controller calls the module, it quickly finds the corresponding implementation according to the module name and instantiates it for operation. Due to the use of the factory pattern, the instantiation of the module is highly abstracted. The controller does not need to understand the internal details of the module and only needs to request the module object according to the name. This dynamic loading mechanism ensures the high scalability of the system: Without modifying the core code of the system, users or developers can add new statistical method modules in the form of plugins at any time, and the system immediately supports this method. In addition, the module management dictionary can also store the mapping rules between the module and the applicable data type or scenario. When called, the system can automatically match the optimal analysis module according to the type or size of the input data. For a dataset of categorical variables, the statistical description module will automatically select the chi-square test path, while for continuous variables, it will select the t-test or analysis of variance path.

[0019] Before the present invention was proposed, there were also some alternative solutions that tried to achieve similar functional goals, but they all had their own limitations compared to each other:

[0020] Solution 1: Statistical software based on a single language. For example, only use R language to develop a graphical user interface (with the help of frameworks such as Shiny) to implement statistical analysis functions, or only use Python (combined with its data science library) to develop an analysis application. The advantage of this type of solution is that the architecture is relatively simple and cross-language calls are avoided. However, the disadvantages are also obvious: using only R will encounter difficulties in building complex GUIs and problems with insufficient front-end interaction. Using only Python is not as good as R in the richness and reliability of statistical professional algorithms. The present invention combines Python+R to take into account both friendly interface interaction and powerful statistical calculations. The two complement each other, which cannot be met simultaneously by a single language solution.

[0021] Solution 2: Cross-language calls with fixed communication modes. There are already some tools that allow Python to call R (for example, simply using rpy2 or using the Rserve service), and developers can also use one of the solutions to implement cross-language analysis. However, fixed use of one calling method means inefficiency in some cases: using rpy2 to process very large data will fail or be extremely slow due to memory limitations, and conversely, using Rserve to process small data will increase overhead. The adaptive communication mechanism of the present invention combines the advantages of both, can be flexibly switched as needed, and greatly improves performance and reliability. In comparison, the fixed mode solution is inferior in terms of versatility and stable performance.

[0022] Solution three: Implementation of single-process multi-threading. Another alternative path is to use multi-threading within a single process to try parallel computing and interface response. However, due to Python's GIL restrictions, multi-threading is difficult to truly execute CPU-intensive tasks in parallel, and the R language itself has limited support for multi-threading. In addition, running R calls in multiple threads in the same process may cause conflicts and easily cause the program to crash. In comparison, the present invention adopts a multi-process architecture to completely avoid these problems and achieve true parallelism and stability. Although the multi-process implementation is relatively complex, it ensures that the calculation and interface each occupy independent resources, and the performance and robustness are better than the single-process multi-threaded solution.

[0023] Solution 4: Non-modular integrated design. Some solutions may write all functions in a large code block or a few modules, which saves the work of module division in implementation. However, this integrated design is very cumbersome when expanding functions. Adding or deleting a function may affect other parts, and the testing and maintenance costs are high. The present invention decouples each functional unit through modular layered design, and replacing or adding a module will not interfere with other functions. The system can evolve as the needs evolve, which is significantly better than the lack of flexibility of non-modular solutions.

[0024] Compared with the prior art, the advantages of the present invention are:

[0025] Cross - language Adaptive Interaction Mechanism: An adaptive data transfer protocol for Python and R to work together is proposed. The system can intelligently judge the data scale and dynamically select embedded calls (rpy2) or network service calls (Rserve). Especially for big data, data sharding transmission and multi - R - process parallel computing mechanisms are introduced, significantly reducing the latency and memory occupancy of cross - language communication, maximizing the use of computing resources while ensuring computational accuracy. This mechanism breaks through the limitations of fixed communication modes, enabling efficient operation of heterogeneous language collaborative analysis in various scenarios.

[0026] Multi - process Architecture and Task Scheduling: A multi - process architecture that separates the GUI from the calculation is designed. Through the main control process + analysis sub - process mode, interface operations and background calculations can be carried out simultaneously, completely solving the problem of interface lag caused by long - term operations. The task scheduling module can manage multiple analysis tasks in parallel and provide process isolation protection, improving the system's reliability. This architecture makes full use of the multi - core advantage, greatly enhancing the execution efficiency of large - scale statistical analysis and the user experience.

[0027] Modular Hierarchical Design and Flexible Expansion: A hierarchical modular statistical analysis process is constructed, breaking down complex data analysis tasks into reusable and replaceable functional modules. Through multi - layer dictionary mapping and the factory pattern, dynamic registration and invocation of modules are achieved, supporting on - demand expansion of analysis functions without modifying the existing system. Standard data interfaces are used between modules to ensure low coupling, facilitating independent development and testing. This design enables the system to have a high degree of customization ability, allowing users to freely combine analysis processes or add new modules according to their needs, ensuring the long - term evolution ability of the system.

[0028] Comprehensive Application of Multiple Design Patterns to Optimize the Architecture: Classic design patterns such as singleton, factory method, adapter, and decorator are cleverly combined in the system architecture. The singleton pattern ensures that the core controller is globally unique, avoiding resource conflicts; the factory pattern and decorator pattern achieve the on - demand production and function enhancement of functional modules, facilitating the extension of new algorithms or the packaging and upgrading of existing algorithms; the adapter pattern ensures the decoupling of data and interface display. The comprehensive application of multiple design patterns makes the system not only maintain the simplicity and high cohesion of the code but also endows each part with the ability of elastic expansion, which is an innovative improvement to the software architecture. Description of the Drawings

[0029] Figure 1 It is the schematic diagram of the overall system architecture of a data analysis software architecture design method proposed by the present invention;

[0030] Figure 2 It is the cross - language data transmission and processing flow chart of a data analysis software architecture design method proposed by the present invention;

[0031] Figure 3Modular hierarchical analysis flowchart of a data analysis software architecture design method proposed by the present invention;

[0032] Figure 4 Block diagrams of the view layer, control layer, and model layer of a data analysis software architecture design method proposed by the present invention. Detailed implementation manners

[0033] Next, the technical solutions in this embodiment will be clearly and completely described in conjunction with the accompanying drawings in this embodiment. Obviously, the described embodiments are only a part of the embodiments of this embodiment, rather than all the embodiments.

[0034] Refer to Figures 1 - 4 , a data analysis software architecture design method, including the following steps:

[0035] S1. Build a statistical analysis system. The statistical analysis system includes a model layer, a view layer, and a controller layer, separating data processing, business logic, and user interface into independent layers with clear responsibilities for each layer and collaborative work;

[0036] S2. Cross-language interaction and adaptive data transmission mechanism: The model layer communicates with the R language environment through an interface. The system integrates two R call methods at the same time: the embedded rpy2 interface call and the server-based Rserve interface call, and adopts an adaptive selection strategy to optimize the interface;

[0037] S3. Multi-process parallel analysis and task scheduling: Introduce a multi-process analysis execution mechanism. The system runs data analysis tasks in independent sub-processes in terms of architecture. The main process is mainly responsible for interface interaction and task control. Use the multiprocessing library of Python to establish a task scheduling manager. When the user initiates an analysis on the interface, the controller creates a new analysis process and passes the corresponding data and parameters to this process to execute the calculation code of the model layer;

[0038] S4. Modular function design and extension: Adopt the modular design concept, divide the statistical analysis process into multiple loosely coupled modules according to different stages and functions, providing good scalability and customization, including multi-level analysis module division, independent module communication, module registration, and dynamic loading.

[0039] In this embodiment, the Model layer: The Model layer is responsible for data storage and analysis processing and is the core computing unit of the system. It covers functions such as data reading and cleaning, field extraction, automatic variable type identification, statistical description calculation, hypothesis testing, and model construction. The Model layer is implemented using the object-oriented and factory method patterns, and corresponding data processing objects are dynamically created according to different data analysis requirements. For example, for different types of data sets or statistical analysis methods, the Model layer can select appropriate algorithm modules through the factory pattern. The Model layer also encapsulates the call logic for R language statistical functions and performs specific statistical calculations through a custom R function interface. After obtaining the analysis results, the Model layer generates a structured result object (such as a data structure containing an analysis result table, visualization chart, and analysis log), and returns the results to the Controller layer.

[0040] The View layer: The View layer provides a graphical user interface and directly interacts with users. This layer is responsible for presenting the analysis results generated by the Model layer in a friendly manner and receiving user input operations (such as parameter settings, analysis commands, etc.). The present invention uses the PySide6 framework to build a cross-platform GUI, and the interface includes components such as a data table display area, a statistical chart drawing area, and a log output area. In order to adapt to diverse data formats and display forms, the View layer introduces the adapter pattern: by writing adapters, different types of underlying data (such as Pandas DataFrame, Numpy arrays, or R return data frames) are converted into forms that can be directly displayed by interface controls (such as table models, image objects, etc.). This design ensures that regardless of the data format returned by the Model layer, the View layer can uniformly process and present it through the adapter. The view interface design focuses on intuitiveness and usability. For example, it provides drop-down menus and forms for users to select statistical methods and set parameters, and triggers the analysis run through buttons. The result display is presented intuitively in the form of charts and tables, facilitating users to understand the analysis output.

[0041] Controller Layer (Controller): The controller is the scheduling center of the system, responsible for responding to user operations and coordinating the model and the view. The controller receives user requests from the view layer (for example, the user selects to execute a certain statistical analysis on the interface), then calls the corresponding analysis function module in the model layer to perform calculations, and then distributes the results returned by the model layer to the view layer for display. To ensure the order and efficiency of the entire system process, the controller is designed as a singleton pattern in the present invention, that is, there is only one controller instance in the entire application for unified management. This can avoid resource waste caused by repeated instantiation and ensure consistent coordination of calls to the model layer and the view layer. The controller also assumes the management role of module loading and extension: through the decorator pattern and the plug-in mechanism, the controller can dynamically load different analysis function modules or attach new features to existing functions at runtime. For example, when the user needs to add a new statistical analysis method, its implementation can be encapsulated into a module and registered with the controller, and the controller can identify and call the new module without modifying the core code. This design greatly enhances the flexible expansion ability of the system.

[0042] Through the above MVC hierarchical architecture, the present invention realizes the decoupling of the interface display and the data processing logic: the view layer and the model layer only interact through the controller and are independent of each other. The advantage of this is that when updating the algorithm or replacing the interface framework in the future, it does not affect other layers, achieving a software structure with high cohesion and low coupling, and improving the maintainability of the system.

[0043] Cross-language interaction and adaptive data transmission mechanism: To make full use of the powerful statistical functions of R language, the model layer of this system communicates with the R language environment through an interface. Specifically, the system integrates two R call methods at the same time: the embedded rpy2 interface and the server-based Rserve interface. The innovation of the present invention lies in designing a cross-language adaptive data transmission protocol, which can intelligently select the optimal communication mechanism between rpy2 and Rserve according to the size of the current analyzed data set and the computing resource status, thus significantly improving the efficiency and stability of the interaction between Python and R.

[0044] Rpy2 interface call: rpy2 allows directly calling R functions and objects in a Python process, which is equivalent to embedding an R interpreter in memory. For medium and small-scale data sets and relatively fast statistical calculations, using rpy2 can avoid starting additional processes or network communication, with low overhead and small call latency, and this system will give priority to using it. In this way, data is directly passed to R through memory sharing, and the results are also returned to Python synchronously, which is suitable for scenarios with moderate data volume and frequent interactions.

[0045] Rserve Interface Call: Rserve provides a service interface through the TCP / IP protocol in an independent R process. When the dataset is large (e.g., exceeding the single-process memory limit) or when the multi-threaded / multi-process computing advantages of R need to be utilized, the system will switch to the Rserve mode. Through Rserve, Python sends data to the R service process for processing, and the results are returned via the network. Although this loosely coupled approach has a certain communication overhead, it is more robust for extremely large data, and the calculations on the R side can be carried out independently without blocking the Python main process. Adaptive Selection Strategy: The present invention dynamically switches between the two call methods by monitoring the size of the dataset to be analyzed and the resource utilization of the current system: when the data volume is small and the system memory is sufficient, the in-memory direct connection method of rpy2 is preferentially used to obtain the fastest response; when the data volume is huge or the memory is tight, Rserve is used instead to share the memory pressure. At the same time, the present invention introduces a data sharding transmission mechanism in the Rserve mode: for extremely large data, instead of transmitting the whole data at once, the data is automatically split into several blocks and transmitted and processed in batches, reducing the amount of data transmitted at a single time. This batch processing mechanism significantly reduces the network transmission latency and the peak memory occupancy. A dynamic R environment management strategy is also designed: the controller dynamically decides whether to start multiple Rserve worker processes to process different data blocks or multiple tasks in parallel according to the current number of CPU cores and task requirements, ensuring that the computing resources are fully utilized without waste. Through the above mechanisms, the collaborative call between Python and R achieves adaptive optimization - achieving near-optimal performance in data analysis tasks of various scales.

[0046] Interface Optimization: To further reduce the cross-language call overhead, the present invention optimizes the parameter passing and result return at the interface layer of the interaction between Python and R. On the one hand, efficient binary data serialization and shared memory technologies (such as directly passing the Numpy array pointer for rpy2) are used to reduce the number of data copies; on the other hand, a unified data result format (such as JSON or DataFrame) is agreed upon to simplify the result parsing. Under this optimization, each time Python calls an R function, it carries the necessary parameters with as little data redundancy as possible, and the R return result is also compressed and packaged into a form that is easy for Python to process. This reduces the additional time consumption of cross-language communication, so that even if the R function is called frequently, the overall performance loss is small.

[0047] Multi - process Parallel Analysis and Task Scheduling: The present invention introduces a multi - process analysis and execution mechanism for time - consuming statistical calculations. In terms of architecture, the system runs data analysis tasks in independent child processes, and the main process is mainly responsible for interface interaction and task control. Specifically, the multiprocessing library of Python is used to establish a task scheduling manager. When a user initiates an analysis on the interface, the controller creates a new analysis process and passes the corresponding data and parameters to this process to execute the calculation code in the model layer. Such a design has the following advantages:

[0048] Avoiding Interface Blocking: Since time - consuming calculations are carried out in child processes, the GUI thread of the main process will not be occupied for a long time, and the user interface can still respond to user operations (for example, allowing the user to browse some of the generated results or prepare for the next analysis). This solves the problem of the interface freezing in traditional single - thread programs and significantly improves the user experience.

[0049] Improving Computational Efficiency: The multi - process mechanism enables the system to utilize multi - core CPUs to execute multiple tasks in parallel. If the user submits multiple analysis tasks simultaneously, the scheduler can start multiple processes concurrently to handle them respectively, thereby shortening the total running time. Additionally, through reasonable inter - process communication, after the child process completes the calculation, it notifies the main process of the result and transmits it back, and the main interface immediately updates to display the result.

[0050] Security and Stability: Using independent processes also enhances the robustness of the system. Even if an analysis process encounters an error or crashes due to abnormal data or extreme situations, it will only affect that child process and will not bring down the entire main application. The main process can capture the abnormal state of the child process and give the user a prompt. This isolation ensures the stability of the system operation and the security of data, preventing problems such as all unsaved data being lost due to a single calculation failure.

[0051] Task Management: The controller contains a simple task management module that can monitor the execution status (in progress, completed, abnormal, etc.) of each analysis process and collect the results. When the task is completed, the results are transmitted back to the main process through a pre - established pipeline or queue. After receiving the results, the controller calls the view to update the result display and stores the results. On the interface, the user can see the list and progress of the currently running tasks and can view the log output in real - time for time - consuming tasks. This enables the user to conveniently manage and track the execution of multiple analysis tasks.

[0052] Modular Function Design and Extension: The present invention adopts the modular design concept, divides the statistical analysis process into multiple loosely - coupled modules according to different stages and functions, and provides good scalability and customizability.

[0053] Multi-level analysis module division: According to the general process of statistical analysis, the system divides the functions into three main levels: data preprocessing, data analysis, and result visualization. Each level is further divided into several independent functional modules. For example, at the data preprocessing level, there are modules such as "missing value handling", "outlier detection", and "variable type conversion"; at the data analysis level, there are modules such as "descriptive statistics", "hypothesis testing", "regression analysis", and "principal component analysis"; at the visualization level, there are modules such as "distribution histogram drawing" and "correlation matrix drawing". Each module performs its own duties and can be developed and tested independently.

[0054] Independent communication of modules: To reduce the coupling degree between modules, each functional module is connected through a uniformly defined data interface. Specifically, the data transmitted between modules uses a standard format (such as a common format file like Pandas DataFrame or CSV), and the modules make agreements on input and output: the data output by the previous module is directly used as the input of the next module. Due to the use of standard data formats, modules do not depend on specific implementations of each other, as long as they follow the interface contract. This means that developers can replace the internal implementation of a certain module without affecting other parts. For example, there can be "outlier handling" modules with multiple different algorithms, and as long as their input and output formats are consistent, they can all be registered and used in the system. Modules are connected and coordinated through a controller. When the user configures the analysis process on the interface (for example, selects to first fill in missing values, then perform regression analysis, and then generate charts), the controller calls the corresponding modules in sequence, transmits data in turn, and completes the entire pipeline.

[0055] Module Registration and Dynamic Loading: The present invention designs a module management mechanism based on the factory pattern and multi-layer dictionary mapping. The system internally maintains a module registration table (which can be understood as a multi-layer nested dictionary structure). The first layer differentiates the levels to which the modules belong (preprocessing / analysis / visualization), and the second layer indexes the specific module implementation objects by module name or function key. When a new module needs to be added, the developer registers the module class or function into this dictionary mapping. When the controller calls a module, it quickly locates the corresponding implementation through the module name and instantiates it for running. Due to the adoption of the factory pattern, the instantiation of the module is highly abstracted. The controller does not need to understand the internal details of the module and only needs to request the module object according to the name. This dynamic loading mechanism ensures the high scalability of the system: without modifying the core code of the system, users or developers can add new statistical method modules in the form of plugins at any time, and the system immediately supports this method. In addition, the module management dictionary can also store the mapping rules between the modules and the applicable data types or scenarios. When called, the system can automatically match the optimal analysis module according to the type or size of the input data. For example, for a dataset of categorical variables, the statistical description module will automatically select the chi-square test path, while for continuous variables, it will select the t-test or analysis of variance path. This intelligent matching improves the accuracy and efficiency of the analysis.

[0056] As described above, the above is only the preferred specific implementation manner of this embodiment, but the protection scope of this embodiment is not limited thereto. Any person skilled in the art within the technical scope disclosed by this embodiment, according to the technical solution and inventive concept of this embodiment, makes equivalent replacements or changes, and all should be covered within the protection scope of this embodiment.

Claims

1. A method for designing a data analysis software architecture, characterized in that It includes the following steps: S1. Build a statistical analysis system. The statistical analysis system includes a model layer, a view layer, and a controller layer, separating data processing, business logic, and user interface into independent layers with clear responsibilities for each layer and collaborative work; S2. Cross-language interaction and adaptive data transmission mechanism: The model layer communicates with the R language environment through an interface, integrating two R call methods: the embedded rpy2 interface call and the server-based Rserve interface call, and adopting an adaptive selection strategy to optimize the interface; S3. Multi-process parallel analysis and task scheduling: Introduce a multi-process analysis execution mechanism, run data analysis tasks in independent sub-processes, and the main process is mainly responsible for interface interaction and task control. Use the Python multiprocessing library to establish a task scheduling manager. When a user initiates an analysis on the interface, the controller creates a new analysis process, and passes the corresponding data and parameters to this process to execute the calculation code of the model layer; S4. Modular function design and extension: Adopt the modular design concept, divide the statistical analysis process into multiple loosely coupled modules according to different stages and functions, providing good scalability and customizability, including multi-level analysis module division, independent module communication, module registration, and dynamic loading.

2. The method for designing a data analysis software architecture according to claim 1, wherein The process of building the statistical analysis system is as follows: Data collection: Ensure that the data source is reliable and relevant to the analysis target. Data cleaning: Process missing values, outliers, and duplicate data to ensure data quality. Data exploration: Initially understand the data characteristics through statistical descriptions and visualization means, discover patterns and anomalies in the data, select appropriate statistical models according to the analysis target and data characteristics, including linear regression and decision trees, and use the selected model for data fitting, including parameter estimation and model training. Evaluate the performance of the model through cross-validation methods to ensure the stability and predictive ability of the model. Interpret the analysis results and use chart tools for result visualization to ensure that the results are easy to understand and communicate. The model layer is responsible for data storage and analysis processing, and is the core computing unit of the system, covering functions such as data reading and cleaning, field extraction, automatic variable type identification, statistical description calculation, hypothesis testing, and model construction. The model layer is implemented using the object-oriented and factory method patterns, dynamically creating corresponding data processing objects according to different data analysis requirements. For different types of data sets or statistical analysis methods, the model layer selects appropriate algorithm module instances through the factory pattern. The model layer has a call logic for R language statistical functions and executes specific statistical calculations through a custom R function interface. After obtaining the analysis results, the model layer generates a structured result object and returns the result to the controller layer.

3. A method for designing a data analysis software architecture according to claim 2, characterized in that, The view layer provides a graphical user interface, interacts directly with users, displays the analysis results generated by the model layer in a friendly manner, and receives user input operations. It uses the PySide6 framework to build a cross-platform GUI. The interface includes components such as a data table display area, a statistical chart drawing area, and a log output area. The adapter pattern is introduced: by writing adapters, different types of underlying data are converted into a form directly displayed by interface controls. The view interface design focuses on intuitiveness and usability, providing a drop-down menu and forms for users to select statistical methods and set parameters, and triggering the analysis run through buttons. The result display is presented intuitively in the form of charts and tables, facilitating users to understand the analysis output.

4. A method for designing a data analysis software architecture according to claim 3, characterized in that, The controller layer is responsible for responding to user operations and coordinating the model and the view. The controller receives user requests from the view layer, calls the corresponding analysis function modules in the model layer to perform calculations accordingly, and then distributes the results returned by the model layer to the view layer for display. To ensure the order and efficiency of the entire system process, the controller is designed as a singleton pattern, that is, there is only one controller in the whole application for unified management, avoiding resource waste caused by repeated instantiation and ensuring consistent coordination of calls to the model layer and the view layer. The controller undertakes the management role of module loading and extension: through the decorator pattern and the plugin mechanism, the controller dynamically loads different analysis function modules at runtime or attaches new features to existing functions. When a user needs to add a new statistical analysis method, its implementation is encapsulated into a module and registered with the controller. The controller identifies and calls this new module without modifying the core code.

5. A method for designing a data analysis software architecture according to claim 4, characterized in that, The specific call of the rpy2 interface is as follows: rpy2 allows directly calling R functions and objects in a Python process, which is equivalent to embedding an R interpreter in memory. For medium and small-scale data sets and relatively fast statistical calculations, using rpy2 avoids starting additional processes or performing network communication, with low overhead and small call latency. Data is directly passed to R through memory sharing, and the results are also synchronously returned to Python.

6. A method for designing a data analysis software architecture according to claim 5, characterized in that, The specific call of the Rserve interface is as follows: Rserve provides a service interface through the TCP / IP protocol in an independent R process. When the data set is large or the advantages of R's multi-threaded / multi-process calculations need to be utilized, the system will switch to the Rserve mode. Through Rserve, Python sends data to the R service process for processing, and the results are returned through the network.

7. A method for designing a data analysis software architecture according to claim 6, characterized in that The specific adaptive selection strategy is as follows: By monitoring the size of the dataset to be analyzed and the resource utilization of the current system, it dynamically switches between two calling methods: When the data volume is small and the system memory is sufficient, the in-memory direct connection method of rpy2 is used to obtain the fastest response; when the data volume is huge or the memory is tight, Rserve is used instead to share the memory pressure. At the same time, a data sharding transmission mechanism is introduced in the Rserve mode: for ultra-large data, instead of transmitting the whole data at once, the data is automatically split into several blocks for batch transmission and processing, reducing the amount of data transmitted at one time, and reducing network transmission latency and peak memory occupancy. A dynamic R environment management strategy is set: the controller will dynamically determine whether to start multiple Rserve worker processes to process different data blocks or multiple tasks in parallel according to the current number of CPU cores and task requirements, ensuring that computing resources are fully utilized without waste. Through the above mechanisms, the collaborative call between Python and R has achieved adaptive optimization, and near-optimal performance can be obtained under various scales of data analysis tasks.

8. A method for designing a data analysis software architecture according to claim 7, characterized in that, The specific interface optimization is as follows: The parameter passing and result return are optimized at the interface layer where Python and R interact. On the one hand, efficient binary data serialization and shared memory technology are used to reduce the number of data copies; On the other hand, a unified data result format is agreed upon to simplify result parsing. Each time Python calls an R function, it carries the necessary parameters and as little data redundancy as possible, and the R return result is also compressed and packaged into a form that is easy for Python to process.

9. A method for designing a data analysis software architecture according to claim 8, characterized in that The specific multi-level analysis module division is as follows: The system divides the functions into three main levels according to the general process of statistical analysis: data preprocessing, data analysis, and result visualization. Each level is further divided into several independent functional modules. At the data preprocessing level, there are "missing value handling", "outlier detection", and "variable type conversion" modules; at the data analysis level, there are "descriptive statistics", "hypothesis testing", "regression analysis", and "principal component analysis" modules; at the visualization level, there are "distribution histogram drawing" and "correlation matrix drawing" modules. Each module performs its own duties and is developed and tested independently.

10. A method for designing a data analysis software architecture according to claim 9, characterized in that, The independent communication of the modules is as follows: Each functional module is connected through a uniformly defined data interface. The data transmitted between modules adopts a standard format. The modules make agreements on input and output: The data output by the previous module is directly used as the input of the next module, adopting the standard data format. The modules do not depend on the specific implementation of each other. As long as the interface contract is followed, the internal implementation of a certain module can be replaced without affecting other parts. There are "outlier handling" modules with multiple different algorithms. As long as their input and output formats are consistent, they are all registered and used in the system. The modules are connected through the controller for scheduling. After the user configures the analysis process on the interface, the controller calls the corresponding modules in sequence, transmits data in turn, and completes the entire pipeline; The module registration and dynamic loading are as follows: Design a module management mechanism based on the factory pattern and multi-layer dictionary mapping. The system internally maintains a module registry. The first layer differentiates the levels to which the modules belong, and the second layer indexes the specific module implementation objects by module name or function key. When a new module needs to be added, the module class or function is registered in this dictionary mapping. When the controller calls a module, it quickly finds the corresponding implementation through the module name and instantiates it for operation. Due to the adoption of the factory pattern, the instantiation of the module is highly abstracted. The controller does not need to understand the internal details of the module, and only needs to request the module object according to the name. This dynamic loading mechanism ensures the high scalability of the system: Without modifying the core code of the system, users or developers can add new statistical method modules in the form of plugins at any time, and the system immediately supports this method. The module management dictionary stores the mapping rules between the modules and the applicable data types or scenarios. When called, the system automatically matches the optimal analysis module according to the type or size of the input data. For a dataset of categorical variables, the statistical description module will automatically select the chi-square test path, while for continuous variables, it will select the t-test or analysis of variance path.

Citation Information

Patent Citations

  • Method and system for executing and equipping application according to use condition

    CN101174219A

  • Software design method and system based on business layering

    CN117632093A

  • Efficient batch processing in a multi-tier application

    EP2495657A1

  • Efficient and intuitive databinding for mobile applications

    WO2016049626A1

Cited By

  • Extensible modular architecture design method and device of industrial large model, computer equipment, storage medium and computer program product

    CN121479305A