A web-based big data analysis method and related equipment
By integrating microservice front-end components and integrating shared data on low-code platforms, dividing and allocating data analysis tasks, centralized scheduling is solved, and efficient data analysis and integration is achieved.
Patent Information
- Application Number
- CN202510061862.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-01-15
AI Technical Summary
Centralized scheduling is difficult to efficiently handle complex data analysis tasks, especially in multi-module systems.
By integrating the microservice front-end components of the subsystem into a low-code platform, integrating shared data according to the subsystem loader, preprocessing and dividing it into multiple subtasks, the distributed intelligent task scheduling mechanism is used for allocation and execution.
It improves the processing efficiency of data analysis tasks and realizes comprehensive analysis and integration of multi-module system data.
Smart Images

Figure CN119474765B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of Internet, and in particular to a web-based big data analysis method and related equipment. Background Art
[0002] In the current information and digital age, enterprises and organizations are increasingly relying on big data analysis to support decision-making and business optimization. In order to achieve efficient big data analysis, systems based on microservice architecture are widely adopted. These systems are usually composed of multiple independent modules, each of which is responsible for processing different data sets or functions. However, since the data storage and management of each module are independent, it becomes difficult to share and integrate data, which further makes it impossible to conduct a comprehensive analysis of the data of each module during the big data analysis process. Some big data analysis systems usually adopt a centralized task scheduling mechanism, and the task division and allocation process lacks flexibility and intelligence. In the case of multiple modules, centralized scheduling is difficult to efficiently handle complex data analysis tasks. Summary of the invention
[0003] The present application provides a web-based big data analysis method and related equipment, which are used to solve the problem in related technologies that centralized scheduling is difficult to efficiently handle complex data analysis tasks.
[0004] The first aspect of the present application provides a web-based big data analysis method, the web-based big data analysis method comprising:
[0005] Integrate the microservice front-end components of the subsystem into the low-code platform;
[0006] Integrate the shared data of the subsystem in the low-code platform according to the subsystem loader;
[0007] Preprocessing the shared data according to preset cleaning rules;
[0008] When receiving a data analysis request from a user, dividing the data analysis task corresponding to the data analysis request into N subtasks according to the preprocessed target shared data; wherein N is an integer greater than or equal to 2;
[0009] The N subtasks are respectively assigned to the corresponding subsystems for data analysis.
[0010] Optionally, in a first implementation of the first aspect of the present application, the step of integrating the microservice front-end component of the subsystem into the low-code platform includes:
[0011] Generate a sandbox in the low-code platform; wherein the sandbox is an independent virtual environment;
[0012] Generate an environment identifier for the microservice front-end component of the subsystem according to the global manager;
[0013] The microservice front-end component and the corresponding environment identifier are assigned to the sandbox.
[0014] Optionally, in a second implementation of the first aspect of the present application, after the step of assigning the microservice front-end component and the corresponding environment identifier to the sandbox, the step further includes:
[0015] Determining a target subsystem according to the data analysis request;
[0016] Obtaining a target environment identifier corresponding to the target subsystem;
[0017] Creating a window container according to the target environment identifier;
[0018] Load the microservice front-end component of the target subsystem into the window container.
[0019] Optionally, in a third implementation of the first aspect of the present application, after the step of integrating the microservice front-end component of the subsystem into the low-code platform, the step further includes:
[0020] Sending the initial data required for data sharing to the microservice front-end component according to a preset configuration file;
[0021] When a user's data query instruction is detected, the data query instruction is sent to the microservice front-end component according to the event bus;
[0022] Receive the shared data sent by the microservice front-end component based on the initial data and the data query instruction.
[0023] Optionally, in a fourth implementation of the first aspect of the present application, the method further includes:
[0024] Determine the target subsystem window that the user needs to query according to the data query instruction;
[0025] Predicting the subsystem window to be accessed based on the target subsystem and the user's historical query records;
[0026] The shared data of the corresponding access subsystem is loaded according to the prediction result of the subsystem window to be accessed.
[0027] Optionally, in a fifth implementation of the first aspect of the present application, when receiving a data analysis request from a user, the step of dividing the data analysis task corresponding to the data analysis request into N subtasks according to the preprocessed target shared data includes:
[0028] Determine a computational logic tree of the data analysis task according to the data analysis request;
[0029] Determining a data format of the target shared data;
[0030] The data analysis task is divided into a plurality of subtasks according to the computational logic tree and the data format.
[0031] Optionally, in a sixth implementation of the first aspect of the present application, after the step of respectively allocating the N subtasks to the corresponding subsystems for data analysis, the step further includes:
[0032] Monitoring the task execution status of the subsystem in real time;
[0033] When the main system detects that a subtask has failed or the result is incomplete, a rescheduling mechanism is triggered;
[0034] The unfinished subtasks are reallocated according to the rescheduling mechanism.
[0035] A second aspect of the present application provides a web-based big data analysis device, the web-based big data analysis device comprising:
[0036] Integration module, used to integrate the microservice front-end components of the subsystem into the low-code platform;
[0037] An integration module, configured to integrate the shared data of the subsystems in the low-code platform according to the subsystem loader;
[0038] A preprocessing module, used for preprocessing the shared data according to preset cleaning rules;
[0039] A division module, for, when receiving a data analysis request from a user, dividing the data analysis task corresponding to the data analysis request into N subtasks according to the preprocessed target shared data; wherein N is an integer greater than or equal to 2;
[0040] The allocation module is used to allocate the N subtasks to the corresponding subsystems for data analysis.
[0041] A third aspect of an embodiment of the present application provides an electronic device, comprising a memory and a processor, wherein the processor is used to execute a computer program stored in the memory, and when the processor executes the computer program, it implements each step of the web-based big data analysis method provided in the first aspect of the embodiment of the present application.
[0042] The fourth aspect of the embodiments of the present application provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, each step of the web-based big data analysis method provided in the first aspect of the embodiments of the present application is implemented.
[0043] In summary, according to a web-based big data analysis method and related equipment provided by the present application, the microservice front-end components of the subsystem are integrated into the low-code platform; the shared data of the subsystem is integrated in the low-code platform according to the subsystem loader; the shared data is preprocessed according to preset cleaning rules; when a user's data analysis request is received, the data analysis task corresponding to the data analysis request is divided into N subtasks according to the preprocessed target shared data; the N subtasks are respectively assigned to the corresponding subsystems for data analysis. Through the implementation of the present application, the shared data of the subsystem is integrated and preprocessed on the low-code platform, and when executing the data analysis task, the data analysis task is divided into multiple subtasks according to the preprocessed target shared data, and the processing efficiency of the data analysis task is effectively improved through the distributed intelligent task scheduling mechanism. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 A schematic diagram of a process flow of a web-based big data analysis method provided in an embodiment of the present application;
[0045] Figure 2 A schematic diagram of a program module of a web-based big data analysis device provided in an embodiment of the present application;
[0046] Figure 3 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0047] In order to make the purpose, features, and advantages of the invention of this application more obvious and easy to understand, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.
[0048] In order to solve the problem that centralized scheduling in related technologies is difficult to efficiently process complex data analysis tasks, the present application embodiment provides a web-based big data analysis method, such as Figure 1 The following is a flow chart of a web-based big data analysis method provided in this embodiment. The web-based big data analysis method includes the following steps:
[0049] Step 110: Integrate the microservice front-end components of the subsystem into the low-code platform.
[0050] Specifically, in this embodiment, the low-code platform is used as the core architecture of the main system, and the microservice architecture is introduced as the organizational form of the subsystem, so as to achieve efficient management and unified integration of multi-system resources. In the low-code platform, the main system integrates the microservice front-end components into the system main interface by dynamically loading custom components.
[0051] In an optional implementation of this embodiment, the step of integrating the microservice front-end component of the subsystem into the low-code platform includes: generating a sandbox in the low-code platform; wherein the sandbox is an independent virtual environment; generating an environment identifier for the microservice front-end component of the subsystem according to the global manager; and assigning the microservice front-end component and the corresponding environment identifier to the sandbox.
[0052] Specifically, in this embodiment, each microservice component will be assigned to an independent virtual environment (sandbox) when loaded. This sandbox uses the browser's memory isolation technology and independent style scope (referring to a set of CSS styles defined for the component, which will only act on the DOM element of the current component, and will not leak to the outside or affect the style of other components) to ensure that the static resources (including HTML structure, CSS style, JavaScript logic) between the subsystems do not interfere with each other. When the components of each subsystem are loaded, the main system will generate a unique environment identifier through a global manager (Resource Orchestrator) and inject it into the sandbox for resource isolation and call tracking. On this basis, the main system and the subsystem communicate through the event bus customized by the low-code platform. The main system is responsible for listening to the operation events triggered by the user, and loading the corresponding subsystem components according to the target of the operation, completing the seamless switching between the main interface and the subsystem interface. This mechanism not only realizes the flexible loading of the microservice front end, but also greatly enhances the maintainability and scalability of the system.
[0053] In an optional implementation of this embodiment, after the step of assigning the microservice front-end component and the corresponding environment identifier to the sandbox, it also includes: determining the target subsystem according to the data analysis request; obtaining the target environment identifier corresponding to the target subsystem; creating a window container according to the target environment identifier; and loading the microservice front-end component of the target subsystem into the window container.
[0054] Specifically, in this embodiment, in order to further optimize the user operation experience, the interface resources of multiple subsystems are integrated into the main system through componentization and multi-windowing to achieve the effect of "multi-task windows" running in parallel. In the traditional micro-frontend architecture, users usually jump to different subsystem interfaces through URLs, but this method cannot display multiple systems at the same time in the same interface, resulting in poor user experience. The core of the multi-window mechanism is the window management module of the low-code platform. When a user triggers a subsystem in the main interface, the main system dynamically creates a window container and loads the microservice front-end component corresponding to the target subsystem into this container. Each window container has independent life cycle management and style scope, and users can freely drag, scale, minimize or close the window. In order to ensure the resource independence between multiple windows, the environment identifier generated by the sandbox mechanism is introduced in the creation process of each window. The window container will be automatically bound to its corresponding sandbox environment to ensure that the data and styles between different windows do not interfere with each other. On this basis, the window management module of the main system also supports the message passing function between windows. For example, a user submits a form data in a window, and another window can receive and process the data in real time to complete the collaborative processing of tasks.
[0055] In an optional implementation of this embodiment, after the step of integrating the microservice front-end component of the subsystem into the low-code platform, it also includes: sending the initial data required for data sharing to the microservice front-end component according to a preset configuration file; when a user's data query instruction is detected, sending the data query instruction to the microservice front-end component according to an event bus; receiving the shared data sent by the microservice front-end component based on the initial data and the data query instruction.
[0056] Specifically, in this embodiment, after the system integration is completed, in order to realize the data interaction between the main system and the subsystem, when the main system loads a subsystem component, the initial data required for its operation will be injected into the subsystem through the configuration file of the custom component. For example, the main system user triggers a data query operation by clicking a button. At this time, the main system will pass the operation instruction to the microservice front-end component corresponding to the target subsystem through the event bus. In order to ensure the security and real-time performance of data transmission, an authentication mechanism based on encrypted tokens can be introduced. Each time the main system sends data to the subsystem, a temporary token will be generated. After receiving the data, the subsystem needs to verify the validity of the token through decryption before continuing to process the data.
[0057] Optionally, when a user initiates a data query operation on the low-code platform, a data query instruction is generated. The instruction contains the specific content and target of the query, such as querying sales data or user behavior records within a specific time period. The query parsing module of the low-code platform will parse the instruction and extract the query content and target subsystem information. Based on the parsing results, the main system determines the target subsystem window that the user needs to query. After the target subsystem window is determined, the main system activates the corresponding subsystem window in the low-code platform through the dynamic component loading mechanism. At the same time, the main system continuously collects the user's query records on the low-code platform, including the target subsystem, query content, timestamp and other information of each query. The main system builds a prediction model through a machine learning algorithm. The input of the model is the user's historical query record, and the output is the prediction result of the subsystem window to be accessed. Commonly used prediction models include but are not limited to predictions based on sequence models. For example, a long short-term memory network (LSTM) is used to predict the subsystem window that the user may query next time. When a user initiates a new data query command, the main system uses the trained prediction model, combined with the target subsystem of the current query and the user's historical query records, to predict other subsystem windows that the user may access. The main system loads the shared data of the corresponding subsystem in advance based on the prediction results of the subsystem window to be accessed.
[0058] Optionally, in modern big data analysis and microservice architecture, in order to achieve efficient data management and analysis, by sharing an independent lightweight data warehouse and using environment identifiers to bind data areas, the problem of data integration and isolation is solved. Specifically, the lightweight data warehouse is a unified data storage center for sharing data between the main system and all subsystems. Compared with traditional databases, lightweight data warehouses are more concise and efficient, suitable for data storage and management of front-end environments. All data is centrally stored in a data warehouse, avoiding the management complexity caused by data dispersion. Although the main system and all subsystems share the same lightweight data warehouse, each subsystem can only access the data area bound to its environment identifier. For example: the environment identifier of subsystem A is EnvA, and it can only access data related to EnvA in the data warehouse, that is, the environment identifier is used to mark the operating environment and data permission range of the subsystem.
[0059] Step 120: Integrate the shared data of the subsystems in the low-code platform according to the subsystem loader.
[0060] Specifically, in this embodiment, the subsystem loader is used to load subsystem-related resource files into the main system. When the shared data of each subsystem (such as SQL database, NoSQL storage, real-time streaming data, etc.) is connected to the main system, it is connected to the resource manager of the main system through the corresponding subsystem loader. The subsystem loader binds the subsystem data according to the environment identifier of each subsystem, so that the data collection and management are completely isolated logically, but can be uniformly queried and operated in the main system. In actual operation, when the user selects a big data analysis task (such as cross-system sales forecast analysis), the main system will dynamically load the adapter module of the relevant subsystem and use the subsystem loader to complete the integration of multiple data sources.
[0061] Step 130: pre-process the shared data according to preset cleaning rules.
[0062] Specifically, in this embodiment, after completing the data source integration, the main system will use the shared lightweight data warehouse to pre-process and clean the shared data, and the main system will uniformly call the predefined cleaning rules (such as deduplication, outlier processing, data type conversion, etc.) for processing. The main system supports dynamic loading and execution of cleaning rules for different business scenarios. For example, when analyzing user behavior data, the main system can automatically clean the repeated user IDs in each subsystem, unify the data format, and remove noise data to ensure the accuracy of subsequent analysis. For example, the time format of shared data is normalized, duplicate data records are cleaned, data uniqueness is ensured, missing values in the data are filled, removed or marked, outliers in the data are detected and processed (such as incorrect input, extreme values, etc.), and field names and field values in different subsystems are mapped to a unified standard to ensure that the types of all data fields are consistent. It can be understood that the execution of cleaning rules depends on the built-in rule engine, which can dynamically load cleaning rules and process data one by one.
[0063] It should be noted that, assuming the given time series data is , the cleaning rule in this embodiment can be expressed by the following formula:
[0064] ,
[0065] in, is the mean of the time series data X, and is defined as , is the standard deviation of the time series data, defined as , is the outlier determination threshold, The value can be adjusted according to the actual application scenario. is the correction coefficient of the outlier, which is used to pull the outlier back to a more reasonable range. , For data points The neighborhood set of Neighborhood set The number of elements of In order to find the average value of the data points in the neighborhood, it can be understood that the judgment formulas in the first and second rows of the formula represent the processing of outliers. The judgment formulas in the first and second rows limit the outliers that exceed the normal range to a reasonable range, thereby reducing the impact of outliers on the analysis results; the judgment formula in the third row of the formula is for processing missing values, using the local trend information of adjacent data to fill in the missing values, thereby avoiding damage to the overall trend of the time series; the judgment formula in the fourth row of the formula is for processing noise data. If Does not meet the outlier condition (i.e. ) and is not a missing value, the original value is retained directly, ensuring the integrity of normal data points and not introducing unnecessary corrections. This formula is a comprehensive cleaning rule that provides an accurate and flexible correction method for outliers, missing values, and noise problems in time series data. By properly setting parameters (such as , ), can adapt to different data characteristics and analysis needs, ensuring that the cleaned data not only truly reflects the original trend but is not disturbed by abnormal data.
[0066] Step 140: when receiving a data analysis request from a user, dividing the data analysis task corresponding to the data analysis request into N subtasks according to the preprocessed target shared data;
[0067] Step 150: Allocate subtasks to subsystems for data analysis.
[0068] Specifically, in this embodiment, when a user initiates a big data analysis request, the main system will parse the task dependencies (the cleaning step will normalize the data structure and content, such as unifying the data format, standardizing field names, merging redundant data, etc. These operations will provide clear dependencies for subsequent distributed computing tasks), and decompose the analysis task into multiple subtasks through the task scheduling module. These subtasks will be distributed to the corresponding subsystems for execution, and the sandbox environment of the subsystems ensures the independence and security of the tasks. The results of the distributed computing will be gradually summarized into the main system, and the main system will merge, count and transform the data through the built-in real-time stream processing module to generate the final analysis results. For example, in the sales forecast analysis scenario, the main system can schedule the sales data of each regional subsystem and generate nationwide sales forecast data through a distributed computing model.
[0069] Optionally, after the data analysis task is completed, the big data analysis results are visualized in multiple views. Users can open multiple windows in the main system interface, each window displays different analysis dimensions or result data. These windows may include but are not limited to the following view types: line charts, bar charts, heat maps, geographic visualization maps, and dynamic dashboards. The window management module of the main system supports real-time interaction and data linkage between windows. For example, when a user selects a specific time range in a window, the display content of other windows will be automatically updated to reflect the data in the selected range. Each window corresponds to an independent data analysis result, and the window sandbox mechanism ensures the independence of view rendering and performance stability.
[0070] In an optional implementation of the present embodiment, when a data analysis request from a user is received, the step of dividing the data analysis task corresponding to the data analysis request into N subtasks according to the preprocessed target shared data includes: determining a computational logic tree of the data analysis task according to the data analysis request; determining a data format of the target shared data; and dividing the data analysis task into N subtasks according to the computational logic tree and the data format.
[0071] Specifically, in this embodiment, the task scheduling module of the main system will parse the user's analysis requirements and convert them into a computational logic tree. The computational logic tree describes the decomposition method and dependency relationship of the tasks. For example, taking the need to analyze sales data as an example: the first-level task: divide the data by region (Region) to perform sales statistics, the second-level task: further subdivide the statistical results by time period, and the third-level task: aggregate the sales of all regions to generate global trends. The basis for task division comes directly from the data preprocessing rules and cleaning results. For example, if the data is partitioned by regional labels, the scheduling module will divide the data into multiple subtasks based on these partitioning rules. The task scheduling module distributes the divided subtasks to the corresponding subsystems through the resource manager in the main system. The task scheduling module will allocate tasks according to the computing resources (CPU, memory, etc.) of each subsystem. For example, a subsystem with stronger computing resources can process larger-scale data or more complex computing logic. After the subtask is distributed to the subsystem, the subsystem executes the task independently through its own sandbox environment and distributed computing framework. After the corresponding data analysis is completed, each subsystem will transmit the data analysis results of the subtask back to the main system. After receiving the data analysis results of each subsystem, the main system will use the real-time stream processing module to summarize the results.
[0072] In an optional implementation of the present embodiment, after the step of respectively allocating the N subtasks to the corresponding subsystems for data analysis, it also includes: real-time monitoring of the task execution status of the subsystem; when the main system detects that the subtask fails or the result is incomplete, triggering a rescheduling mechanism; and reallocating unfinished subtasks according to the rescheduling mechanism.
[0073] Specifically, in this embodiment, each subtask will be assigned a unique task identifier when scheduling. When executing a task, the subsystem will periodically report the task status (such as "in progress", "completed", "failed", etc.) to the main system. The main system monitors the task status and the preset timeout time to determine whether the subtask has failed or has not been completed due to timeout. When the main system detects that the subtask has failed or the result is incomplete, it will trigger the rescheduling mechanism to reallocate and execute the unfinished subtasks. For example, the task retry mechanism: the main system first attempts to reallocate the failed subtask to the atomic system, and the number of retries can be set according to the system configuration. The task migration mechanism: if the number of retries exceeds the preset value, the main system will migrate the task to other available subsystems for execution. When migrating tasks, selection will be made based on the computing resources and current load conditions of the subsystem. Partial result processing: for subtasks that have generated partial results, the system will save their partial results and continue to calculate when reallocating tasks to avoid repeated calculations.
[0074] Optionally, a rescheduling algorithm is used in combination with task retry and migration mechanisms to ensure that tasks can be reassigned and executed in a timely manner after failure. The rescheduling algorithm formula is as follows:
[0075] ,
[0076] in, Represents the rescheduling function, which is used to reassign task T. T represents the task that needs to be rescheduled, S represents the subsystem to which the task was originally assigned, and R represents the current number of retries, with an initial value of 0. Indicates the current retry count of the returned task T. Indicates the maximum number of retries configured by the system. To retry the function, reassign task T to subsystem S. ′ represents a new subsystem, which represents the subsystem to which the task is migrated, and is selected by the task migration mechanism. Represents the migration function, which assigns task T to the new subsystem , it can be understood that if the current retry count R is less than the maximum retry count , then call function, reassigns task T to atomic system S, The function will increase the count of the number of retries R and reallocate the task to the subsystem S for execution. If the current number of retries R has reached or exceeded the maximum number of retries, , then call Function, migrate task T to the new subsystem 'Execute. New subsystem ′ is selected by the task migration mechanism, usually based on the computing resources and load of the subsystem. If task T is successfully completed after retry or migration, save the result and update the task status to "completed". If task T is still not completed after all retries and migrations, record the reason for task failure and trigger an alarm or manual intervention.
[0077] It should be noted that the selection algorithm of the new subsystem is as follows:
[0078] ,
[0079] Where M is the set of all available subsystems, The current load of subsystem S, which indicates the number of tasks currently being processed by the subsystem, It is represented by the computing resources of subsystem S, which indicates the computing power of the subsystem (such as the number of CPU cores, memory size, etc.). The meaning of this formula is to calculate the ratio of the load to the computing resources of each subsystem and select the subsystem with the smallest ratio as the new task migration target. The smaller the ratio, the lighter the load of the subsystem is relative to its computing resources, and it is suitable as the target of task migration.
[0080] Optionally, in big data analysis, users usually need to aggregate and analyze multi-dimensional data and predict it to obtain more comprehensive business insights. For example, in an e-commerce platform, users may need to analyze sales data within a certain period of time, aggregate statistics by region, product type and other dimensions, and predict future sales trends. Therefore, this embodiment includes a multi-dimensional aggregation analysis and prediction algorithm, which generates complex multi-dimensional statistical results and trend predictions based on user query instructions, combined with multi-dimensional data, through aggregation functions and prediction models.
[0081] Specifically, in this embodiment, according to the user query instruction, the target data set is extracted, the data is cleaned and the format is unified, and the data is aggregated and counted in multiple dimensions according to the dimensions specified by the user (such as time, region, product type), and statistical results are generated through aggregation functions (such as sum, average, maximum, minimum, etc.). Based on the aggregation results and historical data, a prediction model is constructed to predict future trends, and time series analysis and machine learning algorithms are applied to generate prediction results. It should be noted that the aggregation function in this embodiment can be expressed as:
[0082] ,
[0083] in, Expressed as in dimension The result of the aggregation calculation on the data set M is: Represents the aggregation dimension specified by the user, such as time, region, product type, etc. Represented as the i-th record in dimension The value on It is an aggregate function, such as sum, average, maximum, minimum, etc.
[0084] The prediction formula can be expressed as:
[0085] ,
[0086] in, is the predicted value at time point t, , , For time point , , The historical values are all aggregated values after the above aggregation function. , , is a model parameter, which indicates the weight of the impact of historical values on the predicted values. is the error term, which represents the random error of the prediction model.
[0087] After obtaining the predicted value, the following formula can be used to assign tasks:
[0088] ,
[0089] ,
[0090] in, Processing data set for subsystem j The computational efficiency of For subsystem j, the data set The result of multi-dimensional aggregation calculation, Processing data set for subsystem j time, The optimal subsystem represents the subsystem with the highest computational efficiency. According to the aggregated computation results and processing time of the subsystem, the computational efficiency of each subsystem is calculated, and the subsystem with the highest computational efficiency is selected for data processing to optimize the computational performance.
[0091] According to a web-based big data analysis method provided by the present application, the microservice front-end components of the subsystem are integrated into the low-code platform; the shared data of the subsystem is integrated in the low-code platform according to the subsystem loader; the shared data is preprocessed according to the preset cleaning rules; when a user's data analysis request is received, the data analysis task corresponding to the data analysis request is divided into N subtasks according to the preprocessed target shared data; the N subtasks are respectively assigned to the corresponding subsystems for data analysis. Through the implementation of the present application, the shared data of the subsystem is integrated and preprocessed on the low-code platform. When executing the data analysis task, the data analysis task is divided into multiple subtasks according to the preprocessed target shared data, and the processing efficiency of the data analysis task is effectively improved through the distributed intelligent task scheduling mechanism.
[0092] Figure 2 A web-based big data analysis device is provided in an embodiment of the present application. The web-based big data analysis device can be used to implement the web-based big data analysis method in the above-mentioned embodiment. Figure 2 As shown, the web-based big data analysis device mainly includes:
[0093] An integration module 10, for integrating the microservice front-end components of the subsystem into the low-code platform;
[0094] An integration module 20, for integrating the shared data of the subsystems in the low-code platform according to the subsystem loader;
[0095] A preprocessing module 30, used for preprocessing the shared data according to preset cleaning rules;
[0096] A division module 40 is used for, when receiving a data analysis request from a user, dividing the data analysis task corresponding to the data analysis request into N subtasks according to the preprocessed target shared data; wherein N is an integer greater than or equal to 2;
[0097] The allocation module 50 is used to allocate the N subtasks to the corresponding subsystems for data analysis.
[0098] In an optional implementation of this embodiment, the integration module is specifically used to: generate a sandbox in the low-code platform; wherein the sandbox is an independent virtual environment; generate an environment identifier for the microservice front-end component of the subsystem according to the global manager; and assign the microservice front-end component and the corresponding environment identifier to the sandbox.
[0099] In an optional implementation of this embodiment, the big data analysis device also includes: a determination module, an acquisition module, a creation module, and a loading module. The determination module is used to determine the target subsystem according to the data analysis request. The acquisition module is used to obtain the target environment identifier corresponding to the target subsystem. The creation module is used to create a window container according to the target environment identifier. The loading module is used to load the microservice front-end component of the target subsystem into the window container.
[0100] In an optional implementation of this embodiment, the big data analysis device further includes: a sending module and a receiving module. The sending module is used to: send the initial data required for data sharing to the microservice front-end component according to a preset configuration file; when a user's data query instruction is detected, the data query instruction is sent to the microservice front-end component according to the event bus. The receiving module is used to: receive the shared data sent by the microservice front-end component based on the initial data and the data query instruction.
[0101] Furthermore, in an optional implementation of this embodiment, the big data analysis device further includes: a prediction module. The determination module is further used to: determine the target subsystem window that the user needs to query according to the data query instruction. The prediction module is used to: predict the subsystem window to be accessed according to the target subsystem and the user's historical query record. The loading module is further used to: load the shared data of the corresponding access subsystem according to the prediction result of the subsystem window to be accessed.
[0102] In an optional implementation of the present embodiment, the partitioning module is specifically used to: determine the computational logic tree of the data analysis task according to the data analysis request; determine the data format of the target shared data; and divide the data analysis task into multiple subtasks according to the computational logic tree and the data format.
[0103] In an optional implementation of this embodiment, the big data analysis device further includes: a monitoring module and a processing module. The monitoring module is used to: monitor the task execution status of the subsystem in real time. The processing module is used to: trigger a rescheduling mechanism when the main system detects that a subtask fails or the result is incomplete; and reallocate unfinished subtasks according to the rescheduling mechanism.
[0104] According to a web-based big data analysis device provided by the present application, the microservice front-end components of the subsystem are integrated into the low-code platform; the shared data of the subsystem is integrated in the low-code platform according to the subsystem loader; the shared data is preprocessed according to preset cleaning rules; when a user's data analysis request is received, the data analysis task corresponding to the data analysis request is divided into N subtasks according to the preprocessed target shared data; the N subtasks are respectively assigned to the corresponding subsystems for data analysis. Through the implementation of the present application, the shared data of the subsystems are integrated and preprocessed on the low-code platform. When executing the data analysis task, the data analysis task is divided into multiple subtasks according to the preprocessed target shared data, and the processing efficiency of the data analysis task is effectively improved through the distributed intelligent task scheduling mechanism.
[0105] According to the application plan provided Figure 3 An electronic device provided in an embodiment of the present application. The electronic device can be used to implement the web-based big data analysis method in the aforementioned embodiment, mainly including:
[0106] The memory 301, the processor 302, and the computer program 303 stored in the memory 301 and executable on the processor 302, the memory 301 and the processor 302 are connected by communication. When the processor 302 executes the computer program 303, the web-based big data analysis method in the aforementioned embodiment is implemented. The number of processors can be one or more.
[0107] The memory 301 may be a high-speed random access memory (RAM) memory, or a non-volatile memory (non-volatile memory), such as a disk memory. The memory 301 is used to store executable program codes, and the processor 302 is coupled to the memory 301 .
[0108] Furthermore, the present application also provides a computer-readable storage medium, which may be provided in the electronic device in the above embodiments. Figure 3 Memory in the illustrated embodiment.
[0109] The computer readable storage medium stores a computer program, and when the program is executed by the processor, the web-based big data analysis method in the aforementioned embodiment is implemented. Furthermore, the computer storable medium can also be a U disk, a mobile hard disk, a read-only memory (ROM), a RAM, a magnetic disk or an optical disk, and other media that can store program codes.
[0110] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0111] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., various media that can store program codes.
[0112] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A web-based big data analysis method, characterized in that: include: Integrate the microservice front-end components of the subsystem into the low-code platform; Integrate the shared data of the subsystem in the low-code platform according to the subsystem loader; Preprocessing the shared data according to preset cleaning rules; When receiving a data analysis request from a user, dividing the data analysis task corresponding to the data analysis request into N subtasks according to the preprocessed target shared data; wherein N is an integer greater than or equal to 2; Allocate the N subtasks to the corresponding subsystems for data analysis; The method further comprises: Aggregate and count the multi-dimensional data of the shared data according to the dimension specified in the user query instruction, and generate an aggregation result according to the aggregation function; build a prediction model based on the aggregation result and corresponding historical data, and generate a prediction result reflecting future trends; allocate computing tasks to the subsystem according to the prediction result; The above aggregation results are calculated by the following formula: , in, Expressed as in dimension The result of the aggregation calculation on the data set M is: Represents the aggregation dimension specified by the user. Represented as the i-th record in dimension The value on is an aggregate function; The above prediction results are calculated by the following formula: , in, is the predicted value at time point t, , , For time point , , The historical values are all the aggregation results after the above aggregation function. , , is a model parameter, which indicates the weight of the impact of historical values on the predicted values. is the error term, which represents the random error of the prediction model; The task allocation is performed using the following formula: , , in, is the computational efficiency of subsystem j in processing dataset M, is the result of multi-dimensional aggregation calculation of data set M by subsystem j, For subsystem j The time to process the dataset M, The optimal subsystem represents the subsystem with the highest computational efficiency. The computational efficiency of each subsystem is calculated based on the aggregated computational results and processing time of the subsystem, and the subsystem with the highest computational efficiency is selected for data processing to optimize the computational performance.
2. The web-based big data analysis method according to claim 1, characterized in that: The step of integrating the microservice front-end components of the subsystem into the low-code platform includes: Generate a sandbox in the low-code platform; wherein the sandbox is an independent virtual environment; Generate an environment identifier for the microservice front-end component of the subsystem according to the global manager; The microservice front-end component and the corresponding environment identifier are assigned to the sandbox.
3. The web-based big data analysis method according to claim 2, characterized in that: After the step of assigning the microservice front-end component and the corresponding environment identifier to the sandbox, the method further includes: Determining a target subsystem according to the data analysis request; Obtaining a target environment identifier corresponding to the target subsystem; Creating a window container according to the target environment identifier; Load the microservice front-end component of the target subsystem into the window container.
4. The web-based big data analysis method according to claim 1, characterized in that: After the step of integrating the microservice front-end components of the subsystem into the low-code platform, it also includes: Sending the initial data required for data sharing to the microservice front-end component according to a preset configuration file; When a user's data query instruction is detected, the data query instruction is sent to the microservice front-end component according to the event bus; Receive the shared data sent by the microservice front-end component based on the initial data and the data query instruction.
5. The web-based big data analysis method according to claim 4, characterized in that: The method further comprises: Determine the target subsystem window that the user needs to query according to the data query instruction; Predicting the subsystem window to be accessed based on the target subsystem and the user's historical query records; The shared data of the corresponding access subsystem is loaded according to the prediction result of the subsystem window to be accessed.
6. The web-based big data analysis method according to claim 1, characterized in that: The step of dividing the data analysis task corresponding to the data analysis request into N subtasks according to the preprocessed target shared data when receiving the data analysis request of the user includes: Determine a computational logic tree of the data analysis task according to the data analysis request; Determining a data format of the target shared data; The data analysis task is divided into N subtasks according to the computational logic tree and the data format.
7. The web-based big data analysis method according to claim 1, characterized in that: After the step of respectively allocating the N subtasks to the corresponding subsystems for data analysis, the method further includes: Monitoring the task execution status of the subsystem in real time; When the main system detects that a subtask has failed or the result is incomplete, a rescheduling mechanism is triggered; The unfinished subtasks are reallocated according to the rescheduling mechanism.
8. A web-based big data analysis device, characterized in that: The web-based big data analysis device comprises: Integration module, used to integrate the microservice front-end components of the subsystem into the low-code platform; An integration module, configured to integrate the shared data of the subsystems in the low-code platform according to the subsystem loader; A preprocessing module, used for preprocessing the shared data according to preset cleaning rules; A division module, for, when receiving a data analysis request from a user, dividing the data analysis task corresponding to the data analysis request into N subtasks according to the preprocessed target shared data; wherein N is an integer greater than or equal to 2; An allocation module, used for respectively allocating the N subtasks to the corresponding subsystems for data analysis; The allocation module is further used to aggregate and count the multi-dimensional data of the shared data according to the dimension specified in the user query instruction, and generate an aggregation result according to the aggregation function; build a prediction model based on the aggregation result and the corresponding historical data, and generate a prediction result reflecting future trends; and allocate the computing tasks of the subsystem according to the prediction result; The above aggregation results are calculated by the following formula: , in, Expressed as in dimension The result of the aggregation calculation on the data set M is: Represents the aggregation dimension specified by the user. Represented as the i-th record in dimension The value on is an aggregate function; The above prediction results are calculated by the following formula: , in, is the predicted value at time point t, , , For time point , , The historical values are all the aggregation results after the above aggregation function. , , is a model parameter, which indicates the weight of the impact of historical values on the predicted values. is the error term, which represents the random error of the prediction model; The task allocation is performed using the following formula: , , in, is the computational efficiency of subsystem j in processing dataset M, is the result of multi-dimensional aggregation calculation of data set M by subsystem j, For subsystem j The time to process the dataset M, The optimal subsystem represents the subsystem with the highest computational efficiency. The computational efficiency of each subsystem is calculated based on the aggregated computational results and processing time of the subsystem, and the subsystem with the highest computational efficiency is selected for data processing to optimize the computational performance.
9. An electronic device, characterized in that: The device comprises a memory and a processor, wherein: The processor is used to execute the computer program stored in the memory; When the processor executes the computer program, the steps in the web-based big data analysis method described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps in the web-based big data analysis method described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Low-code page creation method and device, equipment and medium
CN114860240A
Web front-end application integration system and method based on micro service
CN116842297A
Mutually neutral independent distributed computing and node management method
CN117193987A