Model-based visualized big data stream batch generation management system and method

CN120950543BActive Publication Date: 2026-09-22FIBERHOME TELECOMMUNICATION TECHNOLOGIES CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510982436.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2026-09-22
Estimated Expiration
2045-07-16

AI Technical Summary

Technical Problem

[0005]目前,实现大数据流批一体化,需要创建数据采集任务,数仓各层级计算分析任务,数据调度任务,数据可视化查询配置等,同时大量任务实现又采用有码化方式,流批一体化又靠任务与任务之间的底层配置来构建,实际应用中会造成任务配置项多,底层开发量大,管理效率低下,人工运维成本高昂等问题

Benefits of technology

[0018]本申请实施例提供的技术方案带来的有益效果包括:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120950543B_ABST
    Figure CN120950543B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on model visual big data stream batch generation management system and method, it is related to big data processing technical field.The system, data model metadata information module is used to complete the visual creation of data acquisition model, data warehouse model, business output model;And when data warehouse model is released, the physical of data warehouse model is completed.Data model task module is used to complete task related configuration;And after executing data warehouse model release operation, the corresponding data acquisition task, data warehouse computing task, scheduling task and data output task are generated.Data acquisition module is used to complete the data acquisition processing based on data acquisition model.Data warehouse is used to complete the integrated computing processing based on data warehouse model.Data output module is used to complete the data output processing based on business output model.The application can provide a kind of convenient, efficient, convenient scheme for the creation, integration, management of big data stream batch task, satisfy practical application demand.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of big data processing technology, specifically to a model-based visualization big data batch generation management system and method. Background Technology

[0002] Data plays a vital role not only in the business world but also profoundly impacts all aspects of social interaction and personal life. In the digital age, data is seen as a new form of "currency," its value lying in its ability to support decision-making through analysis and processing, thereby driving business growth and social development.

[0003] However, data itself is merely a carrier of information; its potential value can only be mined and utilized through effective analytical methods. Traditional data analysis methods primarily rely on offline models, which analyze historical data to provide insights based on past data. However, with constantly changing business needs and the increasing importance of real-time decision-making, relying solely on offline model analysis can no longer meet the real-time and accuracy requirements of modern businesses. Therefore, big data stream batch analysis models have emerged, aiming to combine the advantages of real-time data streams and batch data processing to provide more comprehensive and timely data analysis results.

[0004] While big data batch processing models can theoretically meet modern business needs, their practical application faces numerous challenges. For example, the implementation of batch processing models often focuses only on top-level design, and how to implement these top-level models presents unprecedented challenges.

[0005] Currently, achieving integrated big data flow and batch processing requires creating data acquisition tasks, data warehouse calculation and analysis tasks at various levels, data scheduling tasks, and data visualization query configurations. At the same time, many tasks are implemented using coded methods, and integrated flow and batch processing relies on the underlying configuration between tasks. In practical applications, this leads to problems such as numerous task configuration items, large amounts of underlying development work, low management efficiency, and high manual operation and maintenance costs.

[0006] Therefore, how to provide a convenient, efficient, and easy solution for the creation, integration, and management of big data batch processing tasks is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0007] The purpose of this invention is to overcome the shortcomings of the above-mentioned background technology and provide a model-based visualization big data batch generation management system and method, which can provide a convenient, efficient and convenient solution for the creation, integration and management of big data batch tasks, and meet the needs of practical applications.

[0008] To achieve the above objectives, in a first aspect, embodiments of the present invention provide a model-visualized big data batch generation and management system, comprising: a data model metadata information module, a data model task module, a data acquisition module, a data warehouse, and a data output module located in the service layer, and a client located in the presentation layer; The data model metadata information module is used to: respond to client scheduling to complete the visual creation of data acquisition model, data warehouse model, and business output model; and complete the physicalization of data warehouse model when data warehouse model is published; The data model task module is used to: respond to the client's scheduling, complete the task-related configuration; and after executing the data warehouse model release operation, generate corresponding data acquisition tasks, data warehouse computing tasks, scheduling tasks and data output tasks based on the task-related configuration. The data acquisition module is used to: complete data acquisition processing based on the data acquisition model according to the generated data acquisition task; The data warehouse is used to: complete integrated stream and batch computing processing based on the data warehouse model according to the generated data warehouse computing tasks and scheduling tasks; The data output module is used to: complete data output processing based on the business output model according to the generated data output task.

[0009] In conjunction with the first aspect, in one implementation, the data model metadata information module completes the physicalization of the data warehouse model, including: generating specific physical tables or message queue topics based on the data warehouse model metadata.

[0010] In conjunction with the first aspect, in one implementation, the data acquisition model includes a target data model, an exported data model, and an acquisition model mapping relationship; the target data model is used to configure the target acquisition source; the exported data model is used to set exported data fields and construct the data model structure for data warehouse inflow; the acquisition model mapping relationship is a mapping relationship between the target acquisition fields and the exported data model fields. The business output model includes data warehouse acquisition target configuration, business output data model, and output model mapping relationship; the data warehouse acquisition target configuration includes message queue stream data and distributed storage batch data warehouse information; the business output data model is used to construct a business wide table model, and set model index, partition information, and cleanup mechanism for the model table.

[0011] In conjunction with the first aspect, in one implementation, the data model task module completes task-related configuration, including: Scheduling flow configuration related to scheduling tasks: Visually build the data warehouse computing scheduling link by dragging and dropping table model nodes; Data warehouse computing task related computing logic development: If it is a batch computing table, then develop batch task SQL for node information; if it is a stream computing table, then develop stream task SQL for node information; if it is a stream-batch table, then both batch task SQL and stream task SQL for node information are required. Data link configuration related to data acquisition and data output tasks: Set the acquisition target source, exported data model, and acquisition model mapping relationship in the data acquisition model; set the data warehouse acquisition target configuration, business output data model, and output model mapping relationship in the business output model.

[0012] In conjunction with the first aspect, in one implementation, the data acquisition module completes data acquisition processing based on the data acquisition model according to the generated data acquisition task, including: The data acquisition module collects the required computational data according to the generated data acquisition task. During the acquisition process, the target data is transformed into structured data of the ODS layer according to the acquisition model mapping relationship in the set data acquisition model, and the transformed data is imported into the message queue, which serves as the sole entry point for the data warehouse.

[0013] In conjunction with the first aspect, in one implementation, the data warehouse is divided into a real-time data warehouse and an offline data warehouse; the real-time data warehouse relies on a distributed message queue, uses a task scheduler for scheduling, and uses a real-time computing engine to complete the processing of streaming computing tasks; the offline data warehouse relies on a distributed storage database, uses a task scheduler for scheduling, and uses an offline computing engine to complete the processing of batch computing tasks.

[0014] In conjunction with the first aspect, in one implementation, the data output module completes data output processing based on the business output model according to the generated data output task, including: The data output module retrieves the required data warehouse calculation results from the data warehouse according to the generated data output task and the data warehouse collection target configuration in the business output model; it creates a business wide table according to the business wide table model in the business output model; it imports the business fields of the obtained data warehouse calculation results into the business wide table according to the output model mapping relationship in the business output model; and it sets up indexes, partitions, and cleanup mechanisms for the business wide table according to the configuration.

[0015] In conjunction with the first aspect, in one embodiment, the system further includes a visualization database; the visualization database is used to store output data processed by the data output module.

[0016] In conjunction with the first aspect, in one implementation, the system further includes a front-end low-code business display module located in the presentation layer and a back-end OLAP analysis and processing module located in the service layer; the front-end low-code business display module is used to: respond to client scheduling, based on the business output model, call the back-end OLAP analysis and processing module to perform OLAP low-code analysis configuration; and configure visualization components and bind query result datasets through a graphical interface; the back-end OLAP analysis and processing module is used to: call the query engine to parse the OLAP low-code analysis configuration, call the visualization database to perform multidimensional analysis, and render the analysis results to the client interface in real time.

[0017] Secondly, embodiments of the present invention also provide a model-based visualization big data stream batch generation and management method applying the system of the first aspect embodiment, comprising: The data model metadata information module responds to the client's scheduling and completes the visual creation of the data acquisition model, data warehouse model, and business output model; The data model task module responds to the client's scheduling and completes the task-related configuration; After executing the data warehouse model deployment operation, the data model task module triggers an automated process; the automated process includes: The data model metadata information module automatically completes the physicalization of the data warehouse model; Based on the task-related configuration, the data model task module automatically generates corresponding data acquisition tasks, data warehouse computing tasks, scheduling tasks, and data output tasks. The data acquisition module completes data acquisition processing based on the data acquisition model according to the generated data acquisition tasks; the data warehouse completes integrated stream and batch computing processing based on the data warehouse model according to the generated data warehouse computing tasks and scheduling tasks; and the data output module completes data output processing based on the business output model according to the generated data output tasks.

[0018] The beneficial effects of the technical solutions provided in this application include: This application proposes a model-based scheme for constructing integrated batch and stream tasks. It comprehensively considers the creation and association of data tasks across all tables in the data warehouse. Through visual configuration of data acquisition models, data warehouse models, business output models, and task relationship management, it achieves the creation of batch and stream tasks and manages all tasks in association to construct an integrated task. This application uses a configuration model to automatically generate physical storage tables and automatically create acquisition tasks, computation tasks, scheduling tasks, and data output tasks. It eliminates traditional coded operations and cumbersome underlying configurations, reducing operational difficulty and the workload of underlying configuration. This provides a convenient, efficient, and easy-to-use solution for the creation, integration, and management of big data batch and stream tasks, meeting practical application needs. Attached Figure Description

[0019] Figure 1 This is a schematic diagram of the functional modules of the first embodiment of the model-visualized big data batch generation management system of this application; Figure 2 This is a schematic diagram of the data warehouse model in the first embodiment of the system; Figure 3 This is an interface diagram for managing table cell information in the data warehouse model in the first embodiment of the system; Figure 4 This is a schematic diagram of the data acquisition model and data acquisition process in the first embodiment of the system; Figure 5 This is a schematic diagram of the service output model and service output process in the first embodiment of the system; Figure 6 This is an interface diagram of the calculation process generated in the first embodiment of the system; Figure 7 This is an interface diagram showing the target source settings in the first embodiment of the system. Figure 8 This is a schematic diagram of data warehouse model deployment and task generation in the first embodiment of the system; Figure 9 This is a schematic diagram of the functional modules of the second embodiment of the model-visualized big data batch generation management system of this application; Figure 10 This is a schematic diagram of the low-code graphical configuration in the second embodiment of the system; Figure 11 This is a flowchart illustrating the first embodiment of the model-visualized big data batch generation and management method of this application; Figure 12 This is a flowchart illustrating the second embodiment of the model-visualized big data batch generation and management method of this application; Figure 13 This is a schematic diagram illustrating how the various functions are implemented using the method described in this embodiment. Detailed Implementation

[0020] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.

[0021] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0022] In a first aspect, embodiments of this application provide a model-visualized big data batch generation and management system.

[0023] In one embodiment, reference is made to Figure 1 As shown, Figure 1 This is a schematic diagram of the functional modules of the first embodiment of the model-visualized big data batch generation management system of this application. Figure 1 As shown, a model-visualized big data batch generation and management system includes: a data model metadata information module, a data model task module, a data acquisition module, a data warehouse, and a data output module located in the service layer; and a client located in the presentation layer. The service layer is the middle layer in the big data processing architecture, primarily responsible for handling business logic and data operations. It resides between the data source and the presentation layer, providing functions for data processing, business logic processing, and data interaction. The presentation layer is the outermost layer in the big data processing architecture, primarily responsible for interacting with the user, providing the user interface (UI) and user experience (UX). It displays the data provided by the service layer to the user in a user-friendly manner and passes user operation requests to the service layer.

[0024] The data model metadata information module is used to: respond to client scheduling, complete the visual creation of data acquisition models, data warehouse models, and business output models; and complete the physicalization of data warehouse models when they are published.

[0025] It is understandable that building a data warehouse model mainly involves four steps: business modeling, conceptual modeling, logical modeling, and physical modeling. (See also...) Figure 2 As shown, business modeling and conceptual modeling are typically "preparatory work" performed beforehand by operational personnel (such as system development and maintenance personnel). For example, in this "preparatory work," operational personnel will analyze the business based on experience to complete business modeling; they will also complete conceptual modeling based on ER diagrams, dimensional analysis, and indicator analysis, obtaining relevant information. None of the above "preparatory work" needs to be completed within the system. In other words, this system only needs to implement the logical and physical modeling of the data warehouse model. Figure 2As shown, logical modeling can be performed in this system to materialize the conceptual model, generate specific table structures, and fill in the field information within the tables. Metadata information for each level of the table can be entered, making the table information more concrete. Table attributes can be set as stream computing tables, batch computing tables, or stream-batch computing tables. Furthermore, when the data warehouse model is deployed, the physicalization process of the data warehouse model can be completed. This physicalization process is essentially the instantiation of logical modeling information within the data warehouse model, thus completing the physical modeling process. Specifically, when the data warehouse model is deployed, physical modeling is automatically completed, instantiating the logical modeling information and generating specific table information or message queue inflow data information. Calculations and storage are then performed based on the set information. For example, specific physical tables (such as Hive / MySQL) or message queue topics (such as Kafka) can be generated based on the data warehouse model metadata information module. In practical implementation, the backend microservices can interact with the task scheduler to create tasks. These tasks are responsible for creating physical tables. After the task is executed, the execution result is sent back to the backend microservices. The microservices then determine whether the physical table creation was successful based on the callback result, thus determining whether the data warehouse model has been successfully deployed.

[0026] In this embodiment, the data model metadata information module, through the configuration and management of the data warehouse model, can maintain table element information at each level of the data model and generate physical storage at each level. Based on the table element information, batch table creation can be completed on distributed physical storage, and streaming computing topics can be created in a distributed message queue. For table element information management of the data warehouse model, please refer to... Figure 3 As shown.

[0027] See Figure 4 As shown, the data acquisition model in this embodiment includes a target data model, an exported data model, and an acquisition model mapping relationship. The target data model is used to configure the acquisition target source, which includes Excel files, system databases, and device reports. When the acquisition target source is an Excel file, the corresponding column field names need to be set, and the column field information and acquisition path address are used as the acquisition target information. When the acquisition target source is a system database, the system table structure and system database source information need to be entered, and the entered information is used as the target data model. When the acquisition target is a device report, the device object model is directly used as the target data model. The exported data model is used to set the exported data fields and construct the data model structure for the data warehouse. The acquisition model mapping relationship is the mapping relationship between the acquisition target fields and the exported data model fields. During subsequent data acquisition processing, the data acquisition module will transform the acquired target data according to the acquisition model mapping relationship, converting it into ODS layer structured data; and then the transformed data will flow into a message queue. This message queue channel will serve as the sole entry point for data in the data warehouse.

[0028] In this embodiment, for data warehouse data that needs to be displayed for business purposes, a business output model is used to manage the business output of the data warehouse calculation results. See also... Figure 5 As shown, the business output model in this embodiment includes data warehouse acquisition target configuration, business output data model, and output model mapping relationship. The data warehouse acquisition target configuration includes two types: message queue streaming data and distributed storage batch data warehouse information. The acquisition target must be directly attached to an instance of the data warehouse model. For streaming models, the acquisition target can only be attached to one model instance; for batch computing, the acquisition target can be multiple model instances (i.e., batch computing tables can be cascaded across multiple tables, while streaming computing tables can only be retrieved for a single table). The business output data model is used to summarize business data from a business perspective and construct a wide business table model to adapt to high-concurrency query scenarios. Simultaneously, it can also set model indexes, partition information, and cleanup mechanisms for the model tables from a business data perspective. The output model mapping relationship is the mapping relationship established between the data warehouse acquisition target and the business output data model. The field mapping determines the acquisition location and storage location of the acquired field data.

[0029] The data model task module is used to: respond to the client's scheduling, complete the task-related configuration; and after executing the data warehouse model release operation, generate corresponding data acquisition tasks, data warehouse computing tasks, scheduling tasks, and data output tasks based on the task-related configuration.

[0030] It is understood that this embodiment, through the data model task module, can not only manage data acquisition and business output processing, but also manage the data model calculation and scheduling processing for a specific business. When the data model task module manages the data model calculation and scheduling processing for a specific business, it will configure streaming tasks, batch tasks, or streaming-batch tasks on each model table according to the data warehouse model. After configuration, a graphical connection node is used to manage the task scheduling process. Batch processing and streaming processing will each create a specific scheduling task flow; streaming tasks will be executed immediately for real-time calculation; batch tasks will execute batch calculation tasks sequentially according to the scheduling flow when the batch task processing conditions are met. The interface diagram of the generated calculation process is shown below. Figure 6 As shown.

[0031] For example, as an optional implementation, the data model task module completes task-related configurations and may include the following: (1) Scheduling flow configuration related to scheduling tasks: The data warehouse computing scheduling link is built by dragging and dropping table model nodes. Specifically, the table model can be dragged and dropped in the interface to form the information of each node of the computing model, and the information of each node can be connected in an interface to set the data warehouse computing scheduling process.

[0032] (2) Development of computational logic related to data warehouse computing tasks: If it is a batch computing table, then batch task SQL development is performed on the node information; if it is a stream computing table, then stream task SQL development is performed on the node information; if it is a stream batch table, then batch task SQL development and stream task SQL development are required for the node information.

[0033] (3) Data link configuration related to data acquisition and data output tasks: This involves setting the acquisition target source, exported data model, and mapping relationship in the data acquisition model; and setting the data warehouse acquisition target configuration, business output data model, and output model mapping relationship in the business output model. The interface diagram for setting the acquisition target source can be found in [link to interface diagram]. Figure 7 As shown.

[0034] Furthermore, it can be understood that based on the above task-related configurations, corresponding data acquisition tasks, data warehouse computing tasks, scheduling tasks, and data output tasks can be generated. The specific process may include: when the data warehouse model is enabled, corresponding data acquisition tasks will be generated based on the configured data acquisition model; corresponding data output tasks will be generated based on the configured business output model; and corresponding data warehouse computing tasks and scheduling tasks will be generated based on the scheduling flow configuration related to the scheduling tasks and the computing logic development related to the data warehouse computing tasks. Additionally, it can be understood that when the data warehouse model is paused, all generated tasks will be closed; when the data warehouse model is deprecated, all generated tasks will also be deprecated. For details on data warehouse model deployment and task generation, please refer to [link to relevant documentation]. Figure 8 As shown.

[0035] The data acquisition module is used to complete data acquisition and processing based on the data acquisition model according to the generated data acquisition task.

[0036] For example, as an optional implementation, the data acquisition module completes data acquisition processing based on the data acquisition model according to the generated data acquisition task, which may include the following: The data acquisition module collects the required computational data according to the generated data acquisition task. During acquisition, it transforms the target data into structured data at the ODS layer based on the mapping relationship in the set data acquisition model, and imports the transformed data into a message queue that serves as the sole entry point for the data warehouse. Specifically, if the data model metadata information module describes the acquired data as a stream processing table, the acquired data will be used as data to be processed in stream computing; if it describes the acquired data as a batch processing table, the acquired data will be pulled into a distributed repository as data to be processed in batch computing.

[0037] In practical implementation, the data acquisition module can integrate DataX and Flume components in its backend. DataX can collect offline data, while Flume can collect real-time data. In actual applications, the choice of components can be made according to specific circumstances; this embodiment does not impose any specific limitations.

[0038] The data warehouse is used to complete integrated stream and batch computing processing based on the data warehouse model, according to the generated data warehouse computing tasks and scheduling tasks.

[0039] In this embodiment, the data warehouse is divided into a real-time data warehouse and an offline data warehouse. The real-time data warehouse relies on a distributed message queue, uses a task scheduler for scheduling, and employs a real-time computing engine to complete streaming computing tasks. The offline data warehouse relies on a distributed storage database, uses a task scheduler for scheduling, and employs an offline computing engine to complete batch computing tasks. The real-time data warehouse is stored in a distributed message queue, and the offline data warehouse is stored in a distributed storage database.

[0040] For example, in practical implementation, you can use Kafka clusters, Hadoop clusters, Hive clusters, Flink clusters, Dolphinscheduler clusters, and Redis clusters.

[0041] The Kafka cluster, acting as a distributed message queue, is responsible for storing real-time data and transmitting critical information. The HDFS (Hadoop Distributed File System) of the Hadoop cluster, as a file storage system, is responsible for storing massive amounts of offline data. The YARN (Yet Another Resource Negotiator) of the Hadoop cluster is responsible for resource scheduling during massive data processing. The Hive cluster serves as a big data engine for offline computing and analysis. The Flink cluster serves as a big data engine for real-time computing and analysis. The Dolphinscheduler cluster is responsible for scheduling and managing all big data tasks. The Redis cluster implements distributed caching, supporting high-performance queries.

[0042] Data will flow from the Kafka cluster and, depending on whether the scenario is real-time or offline, will be fed into Kafka or HDFS. The DolphInscheduler will then be responsible for scheduling the offline engine (Hive) or the real-time engine (Flink) for data analysis and computation. After obtaining the data analysis results, the necessary data will be appropriately imported into a visualization database (such as PostgreSQL). Furthermore, during data analysis and computation by the offline engine (Hive) or the real-time engine (Flink), the Redis cluster can serve as a computation cache to accelerate computation and improve efficiency; or during queries, the Redis cluster can serve as a query cache to facilitate data retrieval and improve query efficiency.

[0043] The data output module is used to complete data output processing based on the business output model according to the generated data output task.

[0044] For example, as an optional implementation, the data output module completes data output processing based on the business output model according to the generated data output task, which may include the following: The data output module retrieves the required data warehouse calculation results from the data warehouse according to the generated data output task and the data warehouse collection target configuration in the business output model; it creates a business wide table according to the business wide table model in the business output model; it imports the business fields of the obtained data warehouse calculation results into the business wide table according to the output model mapping relationship in the business output model; and it sets indexes, partitions, and cleanup mechanisms for the business wide table according to the configuration.

[0045] Further, see Figure 1 As shown, as an optional implementation, the system also includes a visualization database. This visualization database is used to store the output data processed by the data output module. Specifically, after completing the data output processing based on the business output model, the data output module stores the processed output data in the visualization database, such as... Figure 1 As shown, using a visual database decouples business data querying and application from the data warehouse, allowing focus solely on business-side implementation, which is more conducive to business development.

[0046] It is understood that the system of this embodiment can be applied to (or implemented in) a distributed application platform. For example, the data acquisition module, data model metadata information module, data model task module, and data output module of this system can be integrated into a data processing model microservice of the platform. The data processing model microservice in the platform can be called to realize the big data stream batch generation and management function based on model visualization. Of course, in addition to the microservice that integrates the functional modules of this system (i.e., the data processing model microservice), the platform can also set up basic microservices involving user registration center, gateway, configuration center, and other related microservices. In implementation, the data acquisition module, streaming computing task, and data output module all use distributed message queues for data transmission. The data model task module communicates with the data warehouse through a task scheduling manager, and the data output module communicates with the distributed storage of the data warehouse.

[0047] As can be seen from the above, this embodiment proposes a model-based solution for constructing integrated batch and stream tasks. It comprehensively considers the creation and association of data tasks across all tables in the data warehouse. Through visual configuration of data acquisition models, data warehouse models, business data models, and task relationship management, it achieves the creation of batch and stream tasks and manages all tasks in association to build an integrated task system. This solution uses a configuration model to automatically generate physical storage tables and automatically create acquisition tasks, computation tasks, scheduling tasks, and output tasks, reducing the learning curve and the workload of underlying configuration. It provides a convenient, efficient, and easy-to-use solution for the creation, integration, and management of big data batch and stream tasks, meeting practical application needs.

[0048] Furthermore, in yet another embodiment, referring to Figure 9 As shown, Figure 9 This is a schematic diagram of the functional modules of the second embodiment of the model-visualized big data batch generation management system of this application. Figure 9 As shown, a model-visualized big data batch generation and management system also includes: a front-end low-code business display module located in the presentation layer and a back-end OLAP analysis and processing module located in the service layer.

[0049] The front-end low-code business display module is used to: respond to client scheduling, call the back-end OLAP analysis and processing module for OLAP low-code analysis configuration based on the business output model; and configure visualization components and bind query result datasets through a graphical interface. In practical applications, when configuring OLAP low-code analysis, multiple business output models can be selected, i.e., multi-model combined queries are supported; and the visualization components include cross-tab components, chart components, etc.

[0050] For example, low-code graphical configuration can be shown as in Figure 10. The low-code configuration process can be illustrated as follows: / / Create an analytics page Analysis page: / / Add a cross table component to the page Component.Crosstab: / / Set the data source for the crosstab component.crosstab.data source = query result / / Set row and column fields component.crosstab.row = time field Component.CrossTable.Column = Performance Category Field Component.Crosstab.Value = Performance Value Field / / Add other visualization components to the page, such as charts, etc. Components.Charts: Component.Chart.DataSource = Query Results Component.Chart.Category = Time Field Component.Chart.Value = Performance Value Field The backend OLAP analysis and processing module is used to: call the query engine to parse the OLAP low-code analysis configuration, call the visualization database to perform multidimensional analysis, and render the analysis results to the client interface in real time.

[0051] Understandably, in this embodiment, low-code configuration, rapid visualization, and result output are achieved through a front-end low-code business display module and a back-end OLAP analysis and processing module, enabling data to be presented in charts and graphs. Low-code visualization configuration provides an intuitive and easy-to-use development environment where developers can select and configure various data processing, analysis, and visualization components to build big data solutions that meet their specific needs. In this way, developers can process and analyze big data more efficiently, thereby quickly extracting valuable business insights. Furthermore, due to the characteristics of low-code, this configuration method significantly lowers the technical barrier, allowing more business personnel and analysts to participate in big data processing and analysis.

[0052] Secondly, embodiments of this application provide a model-based visualization big data batch generation and management method that applies the system in the first aspect embodiment.

[0053] In one embodiment, reference is made to Figure 11 As shown, Figure 11 This is a flowchart illustrating the first embodiment of the model-visualized big data batch generation and management method of this application. Figure 11 As shown, a model-visualized big data batch generation and management method includes: Step S10: The data model metadata information module responds to the client's scheduling and completes the visual creation of the data acquisition model, data warehouse model, and business output model. Step S20: The data model task module responds to the client's scheduling and completes the task-related configuration; Step S30: After executing the data warehouse model release operation, the data model task module triggers an automated process; the automated process includes: The data model metadata information module automatically completes the physicalization of the data warehouse model; Based on the task-related configuration, the data model task module automatically generates corresponding data acquisition tasks, data warehouse computing tasks, scheduling tasks, and data output tasks. The data acquisition module completes data acquisition processing based on the data acquisition model according to the generated data acquisition tasks; the data warehouse completes integrated stream and batch computing processing based on the data warehouse model according to the generated data warehouse computing tasks and scheduling tasks; and the data output module completes data output processing based on the business output model according to the generated data output tasks.

[0054] Furthermore, in yet another embodiment, reference is made to Figure 12 As shown, Figure 12 This is a flowchart illustrating the second embodiment of the model-visualized big data batch generation and management method of this application. Figure 12 As shown, a model-visualized big data batch generation and management method also includes: Step S40: The front-end low-code business display module responds to the client's scheduling, calls the back-end OLAP analysis and processing module to perform OLAP low-code analysis configuration based on the business output model, and configures the visualization components and binds the query result dataset through the graphical interface. Step S50: The backend OLAP analysis and processing module calls the query engine to parse the OLAP low-code analysis configuration, calls the visualization database to perform multidimensional analysis, and renders the analysis results to the client interface in real time.

[0055] It should be noted that the various variations and specific examples in the above system embodiments are also applicable to the method in this embodiment. Through the detailed description of the above system, those skilled in the art can clearly understand the various implementation methods of the method in this embodiment. Therefore, for the sake of brevity, they will not be described again here.

[0056] See Figure 13 As shown, Figure 13 This is a schematic diagram illustrating how the methods of this embodiment are used to implement various functions. For example... Figure 13As shown in the embodiments of this application, the model-visualized big data batch generation and management method can complete the creation of model table information upon receiving a request to input model table information; complete the configuration of the input information upon receiving a request to input model collection information; complete the configuration of the output information upon receiving the configuration of the output model; complete the configuration of different batch tasks, stream tasks, and stream-batch tasks when creating tasks for the entire model; and generate unified data collection tasks, data warehouse computing tasks, scheduling tasks, and data output tasks based on the model when issuing control requests for the entire model. This method uses configuration for the data model, which is simple to configure. It provides an efficient data query method during business queries, allows for rapid configuration according to specific needs, is easy to operate, and highly efficient, reducing the workload of underlying configuration. It provides a convenient, efficient, and easy-to-use solution for the creation, integration, and management of big data batch tasks.

[0057] Note: The specific embodiments described above are merely examples and not limitations. Those skilled in the art can combine and integrate some steps and devices from the various embodiments described separately above to achieve the effects of the present invention. Such combined and integrated embodiments are also included in the present invention, but will not be described one by one here.

[0058] The advantages, benefits, and effects mentioned in the embodiments of this invention are merely examples and not limitations. They should not be considered as essential features of each embodiment of this invention. Furthermore, the specific details disclosed in the embodiments of this invention are for illustrative and facilitative purposes only and are not limitations. These details do not restrict the embodiments of this invention from being implemented using these specific details.

[0059] The block diagrams of devices, apparatuses, devices, and systems involved in the embodiments of this invention are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as "comprising," "including," "having," etc., are open-ended terms meaning "including but not limited to," and are used interchangeably with them. The terms "or" and "and" as used in the embodiments of this invention refer to the terms "and / or," and are used interchangeably with them unless the context explicitly indicates otherwise. The term "such as" as used in the embodiments of this invention refers to the phrase "such as but not limited to," and is used interchangeably with it.

[0060] The flowcharts and method descriptions in the embodiments of this invention are merely illustrative examples and are not intended to require or imply that the steps of the various embodiments must be performed in the given order. As those skilled in the art will recognize, the steps in the above embodiments can be performed in any order. Words such as "then," "next," etc., are not intended to limit the order of steps; these words are only used to guide the reader through the description of these methods. Furthermore, any reference to a singular element, such as the use of the articles "a," "one," or "the," is not to be construed as limiting that element to the singular.

[0061] Furthermore, the steps and apparatus in the various embodiments of the present invention are not limited to any one embodiment. In fact, new embodiments can be conceived by combining relevant steps and apparatus in the various embodiments of the present invention with the concepts of the present invention, and these new embodiments are also included within the scope of the present invention.

[0062] The various operations in the embodiments of the present invention can be performed by any suitable means capable of performing the corresponding functions. Such means may include various hardware and / or software components and / or modules, including but not limited to hardware circuits or processors.

[0063] The method of this invention includes one or more actions for implementing the method described above. The methods and / or actions may be interchanged without departing from the scope of the claims. In other words, unless a specific order of actions is specified, the order and / or use of specific actions may be modified without departing from the scope of the claims.

[0064] Those skilled in the art can make various changes, substitutions, and modifications to the technology described herein without departing from the teachings defined by the appended claims. Furthermore, the scope of the claims of this disclosure is not limited to the specific aspects of the processes, machines, manufactures, events, means, methods, and actions described above. Currently existing or later-developed processes, machines, manufactures, events, means, methods, or actions that perform substantially the same function or achieve substantially the same result as the corresponding aspects described herein can be utilized. Therefore, the appended claims include such processes, machines, manufactures, events, means, methods, or actions within their scope.

[0065] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the invention. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of the invention. Therefore, the invention is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features disclosed herein.

[0066] The above description has been given for illustrative and descriptive purposes. Furthermore, this description is not intended to limit the embodiments of the invention to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein. Moreover, anything not described in detail in this specification is prior art well known to those skilled in the art.

Claims

1. A model-visualized big data batch generation management system, characterized in that: The system includes a data model metadata information module, a data model task module, a data acquisition module, a data warehouse, and a data output module located in the service layer, as well as a client located in the presentation layer; The data model metadata information module is used to: respond to client scheduling to complete the visual creation of data acquisition model, data warehouse model, and business output model; and complete the physicalization of data warehouse model when data warehouse model is published; The data model task module is used to: respond to client scheduling and complete task-related configurations; After executing the data warehouse model deployment operation, based on the task-related configuration, corresponding data acquisition tasks, data warehouse computing tasks, scheduling tasks, and data output tasks are generated. The data acquisition module is used to: complete data acquisition processing based on the data acquisition model according to the generated data acquisition task; the data acquisition model includes a target data model, an exported data model, and a mapping relationship between acquisition models; the target data model is used to configure the target source for acquisition; The exported data model is used to set the exported data fields and construct the data model structure for data warehouse inflow; the acquisition model mapping relationship is the mapping relationship between the acquisition target fields and the exported data model fields; The data warehouse is used to: complete integrated stream and batch computing processing based on the data warehouse model according to the generated data warehouse computing tasks and scheduling tasks; The data output module is used to: complete data output processing based on the business output model according to the generated data output task; The business output model includes data warehouse acquisition target configuration, business output data model, and output model mapping relationship; the data warehouse acquisition target configuration includes message queue stream data and distributed storage batch data warehouse information; the business output data model is used to construct a business wide table model, and set model index, partition information, and cleanup mechanism for the model table.

2. The model-visualized big data batch generation management system as described in claim 1, characterized in that, The data model metadata information module completes the physicalization of the data warehouse model, including generating specific physical tables or message queue topics based on the data warehouse model metadata.

3. The model-visualized big data batch generation management system as described in claim 1, characterized in that, The data model task module completes the task-related configuration, including: Scheduling flow configuration related to scheduling tasks: Visually build the data warehouse computing scheduling link by dragging and dropping table model nodes; Data warehouse computing task related computing logic development: If it is a batch computing table, then develop batch task SQL for node information; if it is a stream computing table, then develop stream task SQL for node information; if it is a stream-batch table, then both batch task SQL and stream task SQL for node information are required. Data link configuration related to data acquisition and data output tasks: Set the acquisition target source, exported data model, and acquisition model mapping relationship in the data acquisition model; set the data warehouse acquisition target configuration, business output data model, and output model mapping relationship in the business output model.

4. The model-visualized big data batch generation management system as described in claim 1, characterized in that, The data acquisition module, according to the generated data acquisition task, completes data acquisition processing based on the data acquisition model, including: The data acquisition module collects the required computational data according to the generated data acquisition task. During the acquisition, the target data is transformed into ODS layer structured data according to the acquisition model mapping relationship in the set data acquisition model, and the transformed data is imported into the message queue, which serves as the sole entry point for the data warehouse.

5. The model-visualized big data batch generation management system as described in claim 1, characterized in that: The data warehouse is divided into a real-time data warehouse and an offline data warehouse; The real-time data warehouse relies on a distributed message queue, uses a task scheduler for scheduling, and uses a real-time computing engine to complete the processing of streaming computing tasks. The offline data warehouse relies on a distributed storage database, uses a task scheduler for scheduling, and uses an offline computing engine to complete batch computing tasks.

6. The model-visualized big data batch generation management system as described in claim 1, characterized in that, The data output module completes data output processing based on the business output model according to the generated data output task, including: The data output module, according to the generated data output task and the data warehouse collection target configuration in the business output model, obtains the required data warehouse calculation results from the data warehouse; creates a business wide table according to the business wide table model in the business output model; imports the business fields of the obtained data warehouse calculation results into the business wide table according to the output model mapping relationship in the business output model; and sets up indexing, partitioning, and cleanup mechanisms for the business wide table according to the configuration.

7. The model-visualized big data batch generation management system as described in claim 1, characterized in that: The system also includes a visualization database; the visualization database is used to store the output data processed by the data output module.

8. The model-visualized big data batch generation management system as described in claim 7, characterized in that: The system also includes a front-end low-code business display module located in the presentation layer and a back-end OLAP analysis and processing module located in the service layer; The front-end low-code business display module is used to: respond to client scheduling, call the back-end OLAP analysis and processing module to perform OLAP low-code analysis configuration based on the business output model; and configure visualization components and bind query result datasets through a graphical interface. The backend OLAP analysis and processing module is used to: call the query engine to parse the OLAP low-code analysis configuration, call the visualization database to perform multidimensional analysis, and render the analysis results to the client interface in real time.

9. A model-based visualization-based big data batch generation and management method using the system described in any one of claims 1 to 8, characterized in that, The method includes the following steps: The data model metadata information module responds to the client's scheduling and completes the visual creation of the data acquisition model, data warehouse model, and business output model; The data model task module responds to the client's scheduling and completes the task-related configuration; After executing the data warehouse model deployment operation, the data model task module triggers an automated process; the automated process includes: The data model metadata information module automatically completes the physicalization of the data warehouse model; Based on the task-related configuration, the data model task module automatically generates corresponding data acquisition tasks, data warehouse computing tasks, scheduling tasks, and data output tasks. The data acquisition module completes data acquisition processing based on the data acquisition model according to the generated data acquisition tasks; the data warehouse completes integrated stream and batch computing processing based on the data warehouse model according to the generated data warehouse computing tasks and scheduling tasks; and the data output module completes data output processing based on the business output model according to the generated data output tasks.

Citation Information

Patent Citations

  • Data warehouse visual modeling system and method based on power grid big data

    CN112579563A

  • Decoupling elastic data warehouse architecture

    WO2020220717A1