Real-time Data Analysis and Processing Method, System and Storage Medium Based on Flink and StarRocks
Through the combination of Flink and StarRocks, real-time correlation analysis of multi-source data in the financial industry is realized, and the problems of high data processing complexity and insufficient real-time performance are solved, technology and operation and maintenance complexity are reduced, and data analysis efficiency is improved.
Patent Information
- Application Number
- CN202310762477.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-25
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2043-06-25
AI Technical Summary
In the prior art, the coexistence of multiple databases leads to difficulties in data processing and processing and maintenance in the financial industry, high learning costs and high program complexity, and the correlation analysis of multiple data sources cannot be realized. Real-time financial data analysis time is too long, data writing delay is severe, and the real-time requirements of business functions cannot be met.
The combination of Flink and StarRocks is used to stream analysis and aggregate the original data of the business through Flink, and the correlation aggregation of flow table data and dimension table data is used to store the results in StarRocks, and visual analysis and display it in combination with front-end query requirements.
It realizes simple, efficient and accurate real-time correlation analysis of multi-source business data, reduces the entry threshold for real-time data development, reduces maintenance and server costs, and improves data collection timeliness and analysis and query capabilities.
Smart Images

Figure CN116775768B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data warehouses, and particularly relates to a real-time data analysis and processing method, system and storage medium based on Flink and StarRocks. Background Art
[0002] In the current financial environment, the importance attached to corresponding business data is increasing, and real-time technologies are also evolving continuously. In terms of timeliness, hour-level or minute-level calculations can no longer meet the evolving needs of customer businesses. It has gradually been upgraded from time-window-driven to event-driven. Even for every new piece of data generated, people want to see the data as soon as possible. The data warehouse processing process has changed from an offline or micro-batch process to a real-time streaming process that Flink is good at. In terms of data sources, the overall data representation of a single data source is poor. Currently, people not only hope to analyze and calculate a single data stream but also hope to perform multi-stream calculations by combining multiple data sources. Therefore, every means is taken to make the data representation more abundant. Due to the accumulation of historical data, the continuous expansion of business scale, and the growing demand for data analysis, higher requirements are put forward for the real-time risk analysis engine. Traditional financial industries all have their own relational databases, but they are all applied to core businesses. Therefore, large data clusters are deployed, and big data solutions are used to achieve their data analysis.
[0003] Storing and analyzing data of different financial business scenarios in a way of coexistence of multiple databases makes data processing, processing and maintenance difficult; due to different data processing engines, operation and maintenance personnel and technical personnel need to master the functional codes of multiple technologies to implement, and the learning cost is high; if multi-data sources want to realize the correlation analysis of data, the complexity of the program is extremely high, and even the correlation cannot be realized; moreover, due to the accumulation of historical data and the continuous expansion of business scale, the existing real-time financial data analysis and processing time is too long, unable to meet the real-time requirement that the business function needs to have, the data writing is delayed, and the data backlog is serious; at the same time, most logics use the stored procedure method for logic processing, and the functions are complex. Therefore, some business data needs the participation of original business personnel during the transformation process and the migration to the stream processing process, and the processing efficiency is very low. Summary of the Invention
[0004] The purpose of the present invention is to provide a real-time data analysis and processing method, system and storage medium based on Flink and StarRocks to solve the above problems existing in the prior art.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions:
[0006] In the first aspect, a real-time data analysis and processing method based on Flink and StarRocks is provided, including:
[0007] Collect the original business data upstream and store the collected original business data in the Kafka message queue;
[0008] Use Flink to perform streaming data parsing on the original business data in the message queue to obtain real-time detailed data, and store the real-time detailed data in the Kafka message queue;
[0009] Obtain the real-time tasks of the streaming task platform, where the real-time tasks include several first Flink SQL statements;
[0010] Create a first detailed table, a first temporary table, a first dimension table, and a first result table in the StarRocks database according to the real-time tasks, write the real-time detailed data in the Kafka message queue into the first detailed table in real time, and summarize the real-time detailed data in the first detailed table into the first temporary table according to the set classification metrics in real time to obtain the summarized first temporary table. The first dimension table contains the set first dimension data;
[0011] Associate and aggregate the summarized first temporary table with the first dimension table according to the real-time tasks to obtain the first result data after association and aggregation;
[0012] Write the first result data after association and aggregation into the first result table according to the real-time tasks;
[0013] Obtain the first query instruction from the front end, and perform visual analysis and display on the execution status of the real-time tasks and the first result data written in the first result table according to the first query instruction.
[0014] In a possible design, the method further includes:
[0015] Obtain the batch microsynchronization tasks of the batch task platform, where the batch microsynchronization tasks include several second Flink SQL statements;
[0016] Create a second detailed table, a second temporary table, a second dimension table, and a second result table in the StarRocks database according to the batch microsynchronization tasks, batch synchronize and write the real-time detailed data in the Kafka message queue into the second detailed table at set time intervals or set data volumes, and summarize the real-time detailed data in the second detailed table into the second temporary table according to the set classification metrics to obtain the summarized second temporary table. The second dimension table contains the set second dimension data;
[0017] Associate and aggregate the summarized second temporary table with the second dimension table according to the batch microsynchronization tasks to obtain the second result data after association and aggregation;
[0018] Write the aggregated second result data into the second result table according to the batch microsynchronization tasks;
[0019] Obtain the second query instruction of the front end, and perform visual analysis and display on the execution status of the batch micro-synchronization task and the second result data written in the second result table according to the second query instruction.
[0020] In a possible design, the business raw data includes external business raw data obtained through a third-party interface, internal business raw data stored in a local database, and crawler business raw data crawled on a corresponding website through a crawler.
[0021] In a possible design, using Flink to perform streaming data parsing on the business raw data in the message queue to obtain real-time detail data, including: using Flink to perform formatted parsing on the real-time business raw data in the message queue to obtain data in a set format, and extracting real-time detail data from the data in the set format.
[0022] In a possible design, associating and aggregating the summarized first temporary table with the first dimension table according to the real-time task to obtain the first result data after association and aggregation, including: cleaning the corresponding data in the first temporary table with reference to the first dimension data in the first dimension table according to the real-time task, and performing a join association on the corresponding data in the first temporary table and the first dimension data in the first dimension table to obtain the first result data after association and aggregation.
[0023] In a possible design, associating and aggregating the summarized second temporary table with the second dimension table according to the batch micro-synchronization task to obtain the second result data after association and aggregation, including: cleaning the corresponding data in the second temporary table with reference to the second dimension data in the second dimension table according to the batch micro-synchronization task, and performing a join association on the corresponding data in the second temporary table and the second dimension data in the second dimension table to obtain the second result data after association and aggregation.
[0024] In a second aspect, a real-time data analysis and processing system based on Flink and StarRocks is provided, including a collection unit, an analysis unit, an acquisition unit, a construction unit, an aggregation unit, an input unit, and a query unit, where:
[0025] The collection unit is used to collect the upstream business raw data and store the collected business raw data into the Kafka message queue;
[0026] The analysis unit is used to perform streaming data analysis on the business raw data in the message queue using Flink to obtain real-time detail data, and store the real-time detail data into the Kafka message queue;
[0027] An acquisition unit for acquiring real-time tasks of a streaming task platform, where the real-time tasks include a number of first Flink SQL statements;
[0028] A construction unit for creating a first detail table, a first temporary table, a first dimension table, and a first result table in a StarRocks database according to the real-time tasks, writing the real-time detail data in the Kafka message queue into the first detail table in real time, and summarizing the real-time detail data in the first detail table into the first temporary table according to the set classification metrics in real time to obtain a summarized first temporary table, where the first dimension table contains the set first dimension data;
[0029] An aggregation unit for associatively aggregating the summarized first temporary table and the first dimension table according to the real-time tasks to obtain first result data after associative aggregation;
[0030] An input unit for writing the first result data after associative aggregation into the first result table according to the real-time tasks;
[0031] A query unit for obtaining a first query instruction from the front end and performing visual analysis and display on the execution status of the real-time tasks and the first result data written in the first result table according to the first query instruction.
[0032] In a possible design, the acquisition unit is further configured to acquire batch microsynchronization tasks of a batch task platform, where the batch microsynchronization tasks include a number of second Flink SQL statements;
[0033] The construction unit is further configured to create a second detail table, a second temporary table, a second dimension table, and a second result table in a StarRocks database according to the batch microsynchronization tasks, batch synchronize the real-time detail data in the Kafka message queue into the second detail table at a set time interval or according to a set data volume, and summarize the real-time detail data in the second detail table into the second temporary table according to the set classification metrics to obtain a summarized second temporary table, where the second dimension table contains the set second dimension data;
[0034] The aggregation unit is further configured to associatively aggregate the summarized second temporary table and the second dimension table according to the batch microsynchronization tasks to obtain second result data after associative aggregation;
[0035] The input unit is further configured to write the aggregated second result data into the second result table according to the batch microsynchronization tasks;
[0036] The query unit is further configured to obtain a second query instruction from the front end and perform visual analysis and display on the execution status of the batch microsynchronization tasks and the second result data written in the second result table according to the second query instruction.
[0037] In a third aspect, a real-time data analysis and processing system based on Flink and StarRocks is provided, including:
[0038] A memory for storing instructions;
[0039] A processor for reading the instructions stored in the memory and executing any one of the above-mentioned real-time data analysis and processing methods based on Flink and StarRocks in the first aspect according to the instructions.
[0040] In a fourth aspect, a computer-readable storage medium is provided. Instructions are stored on the computer-readable storage medium. When the instructions run on a computer, the computer is caused to execute any one of the methods in the first aspect. At the same time, a computer program product containing instructions is also provided. When the instructions run on a computer, the computer is caused to execute any one of the real-time data analysis and processing methods based on Flink and StarRocks in the first aspect.
[0041] Advantageous effects: In the present invention, Flink is used to parse and process streaming service data, and based on Flink SQL statements, the aggregation and association of stream table data and corresponding dimension table data are performed. Then, the aggregated data is stored in StarRocks. Finally, according to the query requirements of the front end, the corresponding task processing process and the data in the StarRocks library are visually analyzed and displayed, so as to realize simpler, more efficient, and more accurate real-time correlation analysis and processing of multi-source service data. The present invention performs real-time data analysis and processing based on the open-source Flink and StarRocks, which can effectively reduce the entry threshold of real-time data development and help the rapid development and iteration of real-time services; by using external tables and message middleware as the method of offline and real-time data synchronization, the timeliness and accuracy of data collection can be guaranteed; through the storage characteristics and physical architecture of StarRocks, rapid and stable analysis and query capabilities are provided to adapt to various business analysis scenarios; the present invention only needs one big data engine to support all business data analysis, without the need to introduce other analysis engines, which can effectively reduce the maintenance cost, server cost, as well as the technical and operation complexity. Description of the Drawings
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0043] Figure 1 It is a schematic diagram of the steps of the method in Embodiment 1 of the present invention;
[0044] Figure 2 It is a schematic diagram of the system in Embodiment 2 of the present invention;
[0045] Figure 3 It is a schematic diagram of the system in Embodiment 3 of the present invention. Specific implementation manners
[0046] It should be noted here that the description of these embodiments is used to help understand the present invention, but does not constitute a limitation to the present invention. The specific structural and functional details disclosed herein are only used to describe the exemplary embodiments of the present invention. However, the present invention can be embodied in many alternative forms and should not be construed as limited to the embodiments set forth herein.
[0047] It should be understood that unless otherwise clearly defined and limited, the term "connection" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meaning of the above terms in the embodiments can be understood according to specific situations.
[0048] Specific details are provided in the following description to facilitate a complete understanding of the exemplary embodiments. However, those of ordinary skill in the art should understand that the exemplary embodiments can be implemented without these specific details. For example, the system can be shown in a block diagram to avoid obscuring the example with unnecessary details. In other embodiments, well-known processes, structures, and technologies may not be shown with unnecessary details to avoid obscuring the embodiments.
[0049] Embodiment 1:
[0050] This embodiment provides a real-time data analysis and processing method based on Flink and StarRocks, which can be applied to a corresponding data server, such as Figure 1 As shown, the method includes the following steps:
[0051] S1. Collect the upstream business raw data and store the collected business raw data in the Kafka message queue.
[0052] Specifically, the business raw data includes external business raw data, internal business raw data, and crawler business raw data. Specifically, the external business raw data can be obtained through a third-party interface, the internal business raw data stored in the local database can be retrieved, and the crawler business raw data crawled from the corresponding website through a crawler. The collected business raw data is stored in the Kafka message queue in real time.
[0053] S2. Use Flink to perform streaming data parsing on the business raw data in the message queue, obtain real-time detail data, and store the real-time detail data in the Kafka message queue.
[0054] Specifically, Flink (a distributed streaming data processing engine) can be used to perform streaming data parsing on the business raw data in the message queue. The business raw data has different data formats, such as common JSON nested data, which needs to be extracted and processed according to requirements. Format the parsing of the business raw data to obtain data in a set format. For example, parse and process JSON nested data into row-based data, and extract real-time detail data from the data in the set format.
[0055] S3. Obtain the real-time tasks of the streaming task platform, where the real-time tasks include several first Flink SQL statements.
[0056] Specifically, real-time tasks can be pre-written through the streaming task platform, and then the real-time tasks written by the streaming task platform can be obtained for subsequent real-time data processing. The real-time tasks include several first Flink SQL statements. At the same time, micro-batch synchronization tasks can also be pre-written through the batch task platform, and then the micro-batch synchronization tasks written by the streaming task platform can be obtained for subsequent batch data processing. The micro-batch synchronization tasks include several second Flink SQL statements
[0057] S4. Create a first detail table, a first temporary table, a first dimension table, and a first result table in the StarRocks database according to the real-time tasks. Write the real-time detail data in the Kafka message queue into the first detail table in real time, and summarize the real-time detail data in the first detail table into the first temporary table in real time according to the set classification metrics to obtain the summarized first temporary table. The first dimension table contains the set first dimension data.
[0058] Specifically, after obtaining the real-time tasks, the corresponding first Flink SQL statements in the real-time tasks can be executed to create a first detail table, a first temporary table, a first dimension table, and a first result table in the StarRocks database. Write the real-time detail data in the Kafka message queue into the first detail table in real time, and summarize the real-time detail data in the first detail table into the first temporary table in real time according to the set classification metrics (such as the number of orders, order amount, etc.) to obtain the summarized first temporary table. The first dimension table contains the set first dimension data, and the first dimension data can be sourced from a local database or other externally connected databases.
[0059] Similarly, after obtaining the micro-batch synchronization task, the corresponding second Flink SQL statement in the micro-batch synchronization task can be executed to create a second detail table, a second temporary table, a second dimension table, and a second result table in the StarRocks database. The real-time detail data in the Kafka message queue is batch synchronized and written into the second detail table at a set time interval (such as the data in the queue within 1 minute or 2 minutes) or a set data volume (such as 100 or 200 pieces of data in the queue). The real-time detail data in the second detail table is summarized into the second temporary table according to the set classification indicators to obtain the summarized second temporary table. The second dimension table contains the set second dimension data, and the second dimension data can be sourced from a local database or other externally connected databases.
[0060] S5. According to the real-time task, the summarized first temporary table is associated and aggregated with the first dimension table to obtain the first result data after association and aggregation.
[0061] Specifically in implementation, the corresponding first Flink SQL statement of the real-time task is executed to clean the corresponding data in the first temporary table with reference to the first dimension data in the first dimension table. For example, in the first temporary table data, the gender data is sometimes'male', 'female' and sometimes'man', 'woman', but in the first dimension table, both'male', 'female' and'man', 'woman' are mapped to 1 and 0. The association cleaning is to standardize both to 1 and 0. At the same time, it may also involve various dimension specifications such as arithmetic calculations and exchange rate conversions, as well as performing a join association between the corresponding data in the first temporary table and the first dimension data in the first dimension table to complete the corresponding field information of the first dimension data in the first temporary table data, that is, performing a dimension table join operation. For example, if the first temporary table data does not have area information while the area information is in the first dimension table data, then a join association with the first dimension table data is required to complete the area field information. Finally, the first result data after association and aggregation is obtained.
[0062] Similarly, for the batch processing method, the corresponding second Flink SQL statement of the batch micro-synchronization task can be executed to clean the corresponding data in the second temporary table with reference to the second dimension data in the second dimension table, and perform a join association between the corresponding data in the second temporary table and the second dimension data in the second dimension table to obtain the second result data after association and aggregation.
[0063] S6. According to the real-time task, the first result data after association and aggregation is written into the first result table.
[0064] During specific implementation, after aggregating the first result data, execute the first Flink SQL statement corresponding to the real-time task to write the aggregated first result data into the first result table. Similarly, for the batch processing method, after aggregating the second result data, execute the second Flink SQL statement corresponding to the batch microsynchronization task to write the aggregated second result data into the second result table. The StarRocks database can build different data layers according to requirements, such as the ODS data layer, DWD data layer, APP data layer, etc., to store the analyzed result data in a hierarchical and classified manner, so as to have a clearer control over the result data and facilitate data tracking. For example, data closer to the original business data will be placed in the ODS layer, and data after aggregation and processing will be placed in the DWD layer, and data that needs to provide query applications externally will be placed in the APP layer.
[0065] S7. Obtain the first query instruction from the front end, and perform visual analysis and display on the execution status of the real-time task and the first result data written in the first result table according to the first query instruction.
[0066] During specific implementation, when the front end needs to query the real-time data processing situation, it can send the corresponding first query instruction to the server, and the server performs visual analysis and display on the execution status of the real-time task and the first result data written in the first result table according to the first query instruction. When the front end needs to query the batch data processing situation, it can send the corresponding second query instruction to the server, and the server performs visual analysis and display on the execution status of the batch microsynchronization task and the second result data written in the second result table according to the second query instruction.
[0067] The method of this embodiment is based on open-source Flink and StarRocks for real-time data analysis and processing, which can effectively reduce the entry threshold of real-time data development and help the rapid development and iteration of real-time services; by using external tables and message middleware as the method of offline and real-time data synchronization, the timeliness and accuracy of data collection can be guaranteed; through the storage characteristics and physical architecture of StarRocks, a fast and stable analysis and query ability is provided to adapt to various business analysis scenarios; the present invention only needs one big data engine to support all business data analysis, without the need to introduce other analysis engines, which can effectively reduce the maintenance cost, server cost, as well as the technical and operation complexity.
[0068] Embodiment 2:
[0069] This embodiment provides a real-time data analysis and processing system based on Flink and StarRocks, as Figure 2 shown, including a collection unit, a parsing unit, an obtaining unit, a construction unit, an aggregation unit, an input unit, and a query unit, where:
[0070] The collection unit is used to collect the original business data upstream and store the collected original business data into the Kafka message queue;
[0071] The parsing unit is used to perform streaming data parsing on the original business data in the message queue using Flink to obtain real-time detailed data, and store the real-time detailed data into the Kafka message queue;
[0072] The acquisition unit is used to obtain the real-time tasks of the stream task platform, and the real-time tasks include several first Flink SQL statements;
[0073] The construction unit is used to create a first detailed table, a first temporary table, a first dimension table, and a first result table in the StarRocks database according to the real-time task, write the real-time detailed data in the Kafka message queue into the first detailed table in real time, and summarize the real-time detailed data in the first detailed table into the first temporary table according to the set classification indicators in real time to obtain the summarized first temporary table. The first dimension table contains the set first dimension data;
[0074] The aggregation unit is used to perform associative aggregation on the summarized first temporary table and the first dimension table according to the real-time task to obtain the first result data after associative aggregation;
[0075] The input unit is used to write the first result data after associative aggregation into the first result table according to the real-time task;
[0076] The query unit is used to obtain the first query instruction from the front end, and perform visual analysis and display on the execution status of the real-time task and the first result data written in the first result table according to the first query instruction.
[0077] Further, the acquisition unit is also used to obtain the batch micro-synchronization task of the batch task platform, and the batch micro-synchronization task includes several second Flink SQL statements;
[0078] The construction unit is also used to create a second detailed table, a second temporary table, a second dimension table, and a second result table in the StarRocks database according to the batch micro-synchronization task, batch synchronize and write the real-time detailed data in the Kafka message queue into the second detailed table at a set time interval or set data volume, and summarize the real-time detailed data in the second detailed table into the second temporary table according to the set classification indicators to obtain the summarized second temporary table. The second dimension table contains the set second dimension data;
[0079] The aggregation unit is also used to perform associative aggregation on the summarized second temporary table and the second dimension table according to the batch micro-synchronization task to obtain the second result data after associative aggregation;
[0080] The input unit is further configured to write the aggregated second result data into the second result table according to the batch micro-synchronization task;
[0081] The query unit is further configured to obtain a second query instruction from the front end, and perform visual analysis and display on the execution status of the batch micro-synchronization task and the second result data written in the second result table according to the second query instruction.
[0082] Embodiment 3:
[0083] This embodiment provides a real-time data analysis and processing system based on Flink and StarRocks, as Figure 3 shown. At the hardware level, it includes:
[0084] A data interface for establishing data docking between the processor and an external data terminal;
[0085] A memory for storing instructions;
[0086] A processor for reading the instructions stored in the memory and executing the real-time data analysis and processing method based on Flink and StarRocks in Embodiment 1 according to the instructions.
[0087] Optionally, the device further includes an internal bus. The processor, the memory, and the data interface can be interconnected through the internal bus, and the internal bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc.
[0088] The memory may but is not limited to include Random Access Memory (RAM), Read Only Memory (ROM), Flash Memory, First Input First Output (FIFO), and / or First In Last Out (FILO), etc. The processor may be a general-purpose processor, including Central Processing Unit (CPU), Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), Application Specific Integrated Circuit (ASIC), Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0089] Embodiment 4:
[0090] This embodiment provides a computer-readable storage medium, on which instructions are stored. When the instructions run on a computer, the computer is caused to execute the real-time data analysis and processing method based on Flink and StarRocks in Embodiment 1. Among them, the computer-readable storage medium refers to a carrier for storing data, which may but is not limited to include floppy disks, optical discs, hard disks, flash memories, USB flash drives, and / or Memory Sticks, etc. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable systems.
[0091] This embodiment also provides a computer program product containing instructions. When the instructions run on a computer, the computer is caused to execute the real-time data analysis and processing method based on Flink and StarRocks in Embodiment 1. Among them, the computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable systems.
[0092] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A real-time data analysis and processing method based on Flink and StarRocks, characterized in that, Including: Collect the original business data upstream, and store the collected original business data into the Kafka message queue; Use Flink to perform streaming data parsing on the original business data in the message queue to obtain real-time detailed data, and store the real-time detailed data into the Kafka message queue; Obtain the real-time tasks of the stream task platform, where the real-time tasks include several first Flink SQL statements; Create a first detail table, a first temporary table, a first dimension table, and a first result table in the StarRocks database according to the real-time tasks, write the real-time detailed data in the Kafka message queue into the first detail table in real time, and summarize the real-time detailed data in the first detail table into the first temporary table according to the set classification metrics in real time to obtain the summarized first temporary table. The first dimension table contains the set first dimension data; Associate and aggregate the summarized first temporary table with the first dimension table according to the real-time tasks to obtain the first result data after association and aggregation; Write the first result data after association and aggregation into the first result table according to the real-time tasks; Obtain the first query instruction from the front end, and perform visual analysis and display on the execution status of the real-time tasks and the first result data written in the first result table according to the first query instruction.
2. The real-time data analysis and processing method based on Flink and StarRocks according to claim 1, wherein The method further includes: Obtain the batch micro-synchronization tasks of the batch task platform, where the batch micro-synchronization tasks include several second Flink SQL statements; Create a second detail table, a second temporary table, a second dimension table, and a second result table in the StarRocks database according to the batch micro-synchronization tasks, batch synchronize and write the real-time detailed data in the Kafka message queue into the second detail table at set time intervals or set data volumes, and summarize the real-time detailed data in the second detail table into the second temporary table according to the set classification metrics to obtain the summarized second temporary table. The second dimension table contains the set second dimension data; Associate and aggregate the summarized second temporary table with the second dimension table according to the batch micro-synchronization tasks to obtain the second result data after association and aggregation; Write the second result data after aggregation into the second result table according to the batch micro-synchronization tasks; Obtain the second query instruction from the front end, and perform visual analysis and display on the execution status of the batch micro-synchronization tasks and the second result data written in the second result table according to the second query instruction.
3. The real-time data analysis and processing method based on Flink and StarRocks according to claim 1, wherein The original business data includes external original business data obtained through third-party interfaces, internal original business data stored in local databases, and crawler original business data crawled on corresponding websites through crawlers.
4. The real-time data analysis and processing method based on Flink and StarRocks according to claim 1, characterized in that The step of using Flink to perform streaming data parsing on the original business data in the message queue to obtain real-time detailed data includes: using Flink to perform formatted parsing on the real-time original business data in the message queue to obtain data in a set format, and extracting real-time detailed data from the data in the set format.
5. The real-time data analysis and processing method based on Flink and StarRocks according to claim 1, wherein Associating and aggregating the summarized first temporary table with the first dimension table according to the real-time task to obtain the first result data after association and aggregation, including: cleaning the corresponding data in the first temporary table with reference to the first dimension data in the first dimension table according to the real-time task, and performing a join association between the corresponding data in the first temporary table and the first dimension data in the first dimension table to obtain the first result data after association and aggregation.
6. The real-time data analysis and processing method based on Flink and StarRocks according to claim 2, wherein, Associating and aggregating the summarized second temporary table with the second dimension table according to the batch micro-synchronization task to obtain the second result data after association and aggregation, including: cleaning the corresponding data in the second temporary table with reference to the second dimension data in the second dimension table according to the batch micro-synchronization task, and performing a join association between the corresponding data in the second temporary table and the second dimension data in the second dimension table to obtain the second result data after association and aggregation.
7. A real-time data analysis and processing system based on Flink and StarRocks, characterized in that Including a collection unit, a parsing unit, an acquisition unit, a construction unit, an aggregation unit, an input unit, and a query unit, where: The collection unit is used to collect the upstream business raw data and store the collected business raw data into the Kafka message queue; The parsing unit is used to perform streaming data parsing on the business raw data in the message queue using Flink to obtain real-time detailed data, and store the real-time detailed data into the Kafka message queue; The acquisition unit is used to obtain the real-time tasks of the stream task platform, and the real-time tasks include several first Flink SQL statements; The construction unit is used to create a first detailed table, a first temporary table, a first dimension table, and a first result table in the StarRocks database according to the real-time task, write the real-time detailed data in the Kafka message queue into the first detailed table in real time, and summarize the real-time detailed data in the first detailed table into the first temporary table in real time according to the set classification indicators to obtain the summarized first temporary table, and the first dimension table contains the set first dimension data; The aggregation unit is used to associate and aggregate the summarized first temporary table with the first dimension table according to the real-time task to obtain the first result data after association and aggregation; The input unit is used to write the first result data after association and aggregation into the first result table according to the real-time task; The query unit is used to obtain the first query instruction of the front end, and perform visual analysis and display on the execution situation of the real-time task and the first result data written in the first result table according to the first query instruction.
8. The real-time data analysis and processing system based on Flink and StarRocks according to claim 7, characterized in that The acquisition unit is further used to obtain the batch micro-synchronization tasks of the batch task platform, and the batch micro-synchronization tasks include several second Flink SQL statements; The building unit is further configured to create a second detail table, a second temporary table, a second dimension table, and a second result table in the StarRocks database according to the batch micro-synchronization task, batch-synchronize and write the real-time detail data in the Kafka message queue into the second detail table at a set time interval or a set data volume, and summarize the real-time detail data in the second detail table into the second temporary table according to the set classification metrics to obtain the summarized second temporary table, wherein the second dimension table contains the set second dimension data; The aggregation unit is further configured to perform association aggregation on the summarized second temporary table and the second dimension table according to the batch micro-synchronization task to obtain the second result data after association aggregation; The input unit is further configured to write the aggregated second result data into the second result table according to the batch micro-synchronization task; The query unit is further configured to obtain a second query instruction from the front end, and perform visual analysis and display on the execution status of the batch micro-synchronization task and the second result data written in the second result table according to the second query instruction.
9. A real-time data analysis and processing system based on Flink and StarRocks, characterized in that, Including: A memory for storing instructions; A processor for reading the instructions stored in the memory and executing the real-time data analysis and processing method based on Flink and StarRocks according to any one of claims 1-6.
10. A computer-readable storage medium, characterized in that, Instructions are stored on the computer-readable storage medium, and when the instructions are run on the computer, the computer is caused to execute the real-time data analysis and processing method based on Flink and StarRocks according to any one of claims 1-6.
Citation Information
Patent Citations
Method for constructing real-time data warehouse system based on FlinkDores
CN115033646A
Data real-time analysis method and device, equipment and storage medium
CN116069600A