Stream computing data processing system constructed based on kappa architecture

By adopting the kappa architecture in the streaming computing data processing system, combining kafka+flink components and distributed memory technology, the existing streaming computing system's shortcomings in data cleaning, timeliness and streaming data processing capabilities are solved, and more efficient business operations and risk monitoring are achieved.

CN119988445APending Publication Date: 2025-05-13SHANGHAI RURAL COMML BANK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411992829.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing stream computing system based on the lambda architecture has layer-by-layer cleaning and processing, poor application timeliness, and lack of streaming data processing capabilities, resulting in insufficient utilization of operational data, buried point data and other event data.

Method used

The stream computing data processing system based on the kappa architecture is adopted. Through the cooperation of event acquisition unit, event processing unit, event storage unit and event computing unit, the event acquisition and cleaning of events are completed using the kafka+flink component to perform streaming big data calculations, and combined with distributed memory technology and cluster mechanisms, real-time data processing and storage are realized.

Benefits of technology

The value of streaming data has been fully explored, business operation efficiency and risk monitoring timeliness, streaming data and batch data have been integrated, batch operation efficiency of the data platform has been improved, and streaming data processing gaps have been filled.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988445A_ABST
    Figure CN119988445A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of stream computing data processing, and particularly discloses a stream computing data processing system based on kappa architecture construction, which comprises an event acquisition unit, an event processing unit, an event storage unit, an event computing unit and a management platform, through cooperation of the event collection unit, the event processing unit, the event storage unit and the event calculation unit, collection and cleaning of events are completed through a kafka + flink component, real-time calculation of the events is completed through a streaming big data calculation engine, real-time index storage is performed through a distributed memory technology, and real-time calculation of the events is completed. The use of components is ensured through a distributed cluster mechanism, and real-time calculation data is timely returned to a batch library, so that a Kappa architecture can be applied to data platform establishment, the streaming data value is fully mined, the business operation efficiency and the risk monitoring timeliness are effectively improved, and by integrating the streaming data and the batch data, the risk monitoring efficiency is improved. And the batch processing operation efficiency of the data platform is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of stream computing data processing, and specifically relates to a stream computing data processing system built based on a Kappa architecture. Background Art

[0002] With the popularization of the Internet and mobile devices, the amount of data generated by banking business has increased dramatically. This data needs to be processed and analyzed in real time to support quick decision-making. Stream computing technology can process data streams instantly to meet the banking industry's real-time needs; at the same time, stream computing can help monitor and analyze transaction behaviors in real time, promptly detect and prevent fraud, and strengthen risk control. Through stream computing technology, banks can analyze customer behaviors and preferences in real time, provide personalized services and product recommendations, thereby improving customer satisfaction and loyalty, and optimizing the bank's back-end operating processes, such as automated report generation, real-time monitoring of system performance, etc., to improve overall operational efficiency.

[0003] The currently commonly used stream computing method is based on the lambda architecture. The lambda architecture is generally connected by a batch processing layer, a speed layer and a service layer. The batch processing layer is inherently delayed; the speed layer is responsible for real-time data processing, inputs data in near real time and performs incremental updates, and combines the near real-time update results with the batch processing layer to provide a unified data view; the service layer serves as an access point for accessing services and provides a unified data view.

[0004] However, the bank data platform uses the lambda architecture to separate batch data processing from real-time processing. Although its technology is relatively mature and has a high fault tolerance rate, it requires layer-by-layer cleaning and processing of data, and its application timeliness is poor. At the same time, the platform lacks the ability to process streaming data, and does not make enough use of event data such as operational data and buried data. Therefore, we need to propose a stream computing data processing system built on the kappa architecture to solve the above problems, so that it can apply the kappa architecture to the construction of the data platform, fully tap the value of streaming data, and effectively improve business operation efficiency and risk monitoring timeliness. Summary of the invention

[0005] The purpose of the present invention is to provide a stream computing data processing system built based on the Kappa architecture, which can apply the Kappa architecture to the construction of a data platform, fully tap the value of streaming data, effectively improve business operation efficiency and risk monitoring timeliness, so as to solve the problems raised in the above background technology.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] The stream computing data processing system built based on the Kappa architecture includes: an event collection unit, which completes event collection and cleaning through the Kafka+Flink component;

[0008] An event processing unit, wherein the event processing unit completes real-time calculation of events through a streaming big data computing engine;

[0009] An event storage unit, wherein the event storage unit stores real-time indicators through distributed memory technology;

[0010] An event computing unit, which ensures the use of components through a distributed cluster mechanism and returns real-time computing data to a batch library in a timely manner;

[0011] A management platform is used for operating the event collection unit, the event processing unit, the event storage unit and the event calculation unit on event processing.

[0012] Preferably, the content collected by the kafka+flink component includes file collection, database collection, message queue collection and website interface collection, wherein the database includes a source-side DB2 relational database, an Oracle relational database and a MySQL relational database, and real-time data is pushed to the management platform by reading binlog logs in the source database.

[0013] Preferably, the kafka+flink component includes a java or python API, and during event collection, the source system pushes real-time data to the kafka+flink component through the CDC for Kafka form, the kafka API form, and the kafka synchronization form.

[0014] Preferably, the content processed by the event processing unit includes joining multiple real-time tables, joining real-time tables with offline tables, splicing multiple data streams, custom function processing, data field mapping, abnormal value replacement, string interception and data sorting. The management platform configures the event processing unit's event processing flow in the form of a DAG workflow. Multiple data sources can be dragged and dropped in the processing flow to perform operations such as joining multiple tables, asynchronous search, splicing, and custom function processing. The processed data is transmitted to the supported data source.

[0015] Preferably, the event storage unit includes a temporary storage module and a persistent storage module, the temporary storage module is used to store the memory database and the message middleware, and the persistent storage module is used to store files and relational databases.

[0016] Preferably, the distributed cluster mechanism ensures that the components include an indicator calculation unit, an application management unit, an indicator playback unit and a dynamic reporting unit. The indicator calculation unit contains a variety of calculation operators, including counting, summing, averaging, maximum, minimum, continuous, increasing, decreasing, sorting, fluctuation, skewness and kurtosis; the application management unit has the ability to obtain context and supports statistics in combination with the context. The dynamic reporting unit is used to store the calculated indicator support indicators in a memory database, and the indicator playback unit provides external services through SAS software or Tableau software.

[0017] Preferably, the event calculation unit and the event storage unit constitute a real-time indicator mart, which obtains business system data in real time through the CDC for Kafka service, writes flink jobs according to business caliber to calculate real-time indicators, and stores them in a relational database, which is open to real-time large-screen query through a data service API interface; the management platform, event unit and event processing unit constitute a real-time event center, which uses flink to continuously store data into HDFS, and then uses hive to process data at the end of the day. The event center supports statistical calculation of processed event information to form indicators for the upper-level application system.

[0018] Preferably, the management platform includes a data source management unit, a job development unit, a job monitoring unit, an abnormal alarm unit, a resource management unit and an authority system unit. After the management platform obtains business data in real time, it filters it according to the specified business caliber to form real-time message reminders, reach the downstream application system in real time, and conduct marketing tracking and risk warnings.

[0019] Preferably, it also includes an event application unit, which includes a retail module and a public finance module. The retail module includes a real-time rights issuance unit, an instant product purchase unit and a performance tracking unit, and the public finance module includes an account analysis unit and an assessment unit.

[0020] Preferably, it also includes an event source system unit, which includes a channel end, an account system, a core system and a financial management system.

[0021] The stream computing data processing system based on the Kappa architecture proposed in the present invention has the following advantages compared with the prior art:

[0022] 1. The present invention completes event collection and cleaning through the cooperation of event collection unit, event processing unit, event storage unit and event calculation unit through the kafka+flink component, completes real-time calculation of events through the streaming big data computing engine, stores real-time indicators through distributed memory technology, ensures the use of components through a distributed cluster mechanism, and returns real-time calculation data to the batch library in time. In this way, the Kappa architecture can be applied to the construction of the data platform, fully tapping the value of streaming data, and effectively improving business operation efficiency and risk monitoring timeliness.

[0023] 2. By integrating streaming data and batch data, the present invention can continuously store real-time data into the database, and the database provides real-time interface services and performs data processing at the end of the day, thereby improving the batch processing efficiency of the data platform and filling the gap in streaming data processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 This is a system block diagram of the present invention. DETAILED DESCRIPTION

[0025] The following will be combined with the accompanying drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. The specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0026] The present invention provides Figure 1 The stream computing data processing system based on the kappa architecture shown in the figure includes an event collection unit, an event processing unit, an event storage unit, an event computing unit and a management platform. The event collection unit completes event collection and cleaning through the kafka+flink component, supports multiple data sources, and solves the collection of fragmented event information.

[0027] The content collected by the kafka+flink component includes file collection, database collection, message queue collection and website interface collection, wherein the database includes the source-side DB2 relational database, Oracle relational database and MySQL relational database, and real-time data is pushed to the management platform by reading the binlog log in the source database.

[0028] The kafka+flink component includes a java or python API. When collecting events, the source system pushes real-time data to the kafka+flink component through the CDCforKafka form, the kafka API form, and the kafka synchronization form. Among them, the CDCforKafka form uses database log capture software to read the binlog log of the database to push real-time data to the kafka+flink component. This method has the advantages of high reliability, interface configuration, and re-refresh. In practice, the real-time table of each source system can be sent to the same topic, and the table name field and message status field are used to distinguish them during later processing.

[0029] The Kafka API is suitable for real-time synchronization of flow tables, and attention should be paid to the authentication of the Kafka+Flink components and the push specifications of real-time messages when using it. In practice, the CRO credit data is successfully pushed to the Kafka component of the event center through the Python API.

[0030] Kafka synchronization is when the source system also has a big data platform and Kafka components. That is, Kafka's mirrormaker component is used to perform direct synchronization at the topic level. Its advantages are high reliability and easy configuration.

[0031] The event processing unit completes real-time calculation of events through a streaming big data computing engine to solve the splicing and fusion of event fragment information, data field processing, etc.;

[0032] The content processed by the event processing unit includes joining multiple real-time tables, joining real-time tables with offline tables, splicing multiple data streams, custom function processing, data field mapping, abnormal value replacement, string interception and data sorting. The management platform configures the event processing unit's event processing flow in the form of a DAG workflow. In the processing flow, multiple data sources can be dragged and dropped to perform operations such as joining multiple tables, asynchronous search, splicing, and custom function processing. The processed data can form event information with complete business meanings, such as account movement, contract signing, and transactions. The processed data is transmitted to the supported data sources.

[0033] The event storage unit stores real-time indicators through distributed memory technology, solves the storage problem of event information, and supports temporary storage and persistent storage;

[0034] The event storage unit includes a temporary storage module and a persistent storage module. The temporary storage module is used to store memory databases and message middleware, and the persistent storage module is used to store files and relational databases, wherein the memory databases are such as Redis and Aerospike, and the relational databases are such as MySQL and Oracle.

[0035] The event storage unit also performs real-time to batch data storage processing. The processing flow is that the upstream pushes real-time data to the kafka+flink component. The event center uses flink to continuously store the data into HDFS, and then hive performs data processing at the end of the day. This processing method realizes the interaction between the real-time data warehouse and the offline data warehouse, reflecting the strengths of the kappa architecture event center.

[0036] The event calculation unit ensures the use of components through a distributed cluster mechanism, and returns real-time calculation data to the batch library in a timely manner, performs real-time indicator calculations on processed events, and supports upper-level applications such as marketing and risk control to provide decision support;

[0037] The distributed cluster mechanism ensures that the components include an indicator calculation unit, an application management unit, an indicator playback unit and a dynamic report unit. The indicator calculation unit contains a variety of calculation operators, including counting, summing, averaging, maximum, minimum, continuous, increasing, decreasing, sorting, fluctuation, skewness and kurtosis; the application management unit has the ability to obtain context and supports statistics in combination with context. The dynamic report unit is used to store the calculated indicator support indicators in the memory database. The indicator playback unit provides external services through sas software or tableau software, such as counting, summing, averaging, maximum, minimum, etc., which are conventional calculation operators; complex calculation operators such as continuous, increasing, decreasing, sorting, fluctuation, skewness, kurtosis, CEP, correlation coefficient, etc.; it has the ability to obtain context and supports statistics in combination with context, such as statistics on the maximum transaction time interval of a user in the past 24 hours; it supports the recognition of complex event sequences, such as identifying the number of transactions that are transferred in within 1 minute after logging in by the same user in the past 24 hours, and the number of transactions that are transferred out within 30 seconds after logging in.

[0038] The event calculation unit and the event storage unit constitute a real-time indicator market. The real-time indicator market obtains business system data in real time through the CDCfor Kafka service, writes flink jobs according to business caliber to calculate real-time indicators, and stores them in a relational database. Through the data service API interface, it is open to real-time large-screen query; the management platform, event unit and event processing unit constitute a real-time event center. The real-time event center uses flink to continuously store data in HDFS, and then uses hive to process data at the end of the day. The event center supports statistical calculation of processed event information to form indicators for the upper-level application system.

[0039] The management platform is used for the operation of event processing by the event collection unit, the event processing unit, the event storage unit and the event calculation unit;

[0040] The management platform includes a data source management unit, a job development unit, a job monitoring unit, an abnormal alarm unit, a resource management unit and an authority system unit. After the management platform obtains business data in real time, it filters it according to the specified business caliber to form real-time message reminders, reach downstream application systems in real time, conduct marketing tracking and risk warnings, effectively improve the timeliness of business marketing, timely identify potential risks, and support marketing decision-making. The data source management unit is responsible for collecting, integrating and managing data from different business systems; the job development unit is used to create and execute data processing tasks; the job monitoring unit is used to monitor the execution status and state of data processing tasks in real time; the abnormal alarm unit is used to issue an alarm in time and take corresponding processing measures when an abnormality occurs during data processing or system operation; the resource management unit is used to manage computing resources, storage resources, etc. in the platform to ensure the smooth execution of data processing tasks; the authority system unit is used to manage user permissions and access control in the platform.

[0041] It also includes an event application unit, which includes a retail module and a public finance module. The retail module includes a real-time rights issuance unit, a product instant purchase unit and a performance tracking unit. The real-time rights issuance unit is responsible for issuing various rights to consumers in real time in the retail business, such as points, coupons, discounts, etc. By issuing rights in real time, consumer satisfaction and loyalty can be increased, and sales growth can be promoted; the product instant purchase unit is responsible for processing purchase requests, including verifying consumer identity, checking inventory, calculating prices, generating orders, etc. Through the instant purchase function, the purchase process can be simplified, the purchase efficiency can be improved, and a convenient shopping experience can be provided for consumers; the performance tracking unit is used to track and evaluate the performance of the retail business, collect and analyze sales data, inventory data, customer data, etc. in real time, and help retailers understand business conditions and market trends. Through performance tracking, retailers can formulate more effective sales strategies, inventory management strategies and customer service strategies, thereby improving business efficiency and profitability;

[0042] The public finance module includes a book analysis unit and an assessment unit; the book analysis unit is used to conduct in-depth analysis of the book data of public finance business, and help financial institutions understand business conditions, risk conditions and profitability by extracting, organizing and analyzing various financial data, such as income, expenditure, cost, profit, etc. Through book analysis, financial institutions can formulate more reasonable financial strategies, risk management strategies and business development strategies; the assessment unit is responsible for assessing the work performance of various departments and positions within the financial institution, and evaluating employees' performance, ability, attitude and other aspects according to preset assessment standards and indicators. Through assessment, financial institutions can understand employees' work performance, motivate outstanding employees, and improve overall work efficiency and team cohesion.

[0043] It also includes an event source system unit, which includes a channel end, an account system, a core system and a financial management system. The channel end is an interface or platform for the system to interact with users or external entities. It is responsible for receiving user input, requests or transactions, and passing them to subsequent system units for processing; the account system is a system for managing user funds, assets, liabilities and other financial information, and provides users with a platform for storing, managing and using funds, including deposit, withdrawal, transfer, payment and other functions; the core system is the core processing unit in the event source system unit, which is responsible for processing various transactions, business logic and data storage, receiving requests from the channel end, and verifying, processing and executing according to business rules; the financial management system is a system module specifically used for corporate financial management and investment planning, which provides financial statements, fund management, budget control, investment planning, risk management and other functions to help companies achieve their financial goals.

[0044] Event collection and cleaning are completed through the kafka+flink components, real-time calculation of events is completed through the streaming big data computing engine, real-time indicators are stored through distributed memory technology, the use of components is guaranteed through the distributed cluster mechanism, and real-time calculation data is promptly returned to the batch library. In this way, the Kappa architecture can be applied to the construction of the data platform, fully tapping the value of streaming data, effectively improving business operation efficiency and risk monitoring timeliness; by integrating streaming data and batch data, real-time data can be continuously stored in the database, and the database provides real-time interface services and performs end-of-day data processing, which improves the batch processing efficiency of the data platform and fills the gap in streaming data processing.

[0045] Finally, it should be noted that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, it is still possible for those skilled in the art to modify the technical solutions described in the aforementioned embodiments or to make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A stream computing data processing system built on the Kappa architecture, characterized by: include: An event collection unit, which completes event collection and cleaning through the Kafka+Flink component; An event processing unit, wherein the event processing unit completes real-time calculation of events through a streaming big data computing engine; An event storage unit, wherein the event storage unit stores real-time indicators through distributed memory technology; An event computing unit, which ensures the use of components through a distributed cluster mechanism and returns real-time computing data to a batch library in a timely manner; A management platform is used for operating the event collection unit, the event processing unit, the event storage unit and the event calculation unit on event processing.

2. The stream computing data processing system based on the kappa architecture according to claim 1 is characterized in that: The content collected by the kafka+flink component includes file collection, database collection, message queue collection and website interface collection, wherein the database includes the source-side DB2 relational database, Oracle relational database and MySQL relational database, and real-time data is pushed to the management platform by reading the binlog log in the source database.

3. The stream computing data processing system based on the kappa architecture according to claim 2 is characterized in that: The kafka+flink component includes a java or python API. When collecting events, the source system pushes real-time data to the kafka+flink component through the CDCforKafka form, the kafka API form and the kafka synchronization form.

4. The stream computing data processing system based on the kappa architecture according to claim 3 is characterized in that: The event processing unit processes contents including multiple real-time table joins, real-time table and offline table joins, multiple data stream splicing, custom function processing, data field mapping, abnormal value replacement, string interception and data sorting. The management platform configures the event processing unit's event processing flow in the form of DAG workflow. In the processing flow, multiple data sources can be dragged and dropped to perform operations such as multi-table joins, asynchronous searches, splicing, and custom function processing. The processed data is transmitted to the supported data sources.

5. The stream computing data processing system based on the kappa architecture according to claim 1, characterized in that: The event storage unit includes a temporary storage module and a persistent storage module. The temporary storage module is used to store the memory database and the message middleware, and the persistent storage module is used to store files and relational databases.

6. The stream computing data processing system based on the kappa architecture according to claim 5 is characterized in that: The distributed cluster mechanism ensures that the components include an indicator calculation unit, an application management unit, an indicator playback unit and a dynamic reporting unit. The indicator calculation unit contains a variety of calculation operators, including counting, summing, averaging, maximum, minimum, continuous, increasing, decreasing, sorting, fluctuation, skewness and kurtosis; the application management unit has the ability to obtain context and supports statistics in combination with the context. The dynamic reporting unit is used to store the calculated indicator support indicators in the memory database, and the indicator playback unit provides external services through SAS software or Tableau software.

7. The stream computing data processing system based on the kappa architecture according to claim 6 is characterized in that: The event calculation unit and the event storage unit constitute a real-time indicator market. The real-time indicator market obtains business system data in real time through the CDC for Kafka service, writes flink jobs according to business caliber to calculate real-time indicators, and stores them in a relational database. Through the data service API interface, it is open to real-time large-screen query; the management platform, event unit and event processing unit constitute a real-time event center. The real-time event center uses flink to continuously store data in HDFS, and then uses hive to process data at the end of the day. The event center supports statistical calculation of processed event information to form indicators for the upper-level application system.

8. The stream computing data processing system based on the kappa architecture according to claim 7 is characterized in that: The management platform includes a data source management unit, a job development unit, a job monitoring unit, an abnormal alarm unit, a resource management unit and an authority system unit. After the management platform obtains business data in real time, it filters it according to the specified business caliber to form real-time message reminders, reach downstream application systems in real time, and conduct marketing tracking and risk warnings.

9. The stream computing data processing system based on the kappa architecture according to claim 1, characterized in that: It also includes an event application unit, which includes a retail module and a public finance module. The retail module includes a real-time rights issuance unit, an instant product purchase unit and a performance tracking unit. The public finance module includes an account analysis unit and an assessment unit.

10. The stream computing data processing system based on the kappa architecture according to claim 9, characterized in that: It also includes an event source system unit, which includes a channel end, an account system, a core system and a financial management system.