Multi-source heterogeneous data stream processing method and device based on Flink
Through the Flink-based multi-source heterogeneous data stream processing method, the distributed configuration center and Kafka connector are used to standardize data streams and policy calculation, which solves the problem of low computing efficiency of multi-source heterogeneous data stream strategies in the existing technology, and realizes efficient and flexible data processing and real-time decision support.
Patent Information
- Application Number
- CN202510614102.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-08-26
AI Technical Summary
The existing strategy calculation methods have large coding burdens, difficulty in adjusting rules, weak external data aggregation capabilities and limited application of aggregation results when processing multi-source heterogeneous data streams, resulting in low computing efficiency.
Using a multi-source heterogeneous data stream processing method based on Flink, the distributed configuration center stores preset feature extraction rules and predesign computing strategies, and uses Flink's Kafka connector to obtain data streams from the Kafka message queue, and performs standardized processing and policy calculations. The results are stored in a distributed cache for external services to access.
It realizes flexible and efficient processing of multi-source heterogeneous data streams, and can complete data stream conversion and policy calculation in milliseconds, improving the efficiency of real-time decision support and system flexibility and maintainability.
Smart Images

Figure CN120541103A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data stream processing technology, and specifically to a Flink-based multi-source heterogeneous data stream processing method, a Flink-based multi-source heterogeneous data stream processing device, a computer-readable storage medium, and an electronic device. Background Art
[0002] Policy computing refers to the process of analyzing, processing, and making decisions based on input data according to predefined rules or policies. It is widely used in fields such as real-time data analysis, financial risk control, advertising recommendations, and IoT data analysis. Policy computing involves determining whether data meets certain conditions, filtering, aggregating, and correlating data, and ultimately outputting policy-compliant results. These results can be used to support real-time decisions such as transaction approval, user recommendations, and device alerts.
[0003] Faced with heterogeneous data streams from multiple sources and in multiple formats, existing strategy calculation methods have defects such as heavy coding burden, difficult rule adjustment, weak external data aggregation capabilities, and limited application of aggregation results, resulting in low computational efficiency. Summary of the Invention
[0004] The main purpose of this application is to provide a Flink-based multi-source heterogeneous data stream processing method, a Flink-based multi-source heterogeneous data stream processing device, a computer-readable storage medium, and an electronic device, so as to at least solve the problem of low efficiency of policy calculation of multi-source heterogeneous data streams in the prior art.
[0005] In order to achieve the above-mentioned purpose, according to one aspect of the present application, a method for processing multi-source heterogeneous data streams based on Flink is provided, comprising: obtaining preset feature extraction rules and preset calculation strategies, and storing the preset feature extraction rules and the preset calculation strategies in a distributed configuration center, wherein the preset feature extraction rules include at least one of the following: field mapping rules, data type conversion rules, data cleaning rules and standardized format rules, and the preset calculation strategies include at least one of the following: summation strategy, averaging strategy and statistical distribution strategy; applying Flink's Kafka connector to obtain multi-source heterogeneous data streams from the Kafka message queue, and storing the preset feature extraction rules and the preset calculation strategies in a distributed configuration center. The target feature extraction rules corresponding to the multi-source heterogeneous data stream are obtained, and according to the target feature extraction rules, the multi-source heterogeneous data stream is converted into a standardized feature data stream, wherein the target feature extraction rules are subordinate to a plurality of the preset feature extraction rules; the target calculation strategy corresponding to the standardized feature data stream is obtained from the distributed configuration center, and according to the target calculation strategy, the standardized feature data stream is subjected to strategy calculation to obtain a strategy calculation result, and the strategy calculation result is stored in a distributed cache, so that an external service obtains the strategy calculation result from the distributed cache, wherein the target calculation strategy is subordinate to a plurality of the preset calculation strategies.
[0006] Optionally, according to the target computing strategy, the standardized feature data stream is subjected to policy computing to obtain a policy computing result, including: when there is an aggregation computing requirement for the standardized feature data stream, determining that the target computing strategy includes a window aggregation computing strategy, and the window aggregation computing strategy includes a window aggregation sub-strategy and a computing sub-strategy; obtaining external aggregation data related to the standardized feature data stream from the distributed cache; applying the window aggregation sub-strategy to perform aggregation computing on the external aggregation data and the standardized feature data stream to obtain an aggregation computing result; applying the computing sub-strategy to perform the policy computing on the aggregation computing result to obtain the policy computing result.
[0007] Optionally, the window aggregation sub-strategy is applied to perform aggregation calculation on the external aggregated data and the standardized feature data stream to obtain an aggregation calculation result, including: extracting target aggregated data from the external aggregated data according to the data stream expression filtering configuration in the window aggregation sub-strategy; determining the aggregation time window according to the window aggregation time configuration in the window aggregation sub-strategy; determining the aggregation dimension according to the window aggregation latitude configuration in the window aggregation sub-strategy, the aggregation dimension including the user dimension, the time dimension and the business indicator dimension; performing the aggregation calculation on the target aggregated data and the standardized feature data stream according to the aggregation time window and the aggregation dimension to obtain the aggregation calculation result, wherein the aggregation calculation includes at least one of the following: average value calculation, sum calculation, maximum value calculation and minimum value calculation, and the aggregation calculation result is the result of aggregating the target aggregated data and the standardized feature data stream within the time window according to the aggregation dimension.
[0008] Optionally, a target feature extraction rule corresponding to the multi-source heterogeneous data stream is obtained from the distributed configuration center, including: parsing the identification field of the multi-source heterogeneous data stream; selecting the feature extraction rule mapped to the identification field from the distributed configuration center according to the field mapping rule in the preset feature extraction rule and the identification field of the multi-source heterogeneous data stream; and determining the feature extraction rule as the target feature extraction rule.
[0009] Optionally, each of the preset computing strategies includes a preset identifier, and obtaining a target computing strategy corresponding to the standardized feature data stream from the distributed configuration center includes: obtaining a data stream identifier of the standardized feature data stream, the data stream identifier corresponding one-to-one to the standardized feature data stream; matching the preset identifier corresponding to the data stream identifier from the preset computing strategy; and determining the preset computing strategy corresponding to the preset identifier as the target computing strategy.
[0010] Optionally, after obtaining the target calculation strategy corresponding to the standardized feature data stream from the distributed configuration center, and performing strategy calculation on the standardized feature data stream according to the target calculation strategy to obtain the strategy calculation result, the method further includes: outputting the strategy calculation result to the Kafka message queue so that the downstream system receives the strategy calculation result, wherein the downstream system includes at least one of the following: a database, a real-time analysis system, and a data processing system; obtaining preset feature extraction rules and preset calculation strategies, including: obtaining the preset feature extraction rules and the preset calculation strategies through a restful service.
[0011] Optionally, the multi-source heterogeneous data stream is an established data stream or a newly added data stream, and Flink's Kafka connector is used to obtain the multi-source heterogeneous data stream from the Kafka message queue, including: if the multi-source heterogeneous data stream is the established data stream, then keep its original nested data format unchanged, and convert the established data stream into JSON data stream format through the stream conversion function, and the established data stream is a data stream that has been defined and processed in the system; if the multi-source heterogeneous data stream is the newly added data stream, convert the newly added data stream into a unified and standardized data stream format, and the newly added data stream is a data stream that is introduced into the system for the first time, or its structure and format are different from the established data stream.
[0012] According to another aspect of the present application, a processing device for multi-source heterogeneous data streams based on Flink is provided, comprising: a first acquisition unit, configured to acquire preset feature extraction rules and preset calculation strategies, and store the preset feature extraction rules and the preset calculation strategies in a distributed configuration center, wherein the preset feature extraction rules include at least one of the following: field mapping rules, data type conversion rules, data cleaning rules, and standardized format rules, and the preset calculation strategies include at least one of the following: summation strategies, averaging strategies, and statistical distribution strategies; a second acquisition unit, configured to apply Flink's Kafka connector to acquire multi-source heterogeneous data streams from a Kafka message queue, and store the preset feature extraction rules and the preset calculation strategies in a distributed configuration center. The target feature extraction rules corresponding to the multi-source heterogeneous data stream are obtained, and the multi-source heterogeneous data stream is converted into a standardized feature data stream according to the target feature extraction rules, wherein the target feature extraction rules are subordinate to a plurality of the preset feature extraction rules; a policy calculation unit is used to obtain the target calculation policy corresponding to the standardized feature data stream from the distributed configuration center, and perform policy calculation on the standardized feature data stream according to the target calculation policy to obtain a policy calculation result, and store the policy calculation result in a distributed cache, so that an external service obtains the policy calculation result from the distributed cache, wherein the target calculation policy is subordinate to a plurality of the preset calculation policies.
[0013] According to another aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute any one of the methods for processing multi-source heterogeneous data streams based on Flink.
[0014] According to another aspect of the present application, an electronic device is provided, comprising: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include a method for executing any one of the Flink-based multi-source heterogeneous data stream processing methods.
[0015] Applying the technical solution of the present application, preset feature extraction rules and preset calculation strategies are obtained, and the preset feature extraction rules and preset calculation strategies are stored in a distributed configuration center, wherein the preset feature extraction rules include at least one of the following: field mapping rules, data type conversion rules, data cleaning rules and standardized format rules, and the preset calculation strategies include at least one of the following: summation strategy, averaging strategy and statistical distribution strategy; applying Flink's Kafka connector to obtain multi-source heterogeneous data streams from the Kafka message queue, obtaining target feature extraction rules corresponding to the multi-source heterogeneous data streams from the distributed configuration center, and according to the target feature extraction rules, converting the multi-source heterogeneous data streams into standardized feature data streams, wherein the target feature extraction rules are subordinate to multiple preset feature extraction rules; obtaining a target calculation strategy corresponding to the standardized feature data stream from the distributed configuration center, and performing policy calculation on the standardized feature data stream according to the target calculation strategy to obtain a policy calculation result, and storing the policy calculation result in a distributed cache, so that external services obtain the policy calculation result from the distributed cache, wherein the target calculation strategy is subordinate to multiple preset calculation strategies. In this solution, by pre-setting feature extraction rules and combining them with the dynamic configuration capabilities of the distributed configuration center, it is possible to flexibly and efficiently process multi-source heterogeneous data streams. By obtaining real-time updated multi-source heterogeneous data streams from the Kafka message queue and combining them with Flink's powerful stream processing capabilities, it is possible to complete data stream transformation and policy calculation within milliseconds according to the preset calculation strategy, and obtain instant policy calculation results, thus solving the problem of low efficiency of policy calculation for multi-source heterogeneous data streams in the existing technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The drawings that constitute part of this application are used to provide a further understanding of this application. The illustrative embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation on this application. In the drawings:
[0017] Figure 1 A hardware structure block diagram of a mobile terminal for executing a Flink-based multi-source heterogeneous data stream processing method provided in an embodiment of the present application is shown;
[0018] Figure 2A schematic diagram of a process for processing multi-source heterogeneous data streams based on Flink according to an embodiment of the present application is shown;
[0019] Figure 3 A schematic diagram of a processing flow module of a Flink-based multi-source heterogeneous data stream processing method provided according to an embodiment of the present application is shown;
[0020] Figure 4 A structural block diagram of a Flink-based multi-source heterogeneous data stream processing device provided according to an embodiment of the present application is shown.
[0021] The above drawings include the following reference numerals:
[0022] 102. Processor; 104. Memory; 106. Transmission device; 108. Input / output device. DETAILED DESCRIPTION
[0023] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0024] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0025] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present application described here. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0026] As introduced in the background technology, the policy calculation methods in the prior art have defects such as heavy coding burden, difficult rule adjustment, weak external data aggregation capabilities, and limited application of aggregation results, which lead to low computing efficiency. In order to solve the problem of low efficiency of policy calculation of multi-source heterogeneous data streams in the prior art, the embodiments of the present application provide a Flink-based multi-source heterogeneous data stream processing method, a Flink-based multi-source heterogeneous data stream processing device, a computer-readable storage medium, and an electronic device.
[0027] The technical solutions in the embodiments of the present invention will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present invention.
[0028] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 This is a hardware structure diagram of a mobile terminal for a method of processing multi-source heterogeneous data streams based on Flink according to an embodiment of the present invention. Figure 1 As shown, the mobile terminal may include one or more ( Figure 1 Only one is shown) a processor 102 (the processor 102 may include but is not limited to a microprocessor MCU or a programmable logic device FPGA and other processing devices) and a memory 104 for storing data, wherein the mobile terminal may also include a transmission device 106 and an input and output device 108 for communication functions. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the mobile terminal. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.
[0029] The memory 104 can be used to store computer programs, such as software programs and modules of application software, such as the computer program corresponding to the device information display method in the embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implementing the above-mentioned method. The memory 104 may include a high-speed random access memory and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories can be connected to the mobile terminal via a network. Examples of the above-mentioned networks include but are not limited to the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by the mobile terminal's communication provider. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices via a base station to communicate with the Internet. In one example, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0030] This embodiment provides a Flink-based method for processing multi-source heterogeneous data streams running on a mobile terminal, a computer terminal, or a similar computing device. It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Moreover, although a logical order is shown in the flowcharts, in some cases, the steps shown or described may be executed in a different order than shown.
[0031] Figure 2 This is a flowchart of a method for processing multi-source heterogeneous data streams based on Flink according to an embodiment of the present application.
[0032] like Figure 2 As shown, the method includes the following steps:
[0033] Step S201: Obtain preset feature extraction rules and preset calculation strategies, and store the preset feature extraction rules and the preset calculation strategies in a distributed configuration center, wherein the preset feature extraction rules include at least one of the following: field mapping rules, data type conversion rules, data cleaning rules, and standardized format rules; and the preset calculation strategies include at least one of the following: summation strategy, averaging strategy, and statistical distribution strategy;
[0034] Specifically, the preset feature extraction rules are a series of rules defined according to data preprocessing requirements, including field mapping, data type conversion, data cleaning, and standardized formats. Field mapping rules ensure that the same fields in data streams from different sources can be uniformly identified; data type conversion rules handle field type differences to ensure that data can be used for unified calculation logic; data cleaning rules are responsible for cleaning invalid or abnormal data and improving data quality; standardized format rules convert data streams into a unified, standardized format to facilitate unified processing of strategy calculations. The preset calculation strategy defines how to calculate and analyze feature data streams, including summation strategies, averaging strategies, and statistical distribution strategies. The summation strategy is used to calculate the sum of specific fields; the averaging strategy calculates the average value; and the statistical distribution strategy analyzes the distribution of data, such as frequency distribution and quantiles.
[0035] The preset feature extraction rules and calculation strategies are stored in a distributed configuration center (such as Apache Apollo) to enable centralized management, dynamic modification, and real-time push of these rules and strategies. The distributed configuration center serves as the "brain" of these strategies and rules. It allows these rules and strategies to be shared across multiple servers and updated and pushed to running Flink data processing tasks in real time, enabling dynamic adjustments to strategy calculations without requiring redeployment. Flink data processing tasks can read these rules and strategies in real time, reducing the system's reliance on static code and improving its flexibility and maintainability.
[0036] By storing pre-set feature extraction rules and calculation strategies in a distributed configuration center, we achieve centralized management, real-time updates, and flexible application of strategy configurations, significantly improving the responsiveness and business adaptability of the real-time strategy calculation system. This design pattern also adheres to the principle of configuration externalization in microservices architectures, making the system more robust and easier to maintain.
[0037] Step S202: Apply Flink's Kafka connector to obtain multi-source heterogeneous data streams from the Kafka message queue, obtain target feature extraction rules corresponding to the multi-source heterogeneous data streams from the distributed configuration center, and convert the multi-source heterogeneous data streams into standardized feature data streams according to the target feature extraction rules, wherein the target feature extraction rules are subordinate to the plurality of preset feature extraction rules.
[0038] Specifically, in this embodiment, the Flink framework is used to process multi-source heterogeneous data streams. Flink provides a powerful set of APIs that can process unbounded and bounded data streams. Among them, the Kafka connector is a bridge for Flink to interact with the Kafka message queue, allowing Flink to seamlessly read data from Kafka and also write processed data to Kafka. The Kafka message queue is a distributed, partitioned, multi-subscriber messaging system that is widely used to build real-time data pipelines and stream processing applications. It can process large amounts of data streams and ensure the order and reliability of messages. In this embodiment, Kafka serves as the entry point for receiving and storing multi-source heterogeneous data streams. These data streams may come from different data sources, such as databases, log files, sensor data, etc., and the data structure and format of each data source are different.
[0039] Flink's Kafka connector subscribes to Kafka topics, feeding multi-source, heterogeneous data streams into Flink data processing tasks as streaming data. Multi-source, heterogeneous data streams are read as messages, each containing a timestamp and data content. These data streams originate from diverse sources, such as different databases, sensors, or logging systems. Each source may have different data formats, field names, or data types. For example, one data stream may contain structured JSON data, while another may contain database query results. To process these multi-source, heterogeneous data streams, different preprocessing logic must be applied based on the characteristics of each data stream. This preprocessing logic is encapsulated as the pre-set feature extraction rules described above and stored in a distributed configuration center. The distributed configuration center is a system that centrally manages and pushes configurations in real time, such as Apache Apollo, ensuring that configuration information remains consistent across multiple service instances and can be updated in real time.
[0040] Before processing multi-source heterogeneous data streams, Flink data processing tasks retrieve target feature extraction rules that match the current multi-source heterogeneous data streams from the distributed configuration center. The Flink data processing tasks preprocess the multi-source heterogeneous data streams based on the target feature extraction rules, converting them into a unified, standardized feature data stream. This process typically includes data cleansing, field format and type conversion, and data structure standardization, enabling consistent processing and analysis of diverse data sources.
[0041] Step S202 effectively unifies data streams from various sources and formats, converting them into standardized feature data streams to facilitate subsequent policy calculations and decision support. This not only improves the automation of data processing but also enhances the system's flexibility and scalability, as feature extraction rules and calculation strategies can be dynamically adjusted without modifying or redeploying Flink data processing tasks.
[0042] Step S203: Obtain the target computing strategy corresponding to the standardized feature data stream from the distributed configuration center, and perform strategy calculation on the standardized feature data stream according to the target computing strategy to obtain the strategy calculation result, and store the strategy calculation result in the distributed cache so that the external service can obtain the strategy calculation result from the distributed cache, wherein the target computing strategy belongs to a plurality of the preset computing strategies.
[0043] Specifically, after standardizing multi-source heterogeneous data, the next step is strategy calculation. First, the Flink data processing task retrieves a calculation strategy that matches the current standardized feature data stream from the distributed configuration center. These predefined calculation strategies guide how to perform calculations and analysis on the data stream, including summation strategies, averaging strategies, and statistical distribution strategies. These strategies can be flexibly selected and applied based on specific business scenarios. These predefined calculation strategies cover a variety of computing needs. For example, a summation strategy calculates the sum of a business metric within a specific time window; an averaging strategy calculates the average transaction volume for a user group; and a statistical distribution strategy analyzes data distribution, such as calculating data quantiles. Because these strategies are stored in the distributed configuration center, they can be dynamically updated and applied in real time without modifying the Flink data processing task code. This means that adjustments to business logic can be quickly responded to without requiring application redeployment.
[0044] Based on the obtained target calculation strategy, the Flink data processing task performs the corresponding strategy calculation on the standardized feature data stream. This involves steps such as data filtering, aggregation, calculation, and analysis to ensure that the calculation results meet business requirements and policy definitions. After the strategy calculation is completed, the strategy calculation result is obtained. The strategy calculation result contains the output of the calculation of the standardized feature data stream according to the target calculation strategy, which can be the calculated value of one or more key business indicators.
[0045] Policy calculation results are stored in a distributed cache (such as Redis), a high-performance in-memory data structure designed for fast data read and write access, ensuring the real-time and high availability of policy calculation results. As a real-time storage and access point for policy calculation results, the distributed cache not only provides high-speed data access but also handles highly concurrent requests, ensuring the timeliness of calculation results and high system availability. External services (such as decision support systems and real-time analytics tools) can access the distributed cache to obtain the latest policy calculation results without directly interacting with Flink data processing tasks. This reduces inter-system communication latency and improves the overall system's responsiveness and efficiency.
[0046] During runtime, Flink data processing tasks retrieve computation strategies from a distributed configuration center. These strategies are associated with standardized feature data streams, ensuring targeted and accurate computation. Flink data processing tasks then perform specific computations on the data streams based on the strategies. Dynamically acquiring and applying computation strategies enables the system to flexibly adapt to changing business needs. For example, when a strategy needs to be adjusted, only the strategy definition in the configuration center needs to be updated, without having to modify and redeploy the Flink application. This significantly improves system maintainability and responsiveness. Strategy computation results are stored in a distributed cache, allowing external services to access these results in real time through a cached API without having to wait for Flink data processing tasks to execute or query the database. This design reduces data access latency and improves the system's real-time interactive capabilities. This is particularly important for scenarios requiring rapid decision-making, such as real-time risk assessment in financial transactions, providing immediate feedback and support.
[0047] Through step S203, real-time strategy calculation of multi-source heterogeneous data streams based on Flink is implemented, and distributed cache enables external services to obtain calculation results in real time, thereby enhancing the system's real-time decision-making capabilities and business flexibility.
[0048] This embodiment combines the stream processing capabilities of Apache Flink, the efficient data transmission of Kafka, and the dynamic configuration management and high-speed data access capabilities of distributed configuration centers (such as Apache Apollo) and distributed caches (such as Redis), and can automatically adjust data processing logic to adapt to different data sources and business needs. At the same time, by storing policy calculation results in a distributed cache in real time, it ensures that external services can immediately access the latest calculation results, thereby significantly improving the real-time and efficiency of decision support while ensuring data processing accuracy and consistency, thereby solving the problem of low efficiency of policy calculation for multi-source heterogeneous data streams in the existing technology.
[0049] During the specific implementation process, according to the above-mentioned target calculation strategy, the above-mentioned standardized feature data stream is subjected to strategy calculation to obtain a strategy calculation result, including: when there is an aggregation calculation requirement for the above-mentioned standardized feature data stream, determining that the above-mentioned target calculation strategy includes a window aggregation calculation strategy, and the above-mentioned window aggregation calculation strategy includes a window aggregation sub-strategy and a calculation sub-strategy; obtaining external aggregation data related to the above-mentioned standardized feature data stream from the above-mentioned distributed cache; applying the above-mentioned window aggregation sub-strategy to perform aggregation calculation on the above-mentioned external aggregation data and the above-mentioned standardized feature data stream to obtain an aggregation calculation result; applying the above-mentioned calculation sub-strategy to perform the above-mentioned strategy calculation on the above-mentioned aggregation calculation result to obtain the above-mentioned strategy calculation result.
[0050] Specifically, first, when there is a need for aggregation calculations in the standardized feature data stream, the Flink data processing task will analyze the target calculation strategy to confirm whether it includes a window aggregation calculation strategy. The window aggregation calculation strategy here usually involves statistical analysis of data within a period of time, including aggregation operations such as summation, averaging, maximum or minimum values. The window aggregation calculation strategy consists of a window aggregation sub-strategy and a calculation sub-strategy. The window aggregation sub-strategy is responsible for defining the time window and dimension for data aggregation and determining which data will be aggregated together for analysis. This usually involves grouping data streams, setting time windows (such as sliding windows, session windows, etc.), and how to store intermediate results after aggregation for subsequent use by the calculation sub-strategy. The calculation sub-strategy then executes further calculation instructions based on the aggregate calculation results generated by the window aggregation sub-strategy, ultimately generating the strategy calculation results. These calculations can be complex mathematical operations, statistical analysis, or other business logic calculations.
[0051] Before performing windowed aggregations, Flink data processing tasks retrieve external aggregate data related to the standardized feature data stream from the distributed cache. This step is designed to incorporate information from historical data or other relevant data sources to enhance the comprehensiveness and accuracy of the aggregation. External aggregate data includes previous aggregation results within the same or different windows, providing context for the current computation. A windowed aggregation sub-strategy is then applied to aggregate the standardized feature data stream and the external aggregate data. Using Flink's window operators, such as Window or SessionWindow, the data stream is grouped and aggregated by time windows and specified dimensions to produce the aggregated results. During the aggregation process, Flink leverages its stream processing capabilities to efficiently process large amounts of real-time data streams while also incorporating historical data to generate the most up-to-date aggregated results. Finally, the calculation sub-strategy performs further computations on the aggregated results. For example, the average strategy first performs a summation operation and then divides it by the number of data points within the window to obtain the average. The result of this stage is the final policy calculation result, reflecting the business value or metric calculated for the standardized feature data stream within the specified time window according to the predefined rules and strategies.
[0052] This process not only enables aggregate calculations on standardized feature data streams, but also enables the integration of external aggregate data for more comprehensive policy calculations, resulting in refined policy calculation results. The results are stored in a distributed cache, ensuring that external services have real-time access to the latest aggregate calculation information, supporting real-time decision-making and analysis needs.
[0053] Furthermore, the above-mentioned window aggregation sub-strategy is applied to perform aggregation calculation on the above-mentioned external aggregated data and the above-mentioned standardized feature data stream to obtain the aggregation calculation result, including: extracting the target aggregated data from the above-mentioned external aggregated data according to the data stream expression filtering configuration in the above-mentioned window aggregation sub-strategy; determining the aggregation time window according to the window aggregation time configuration in the above-mentioned window aggregation sub-strategy; determining the aggregation dimension according to the window aggregation latitude configuration in the above-mentioned window aggregation sub-strategy, and the above-mentioned aggregation dimension includes the user dimension, the time dimension and the business indicator dimension; performing the above-mentioned aggregation calculation on the above-mentioned target aggregated data and the above-mentioned standardized feature data stream according to the above-mentioned aggregation time window and the above-mentioned aggregation dimension to obtain the above-mentioned aggregation calculation result, wherein the above-mentioned aggregation calculation includes at least one of the following: average value calculation, sum calculation, maximum value calculation and minimum value calculation, and the above-mentioned aggregation calculation result is the result of aggregating the above-mentioned target aggregated data and the above-mentioned standardized feature data stream within the above-mentioned time window according to the above-mentioned aggregation dimension.
[0054] Specifically, the window aggregation sub-strategy includes a data stream expression filtering configuration, which is used to filter out a subset of data that meets specific conditions from external aggregate data. For example, if the goal is to analyze the transaction behavior of a specific user group, the filtering configuration can be set to select only the data of these users. The expression filtering configuration is usually based on the field values of the data stream and is implemented using the expression engine in the Flink framework (such as Apache Flink's built-in expression engine). By applying the above filtering configuration, Flink data processing tasks can accurately filter and extract target aggregate data from external aggregate data (stored in distributed caches such as Redis). These data are the basis for the next step of aggregation calculation, ensuring the targeted and efficient calculation.
[0055] The window aggregation time configuration in the window aggregation sub-strategy defines the time range for aggregation calculations. The time window can be fixed (such as every 5 minutes) or sliding (such as sliding 1 minute every 5 minutes), or event-driven (such as user sessions). The time window is determined in order to perform aggregate analysis of data within a specific time period to reflect the short-term or long-term trends of the business. The aggregation dimension configuration determines how the data is grouped and aggregated. Common aggregation dimensions include user dimension, time dimension, and business indicator dimension, such as user ID, transaction timestamp, and transaction amount. These dimensions enable more detailed data analysis, such as calculating the average transaction amount of a specific user in a specific time period, or calculating the maximum value of a business indicator in different time windows.
[0056] Based on the specified time window and aggregation dimension, Flink data processing tasks aggregate the target aggregate data and the standardized feature data stream. Calculation types include average, sum, maximum, or minimum, depending on business requirements and policy definitions. This process can be implemented using Flink's window operators (such as TumblingWindow or SlidingWindow). Flink automatically handles the data stream grouping and window aggregation logic based on the configuration to generate the aggregate calculation results.
[0057] By configuring data stream expression filters, it's possible to extract data subsets tailored to specific business needs from massive amounts of external aggregated data, avoiding the processing of irrelevant data and improving the relevance and efficiency of calculations. Dynamic configuration of the aggregation time window allows for flexible adjustment of the data analysis timeframe based on different business scenarios, effectively supporting both short-term, immediate monitoring and long-term trend analysis. Setting aggregation dimensions allows for multi-dimensional data analysis, providing insights into business performance from diverse perspectives, including detailed data analysis results at the user, time, and business metric levels. Aggregate calculation results are generated and stored in real-time in a distributed cache, ensuring that external services can instantly access and apply these results, supporting real-time decision-making and analysis. In summary, this process enables efficient, dynamic, and accurate aggregation calculations of data streams, providing powerful data support for complex business scenarios while also improving the real-time and accuracy of decision-making.
[0058] This application can adopt a rule learning model based on machine learning, which can automatically analyze historical data streams and policy execution results to explore potential rules and patterns. Through training, the rule learning model can predict which rules are more effective in specific situations, thereby intelligently adjusting the priority and scope of application of the rules. Based on the results of rule learning, it can adaptively optimize the policy calculation process and dynamically adjust the filtering conditions, window size, aggregation dimensions and other parameters in the policy to achieve the best computing effect. For example, for frequently triggered policies, its execution path can be optimized to reduce unnecessary calculations; for low-frequency but high-risk policies, resource allocation can be improved to ensure timely response. Intelligent rule learning and adaptive optimization further enrich the functionality and security of the policy calculation method based on the Flink framework, and enhance the adaptability and optimization capabilities of complex policy calculations.
[0059] In order to improve the efficiency and effectiveness of data processing, in some embodiments of the present application, target feature extraction rules corresponding to the above-mentioned multi-source heterogeneous data stream are obtained from the above-mentioned distributed configuration center, including: parsing the identification field of the above-mentioned multi-source heterogeneous data stream; according to the above-mentioned field mapping rules in the above-mentioned preset feature extraction rules and the above-mentioned identification field of the above-mentioned multi-source heterogeneous data stream, selecting the feature extraction rule mapped by the above-mentioned identification field from the above-mentioned distributed configuration center; and determining the above-mentioned feature extraction rule as the above-mentioned target feature extraction rule.
[0060] Specifically, in the data stream processing process, the multi-source, heterogeneous data streams received from the Kafka message queue are first parsed to identify fields that can identify the source and characteristics of the data stream. These identification fields can be identifiers of the data stream source (such as different data source IDs), data type identifiers (such as file types like JSON and CSV), or specific business tags (such as transaction data or user behavior data). Parsing the identification fields provides the basis for subsequent rule matching and application. The preset feature extraction rules include field mapping rules, which guide how to map fields in the multi-source, heterogeneous data streams to a unified feature field structure. Based on the parsed identification fields and the field mapping rules in the preset feature extraction rules, the Flink data processing task searches and selects a feature extraction rule that matches the identification field in the distributed configuration center. For example, if the parsed identification field indicates that the data stream originates from user transaction data, the task searches for feature extraction rules related to user transaction data processing.
[0061] The distributed configuration center contains multiple preset feature extraction rules, each corresponding to a different data source and business scenario. The above steps allow you to locate and determine the feature extraction rule that best matches the current data stream characteristics from these preset rules. This is known as the target feature extraction rule. This rule guides the Flink data processing task in preprocessing the data stream, ensuring that the data stream is converted into a unified, standardized feature data stream to facilitate subsequent policy calculations.
[0062] By parsing the identification fields of the data stream and matching the feature extraction rules according to the field mapping rules, the source and type of the data stream can be automatically identified, thereby selecting the most appropriate preprocessing method, avoiding manual intervention, and improving the automation level of the processing flow. The acquisition of target feature extraction rules is performed dynamically at runtime, which means that even if the data source or business scenario changes, it can be quickly adjusted and new rules can be selected for data conversion, enhancing flexibility and adaptability. In short, the above process realizes the automated and standardized preprocessing of multi-source heterogeneous data streams, providing input data with a unified format, rich content and clear structure for subsequent policy calculations, facilitating subsequent policy calculations, and also providing a standardized method for data storage and querying, improving the efficiency and effectiveness of data processing.
[0063] In other embodiments of the present application, each of the above-mentioned preset computing strategies includes a preset identifier, and obtaining a target computing strategy corresponding to the above-mentioned standardized feature data stream from the above-mentioned distributed configuration center includes: obtaining a data stream identifier of the above-mentioned standardized feature data stream, the above-mentioned data stream identifier corresponding one-to-one to the above-mentioned standardized feature data stream; matching the above-mentioned preset identifier corresponding to the above-mentioned data stream identifier from the above-mentioned preset computing strategy; and determining the above-mentioned preset computing strategy corresponding to the above-mentioned preset identifier as the above-mentioned target computing strategy.
[0064] Specifically, first, you need to obtain a unique identifier for the current standardized feature data stream. This identifier uniquely identifies and distinguishes different data stream sources. In Flink data processing tasks, the data stream identifier is a metadata attribute of the data stream that is preserved when the data stream is converted to a standardized feature data stream. This identifier is crucial for subsequently selecting the appropriate computation strategy. When preset computation strategies are stored in the distributed configuration center, each strategy is assigned a preset identifier. These identifiers are used to quickly locate and match specific data stream requirements at runtime. When a Flink data processing task searches for a computation strategy from the distributed configuration center, it matches the preset identifiers in the preset computation strategies against the data stream identifiers of the standardized feature data stream, finding the computation strategy that best matches the characteristics of the current data stream.
[0065] Once a preset identifier matching the data stream identifier is found, the corresponding preset calculation strategy is determined as the target calculation strategy. This strategy guides the Flink data processing task to perform specific calculations on the standardized feature data stream, such as summation, averaging, maximum, or minimum values. By matching identifiers, the most appropriate strategy is automatically selected without manual intervention, ensuring targeted and efficient calculations.
[0066] By matching data stream identifiers with pre-set identifiers, the most relevant computation strategy for the current data stream is automatically selected, preventing policy inefficiency or misuse, and improving the accuracy and specificity of data processing. Pre-set computation strategies are stored in a distributed configuration center and retrieved through identifier matching. This means computation strategies can be dynamically updated and applied in real time without modifying the Flink data processing task code, enhancing flexibility and maintainability. Automatically matching and determining the target computation strategy eliminates the need for in-depth understanding of data stream details or the processing logic of Flink data processing tasks, simplifying the policy computation process, reducing its complexity and entry level, and improving its efficiency. Furthermore, the target computation strategy is selected based on the real-time characteristics of the data stream, ensuring the timeliness of computation results and supporting real-time data analysis and decision support scenarios, such as real-time transaction monitoring and real-time analysis of IoT device status. This process not only enables precise processing of standardized feature data streams, but also improves the real-time and flexibility of data processing through dynamic policy matching and application.
[0067] In order to ensure the immediacy and availability of the results, a real-time data source is provided for the downstream system. After obtaining the target calculation strategy corresponding to the above-mentioned standardized feature data stream from the above-mentioned distributed configuration center, and performing strategy calculation on the above-mentioned standardized feature data stream according to the above-mentioned target calculation strategy, and obtaining the strategy calculation result, the above-mentioned method also includes: outputting the above-mentioned strategy calculation result to the above-mentioned Kafka message queue so that the downstream system receives the above-mentioned strategy calculation result, wherein the above-mentioned downstream system includes at least one of the following: a database, a real-time analysis system and a data processing system.
[0068] Specifically, after policy calculation is performed on the standardized feature data stream and the results are obtained, they are output to the Kafka message queue. As a high-throughput distributed publish-subscribe messaging system, Kafka can efficiently and reliably process and transmit large amounts of real-time data. Outputting policy calculation results to Kafka ensures the immediacy and availability of the results, providing a real-time data source for downstream systems. Downstream systems are systems that receive policy calculation results and further process or store them. These systems can be databases for persistent storage of results, real-time analytics systems for immediate analysis and alerting, or data processing systems for data reprocessing or integration. Outputting policy calculation results to the Kafka message queue enables downstream systems to receive and process them in real time, providing direct data support for real-time decision-making and analysis, and shortening the time delay from data processing to decision execution.
[0069] In some further embodiments of the present application, obtaining the preset feature extraction rules and the preset calculation strategy includes: obtaining the above-mentioned preset feature extraction rules and the above-mentioned preset calculation strategy through a restful service.
[0070] Specifically, the acquisition of preset feature extraction rules and preset calculation strategies is achieved through RESTful services. RESTful is a design style and development method for network applications. Based on the HTTP protocol, data can be exchanged in XML or JSON format. The configuration page stores the configured preset feature extraction rules and preset calculation strategies in the Apollo distributed configuration center in a RESTful manner. This means that the modification, addition or deletion of rules and strategies can be performed without restarting the service, and can be completed through a simple HTTP request, which greatly facilitates the dynamic management and adjustment of rules. Interact with Apollo through the Java API to obtain preset feature extraction rules and preset calculation strategies. Rules can be pulled from the remote server at any time as parameters without hard coding in the application, which reduces the coupling between systems, enables faster adaptation to business changes, reduces system downtime caused by rule updates, and improves system stability and user experience.
[0071] Uploading and updating rules and policies through RESTful services enables dynamic management of rules and policies, allowing timely policy adjustments even during peak business periods without worrying about the impact of system restarts. Using a Java API to retrieve rules and policies from Apollo reduces network I / O time consumption and ensures efficient policy calculation. Real-time updates of rules and policies also ensure high system stability, enabling rapid response to business needs and market changes.
[0072] The above-mentioned multi-source heterogeneous data stream is an existing data stream or a newly added data stream, and Flink's Kafka connector is used to obtain the multi-source heterogeneous data stream from the Kafka message queue, including: if the above-mentioned multi-source heterogeneous data stream is the above-mentioned established data stream, then keep its original nested data format unchanged, and convert the above-mentioned established data stream into JSON data stream format through the stream conversion function. The above-mentioned established data stream is a data stream that has been defined and processed in the system; if the above-mentioned multi-source heterogeneous data stream is the above-mentioned newly added data stream, convert the above-mentioned newly added data stream into a unified and standardized data stream format. The above-mentioned newly added data stream is a data stream that is introduced into the above-mentioned system for the first time, or its structure and format are different from the above-mentioned established data stream.
[0073] Specifically, established data streams refer to data streams that have been previously defined and processed in the system. These data streams, defined by the business system's custom format, are not standard data formats and require additional processing. These data streams use nested data formats (for example, complex JSON or XML structures) and contain multi-level data information. Maintaining the original nested data format when processing established data streams avoids information loss or format inconsistencies during data conversion. While the original format of the established data stream is preserved, it must be converted to the system-standard JSON data stream format before feature extraction and policy calculation. This conversion utilizes Flink's stream transformation capabilities, ensuring that the data stream can be recognized and processed by other components in the system (such as the feature extraction module and the policy calculation module). By uniformly converting nested data formats to JSON, data can be parsed and used more efficiently, while also facilitating dynamic rule matching and application. After conversion to the JSON data stream format, the data is then converted to the standardized format required for pre-defined policy calculation using pre-defined feature extraction rules.
[0074] New data streams refer to data streams that are introduced into the system for the first time, or data streams whose structure and format are significantly different from established data streams. For this type of data stream, its structure and format are accessed in a standardized manner and do not require additional processing. It is converted into the standardized format required for preset policy calculations through preset feature extraction rules, ensuring that all data streams have the same format and structure after entering the system, facilitating subsequent data processing and rule application. Converting new data streams into a unified and standardized data stream format is a key step in achieving system scalability and data processing flexibility. Through this conversion, new data streams can seamlessly connect to various functions in the system, such as feature extraction, policy calculation, etc., while also ensuring the compatibility and consistency of data streams in the entire system, reducing anomalies and errors in data processing.
[0075] By maintaining the nested data format of established data streams and subsequently converting the format to JSON, it is possible to quickly adapt to the processing needs of known data streams without requiring extensive modifications to the existing processing logic, reducing the complexity of system maintenance and upgrades. For newly added data streams, standardized data stream format conversion allows for flexible access and processing without requiring large-scale adjustments to the existing system architecture. This provides a solid foundation for system scalability and adaptability, enabling faster response to business changes and increased data sources.
[0076] The above process ensures consistent processing logic and rule application across all data streams, reducing variability and improving accuracy and efficiency. This mechanism enables efficient and consistent processing of heterogeneous data streams from multiple sources, providing a solid data foundation for real-time data analysis and policy calculation, while also improving the system's overall operational efficiency and business adaptability.
[0077] In this application, an intelligent scheduling framework is introduced that can automatically detect the resource requirements and load conditions of Flink data processing tasks and dynamically adjust the running resources (such as CPU, memory) of Flink data processing tasks. This framework uses the monitoring data of the Flink cluster, combined with prediction algorithms (such as ARIMA, Prophet, etc.) to predict the load conditions of future Flink data processing tasks and expand or shrink resources in advance. The intelligent scheduling framework can automatically send resource adjustment instructions to the Flink cluster manager based on the prediction results, thereby realizing the elastic expansion and contraction of Flink data processing tasks. This mechanism ensures that Flink data processing tasks always run under the optimal resource configuration, and there will be no backlog of tasks due to insufficient resources, nor will there be waste due to excess resources.
[0078] Through the integration and optimization of a multi-level rule engine, we can better handle a variety of policy calculation needs, from small numerical comparisons to large-scale machine learning decisions, and achieve timely and accurate responses. This not only improves the comprehensiveness of policy calculations, but also greatly optimizes computing efficiency and resource utilization. Intelligent scheduling and dynamic resource adjustment mechanisms enable automatic adaptation to load changes, dynamically adjusting resources without manual intervention, ensuring high system availability and stability. Especially in the face of sudden high data flow pressures, we can respond quickly and avoid data processing bottlenecks and delays.
[0079] In order to enable those skilled in the art to more clearly understand the technical solution of the present application, the implementation process of the Flink-based multi-source heterogeneous data stream processing method of the present application will be described in detail below with reference to specific embodiments.
[0080] This embodiment relates to a specific Flink-based method for processing multi-source heterogeneous data streams. The specific process is as follows:
[0081] 1) Receive data stream feature extraction rules and policies through the RESTful service, and save the submitted configuration information in JSON format to the distributed configuration center Apollo.
[0082] On the management side, feature extraction rules and policies are configured separately. The initial policy state is stopped. Once activated, they are submitted to Apollo, the distributed configuration center, via a RESTful service. Incoming data streams are first converted into standardized feature data streams using feature extraction rules. The feature data streams are then processed through a series of rule-based processing. First, the rule filters are configured to filter data streams that match the rules, and then the rule aggregation configuration is used to aggregate the data streams. Apollo is a distributed configuration center that pushes configuration changes to applications in real time. It also provides a convenient visual management interface, enabling dynamic policy changes to take effect in real time.
[0083] 2) Obtain multi-source heterogeneous data streams from the message queue.
[0084] Multi-source heterogeneous data streams are received through Kafka, a high-performance, low-latency distributed message queue. Accessing multi-source heterogeneous data streams is primarily categorized into existing data stream access and new data stream access. Existing data stream access maintains the existing data's nested format, converts it into a JSON data stream format, and extracts it using customized extraction methods within feature extraction rules. This reduces service access costs while also decoupling internal data streams from services. New data stream access uses a unified, standardized data stream format and the universal extraction methods within feature extraction rules.
[0085] 3) Obtain the feature extraction rules corresponding to the data stream through the Apollo client, convert the multi-source heterogeneous data stream into a normalized feature data stream, and output it to the message queue.
[0086] The Apollo client can obtain configuration data in real time. At the same time, there is a mapping relationship between the Apollo configuration key and the data stream. The feature extraction rules that conform to the data stream on Apollo can be directly obtained through the data stream mapping field. The feature extraction rules are used to convert the access data into data in a standardized feature format, so that the data streams processed by the rule stream are all standardized and unified format data, realizing decoupling from the business. The feature extraction rules define the mapping relationship and extraction method between the access data and the standardized feature data. The extraction methods include direct extraction, JSON path extraction, and custom extraction methods: JSON path extraction is implemented by combining Java open source JSON tools, and can realize arbitrary access path Java property extraction without coding; custom extraction meets the open-closed principle by combining factory methods, reducing development and maintenance costs. The data format of the standardized feature data stream is a dictionary format, the dictionary key is the feature parameter name, the dictionary value is the feature structure, and the feature structure includes feature key, feature value, and feature type.
[0087] 4) Get the feature data stream from the message queue.
[0088] Obtain feature data streams from Kafka in dictionary format. Standardizing the data stream format decouples policy calculation from business, facilitating the provision of general policy calculation capabilities.
[0089] 5) Obtain the corresponding strategy for the data stream through the Apollo client, process the data stream, obtain the aggregated data from the external distributed storage, write the aggregated results to the distributed storage in real time, and output the data stream processing results to the message queue.
[0090] The Apollo client obtains the policy that matches the data flow on Apollo in real time based on the data flow mapping field.
[0091] Window aggregation rules in a policy are used to aggregate data streams that meet specific conditions in real time. Window aggregation rules define the data stream expression filter, window aggregation time, groupby, and storage. The filter is used to select data streams that meet the window aggregation rule and is implemented using the open-source Java expression engine, aviator. External window data storage locations are determined through the storage, groupby, and window configurations, and code-free implementation using the Redis API is also possible. Window aggregation with external data is implemented using storage, groupby, and window. Storage determines the data storage method for window aggregations, typically implemented using the distributed, high-performance Redis cache. Groupby supports multiple aggregation field configurations based on specific rules and determines the storage location of window aggregate data in the Redis dictionary. Window determines the time range of window aggregation data, corresponding to the data range in the Redis dictionary. The window format includes the window type, window time unit, and window time. The corresponding Redis dictionary field field is a date in time unit, and the dictionary field value is the aggregated data in time unit.
[0092] The result of data stream processing is a policy event data structure, including features, policies, and the processing status of each rule in the policy.
[0093] 6) External services obtain aggregated results by accessing distributed storage in real time.
[0094] It should be noted that the above aggregation results may include aggregation calculation results, or may include aggregation calculation results and policy calculation results.
[0095] The above processing flow mainly includes seven modules, see Figure 3 , as follows:
[0096] 1) Policy input module: Receives feature extraction rules and policies through restful services and saves the received data to the distributed configuration center Apollo. 2) Access data stream acquisition module: Obtains multi-source heterogeneous data streams from the message queue through the Flink Kafka connector. 3) Feature extraction rule acquisition module: Obtains feature extraction rules corresponding to the data stream through the Apollo client. 4) Feature extraction module: Converts multi-source heterogeneous data streams into normalized data streams through Flink and outputs them to the message queue. 5) Feature data stream acquisition module: Obtains normalized feature data streams from the message queue through the Flink Kafka connector. 6) Policy acquisition module: Obtains policies corresponding to the data stream through the Apollo client. 7) Policy calculation module: Processes the data stream through Flink, obtains aggregated data from external distributed storage, writes the aggregation results to the distributed storage in real time, and outputs the data stream processing results to the message queue.
[0097] The above aggregation results can include aggregate calculation results, or aggregate calculation results and policy calculation results. External services obtain aggregation results by accessing distributed storage data in real time.
[0098] By defining standardized access processes, abstract feature structures, and feature extraction rules, and relying on the distributed configuration center Apollo and the distributed message queue Kafka, incoming data streams are converted into standard feature data streams, enabling scalable access to multi-source heterogeneous data streams. By abstracting the policy structure, relying on the distributed stream processing engine Flink for real-time policy calculations and the distributed cache Redis for real-time storage of aggregation results, this ensures system flexibility, scalability, policy calculation performance, and the timeliness of the application of aggregation results.
[0099] Policy computing typically serves multiple scenarios, and the data structures in different scenarios may differ. Connecting to different scenarios separately is too costly, and heterogeneous data structures are also detrimental to the generalization of policy computing capabilities. This application defines a set of standard business access processes, uses the distributed message queue Kafka to decouple business access, abstracts feature extraction rules and features, and enables low-cost, scalable access to multi-source heterogeneous data sources. Policy aggregation rules also rely on external aggregated data to aggregate data. This application uses abstract policies, relies on the Flink distributed high-performance stream processing engine, and the distributed high-performance cache Redis, to flexibly support policy aggregation rule data aggregation in real time, and make aggregated data available to external services in real time.
[0100] The following is a specific embodiment of the present application.
[0101] Background: In banking systems, real-time transaction monitoring is crucial for fraud prevention, risk management, and optimizing the customer experience. Banks process transaction data streams from a wide range of sources, including online banking transactions, ATM operations, and credit card payments. These data also come in a variety of formats, such as JSON, CSV, and specialized internal bank formats. Furthermore, banks need to implement dynamically adaptable monitoring strategies to respond to evolving market conditions and customer demands.
[0102] Implementation Plan: First, a standardized business access process was defined. Before entering the system, all transaction data streams, regardless of their original data format, were converted into a unified JSON data stream format. This process ensured data format consistency, facilitating subsequent processing and analysis. Feature extraction rules were configured in the Apollo distributed configuration center. These rules defined how to extract key features, such as user ID, transaction amount, and transaction time, from transaction data from various sources. When a Flink data processing task ingested a transaction data stream from a Kafka message queue, it retrieved the corresponding feature extraction rules from Apollo based on the data stream's identifier fields (such as the data source ID). For existing data streams, the original nested data format was maintained and the streams were converted to JSON format. For newly added data streams, they were converted to a unified JSON format according to a pre-defined normalization process. Apollo defined multiple pre-defined computation strategies, each with a pre-defined identifier that guided the Flink data processing task on how to process the standardized feature data streams. Strategies included filtering rules, window aggregation rules, and access logic for external distributed storage.
[0103] When a standardized feature data stream enters the system, the Flink data processing task retrieves the corresponding preset calculation strategy from Apollo based on the data stream identifier. The strategy calculation process first filters the data using expression rules. Then, based on the configuration of the window aggregation sub-strategy (such as window time and aggregation dimension), it aggregates the data stream and external aggregate data (such as historical user transaction records) in real time, and stores the aggregation results in Redis in real time.
[0104] After the calculation is complete, the policy calculation results are output to the Kafka message queue. Downstream systems such as databases, real-time analysis systems, or data processing systems can receive and further process these results in real time. External services (such as a bank's risk management platform) can access the data in Redis in real time to obtain the latest aggregated results for real-time transaction risk assessment, such as real-time monitoring of abnormal user behavior.
[0105] Through the above methods, banks can monitor and analyze transaction data streams in real time, identify and prevent potential fraud, optimize business processes, and enhance the customer experience. Specifically, standardized business access processes and non-coding feature extraction rules enable the system to efficiently process multi-source heterogeneous data streams while easily adapting to the access of new data streams, reducing development and maintenance costs. Relying on Flink's high-performance stream processing capabilities and distributed storage Redis, it can perform policy calculations and data aggregation in real time, supporting real-time decision-making and risk control. Aggregation results are stored in Redis in real time, and external services such as risk management platforms can access these results instantly, improving the timeliness and accuracy of risk warnings and enhancing the bank's risk management capabilities.
[0106] The embodiment of the present application also provides a processing device for multi-source heterogeneous data streams based on Flink. It should be noted that the processing device for multi-source heterogeneous data streams based on Flink in the embodiment of the present application can be used to execute the processing method for multi-source heterogeneous data streams based on Flink provided in the embodiment of the present application. The device is used to implement the above-mentioned embodiments and preferred implementation methods, and the details that have been explained will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware for a predetermined function. Although the device described in the following embodiments is preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceived.
[0107] The following introduces a Flink-based multi-source heterogeneous data stream processing device provided in an embodiment of the present application.
[0108] Figure 4 Schematic diagram of a multi-source heterogeneous data stream processing device based on Flink according to an embodiment of the present application. Figure 4As shown, the device includes a first acquisition unit 10, a second acquisition unit 20 and a strategy calculation unit 30. The first acquisition unit is used to obtain preset feature extraction rules and preset calculation strategies, and store the preset feature extraction rules and the preset calculation strategies in a distributed configuration center, wherein the preset feature extraction rules include at least one of the following: field mapping rules, data type conversion rules, data cleaning rules and standardized format rules, and the preset calculation strategies include at least one of the following: summation strategy, averaging strategy and statistical distribution strategy; the second acquisition unit is used to apply Flink's Kafka connector to obtain multi-source heterogeneous data streams from the Kafka message queue, and obtain target features corresponding to the multi-source heterogeneous data streams from the distributed configuration center. The target feature extraction rule is used to convert the multi-source heterogeneous data stream into a standardized feature data stream according to the target feature extraction rule, wherein the target feature extraction rule is subordinate to a plurality of the preset feature extraction rules; the policy calculation unit is used to obtain the target calculation policy corresponding to the standardized feature data stream from the distributed configuration center, and perform policy calculation on the standardized feature data stream according to the target calculation policy to obtain the policy calculation result, and store the policy calculation result in the distributed cache, so that the external service obtains the policy calculation result from the distributed cache, wherein the target calculation policy is subordinate to a plurality of the preset calculation policies.
[0109] This embodiment combines the stream processing capabilities of Apache Flink, the efficient data transmission of Kafka, and the dynamic configuration management and high-speed data access capabilities of distributed configuration centers (such as Apache Apollo) and distributed caches (such as Redis), and can automatically adjust data processing logic to adapt to different data sources and business needs. At the same time, by storing policy calculation results in a distributed cache in real time, it ensures that external services can immediately access the latest calculation results, thereby significantly improving the real-time and efficiency of decision support while ensuring data processing accuracy and consistency, thereby solving the problem of low efficiency of policy calculation for multi-source heterogeneous data streams in the existing technology.
[0110] In the specific implementation process, the above-mentioned strategy calculation unit includes a first determination module, a first acquisition module, an aggregation calculation module and a strategy calculation module. The first determination module is used to determine that the above-mentioned target calculation strategy includes a window aggregation calculation strategy when there is an aggregation calculation demand for the above-mentioned standardized feature data stream, and the above-mentioned window aggregation calculation strategy includes a window aggregation sub-strategy and a calculation sub-strategy; the first acquisition module is used to obtain external aggregation data related to the above-mentioned standardized feature data stream from the above-mentioned distributed cache; the aggregation calculation module is used to apply the above-mentioned window aggregation sub-strategy to aggregate the above-mentioned external aggregation data and the above-mentioned standardized feature data stream to obtain an aggregation calculation result; the strategy calculation module is used to apply the above-mentioned calculation sub-strategy to perform the above-mentioned strategy calculation on the above-mentioned aggregation calculation result to obtain the above-mentioned strategy calculation result.
[0111] This process not only enables aggregate calculations on standardized feature data streams, but also enables the integration of external aggregate data for more comprehensive policy calculations, resulting in refined policy calculation results. The results are stored in a distributed cache, ensuring that external services have real-time access to the latest aggregate calculation information, supporting real-time decision-making and analysis needs.
[0112] Furthermore, the above-mentioned aggregation calculation module includes an extraction submodule, a first determination submodule, a second determination submodule, and an aggregation calculation submodule. Among them, the extraction submodule is used to extract the target aggregation data from the above-mentioned external aggregation data according to the data flow expression filtering configuration in the above-mentioned window aggregation sub-strategy; the first determination submodule is used to determine the aggregation time window according to the window aggregation time configuration in the above-mentioned window aggregation sub-strategy; the second determination submodule is used to determine the aggregation dimension according to the window aggregation latitude configuration in the above-mentioned window aggregation sub-strategy, and the above-mentioned aggregation dimension includes the user dimension, time dimension and business indicator dimension; the aggregation calculation submodule is used to perform the above-mentioned aggregation calculation on the above-mentioned target aggregation data and the above-mentioned standardized feature data stream according to the above-mentioned aggregation time window and the above-mentioned aggregation dimension, and obtain the above-mentioned aggregation calculation result, wherein the above-mentioned aggregation calculation includes at least one of the following: average value calculation, sum calculation, maximum value calculation and minimum value calculation, and the above-mentioned aggregation calculation result is the result of aggregating the above-mentioned target aggregation data and the above-mentioned standardized feature data stream within the above-mentioned time window according to the above-mentioned aggregation dimension.
[0113] By configuring data stream expression filters, it's possible to extract data subsets tailored to specific business needs from massive amounts of external aggregated data, avoiding the processing of irrelevant data and improving the relevance and efficiency of calculations. Dynamic configuration of the aggregation time window allows for flexible adjustment of the data analysis timeframe based on different business scenarios, effectively supporting both short-term, immediate monitoring and long-term trend analysis. Setting aggregation dimensions allows for multi-dimensional data analysis, providing insights into business performance from diverse perspectives, including detailed data analysis results at the user, time, and business metric levels. Aggregate calculation results are generated and stored in real-time in a distributed cache, ensuring that external services can instantly access and apply these results, supporting real-time decision-making and analysis. In summary, this process enables efficient, dynamic, and accurate aggregation calculations of data streams, providing powerful data support for complex business scenarios while also improving the real-time and accuracy of decision-making.
[0114] To improve the efficiency and effectiveness of data processing, in some embodiments of the present application, the second acquisition unit includes a parsing module, a selection module, and a second determination module. The parsing module is configured to parse the identification field of the multi-source heterogeneous data stream; the selection module is configured to select, from the distributed configuration center, a feature extraction rule mapped to the identification field based on the field mapping rule in the preset feature extraction rule and the identification field of the multi-source heterogeneous data stream; and the second determination module is configured to determine the feature extraction rule as the target feature extraction rule.
[0115] By parsing the identification fields of the data stream and matching the feature extraction rules according to the field mapping rules, the source and type of the data stream can be automatically identified, thereby selecting the most appropriate preprocessing method, avoiding manual intervention, and improving the automation level of the processing flow. The acquisition of target feature extraction rules is performed dynamically at runtime, which means that even if the data source or business scenario changes, it can be quickly adjusted and new rules can be selected for data conversion, enhancing flexibility and adaptability. In short, the above process realizes the automated and standardized preprocessing of multi-source heterogeneous data streams, providing input data with a unified format, rich content and clear structure for subsequent policy calculations, facilitating subsequent policy calculations, and also providing a standardized method for data storage and querying, improving the efficiency and effectiveness of data processing.
[0116] In some other embodiments of the present application, each of the above-mentioned preset computing strategies includes a preset identifier, and the above-mentioned strategy computing unit includes a second acquisition module, a matching module, and a third determination module. The second acquisition module is configured to acquire a data stream identifier of the above-mentioned standardized feature data stream, wherein the data stream identifier corresponds one-to-one with the above-mentioned standardized feature data stream; the matching module is configured to match the above-mentioned preset identifier corresponding to the above-mentioned data stream identifier from the above-mentioned preset computing strategies; and the third determination module is configured to determine the above-mentioned preset computing strategy corresponding to the above-mentioned preset identifier as the above-mentioned target computing strategy.
[0117] By matching data stream identifiers with pre-set identifiers, the most relevant computation strategy for the current data stream is automatically selected, preventing policy inefficiency or misuse, and improving the accuracy and specificity of data processing. Pre-set computation strategies are stored in a distributed configuration center and retrieved through identifier matching. This means computation strategies can be dynamically updated and applied in real time without modifying the Flink data processing task code, enhancing flexibility and maintainability. Automatically matching and determining the target computation strategy eliminates the need for in-depth understanding of data stream details or the processing logic of Flink data processing tasks, simplifying the policy computation process, reducing its complexity and entry level, and improving its efficiency. Furthermore, the target computation strategy is selected based on the real-time characteristics of the data stream, ensuring the timeliness of computation results and supporting real-time data analysis and decision support scenarios, such as real-time transaction monitoring and real-time analysis of IoT device status. This process not only enables precise processing of standardized feature data streams, but also improves the real-time and flexibility of data processing through dynamic policy matching and application.
[0118] In order to ensure the immediacy and availability of the results and provide a real-time data source for the downstream system, the above-mentioned device also includes an output unit, which is used to obtain the target calculation strategy corresponding to the above-mentioned standardized feature data stream from the above-mentioned distributed configuration center, and perform strategy calculation on the above-mentioned standardized feature data stream according to the above-mentioned target calculation strategy. After obtaining the strategy calculation result, the above-mentioned strategy calculation result is output to the above-mentioned Kafka message queue so that the downstream system receives the above-mentioned strategy calculation result, wherein the above-mentioned downstream system includes at least one of the following: a database, a real-time analysis system and a data processing system.
[0119] Outputting the policy calculation results to the Kafka message queue enables downstream systems to receive and process these results in real time, providing direct data support for real-time decision-making and analysis, and shortening the time delay from data processing to decision execution.
[0120] In some further embodiments of the present application, the first acquisition unit includes obtaining the preset feature extraction rules and the preset calculation strategy via a RESTful service. Obtaining the preset feature extraction rules and the preset calculation strategy via a RESTful service makes rule management more flexible and efficient. This dynamic configuration mechanism ensures that rules can respond to changes in business needs in a timely manner, reducing the complexity and risk of rule updates.
[0121] The multi-source heterogeneous data stream is an established data stream or a newly added data stream, and the second acquisition unit includes: a first conversion module and a second conversion module. The first conversion module is used to, if the multi-source heterogeneous data stream is the established data stream, maintain its original nested data format unchanged and convert the established data stream into a JSON data stream format through a stream conversion function. The established data stream is a data stream that has already been defined and processed in the system; the second conversion module is used to, if the multi-source heterogeneous data stream is the newly added data stream, convert the newly added data stream into a unified and standardized data stream format. The newly added data stream is a data stream that is introduced into the system for the first time, or whose structure and format are different from the established data stream.
[0122] The above process ensures consistent processing logic and rule application across all data streams, reducing variability and improving accuracy and efficiency. This mechanism enables efficient and consistent processing of heterogeneous data streams from multiple sources, providing a solid data foundation for real-time data analysis and policy calculation, while also improving the system's overall operational efficiency and business adaptability.
[0123] The Flink-based multi-source heterogeneous data stream processing device includes a processor and memory. The first acquisition unit, second acquisition unit, policy calculation unit, etc. are all stored as program units in the memory. The processor executes these program units stored in the memory to implement the corresponding functions. The above modules are all located in the same processor; alternatively, the above modules can be located in different processors in any combination.
[0124] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.
[0125] An embodiment of the present invention provides a computer-readable storage medium, which includes a stored program. When the program is run, the device containing the computer-readable storage medium is controlled to execute the Flink-based multi-source heterogeneous data stream processing method.
[0126] An embodiment of the present invention provides a processor, which is used to run a program, wherein the program executes the Flink-based multi-source heterogeneous data stream processing method when running.
[0127] An embodiment of the present invention provides a device comprising a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps of the Flink-based multi-source heterogeneous data stream processing method. The device herein can be a server, a PC, a PAD, a mobile phone, or the like.
[0128] The present application also provides a computer program product, which, when executed on a data processing device, is suitable for executing a program that initializes the steps of the above-mentioned Flink-based multi-source heterogeneous data stream processing method.
[0129] Obviously, those skilled in the art will appreciate that the various modules or steps of the present invention described above can be implemented using a general-purpose computing device, can be centralized on a single computing device, or can be distributed across a network of multiple computing devices. They can be implemented using program code executable by the computing device, and thus, can be stored in a storage device and executed by the computing device. In some cases, the steps shown or described herein can be performed in a different order than that shown, or can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.
[0130] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0131] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1A device that provides the functions specified in a block or multiple blocks.
[0132] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0133] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0134] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0135] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0136] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0137] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0138] The above description is merely a preferred embodiment of the present application and is not intended to limit the present application. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present application shall be included within the scope of protection of the present application.
Claims
1. A method for processing multi-source heterogeneous data streams based on Flink, characterized in that: include: Obtaining preset feature extraction rules and preset calculation strategies, and storing the preset feature extraction rules and the preset calculation strategies in a distributed configuration center, wherein the preset feature extraction rules include at least one of the following: field mapping rules, data type conversion rules, data cleaning rules, and standardized format rules; and the preset calculation strategies include at least one of the following: summation strategy, averaging strategy, and statistical distribution strategy; Apply Flink's Kafka connector to obtain multi-source heterogeneous data streams from the Kafka message queue, obtain target feature extraction rules corresponding to the multi-source heterogeneous data streams from the distributed configuration center, and convert the multi-source heterogeneous data streams into standardized feature data streams according to the target feature extraction rules, wherein the target feature extraction rules are subordinate to the multiple preset feature extraction rules; Obtain a target computing strategy corresponding to the standardized feature data stream from the distributed configuration center, and perform strategy calculation on the standardized feature data stream according to the target computing strategy to obtain a strategy calculation result, and store the strategy calculation result in a distributed cache so that an external service can obtain the strategy calculation result from the distributed cache, wherein the target computing strategy belongs to a plurality of preset computing strategies.
2. The method according to claim 1, characterized in that According to the target calculation strategy, the strategy calculation is performed on the standardized feature data stream to obtain a strategy calculation result, including: In a case where there is an aggregation computing requirement for the standardized feature data stream, determining that the target computing strategy includes a window aggregation computing strategy, wherein the window aggregation computing strategy includes a window aggregation sub-strategy and a computing sub-strategy; Obtaining external aggregated data related to the standardized feature data stream from the distributed cache; Applying the window aggregation sub-strategy to aggregate and calculate the external aggregate data and the standardized feature data stream to obtain an aggregate calculation result; Apply the calculation sub-strategy to perform the policy calculation on the aggregate calculation result to obtain the policy calculation result.
3. The method according to claim 2, characterized in that Applying the window aggregation sub-strategy to aggregate and calculate the external aggregate data and the standardized feature data stream to obtain an aggregate calculation result, including: Extracting target aggregate data from the external aggregate data according to the data flow expression filtering configuration in the window aggregation sub-strategy; Determine the aggregation time window according to the window aggregation time configuration in the window aggregation sub-strategy; Determine aggregation dimensions according to the window aggregation latitude configuration in the window aggregation sub-strategy, where the aggregation dimensions include user dimension, time dimension, and business indicator dimension; According to the aggregation time window and the aggregation dimension, the aggregation calculation is performed on the target aggregation data and the standardized feature data stream to obtain the aggregation calculation result, wherein the aggregation calculation includes at least one of the following: average value calculation, sum calculation, maximum value calculation and minimum value calculation, and the aggregation calculation result is the result of aggregating the target aggregation data and the standardized feature data stream within the time window according to the aggregation dimension.
4. The method according to claim 1, wherein Obtaining target feature extraction rules corresponding to the multi-source heterogeneous data stream from the distributed configuration center includes: Parsing the identification field of the multi-source heterogeneous data stream; According to the field mapping rule in the preset feature extraction rule and the identification field of the multi-source heterogeneous data stream, selecting a feature extraction rule mapped to the identification field from the distributed configuration center; The feature extraction rule is determined as the target feature extraction rule.
5. The method according to claim 1, wherein Each of the preset computing strategies includes a preset identifier, and obtaining a target computing strategy corresponding to the standardized feature data stream from the distributed configuration center includes: Acquire a data stream identifier of the standardized feature data stream, wherein the data stream identifier corresponds one-to-one to the standardized feature data stream; Matching the preset identifier corresponding to the data flow identifier from the preset calculation strategy; The preset computing strategy corresponding to the preset identifier is determined as the target computing strategy.
6. The method according to claim 1, characterized in that After obtaining a target calculation strategy corresponding to the standardized feature data stream from the distributed configuration center and performing strategy calculation on the standardized feature data stream according to the target calculation strategy to obtain a strategy calculation result, the method further includes: outputting the strategy calculation result to the Kafka message queue so that a downstream system receives the strategy calculation result, wherein the downstream system includes at least one of the following: a database, a real-time analysis system, and a data processing system; Obtaining preset feature extraction rules and preset calculation strategies, including: obtaining the preset feature extraction rules and the preset calculation strategies through a restful service.
7. The method according to claim 1, characterized in that The multi-source heterogeneous data stream is an existing data stream or a new data stream. The Kafka connector of Flink is used to obtain the multi-source heterogeneous data stream from the Kafka message queue, including: If the multi-source heterogeneous data stream is the established data stream, its original nested data format is kept unchanged, and the established data stream is converted into the JSON data stream format through the stream conversion function. The established data stream is a data stream that has been defined and processed in the system; If the multi-source heterogeneous data stream is the newly added data stream, the newly added data stream is converted into a unified and standardized data stream format. The newly added data stream is introduced into the system for the first time, or its structure and format are different from the established data stream.
8. A multi-source heterogeneous data stream processing device based on Flink, characterized in that: include: A first acquisition unit is configured to acquire a preset feature extraction rule and a preset calculation strategy, and store the preset feature extraction rule and the preset calculation strategy in a distributed configuration center, wherein the preset feature extraction rule includes at least one of the following: a field mapping rule, a data type conversion rule, a data cleaning rule, and a standardized format rule; and the preset calculation strategy includes at least one of the following: a summation strategy, an averaging strategy, and a statistical distribution strategy; a second acquisition unit, configured to apply Flink's Kafka connector to acquire a multi-source heterogeneous data stream from a Kafka message queue, acquire a target feature extraction rule corresponding to the multi-source heterogeneous data stream from the distributed configuration center, and convert the multi-source heterogeneous data stream into a standardized feature data stream according to the target feature extraction rule, wherein the target feature extraction rule is subordinate to the plurality of preset feature extraction rules; A policy calculation unit is used to obtain a target calculation policy corresponding to the standardized feature data stream from the distributed configuration center, and perform policy calculation on the standardized feature data stream according to the target calculation policy to obtain a policy calculation result, and store the policy calculation result in a distributed cache so that an external service can obtain the policy calculation result from the distributed cache, wherein the target calculation policy belongs to a plurality of preset calculation policies.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the Flink-based multi-source heterogeneous data stream processing method according to any one of claims 1 to 7.
10. An electronic device, characterized in that: include: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include a method for executing the Flink-based multi-source heterogeneous data stream processing method according to any one of claims 1 to 7.