Data processing method and device, electronic equipment, storage medium and program product

By introducing a broadcast state mechanism and a configuration sensing source module into the stream processing system, the data processing configuration is dynamically updated, solving the downtime problem of the stream processing system when the configuration changes, and realizing dynamic updates and efficient operation and maintenance of the stream processing system.

CN120705157APending Publication Date: 2025-09-26BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510788251.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing stream processing systems require stopping and restarting when data processing configurations change, and cannot achieve dynamic updates.

Method used

By introducing a broadcast state mechanism into the stream processing system, the configuration awareness source module dynamically obtains the data processing configuration in the target storage space and broadcasts it to the computing instance, thereby realizing the dynamic updating of the data processing configuration.

Benefits of technology

The system can dynamically update data processing configurations without stopping the stream processing system, improving system availability and business continuity, reducing maintenance time, lowering error risk, and enhancing system scalability and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705157A_ABST
    Figure CN120705157A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method and device, electronic equipment, a storage medium and a program product, and relates to the technical field of data processing.The method comprises the steps that first data processing configuration in a target storage space is obtained, and the target storage space is configured to be data processing configuration for storing data streams; generating broadcast information based on the first data processing configuration, and broadcasting the broadcast information to at least one calculation instance, the at least one calculation instance being configured to update locally stored data processing configuration based on the broadcast information to obtain a second data processing configuration of the calculation instance; and in response to the received first data stream, processing the first data stream based on the at least one calculation instance to obtain a data processing result, and the at least one calculation instance is further configured to process the allocated data based on the second data processing configuration. According to the invention, the dynamic updating problem of the data processing configuration of the stream processing system can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technology, and in particular to a data processing method, device, electronic device, storage medium, and program product. Background Art

[0002] Currently, stream processing systems statically load data processing configurations during job initialization. Subsequently, data streams are processed according to these static configurations to produce data processing results. However, adjusting the data processing configuration with this data processing approach requires stopping and restarting the stream processing system. Therefore, dynamically updating the data processing configuration in stream processing systems has become an urgent problem.

[0003] Public content

[0004] In view of this, the present disclosure provides a data processing method, apparatus, electronic device, storage medium, and program product to solve the problem of dynamic update of data processing configuration of a stream processing system.

[0005] In a first aspect, the present disclosure provides a data processing method, the method comprising:

[0006] Obtaining a first data processing configuration in a target storage space, wherein the target storage space is configured to store a data processing configuration of a data stream;

[0007] Generate broadcast information based on the first data processing configuration, and broadcast the broadcast information to at least one computing instance, wherein the at least one computing instance is configured to update a locally stored data processing configuration based on the broadcast information to obtain a second data processing configuration of the computing instance;

[0008] In response to the received first data stream, the first data stream is processed based on the at least one computing instance to obtain a data processing result. The at least one computing instance is also configured to process the allocated data based on the second data processing configuration.

[0009] In a second aspect, the present disclosure provides a data processing device, the device comprising:

[0010] a data acquisition module, configured to acquire a first data processing configuration in a target storage space, wherein the target storage space is configured to store a data processing configuration of a data stream;

[0011] a data broadcast module, configured to generate broadcast information based on the first data processing configuration, and broadcast the broadcast information to at least one computing instance, wherein the at least one computing instance is configured to update a locally stored data processing configuration based on the broadcast information to obtain a second data processing configuration of the computing instance;

[0012] A data processing module is used to process the first data stream received based on the at least one computing instance to obtain a data processing result, and the at least one computing instance is also configured to process the allocated data based on the second data processing configuration.

[0013] In a third aspect, the present disclosure provides an electronic device comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the data processing method of the first aspect or any corresponding embodiment thereof by executing the computer instructions.

[0014] In a fourth aspect, the present disclosure provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the data processing method of the first aspect or any corresponding embodiment thereof.

[0015] In a fifth aspect, the present disclosure provides a computer program product, comprising computer instructions for causing a computer to execute the data processing method of the first aspect or any corresponding embodiment thereof.

[0016] The data processing method provided by the embodiment of the present disclosure is that the stream processing system obtains the first data processing configuration from the target storage space, and then broadcasts the first data processing configuration to the computing instance in the stream processing system in a broadcast manner. Therefore, when the data processing configuration in the target storage space changes, the stream processing system can synchronize the changed data processing configuration to the computing instance. Furthermore, when the stream processing system receives the first data stream, it can process the first data stream through the computing instance and the effective second data processing configuration to obtain a data processing result. Since the computing instance processes the allocated data based on the second data processing configuration of the local cache, when the data processing configuration changes, there is no need to stop and restart the stream processing system, thereby enabling dynamic update of the data processing configuration of the stream processing system.

[0017] The beneficial effects of the data processing device, electronic device, storage medium, and program product correspond to the beneficial effects of the data processing method and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0019] Figure 1 is a schematic diagram of an optional application scenario according to an embodiment of the present disclosure;

[0020] Figure 2 is a flow chart of a data processing method according to an embodiment of the present disclosure;

[0021] Figure 3 is a schematic diagram of a data processing system according to an embodiment of the present disclosure;

[0022] Figure 4 is a schematic diagram of a framework of another stream processing system according to an embodiment of the present disclosure;

[0023] Figure 5 is a schematic diagram of Kafka lag accumulation and operation delay using the data processing method of the present disclosure;

[0024] Figure 6 This is a schematic diagram of Kafka Lag accumulation and operation delay without using the data processing method disclosed herein;

[0025] Figure 7 is a structural block diagram of a data processing device according to an embodiment of the present disclosure;

[0026] Figure 8 It is a structural block diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0027] To make the purpose, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of the present disclosure.

[0028] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0029] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.

[0030] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0031] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0032] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.

[0033] Currently, stream processing systems statically load data processing configurations during job initialization. Subsequently, they process the data stream based on these static configurations to produce the data processing results. While some stream processing systems can leverage external configuration services for parameter adjustments, these adjustments are often limited to simple parameters and are difficult to adapt to dynamic structural changes in the data processing pipeline.

[0034] The following are some of the main problems with stream processing systems in related technologies:

[0035] 1. Inflexible indicator management: Multi-scenario and multi-table operations such as Customer Relationship Management (CRM) make it difficult to integrate new indicators or verify task completion.

[0036] 2. Real-time indicator processing lacks dynamism: the calculation logic of related technologies is rigid, windows and scripts are difficult to adjust online, and changes require shutdown and restart.

[0037] 3. Data skew and high-cardinality aggregation bottlenecks: High-cardinality dimension keyBy operations can easily lead to uneven operator load, affecting throughput and latency. KeyBy operations are used to aggregate identical data into the same partition.

[0038] 4. Complex evolution of deduplication, rollback, and database object collections (Schema): Dynamic adjustments to multiple dimensions, such as deduplication strategies, aggregation dimensions, and target table structures, must be completed without interrupting tasks.

[0039] 5. Multiple aggregation types coexist: Multiple aggregation methods such as sum (SUM), total (COUNT), average (AVG), and count distinct (COUNT DISTINCT) must be supported simultaneously, and accuracy and consistency must be guaranteed.

[0040] In view of this, according to an embodiment of the present disclosure, a data processing method embodiment is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0041] As an optional application scenario of the embodiment of the present disclosure, Figure 1 As shown, the entire data processing scenario includes client 1, target storage space 2, and stream processing system 3. Client 1 is used to configure the data processing configuration of the data stream and send the configured data processing configuration to target storage space 2 for storage. Stream processing system 3 obtains the data processing configuration from target storage space 2 to process the received data stream.

[0042] Open-source distributed stream processing frameworks, such as Apache Flink, provide a fundamental mechanism called broadcast state. This embodiment leverages broadcast state to build an end-to-end stream processing system for dynamic task adjustment. Specifically, for complex configurations involving external storage and management, such as metric definitions and database object schema mappings, a specific architectural solution is developed to integrate external target storage, runtime metric processing result management within the stream processing system, and logic at the compute instance (e.g., operator) level.

[0043] In this embodiment, a data processing method is provided, which can be used in a stream processing system. Figure 2 is a flow chart of a data processing method according to an embodiment of the present disclosure, such as Figure 2 As shown, the process includes the following steps:

[0044] Step S201: Acquire a first data processing configuration in a target storage space, where the target storage space is configured to store a data processing configuration for a data stream.

[0045] The target storage space is an external configuration database that can be dynamically updated. The target storage space is configured to persistently store various configurations that define data processing behaviors, namely, data processing configurations.

[0046] Optionally, the first data processing configuration includes one or more of pre-declared data processing rules, data mapping patterns, and serializable code snippets. Among them, data processing rules can be defined in data formats such as JSON, YAML, XML, etc., and data processing rules include filtering conditions, indicator calculation expressions, regular expressions, etc. The data mapping pattern includes field mapping tables defined in one or more ways, such as field mapping tables defined by Avro, field mapping tables defined by Protobuf Schema, and customized field mapping tables. Serializable code snippets include one or more types of scripts, such as Groovy scripts, AviatorScript scripts, etc. In addition, other data processing configurations can be selected according to actual conditions, which are not limited here.

[0047] Optionally, the target storage space can adopt one or more of a relational database, a non-relational database (such as a NoSQL database), a distributed key-value storage area, and a preset configuration management center. Other storage methods can also be selected according to actual conditions, which are not limited here.

[0048] Specifically, a configuration-aware source module can be set up in the stream processing system to actively or passively detect configuration changes in the target storage space through Change Data Capture (CDC) or message queue notifications of the target storage space. Specifically, the configuration-aware source module can periodically poll or event-drivenly pull the latest data processing configuration from the target storage space to obtain the first data processing configuration.

[0049] Alternatively, for a stream processing system using the Flink framework, a Flink RichSourceFunction can be used to construct a configuration-aware source module to obtain the first data processing configuration in the target storage space. The configuration-aware module can be pluggable to accommodate different target storage spaces.

[0050] Step S202: Generate broadcast information based on the first data processing configuration, and broadcast the broadcast information to at least one computing instance. The at least one computing instance is configured to update the locally stored data processing configuration based on the broadcast information to obtain a second data processing configuration of the computing instance.

[0051] Specifically, the first data processing configuration is serialized into a unified configuration object or event stream through the configuration-aware source module, and the configuration object or event stream is converted into a broadcast stream to obtain broadcast information. The broadcast information is then broadcast to at least one computing instance using a broadcast operator (such as Flink's broadcast operator). The stream processing system's checkpoint mechanism ensures the consistency and fault tolerance of the broadcast state. Even after failure recovery, the broadcast operator can be restored to the correct configuration state to ensure the accuracy of the broadcast information.

[0052] A compute instance is an operator in a stream processing system that processes job logic. Upon receiving a broadcast message, the compute instance updates its cached copy of the data processing configuration and / or the broadcast state (BroadcastState) maintained by the stream processing system based on the broadcast message. For example, a compute instance updates its locally cached data processing configuration using the processBroadcastElement() method.

[0053] For example, the computing instance updates the rule engine, indicator calculator, field mapper, target schema definition, etc. in the local cache according to the first data processing configuration.

[0054] Step S203: In response to the received first data stream, the first data stream is processed based on at least one computing instance to obtain a data processing result. The at least one computing instance is further configured to process the allocated data based on a second data processing configuration.

[0055] Among them, the stream processing system can obtain the first data stream through operators such as "processElement()" and "onTimer()".

[0056] Optionally, the first data stream is input data (deduplicatedStream) of the stream processing system after basic deduplication.

[0057] Specifically, when the stream processing system receives the first data stream, the computing instance performs corresponding data processing tasks on the first data stream based on the currently effective second data processing configuration. For example, it can dynamically apply new data cleansing rules, recalculate derived indicators based on updated indicator formulas, dynamically adjust the structure and field population logic of the output row data (RowData) based on the changed schema, or dynamically switch data routing strategies.

[0058] The data processing method provided in this embodiment is that the stream processing system obtains the first data processing configuration from the target storage space, and then broadcasts the first data processing configuration to the computing instance in the stream processing system in a broadcast manner. Therefore, when the data processing configuration in the target storage space changes, the stream processing system can synchronize the changed data processing configuration to the computing instance. Furthermore, when the stream processing system receives the first data stream, it can process the first data stream through the computing instance and the effective second data processing configuration to obtain a data processing result. Since the computing instance processes the allocated data based on the second data processing configuration of the local cache, when the data processing configuration changes, there is no need to stop and restart the stream processing system, thereby enabling dynamic updating of the data processing configuration of the stream processing system.

[0059] In some optional implementations, the above step S201 includes: in response to a configuration change message of the target storage space, acquiring an updated data processing configuration in the target storage space to obtain a first data processing configuration.

[0060] It can be understood that when the stream processing system starts, the configuration-aware source module of the stream processing system loads the initial data processing configuration from the target storage and distributes the initial data processing configuration to the computing instance through the configuration broadcast operator. The computing instance processes the data stream based on the initial data processing configuration. When the data processing configuration in the target storage space changes, for example, the user modifies the calculation formula of an analysis indicator, the configuration-aware source module perceives the change in the data processing configuration in the target storage space through polling or message queue notification, and obtains the updated data processing configuration from the target storage space as the first data processing configuration. Then, the first data processing configuration is injected as a new configuration event into the broadcast stream related to the data processing configuration, so that the computing instance receives the broadcast information and updates the locally cached data processing configuration. Starting from the next data element or the next data processing cycle, the computing instance uses the new second data processing configuration to execute its job logic. The entire process does not require stopping or restarting the stream processing system's job, thereby realizing hot plugging of data processing configuration (i.e., replacing the target storage space and data processing configuration through the plug-in of the configuration-aware source module connected to the target storage space) and real-time adjustment of data processing configuration.

[0061] In the data processing method provided by this embodiment, the stream processing system responds to the configuration change information of the target storage space and automatically obtains the updated data processing configuration. Therefore, it can dynamically update the data processing configuration of the computing instance following the change of the data processing configuration of the target storage space to instantly adjust the data processing logic of the stream processing system.

[0062] In some optional implementations, at least one computing instance is configured to have at least one first computing instance and at least one second computing instance. The above step S203 includes:

[0063] Step a1: Generate a hash key for the data content in the first data stream.

[0064] Specifically, a dynamic key generation function is used to generate a hash key for data distribution based on attribute information of the data content and a random salt value, so that the data content of the first data stream is distributed as evenly as possible to the parallel computing instances using the hash key.

[0065] Step a2: Based on the hash key, the data content in the first data stream is assigned to at least one first computing instance, and the logical key, field information and target analysis indicators of the assigned data content are analyzed using the first computing instance to obtain a second data stream. The logical key is generated based on the attribute information of the data content; the attribute information includes indicator dimensions and deduplication information.

[0066] The deduplication information includes the deduplication dimension and the time dimension (truncatedTime), which can be adjusted according to actual conditions.

[0067] In practical applications, the ArithmeticProcessFunction of the stream processing system can be used as the first computing instance. When the first computing instance receives the assigned data content, it determines the dynamic expression of the expression evaluation engine (Aviator) through the local second data processing configuration, and uses the dynamic expression to perform preliminary calculations on the assigned data content, for example, calculating derived fields. Among them, the first data processing configuration can be divided into dataset configuration (FlinkDatasetConf) and indicator configuration (FlinkMetricConf) according to the configuration content. The dataset configuration includes the calculation field configuration, field type, field information, dataset dimension, time aggregation dimension, time aggregation granularity, deduplication information, etc. The indicator configuration includes indicator dimension, indicator type, and indicator formula, etc. The first computing instance dynamically loads the dataset configuration and indicator configuration in the second data processing configuration, queries the attribute information required to configure the logical key for each data content, and constructs a logical key for each assigned data content based on the attribute information. The logical key is used by the downstream second computing instance to deduplicate the field information and accurately aggregate the indicator processing results of the analysis indicator.

[0068] In the indicator configuration, indicator identifiers for different analysis indicators are configured for different indicator dimensions. The target analysis indicator to be executed by the downstream second computing instance can be determined based on the indicator dimensions of the data content. The fields and field values ​​in the data content are extracted to obtain field information in the form of a key and value map.

[0069] It can be understood that the second data stream includes the logical key, field information and target analysis indicators of the data content corresponding to the first computing instance.

[0070] Exemplarily, the second data stream Tuple1 may be output in the form of “<logical key, List<indicator identifier of analysis indicator>, Map<field, field value>>”.

[0071] Step a3: assign the second data stream to at least one second computing instance, and use the second computing instance to hierarchically aggregate the indicator processing results of the target analysis indicators in the assigned second data stream to obtain a data processing result.

[0072] Specifically, the second data stream can be assigned to a second computing instance according to the logical key in the second data stream or the indicator dimension extracted according to the logical key, and the second computing instance can be used to perform local aggregation / global aggregation on the indicator processing results of the target analysis indicators in the assigned second data stream to obtain the data processing results.

[0073] The data processing method provided in this embodiment distributes the data content of the first data stream to the parallel first computing instance using a hash key of the data content in the first data stream, thereby resolving the problem of uneven physical load in the stream processing system. Furthermore, the second data stream output by the first computing instance is assigned to the parallel second computing instance for hierarchical aggregation of the indicator processing results of the analysis indicators, thereby reducing the computational complexity of a single indicator processing result aggregation and improving data processing efficiency.

[0074] In some optional implementations, the above step a1 includes:

[0075] Step a11: Acquire attribute information of the data content and the target salt value of the data content.

[0076] The target salt value may be a salt value randomly generated based on the data content.

[0077] Optionally, the attribute information includes indicator dimensions and deduplication information, where the indicator dimensions represent the combination of actual dimensions of the analysis indicator corresponding to the data content (actualDimensionKey). The deduplication information includes the deduplication dimension and the time dimension (truncatedTime). The deduplication dimension can be represented by a precise deduplication business key (deduplicationKey).

[0078] Step a12: performing a hash operation on the attribute information and the target salt value to obtain a hash key of the data content.

[0079] Specifically, the attribute information and the target salt value are combined, and a hash operation is performed on the combined result to obtain a hash key of the data content.

[0080] The data processing method provided in this embodiment performs a hash operation on the attribute information of the data content and the target salt value to obtain the hash key of the data content. Therefore, the target salt value can be used to break up the hash keys of data content of the same type, thereby evenly distributing the data content to parallel computing instances, solving the problem of uneven physical load of the stream processing system.

[0081] In some optional implementations, the step a2 of analyzing the logical key, field information, and target analysis indicator of the allocated data content using the first computing instance to obtain the second data stream includes:

[0082] Step a21: Analyze the attribute information of the allocated data content using the first computing instance and its second data processing configuration to generate a logical key for the data content.

[0083] Specifically, the first computing instance searches for attribute information of the data content in the allocated data content according to the second data processing configuration, and generates a logical key of the data content according to the attribute information.

[0084] Specifically, different information in the attribute information is merged to obtain a logical key.

[0085] For example, the logical key rawkey = actualDimensionKey "#" + truncatedTime + "#" + deduplicationKey, where actualDimensionKey is the indicator dimension, truncatedTime is the time dimension in the deduplication information, and deduplicationKey is the deduplication dimension in the deduplication information.

[0086] Step a22: Analyze the analysis indicators associated with the indicator dimension using the first computing instance and its second data processing configuration to obtain target analysis indicators.

[0087] Specifically, the second data processing configuration stores indicator identifiers of analysis indicators under different indicator dimensions. The first computing instance determines the indicator identifiers of analysis indicators under the indicator dimension of the data content according to the second data processing configuration to obtain the target analysis indicator.

[0088] Step a23: Analyze the fields and the field values ​​of the fields in the allocated data content using the first computing instance and its second data processing configuration to obtain field information to obtain a second data stream.

[0089] Specifically, the first computing instance extracts fields and field values ​​in the data content according to the second data processing configuration to obtain field information.

[0090] The data processing method provided in this embodiment utilizes a first computing instance and its second data processing configuration to analyze attribute information of data content to obtain a logical key. Therefore, the logical key can be used to further segment data content with different attribute information. Furthermore, the first computing instance is used to analyze the target analysis indicators and field information corresponding to the data content. Therefore, the data content can be dynamically distributed based on the logical key, thereby improving data processing efficiency.

[0091] In some optional embodiments, at least one second computing instance is configured to have at least one third computing instance and at least one fourth computing instance. Step a3 includes:

[0092] Step a31: Allocate the second data stream to at least one third computing instance based on the logical key in the second data stream.

[0093] Optionally, LocalDedupAggFunction of the stream processing system is used as the third computing instance, and the third computing instance is used to maintain the local aggregation result corresponding to the logical key.

[0094] Specifically, the logical key in the second data stream output by the first computing instance is used as the basis of keyBy, and the second data stream is allocated to the parallel third computing instance.

[0095] Step a32: Use the third computing instance to locally aggregate the indicator processing results of the target analysis indicators in the assigned second data stream to obtain the local aggregation results of the target analysis indicators to obtain the third data stream. The third data stream includes the indicator dimensions of the logical key corresponding to the third computing instance and the local aggregation results of the target analysis indicators.

[0096] Specifically, the third computing instance performs local aggregation of the corresponding aggregation type on the indicator processing results of the target analysis indicator of different aggregation types based on the local second data processing configuration, thereby obtaining the local aggregation result of the target analysis indicator. The third computing instance can dynamically obtain the aggregation type of the target analysis indicator from the second data processing configuration using FlinkMetricConf.

[0097] For example, the third computing instance performs corresponding aggregations, such as addition and comparison, on the target analysis metrics of the SUN, MIN, and MAX types directly within the AggregationBundle of the stream processing system. For target analysis metrics of the COUNT, AVG, and COUNT_DISTINCT types, the third computing instance retains the deduplicated data or tags within the AggregationBundle as the local aggregation results.

[0098] Exemplarily, the third data stream Tuple2 may be output in the form of “<logical key, local aggregation result>”.

[0099] Step a33: assign the third data stream to at least one fourth computing instance, and use the fourth computing instance to globally aggregate the local aggregation results of the target analysis indicators in the assigned third data stream to obtain a data processing result.

[0100] Optionally, the AggMetricProcessFunction of the stream processing system is used as the fourth computing instance. The AggMetricProcessFunction is a keyedBroadcastProcessFunction. The fourth computing instance is used to maintain the global aggregation results for each indicator dimension. These global aggregation results can be designed for different aggregation types. For example, for the indicator dimension of the MapState aggregation type, MapState is used.<Long,MetricValue> Stores global aggregation results, indicator dimensions of ValueState aggregation type, using ValueState <Map<Long,HyperLogLogPlus> Store the global aggregation results (i.e. HLL objects).

[0101] The data processing method provided in this embodiment distributes the computational load of high-cardinality aggregation tasks to each computing instance of the stream processing system through hash keys and logical keys, significantly alleviating data skew. At the same time, the third computing instance is used to locally aggregate the indicator processing results of the target analysis indicators in the second data stream, and then the fourth computing instance is used to globally aggregate the local aggregation results in the third data stream output by the third computing instance. Therefore, the aggregation efficiency can be improved, and multiple aggregation types can be processed simultaneously, for example, the cardinality estimation of the algorithm for cardinality estimation (HyperLogLogPlus) to support complex and diverse indicator aggregation, thereby meeting the indicator requirements of complex business scenarios. Moreover, the phased indicator processing allows each stage to be independently optimized and expanded, and the overall architecture has better adaptability to changes in data volume and indicator complexity, which can enhance the scalability and robustness of the stream processing system.

[0102] In some optional implementations, the above step a32 includes:

[0103] Step a321: Use the third computing instance to extract the indicator dimension and deduplication information from the logical key of the allocated second data stream.

[0104] Specifically, within the third computing instance, the third computing instance parses the logical key in the second data stream to obtain indicator dimensions and deduplication information.

[0105] Step a322: Deduplication is performed on the field information in the allocated second data stream using the third computing instance, deduplication information, and the corresponding second data processing configuration to obtain deduplication-processed field information.

[0106] Specifically, the field information in the third data stream is screened according to the deduplication dimension, and the latest field value in the field information is retained according to the time dimension to obtain the field information after deduplication processing.

[0107] Step a323, using the third computing instance, the deduplicated field information and the corresponding second data processing configuration, locally aggregate the indicator processing results of the target analysis indicators in the allocated second data stream to obtain the local aggregation results of the target analysis indicators to obtain the third data stream.

[0108] Specifically, the third computing instance is used to determine the aggregation type of the target analysis indicator in the local second data processing configuration, and based on the field information after deduplication processing and the aggregation type of the target analysis indicator, the indicator processing results of the target analysis indicator are locally aggregated with the corresponding aggregation type to obtain the local aggregation results of the target analysis indicator.

[0109] Specifically, the third computing instance uses a timer (onTimer) and a batch mechanism to periodically (for example, every minute) output the aggregation results in the third data stream.

[0110] The data processing method provided in this embodiment first deduplicates the field information according to the deduplication information, and then uses the deduplicated field information to locally aggregate the indicator processing results of the target analysis indicator. Therefore, it can ensure that in a distributed environment, the field information of the same deduplication dimension is accurately processed once in the local aggregation of the target analysis indicator of the same logical key or is processed according to the latest timestamp, thereby improving the accuracy of local aggregation.

[0111] In some optional implementations, the above step a33 includes:

[0112] Step a331: Allocate the third data stream to at least one fourth computing instance based on the indicator dimension in the third data stream.

[0113] Optionally, the indicator dimension in the third data stream is used as the basis for keyBy and distributed to the parallel fourth computing instance.

[0114] Step a332, using the fourth computing instance to perform global aggregation on the local aggregation results of the same indicator dimension in the assigned third data stream, to obtain the first global aggregation result of the target analysis indicator and the second global aggregation result of the indicator dimension, to obtain the data processing result.

[0115] Specifically, the fourth computing instance receives the third data stream from different third computing instances, and performs aggregation on the local aggregation results of the same indicator dimension according to the local second data processing configuration to merge the local aggregation results of the same indicator dimension to obtain the data processing result.

[0116] For example, merge the indicator processing results of SUN type target analysis indicators, select MIN or MAX aggregation type for global aggregation, merge HLL cardinality estimation, calculate AVG of local aggregation results, etc.

[0117] Exemplarily, the fourth data stream Tuple3 may be output in the form of “<logical key, Map<indicator dimension, second global aggregation result>, Map<target analysis indicator, first global aggregation result>”.

[0118] The data processing method provided in this embodiment allocates the third data stream to a parallel fourth computing instance based on the indicator dimension in the third data stream, and globally aggregates the local aggregation results of the same indicator dimension. Therefore, through the pipeline operation of the three stages of the first computing instance, the third computing instance, and the fourth computing instance, the size of the network transmission and global aggregation data can be reduced, and the aggregation efficiency can be improved.

[0119] As a specific application example, see Figure 3 , the stream processing system using the data processing method disclosed herein can be configured as a target storage space, a configuration perception source module, and a dynamic processing module. Figure 4 The target storage space stores the data processing configuration for the data stream, such as dataset configuration and indicator configuration. When the configuration-aware source module detects a change in the dataset configuration, it synchronizes the latest dataset configuration to various types of operators and compute instances in the dynamic processing module through a first broadcast operator. When the configuration-aware source module detects a change in the indicator configuration, it synchronizes the latest indicator configuration to various types of operators and compute instances in the dynamic processing module through a second broadcast operator. After the stream processing system constructs a data processing task, it receives incoming information from the presentation middleware (Dorado). The data processing task initializes the query for dataset information. The source data operator of the stream processing system receives message queue messages, determines the data source type, initializes the data source's TCC (Try-Confirm-Cancel) configuration information, sets the initial consumption point to consume the data from the data source, and transmits the data to the input operator of the dynamic processing module in the form of a data stream. After receiving the data stream, the input operator of the dynamic processing module parses the data stream's field information according to the data processing configuration and adds a time dimension and a deduplication dimension to facilitate processing of the data content in the data stream according to these time and deduplication dimensions. The input operator sends the processed data stream to the deduplication operator, which performs basic deduplication on the data sent by the input operator based on the deduplication dimension and time dimension in the second data processing configuration, generating a first data stream. A hash key is then generated for the data content in the first data stream, and the data content in the first data stream is distributed to the first compute instance according to the hash key. The first compute instance imports indicator operations to calculate basic indicators and performs derived indicator operations using Aviator dynamic expressions. The indicator processing results are added to the first data stream for downstream aggregate indicator calculations. Simultaneously, the first compute instance uses Aviator to filter data that does not meet the preset aggregate indicator filter conditions, generating a second data stream. The dynamic processing module distributes the second data stream output by the first compute instance to a third compute instance based on the logical key in the second data stream, performing indicator aggregation of different aggregation types, generating a third data stream. The third data stream output by the third compute instance is then distributed to a fourth compute instance based on the indicator dimensions in the third data stream. The fourth compute instance performs global aggregation on the indicators based on the indicator dimensions, generating the data processing results.

[0120] The data processing method disclosed in this disclosure has the following key advantages: First, it supports incremental updates and parsing of nested objects, lists, script expressions, and the like in the broadcast state, enabling deep integration with complex external configurations. This allows operators and compute instances in the stream processing system to dynamically update data processing configurations, such as compute scripts, window parameters, and output schemas, during operation. Consequently, changes to business logic, indicator caliber, data planning, or target schemas can be made without interrupting the real-time tasks of the stream processing system, significantly improving the availability and business continuity of the stream processing system and enabling configuration updates with near-zero downtime. Second, the disclosure utilizes an automated data processing configuration update process, which can shorten data processing configuration update time based on the frequency of data processing configuration pulls, enabling businesses to quickly respond to external changes and improving operational efficiency and agility. Third, the decoupling of data processing configuration from code makes data processing configuration management more centralized and transparent, effectively reducing code maintenance complexity. Furthermore, new types of data processing configurations or processing logic can be added by extending the configuration model and compute instances, thereby enhancing the maintainability and scalability of the stream processing system. Fourth, dynamic, controlled updates to data processing configurations reduce the likelihood of errors introduced by manual operations, especially for complex data mapping and transformation rules, effectively reducing the risk of data errors. Fifth, the checkpoint and broadcast state mechanisms of the stream processing system ensure that the configuration state of data processing remains consistent even during data processing configuration updates or after failure recovery, thereby ensuring data processing consistency. Sixth, relevant personnel can directly and finely adjust the processing method of real-time data streams by modifying external configurations (i.e., the data processing configuration in the target storage space), without relying on the immediate intervention of the R&D team, thus achieving refined real-time intervention capabilities.

[0121] See also Figure 5 and Figure 6 , which are respectively the MySQL Lag accumulation and operation delay when the data processing method of the present invention is adopted and when the data processing method of the present invention is not adopted. It can be seen that after adopting the data processing method of the present invention, the MySQL Lag accumulation and operation delay are effectively solved.

[0122] In this embodiment, a data processing device is also provided, which is used to implement the above-mentioned embodiments and preferred embodiments. Details that have been described will not be repeated here. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.

[0123] This embodiment provides a data processing device, such as Figure 7 Shown, including:

[0124] A data acquisition module 701 is configured to acquire a first data processing configuration in a target storage space, where the target storage space is configured to store a data processing configuration for a data stream;

[0125] a data broadcast module 702 configured to generate broadcast information based on the first data processing configuration and broadcast the broadcast information to at least one computing instance, wherein the at least one computing instance is configured to update a locally stored data processing configuration based on the broadcast information to obtain a second data processing configuration of the computing instance;

[0126] The data processing module 703 is used to process the first data stream based on at least one computing instance in response to the received first data stream to obtain a data processing result. The at least one computing instance is also configured to process the allocated data based on a second data processing configuration.

[0127] In some optional implementations, the data acquisition module 701 includes:

[0128] The data acquisition unit is configured to acquire the updated data processing configuration in the target storage space in response to the configuration change message of the target storage space, so as to obtain the first data processing configuration.

[0129] In some optional implementations, at least one computing instance is configured to have at least one first computing instance and at least one second computing instance. The data processing module 703 includes:

[0130] a first processing unit, configured to generate a hash key for data content in the first data stream;

[0131] a second processing unit, configured to assign data content in the first data stream to at least one first computing instance based on a hash key, and analyze, using the first computing instance, a logical key, field information, and target analysis indicators of the assigned data content to obtain a second data stream, wherein the logical key is generated based on attribute information of the data content; the attribute information includes indicator dimensions and deduplication information;

[0132] The third processing unit is used to assign the second data stream to at least one second computing instance, and use the second computing instance to hierarchically aggregate the indicator processing results of the target analysis indicators in the assigned second data stream to obtain a data processing result.

[0133] In some optional implementations, the first processing unit includes:

[0134] An information acquisition subunit, used to obtain attribute information of the data content and a target salt value of the data content;

[0135] The hash operation subunit is used to perform a hash operation on the attribute information and the target salt value to obtain the hash key of the data content.

[0136] In some optional embodiments, the second processing unit includes:

[0137] a first processing subunit, configured to analyze attribute information of the allocated data content using the first computing instance and its second data processing configuration to generate a logical key of the data content;

[0138] A second processing subunit is configured to analyze the analysis indicator associated with the indicator dimension using the first computing instance and its second data processing configuration to obtain a target analysis indicator;

[0139] The third processing subunit is used to analyze the fields and field values ​​in the allocated data content using the first computing instance and its second data processing configuration to obtain field information to obtain a second data stream.

[0140] In some optional embodiments, at least one second computing instance is configured to have at least one third computing instance and at least one fourth computing instance. The third processing unit includes:

[0141] a fourth processing subunit, configured to assign the second data stream to at least one third computing instance based on the logical key in the second data stream;

[0142] a fifth processing subunit, configured to locally aggregate the indicator processing results of the target analysis indicators in the allocated second data stream using the third computing instance to obtain the locally aggregated results of the target analysis indicators, so as to obtain a third data stream, the third data stream including the indicator dimension of the logical key corresponding to the third computing instance and the locally aggregated results of the target analysis indicators;

[0143] The sixth processing sub-unit is used to assign the third data stream to at least one fourth computing instance, and use the fourth computing instance to globally aggregate the local aggregation results of the target analysis indicators in the assigned third data stream to obtain a data processing result.

[0144] In some optional embodiments, the fifth processing sub-unit is specifically used to: use the third computing instance to extract indicator dimensions and deduplication information in the logical key of the assigned second data stream; use the third computing instance, deduplication information and the corresponding second data processing configuration to deduplicate the field information in the assigned second data stream to obtain the deduplicated field information; use the third computing instance, the deduplication field information and the corresponding second data processing configuration to locally aggregate the indicator processing results of the target analysis indicators in the assigned second data stream to obtain the local aggregation results of the target analysis indicators to obtain the third data stream.

[0145] In some optional embodiments, the sixth processing sub-unit is specifically used to: assign the third data stream to at least one fourth computing instance based on the indicator dimension in the third data stream; use the fourth computing instance to globally aggregate the local aggregation results of the same indicator dimension in the assigned third data stream to obtain a first global aggregation result of the target analysis indicator and a second global aggregation result of the indicator dimension to obtain a data processing result.

[0146] The data processing device provided by the embodiment of the present disclosure can execute the data processing method provided by any embodiment of the present disclosure, and has the functional modules and beneficial effects corresponding to the execution method. The data processing device of the present disclosure obtains the first data processing configuration from the target storage space by the stream processing system, and then broadcasts the first data processing configuration to the computing instance in the stream processing system in a broadcast manner. Therefore, when the data processing configuration in the target storage space changes, the stream processing system can synchronize the changed data processing configuration to the computing instance. Furthermore, when the stream processing system receives the first data stream, it can process the first data stream through the computing instance and the effective second data processing configuration to obtain a data processing result. Since the computing instance processes the allocated data based on the second data processing configuration of the local cache, when the data processing configuration changes, there is no need to stop and restart the stream processing system, thereby enabling the dynamic update of the data processing configuration of the stream processing system.

[0147] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0148] Figure 8 This is a structural block diagram of an electronic device provided in an embodiment of the present disclosure.

[0149] The following specific reference Figure 8 , which shows a block diagram of the structure of the electronic device suitable for implementing the embodiments of the present disclosure. The electronic device may include a processor (e.g., a central processing unit, a graphics processing unit, etc.) 801, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 802 or the program loaded from the memory 808 into the random access memory (RAM) 803. Various programs and data required for the operation of the electronic device are also stored in the RAM 803. The processor 801, ROM 802, and RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0150] Typically, the following devices may be connected to the I / O interface 805: an input device 806 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 807 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 808 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 809. The communication device 809 may allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Although Figure 8 An electronic device having various devices is shown, but it should be understood that it is not required to implement or possess all of the devices shown, and more or fewer devices may be implemented or possessed instead.

[0151] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 809, or installed from the memory 808, or installed from the ROM 802. When the computer program is executed by the processor 801, the above-mentioned functions defined in the data processing method of the embodiment of the present disclosure are performed.

[0152] Figure 8 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0153] The embodiments of the present disclosure also provide a computer-readable storage medium. The above-mentioned method according to the embodiments of the present disclosure can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the data processing method shown in the above embodiment is implemented.

[0154] A portion of the present disclosure may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present disclosure through the operation of the computer. Those skilled in the art should understand that the form in which the computer program instruction exists in a computer-readable medium includes but is not limited to a source file, an executable file, an installation package file, etc. Accordingly, the way in which the computer program instruction is executed by the computer includes but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium that can be accessed by the computer.

[0155] Although the embodiments of the present disclosure have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A data processing method, characterized in that: The method comprises: Obtaining a first data processing configuration in a target storage space, wherein the target storage space is configured to store a data processing configuration of a data stream; Generate broadcast information based on the first data processing configuration, and broadcast the broadcast information to at least one computing instance, wherein the at least one computing instance is configured to update a locally stored data processing configuration based on the broadcast information to obtain a second data processing configuration of the computing instance; In response to the received first data stream, the first data stream is processed based on the at least one computing instance to obtain a data processing result. The at least one computing instance is also configured to process the allocated data based on the second data processing configuration.

2. The method according to claim 1, characterized in that The obtaining of the first data processing configuration in the target storage space includes: In response to the configuration change message of the target storage space, the updated data processing configuration in the target storage space is acquired to obtain the first data processing configuration.

3. The method according to claim 1, characterized in that The at least one computing instance is configured to have at least one first computing instance and at least one second computing instance; and the processing of the first data stream based on the at least one computing instance to obtain a data processing result includes: generating a hash key for the data content in the first data stream; Assigning data content in the first data stream to the at least one first computing instance based on the hash key, and analyzing the logical key, field information, and target analysis indicators of the assigned data content using the first computing instance to obtain a second data stream, wherein the logical key is generated based on attribute information of the data content; the attribute information includes indicator dimensions and deduplication information; The second data stream is assigned to the at least one second computing instance, and the second computing instance is used to hierarchically aggregate the indicator processing results of the target analysis indicators in the assigned second data stream to obtain the data processing results.

4. The method according to claim 3, characterized in that Generating a hash key of the data content in the first data stream includes: Obtaining attribute information of the data content and a target salt value of the data content; A hash operation is performed on the attribute information and the target salt value to obtain a hash key of the data content.

5. The method according to claim 3, characterized in that The using the first computing instance to analyze the assigned logical key, field information, and target analysis indicator of the data content to obtain a second data stream includes: Analyzing attribute information of the assigned data content using the first computing instance and the second data processing configuration thereof to generate a logical key for the data content; Analyzing the analysis indicator associated with the indicator dimension using the first computing instance and the second data processing configuration thereof to obtain the target analysis indicator; The first computing instance and the second data processing configuration thereof are used to analyze the fields in the allocated data content and the field values ​​of the fields to obtain the field information, so as to obtain the second data stream.

6. The method according to claim 3, characterized in that The at least one second computing instance is configured to have at least one third computing instance and at least one fourth computing instance; the assigning the second data stream to the at least one second computing instance, and using the second computing instance to hierarchically aggregate the indicator processing results of the target analysis indicator in the assigned second data stream to obtain the data processing result, including: assigning the second data stream to the at least one third computing instance based on the logical key in the second data stream; Using the third computing instance to locally aggregate the indicator processing results of the target analysis indicator in the second data stream allocated thereto, to obtain the local aggregation results of the target analysis indicator, so as to obtain a third data stream, the third data stream including the indicator dimension of the logical key corresponding to the third computing instance and the local aggregation results of the target analysis indicator; The third data stream is assigned to the at least one fourth computing instance, and the fourth computing instance is used to globally aggregate the local aggregation results of the target analysis indicators in the assigned third data stream to obtain the data processing result.

7. The method according to claim 6, characterized in that The using the third computing instance to locally aggregate the indicator processing results of the target analysis indicator in the allocated second data stream to obtain the local aggregation results of the target analysis indicator to obtain a third data stream includes: extracting the indicator dimension and the deduplication information from the assigned logical key of the second data stream using the third computing instance; Using the third computing instance, the deduplication information, and the corresponding second data processing configuration, perform deduplication processing on the field information in the allocated second data stream to obtain deduplication processed field information; Using the third computing instance, the deduplicated field information and the corresponding second data processing configuration, the indicator processing results of the target analysis indicators in the assigned second data stream are locally aggregated to obtain the local aggregation results of the target analysis indicators to obtain the third data stream.

8. The method according to claim 6, characterized in that The assigning the third data stream to the at least one fourth computing instance, and using the fourth computing instance to globally aggregate the local aggregation results of the target analysis indicator in the assigned third data stream to obtain the data processing result, includes: Allocating the third data stream to the at least one fourth computing instance based on the indicator dimension in the third data stream; The fourth computing instance is used to globally aggregate the local aggregation results of the same indicator dimension in the third data stream assigned thereto, to obtain a first global aggregation result of the target analysis indicator and a second global aggregation result of the indicator dimension, so as to obtain the data processing result.

9. A data processing device, characterized in that: The device comprises: a data acquisition module, configured to acquire a first data processing configuration in a target storage space, wherein the target storage space is configured to store a data processing configuration of a data stream; a data broadcast module, configured to generate broadcast information based on the first data processing configuration, and broadcast the broadcast information to at least one computing instance, wherein the at least one computing instance is configured to update a locally stored data processing configuration based on the broadcast information to obtain a second data processing configuration of the computing instance; A data processing module is used to process the first data stream received based on the at least one computing instance to obtain a data processing result, and the at least one computing instance is also configured to process the allocated data based on the second data processing configuration.

10. An electronic device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the data processing method according to any one of claims 1 to 8 by executing the computer instructions.

11. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the data processing method according to any one of claims 1 to 8.

12. A computer program product, characterized in that The method comprises computer instructions for causing a computer to execute the data processing method according to any one of claims 1 to 8.