Data distribution method and device based on stream processing framework and electronic equipment
By automatically generating the mapping relationship between the resource slot and the database address and establishing a connection when the stream processing framework is started, the problem of Flink not supporting multi-database distribution is solved, and processing performance and distribution efficiency are improved.
Patent Information
- Application Number
- CN202311558385.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-21
- Publication Date
- 2025-05-23
AI Technical Summary
Flink itself does not support the distribution of data to multiple databases, resulting in low data distribution efficiency and poor performance.
By automatically generating a mapping relationship between the resource slot and the database address when the stream processing framework is started, and establishing a connection based on this, flexible data distribution is achieved.
It improves the processing performance and data distribution efficiency of the stream processing framework, supports flexible distribution to multiple database addresses, and reduces maintenance costs.
Smart Images

Figure CN120030069A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of big data technology, and more specifically to a data distribution method, device and electronic device based on a stream processing framework. Background Art
[0002] With the continuous development of big data technology, distributed data processing frameworks can be used to process a large amount of data generated by various business scenarios. For example, the stream processing framework Flink can distribute the data generated by various business scenarios to the corresponding database.
[0003] In the process of implementing the inventive concept of the present disclosure, the inventors found that there are at least the following technical problems in the prior art: the stream processing framework Flink itself does not support distributing data to multiple databases, resulting in low data distribution efficiency and poor performance. Summary of the invention
[0004] In view of the above problems, the present disclosure provides a data distribution method, device and electronic device based on a stream processing framework.
[0005] According to a first aspect of the present disclosure, a data distribution method based on a stream processing framework is provided, comprising: in response to receiving a start request for the stream processing framework, generating a mapping relationship between at least one resource slot in the stream processing framework and at least one database address, wherein the resource slot is used to execute at least one data distribution task of the stream processing framework; based on the mapping relationship, establishing a connection between at least one resource slot and at least one database address; and in response to receiving a request for distributing target data, distributing the target data to a target database address in at least one database address based on the connection.
[0006] According to an embodiment of the present disclosure, in response to receiving a startup request for a stream processing framework, a mapping relationship between at least one resource slot in the stream processing framework and at least one database address is generated, including: in response to receiving a startup request for the stream processing framework, obtaining target configuration parameters from a data interaction class of the stream processing framework, wherein the target configuration parameters are user-defined new parameters; determining the number and database addresses of databases related to the stream processing framework based on the target configuration parameters to obtain at least one database address; and generating a mapping relationship between at least one resource slot of the stream processing framework and at least one database address.
[0007] According to an embodiment of the present disclosure, the method further includes: when the number of databases and / or the database addresses change, updating the target configuration parameters by restarting the stream processing framework.
[0008] According to an embodiment of the present disclosure, generating a mapping relationship between at least one resource slot of a stream processing framework and at least one database address includes: based on a mapping strategy, calculating a calculation result of each at least one resource slot according to a first number of each at least one resource slot; generating a mapping relationship between at least one resource slot of the stream processing framework and at least one database address according to a calculation result of each at least one resource slot and a second number of at least one database address.
[0009] According to an embodiment of the present disclosure, a connection is established between at least one resource slot and at least one database address based on a mapping relationship, including: establishing M connections between the resource slots and database addresses where a mapping relationship exists, wherein the M connections include a master connection and (M-1) slave connections, and M is a positive integer greater than 1.
[0010] According to an embodiment of the present disclosure, the method further includes: managing M connections through a connection pool. Managing M connections through a connection pool further includes: when the stream processing framework is in a running state, detecting the connection state of the master connection; and in response to detecting that the connection state of the master connection is abnormal, taking any slave connection with a normal connection state among the (M-1) slave connections as a new master connection.
[0011] According to an embodiment of the present disclosure, in response to receiving a request for distributing target data, the target data is distributed to a target database address in at least one database address based on a connection, including: in response to receiving a request for distributing target data, determining a target resource slot for performing a distribution task related to the target data; determining a target database address corresponding to the target resource slot based on a mapping relationship; and calling a distribution function of a stream processing framework to distribute the target data to the target database address based on a primary connection between the target resource slot and the target database address.
[0012] A second aspect of the present disclosure provides a data distribution device based on a stream processing framework, including: a generation module, which is used to generate a mapping relationship between at least one resource slot and at least one database address in the stream processing framework in response to receiving a start request for the stream processing framework, wherein the resource slot is used to execute at least one data distribution task of the stream processing framework; a connection module, which is used to establish a connection between at least one resource slot and at least one database address based on the mapping relationship; and a distribution module, which is used to distribute the target data to a target database address in at least one database address based on the connection in response to receiving a request for distributing the target data.
[0013] The third aspect of the present disclosure provides an electronic device, comprising: one or more processors; and a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors execute the above-mentioned data distribution method based on the stream processing framework.
[0014] The fourth aspect of the present disclosure also provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, causes the processor to execute the above-mentioned data distribution method based on the stream processing framework.
[0015] The fifth aspect of the present disclosure also provides a computer program product, including a computer program, which implements the above-mentioned data distribution method based on the stream processing framework when executed by a processor.
[0016] The embodiments of the present disclosure automatically generate a mapping relationship between at least one resource slot and at least one database address in the stream processing framework when starting the stream processing framework; based on the mapping relationship, a connection is established between at least one resource slot and at least one database address, without the need to pre-set the mapping relationship and the connection in an encoding manner, thereby achieving a flexible mapping between the resource slot and the database address, facilitating the flexible distribution of data to multiple database addresses based on the connection established according to the mapping relationship, and improving the processing performance and data distribution efficiency of the stream processing framework. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The above contents and other purposes, features and advantages of the present disclosure will become more apparent through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, in which:
[0018] Figure 1 A schematic diagram of a data distribution architecture based on a stream processing framework according to an embodiment of the present disclosure is schematically shown;
[0019] Figure 2 A flow chart of a data distribution method based on a stream processing framework according to an embodiment of the present disclosure is schematically shown;
[0020] Figure 3 A schematic diagram schematically shows the relationship between target configuration parameters and database addresses according to an embodiment of the present disclosure;
[0021] Figure 4 A schematic diagram of a mapping relationship between a resource slot and a database address according to a specific embodiment of the present disclosure is shown;
[0022] Figure 5 The following schematically shows an application scenario diagram of a data distribution method based on a stream processing framework according to a specific embodiment of the present disclosure;
[0023] Figure 6A structural block diagram of a data distribution device based on a stream processing framework according to an embodiment of the present disclosure is schematically shown; and
[0024] Figure 7 A block diagram of an electronic device suitable for a data distribution method based on a stream processing framework according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION
[0025] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present disclosure. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.
[0026] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise", "include", etc. used herein indicate the existence of the features, steps, operations and / or components, but do not exclude the existence or addition of one or more other features, steps, operations or components.
[0027] All terms (including technical and scientific terms) used herein have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification, and should not be interpreted in an idealized or overly rigid manner.
[0028] When using expressions such as "at least one of A, B, and C, etc.", they should generally be interpreted according to the meaning of the expression commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0029] In the embodiments of the present disclosure, the collection, updating, analysis, processing, use, transmission, provision, disclosure, storage, etc. of the data involved (for example, including but not limited to user personal information) are in compliance with the provisions of relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. In particular, necessary measures are taken for user personal information to prevent illegal access to user personal information data and maintain the security of user personal information, network security, and national security.
[0030] In the embodiments of the present disclosure, the user's authorization or consent is obtained before obtaining or collecting the user's personal information.
[0031] In the data distribution scenario in the big data field, due to the different data types, roles, functions, etc., when executing the distribution task through the stream processing framework Flink, it is necessary to distribute the received data to multiple databases, or different data shards of the same database. For example, the product information data, transaction data, logistics data, etc. of the shopping platform are distributed to different databases through the stream processing framework Flink; the logistics order number, cargo type, cargo storage warehouse and other information in the logistics platform can also be distributed to different databases through the stream processing framework Flink.
[0032] However, the stream processing framework Flink itself does not directly support writing data to multiple JDBC (JavaDataBase Connectivity) URLs (Uniform Resource Locator), that is, the built-in function of the database address, and does not support distributing data to multiple databases, thereby limiting Flink's data distribution performance and efficiency. In the related art, although Flink can be coded to support writing data to multiple databases. However, most of the above-mentioned modification methods pre-write the data distribution logic into Flink by coding, so that it performs data distribution tasks according to the pre-written logic. The above-mentioned modification method not only cannot flexibly adjust the storage address of the data during the data distribution process, but also needs to be maintained by modifying the coding when the database is updated and upgraded, which also limits the data distribution performance and efficiency.
[0033] An embodiment of the present disclosure provides a data distribution method based on a stream processing framework, comprising: in response to receiving a start request for the stream processing framework, generating a mapping relationship between at least one resource slot in the stream processing framework and at least one database address, wherein the resource slot is used to execute at least one data distribution task of the stream processing framework; based on the mapping relationship, establishing a connection between at least one resource slot and at least one database address; and in response to receiving a request for distributing target data, distributing the target data to a target database address in at least one database address based on the connection.
[0034] Figure 1 A schematic diagram of a data distribution architecture based on a stream processing framework according to an embodiment of the present disclosure is schematically shown.
[0035] like Figure 1As shown, the data distribution architecture 100 according to this embodiment may include a server cluster 101, a first database 102, a second database 103, and a third database 104. The network is used to provide a medium for communication links between the server cluster 101, the first database 102, the second database 103, and the third database 104. The network may include various connection types, such as wired, wireless communication links, or optical fiber cables, etc.
[0036] The server cluster 101 is provided with a stream processing framework for real-time processing of stream data from multiple data sources. The data operation modules in the stream processing framework can be provided in multiple servers of the server cluster 101 so as to realize the data processing function through the cooperation of multiple servers.
[0037] The first database 102, the second database 103 and the third database 104 may be a variety of databases with storage functions. For example, a relational database, a non-relational database, a network database, a hierarchical database. The first database 102, the second database 103 and the third database 104 may be databases of the same type or databases of different types.
[0038] For example, the first database 102 , the second database 103 , and the third database 104 may all be relational databases, and the first database 102 , the second database 103 , and the third database 104 may be an Oracle database, a SQL Server database, and a MySQL database in relational databases, respectively.
[0039] According to an embodiment of the present disclosure, the stream processing framework can obtain the stream data to be processed from other data sources, and distribute the stream data to be processed to one of the first database 102, the second database 103, and the third database 104. After distributing the stream data to one of the first database 102, the second database 103, and the third database 104, the database can feed back the storage result to the stream processing framework.
[0040] Alternatively, one of the first database 102, the second database 103 and the third database 104 can be used as a data source, and the stream processing framework obtains stream data from the database serving as the data source, and distributes the obtained stream data to other databases among the first database 102, the second database 103 and the third database except the data source.
[0041] It should be noted that the data distribution method based on the stream processing framework provided in the embodiment of the present disclosure can generally be executed by the server cluster 101. Accordingly, the data distribution device based on the stream processing framework provided in the embodiment of the present disclosure can generally be set in the server cluster 101. The data distribution method based on the stream processing framework provided in the embodiment of the present disclosure can also be executed by a server or server cluster that is different from the server cluster 101 and can communicate with the first database 102, the second database 103 and the third database 104 and / or the server cluster 101. Accordingly, the data distribution device based on the stream processing framework provided in the embodiment of the present disclosure can also be set in a server or server cluster that is different from the server cluster 101 and can communicate with the first database 102, the second database 103 and the third database 104 and / or the server cluster 101.
[0042] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of databases, networks and server clusters may be provided as required.
[0043] The following will be based on Figure 1 The scene described by Figure 2 to Figure 6 The data distribution method based on the stream processing framework of the disclosed embodiment is described in detail.
[0044] Figure 2 The flowchart of the data distribution method based on the stream processing framework according to the embodiment of the present disclosure is schematically shown.
[0045] like Figure 2 As shown, the data distribution method 200 based on the stream processing framework includes operations S210 to S230.
[0046] In operation S210, in response to receiving a start request for a stream processing framework, a mapping relationship between at least one resource slot in the stream processing framework and at least one database address is generated, wherein the resource slot is used to execute at least one data distribution task of the stream processing framework.
[0047] According to an embodiment of the present disclosure, a stream processing framework can execute multiple data distribution tasks in real time. A data distribution task refers to a task in which the stream processing framework receives stream data from a data source and distributes the stream data to a database address related to the stream data. Stream data can be understood as a set of sequential, large, fast, and continuous data sequences, and stream data can be regarded as a dynamic data set that grows infinitely over time.
[0048] According to an embodiment of the present disclosure, the implementation of the data distribution task depends on the resource slot in the stream processing framework. In the stream processing framework Flink, each task manager (TaskManager) is allocated one or more resource slots for executable tasks. A resource slot can execute one or more data distribution tasks, and each resource slot can be considered as a computing unit that performs a concurrent calculation. The stream processing framework may include one or more resource slots.
[0049] According to an embodiment of the present disclosure, in the process of starting a stream processing framework, a mapping relationship between at least one resource slot and at least one database address may be automatically generated, and a connection may be established based on the generated mapping relationship.
[0050] According to an embodiment of the present disclosure, when starting a stream processing framework, the server or server cluster where the stream processing framework is located can receive a startup request for starting the stream processing framework, and then call a data class or function of the stream processing framework to generate a mapping relationship between the stream data processing framework and at least one data.
[0051] According to an embodiment of the present disclosure, a mapping relationship is used to represent the association between a resource slot and a database address. For example, there is a mapping relationship between resource slot 1 and database address 1, which represents that all stream data processed by resource slot 1 needs to be distributed to database address 1.
[0052] According to an embodiment of the present disclosure, a database address may represent the address of an independent database; it may also represent the address of a database shard in a database cluster used to execute part of the tasks or store part of the data.
[0053] In operation S220, a connection is established between at least one resource slot and at least one database address based on the mapping relationship.
[0054] According to an embodiment of the present disclosure, after generating a mapping relationship between at least one resource slot and at least one database address, it is necessary to establish a connection between the resource slot and the database address with the mapping relationship so that the resource slot distributes the stream data to the corresponding database address through the established connection.
[0055] According to an embodiment of the present disclosure, a connection may be understood as a channel for data interaction between a stream processing framework and a database.
[0056] According to an embodiment of the present disclosure, a connection can be established between a resource slot and a database address with a mapping relationship through JDBC (Java Database Connectivity) of a stream processing framework. JDBC includes an application programming interface and a data interaction class for interacting with a database in the Java programming language. By calling the interface and data interaction class in JDBC, a connection between a database to which the database address belongs and a resource slot can be achieved.
[0057] According to an embodiment of the present disclosure, after a connection is established between a resource slot of a stream processing framework and a database address, the stream processing framework is started, at which point the stream processing framework can obtain data from at least one data source and perform data distribution tasks.
[0058] In operation S230 , in response to receiving the request for distributing the target data, the target data is distributed to a target database address among the at least one database address based on the connection.
[0059] According to an embodiment of the present disclosure, the target data may be one or more stream data from one or more data sources. The target data may be data related to multiple business functions. For example, the target data may be real-time logistics information, resource consumption information, office software information, shopping platform recommendation information, etc.
[0060] According to an embodiment of the present disclosure, after receiving a request for distributing target data sent by one or more data sources, the stream processing framework can allocate target resources to corresponding resource slots, and the resource slots perform predetermined processing on the target data, and based on the connection established at startup, distribute the processed target data to the target database address.
[0061] In an embodiment of the present disclosure, when starting the stream processing framework, a mapping relationship between at least one resource slot and at least one database address in the stream processing framework is automatically generated; based on the mapping relationship, a connection is established between at least one resource slot and at least one database address, without the need to pre-set the mapping relationship and the connection in an encoding manner, thereby achieving a flexible mapping between the resource slot and the database address, facilitating flexible distribution of data to multiple database addresses based on the connection established according to the mapping relationship, thereby improving the processing performance and data distribution efficiency of the stream processing framework.
[0062] In the embodiments of the present disclosure, the collection, use, storage, sharing and transfer of target data are in compliance with the relevant laws and regulations, and the user has been informed and the user's consent or authorization has been obtained. When used, the user information is de-identified and / or anonymized and / or encrypted.
[0063] According to an embodiment of the present disclosure, in response to receiving a startup request for a stream processing framework, a mapping relationship between at least one resource slot in the stream processing framework and at least one database address is generated, including: in response to receiving a startup request for the stream processing framework, obtaining target configuration parameters from a data interaction class of the stream processing framework, wherein the target configuration parameters are user-defined new parameters; determining the number and database addresses of databases related to the stream processing framework based on the target configuration parameters to obtain at least one database address; and generating a mapping relationship between at least one resource slot of the stream processing framework and at least one database address.
[0064] According to an embodiment of the present disclosure, the data exchange class of the stream processing framework Flink is a JDBC Sink class, also known as a JdbcSink operator. The JDBC Sink class provides function methods and variable parameters that support functions such as database connection, writing, and distributed transactions.
[0065] According to an embodiment of the present disclosure, the target configuration parameter is a user-defined parameter newly added to write information related to the database address into the stream processing framework Flink. The target configuration parameter may include the number of databases and database addresses related to the stream processing framework.
[0066] According to an embodiment of the present disclosure, the database address associated with the stream processing framework refers to the address corresponding to the database or database shard to be stored after the stream processing framework obtains the stream data from the data source. Among them, the database address can be an address determined according to actual business needs. The number of databases associated with the stream processing framework represents the number of the above database addresses.
[0067] For example, database address 1 and database address 2 are predetermined according to business functions; database address 3 and database address 4 are predetermined according to the purpose of stream data, and the number of databases is 4.
[0068] According to an embodiment of the present disclosure, in response to receiving a start request for a stream processing framework, after obtaining a target configuration parameter from a data interaction class of the stream processing framework, the number of databases and database addresses related to the stream processing framework Flink can be directly obtained from the target configuration parameter. After obtaining at least one database address from the target configuration parameter, a mapping relationship can be generated between the at least one database address and at least one resource slot in the stream processing framework.
[0069] According to an embodiment of the present disclosure, target configuration parameters may be added by extending the JDBC Sink class of Flink.
[0070] According to a specific embodiment of the present disclosure, target configuration parameters can be added by adding parameters to the constructor. Specifically, there may be no constructor or a constructor in the JDBC Sink class. The constructor is used to assign initial values to object member variables in the statement that creates the object. When there is no constructor in the JDBC Sink class, the compiler will add a constructor with no parameters and an empty method body by default. When there is a constructor in the JDBC Sink class, the constructor may include one or more parameters.
[0071] If there is no constructor, you can add a constructor that includes the target configuration parameters to add new target configuration parameters to the JDBC Sink class. If there is a constructor, add the target configuration parameters as new parameters to the original constructor to add new target configuration parameters to the JDBC Sink class.
[0072] According to another specific embodiment of the present disclosure, target configuration parameters can also be added through the setter method. The setter method is essentially an instantiation method, which is accessed as a property outside the JDBC Sink class, allowing the creation of read-only properties and write-only properties. When starting and running the stream processing framework Flink, the target configuration parameters have read-only properties so that the database address and number of databases can be obtained from the target configuration parameters. When the stream processing framework Flink stops running or is shut down, the target configuration parameters can be modified to update the database address.
[0073] In the embodiment of the present disclosure, by adding target configuration parameters, only multiple database addresses and the number of databases are introduced into the stream processing framework. Therefore, when the stream processing framework is started, the current database related to the stream processing framework is obtained by reading the parameter value of the target configuration parameter. When the stream processing framework is started, no other code logic needs to be executed, thereby avoiding increasing the complexity of the stream data processing framework and thus avoiding affecting the performance of the stream processing framework. In addition, the flexible writing method realized by the target configuration parameters can meet the needs of data distribution and concurrent writing, and improve writing performance and scalability.
[0074] Figure 3 The figure schematically shows the relationship between the target configuration parameters and the database address according to the embodiment of the present disclosure.
[0075] like Figure 3 As shown, the target configuration parameter 3011 of the stream processing framework 301 may include database addresses of N databases, such as database address 1 302_1, database address 2 302_2, database address 3 302_3, ... database address N 302_N.
[0076] According to an embodiment of the present disclosure, in the case where the number of databases and / or the database address change, the target configuration parameters are updated by restarting the stream processing framework.
[0077] According to an embodiment of the present disclosure, a change in the number of databases may refer to expanding the databases by adding new databases. A change in the database address may refer to upgrading and transforming the database, that is, data migration occurs, or a certain database fails, and the corresponding stream data is stored in a standby database. A change in both the number of databases and the database address may refer to re-performing a data sharding operation on the database cluster.
[0078] According to an embodiment of the present disclosure, since the target configuration parameters include the database address and the number of databases, in the case where the number of databases and / or the database address change, by restarting the stream processing framework, the parameter values of the target configuration parameters can be updated in Flink. Since generating the mapping relationship between at least one resource slot and at least one database address is an operation that needs to be performed each time the stream processing framework is started, when restarting Flink, the stream processing architecture can generate a new mapping relationship according to the updated target configuration parameters, so as to realize flexible updating of the connection between the resource slot and the database address as the number of databases and / or the database address change.
[0079] For example, in response to receiving a restart request for the stream processing framework, obtain the updated target configuration parameters from the data interaction class of the stream processing framework; according to the updated target configuration parameters, determine the number of databases and the database address related to the stream processing framework to obtain at least one updated database address; generate the mapping relationship between at least one resource slot of the stream processing framework and at least one updated database address.
[0080] In the related art, generally, multiple database addresses, that is, JDBC URLs, are written in Flink through hard-coded operations. For example, Side Output channels corresponding to each database address are created in advance through coding, data is sent to different Side Output channels according to specific conditions, and the data is written into different JDBC URLs through the Side Output channels. Or, the data is written into the Elasticsearch intermediate layer, and Elasticsearch is used to implement data distribution. Or, by using a custom Sink function, the relationship between the resource slot and the JDBC URL is written into Flink in advance. Or, other data processing tools are used to receive the data stream output by Flink, such as Apache NiFi, Apache Camel, etc., and the data processing tools send the data to multiple JDBC URLs.
[0081] However, Side Output channels, Elasticsearch Connectors, custom Sink functions, and data processing tools all predefine data transmission streams through coding, which increases the complexity of the stream processing framework, easily introduces configuration errors and processing logic errors, and thus increases code maintenance costs. In addition, predefine data transmission streams through coding not only introduces additional data transmission delays, such as in high-load conditions in distributed environments, which will lead to reduced efficiency of data distribution tasks and reduced performance of the stream processing framework; moreover, when executing data distribution tasks, the same data may be written to multiple database addresses, leading to data consistency issues.
[0082] In the embodiments of the present disclosure, by adding target configuration parameters, only multiple database addresses and the number of databases are introduced into the stream processing framework, so that the parameter values of the target configuration parameters can be read only by starting the initialization operation of the stream processing framework, without executing other code logic, thereby avoiding increasing the complexity of the stream processing framework. In addition, since the mapping relationship between the resource slot and the database address is generated by the stream processing framework itself, when the database address and / or the number of databases changes, there is no need to perform any code maintenance operations on the stream processing framework, thereby reducing maintenance costs and maintenance workload.
[0083] According to an embodiment of the present disclosure, a mapping relationship between at least one resource slot of a stream processing framework and at least one database address is generated, including: based on a mapping strategy, calculating a calculation result of each at least one resource slot according to a first number of each at least one resource slot; generating a mapping relationship between at least one resource slot of the stream processing framework and at least one database address according to a calculation result of each at least one resource slot and a second number of at least one database address.
[0084] According to an embodiment of the present disclosure, each resource slot corresponds to a first number for uniquely identifying the resource slot, and each database address corresponds to a second number for uniquely identifying the database address.
[0085] For example, the first number of resource slot 1 is 001, the first number of resource slot 2 is 002, the first number of resource slot 3 is 003, and the first number of resource slot 4 is 004. The second number of database address 1 is DB01, the second number of database address 2 is DB02, and the second number of database address 3 is DB03.
[0086] According to an embodiment of the present disclosure, a mapping strategy includes a calculation rule and a mapping rule, wherein the calculation rule is used to determine a method for obtaining a calculation result according to a first number of a resource slot, and the mapping rule is used to determine a mapping relationship between the calculation result and the second number. The mapping strategy can be determined according to actual needs.
[0087] According to an embodiment of the present disclosure, a mapping relationship between at least one resource slot of a stream processing framework and at least one database address is generated based on the calculation results of each of the at least one resource slots and the second number of the at least one database address, including: based on a mapping rule, the calculation result is matched with the second number to obtain a mapping relationship between the resource slot corresponding to the calculation result and the database address represented by the second number.
[0088] According to an embodiment of the present disclosure, the calculation rule may be to poll and calculate the remainder obtained by dividing the first number by a predetermined value. The mapping rule is to determine the mapping relationship between the resource slot and the database address according to the numerical interval in which the remainder is located. For example, when the number of databases is 5, after dividing the first number by 5, a mapping relationship is established between the resource slot with a remainder of 1 and the database address 1; a mapping relationship is established between the resource slot with a remainder of 2 and the database address 2...a mapping relationship is established between the resource slot with a remainder of 5 and the database address 5.
[0089] According to an embodiment of the present disclosure, the calculation rule may be based on a load balancing strategy, and the flow of each resource slot may be evenly distributed to multiple database addresses according to the first number. In this case, the mapping rule is random mapping. Load balancing can avoid performance degradation caused by excessive load on a database.
[0090] According to an embodiment of the present disclosure, the calculation rule may also be to perform a hash calculation on the first number, and the mapping rule may be to map resource slots having the same calculation result after the hash calculation to a database address. Alternatively, the calculation rule may also be to process the first number using a hash algorithm, and the mapping rule may be to map resource slots having the same calculation result after the hash algorithm calculation to a database address.
[0091] In the embodiments of the present disclosure, a mapping strategy can be used to calculate the calculation result of the first number and establish a connection between the calculation result and the second number, expand the applicable scenarios of Flink based on a flexible mapping strategy, and improve the writing speed and throughput of the system. In addition, since the mapping strategy can uniquely map a resource slot to a database address, it ensures that the same data is written to the same database address, avoiding duplicate writing or loss of data, and avoiding data consistency problems that may occur during write operations.
[0092] In addition, this connection management and maintenance mechanism can improve the efficiency and stability of writing, while reducing the overhead of creating and destroying JDBC connections and improving the performance of the Flink framework.
[0093] According to an embodiment of the present disclosure, generating a mapping relationship between at least one resource slot and at least one database address in a stream processing framework, and establishing a connection between at least one resource slot and at least one database address based on the mapping relationship can be implemented by the open() method of the Sink class in the stream processing framework.
[0094] According to an embodiment of the present disclosure, based on the connection, distributing the target data to the target database address in at least one database address can be implemented through the invoke() method of the Sink class in the stream processing framework.
[0095] Figure 4 The figure schematically shows a mapping relationship between resource slots and database addresses according to a specific embodiment of the present disclosure.
[0096] like Figure 4 As shown, the mapping relationship between the resource slot and the database address can be one-to-one, that is, there is a mapping relationship between one resource slot and only one database address. For example, there is a mapping relationship between database address 1402_1 and resource slot 1 401_1, there is a mapping relationship between database address 2 402_2 and resource slot 2 401_2, and there is a mapping relationship between database address 3 402_3 and resource slot 1 401_3.
[0097] The mapping relationship between resource slots and database addresses can also be many-to-one, that is, there is a mapping relationship between one resource slot and one database address, and multiple resource slots can have a mapping relationship with the same database address. For example, database address N 402_N has a mapping relationship with resource slots m 401_m, ... resource slots M 401_M at the same time.
[0098] According to an embodiment of the present disclosure, based on a mapping relationship, a connection is established between at least one resource slot and at least one database address, including: establishing M connections between the resource slots and database addresses where a mapping relationship exists, wherein the M connections include a master connection and (M-1) slave connections, and M is a positive integer greater than 1.
[0099] According to an embodiment of the present disclosure, if only one connection is established between a resource slot and a database address with a mapping relationship, when the connection fails, the stream processing framework will not only be unable to execute the data distribution task, but will also need to consume additional computing resources to re-establish the connection and re-execute the data distribution task, affecting efficiency and performance.
[0100] According to an embodiment of the present disclosure, when establishing a connection, multiple connections can be established simultaneously between resource slots and database addresses that have a mapping relationship. The main connection among the multiple connections is used to transmit data, and the slave connection serves as a backup connection for the main connection.
[0101] According to an embodiment of the present disclosure, when establishing a connection based on a mapping relationship, multiple connections are simultaneously established between a resource slot and a database address where a mapping relationship exists, and one of the connections is used as the main connection and the other connections are used as slave connections, thereby improving the fault tolerance of the stream processing framework and thereby ensuring the reliability and consistency of data writing.
[0102] According to an embodiment of the present disclosure, the data distribution method based on the stream processing framework also includes: managing M connections through a connection pool. Managing M connections through a connection pool also includes: when the stream processing framework is in operation, detecting the connection state of the master connection; and in response to detecting that the connection state of the master connection is abnormal, taking any one of the (M-1) slave connections whose connection state is normal as a new master connection.
[0103] According to an embodiment of the present disclosure, a connection pool is used to manage all connections between at least one resource slot and at least one database address when starting a stream processing framework Flink. For each of the M connections under a mapping relationship, the connection pool can detect the connection status of the main connection among the M connections in real time. When the connection status of the main connection is normal, the connection obtained by the stream processing framework through the invoke() method is the main connection.
[0104] According to an embodiment of the present disclosure, when it is detected that the connection status of the main connection is abnormal, regardless of whether there is a data distribution task at this time, any slave connection with a normal connection status among the (M-1) slave connections will be used as a new main connection, so that the stream processing framework can obtain the updated main connection through the invoke() method, and distribute the target data to the target database address based on the updated main connection.
[0105] According to an embodiment of the present disclosure, for a primary connection whose connection status is abnormal, the connection pool can repair the primary connection in the background without affecting the data distribution task.
[0106] In the embodiments of the present disclosure, by integrating a connection pool in the stream processing framework Flink, the validity of the connection during the execution of the data distribution task can be determined, thereby avoiding the waste of resources caused by re-executing the data distribution task due to connection failure.
[0107] According to an embodiment of the present disclosure, in response to receiving a request for distributing target data, the target data is distributed to a target database address in at least one database address based on a connection, including: in response to receiving a request for distributing target data, determining a target resource slot for performing a distribution task related to the target data; determining a target database address corresponding to the target resource slot based on a mapping relationship; and calling a distribution function of a stream processing framework to distribute the target data to the target database address based on a primary connection between the target resource slot and the target database address.
[0108] According to an embodiment of the present disclosure, the stream processing framework may determine a target resource slot for processing target data and a first number of the target resource slot according to the type or attribute information of the target data.
[0109] For example, if the type of the target data is business, the resource slot for processing the target data of the business type may be determined as 002; if the type of the target data is basic attribute, the resource slot for processing the target data of the basic attribute type may be determined as 004.
[0110] According to an embodiment of the present disclosure, since the generation of the mapping relationship between the resource slot and the database address has been completed when the stream processing framework is started, as long as the stream processing framework is not closed, the stream processing framework can determine the target database address corresponding to the target resource slot according to the mapping relationship.
[0111] According to an embodiment of the present disclosure, even if a restart operation occurs in the stream processing framework, as long as the target configuration parameters and the mapping strategy do not change, the mapping relationship regenerated by the stream processing framework will not change, and the stream processing framework can still determine the target database address corresponding to the target resource slot according to the mapping relationship.
[0112] According to an embodiment of the present disclosure, after determining the target resource slot and the target database address, the method of the distribution function invoke() may be called to obtain the main connection between the target resource slot and the target database address, and the target data may be distributed to the target database address based on the obtained main connection.
[0113] According to an embodiment of the present disclosure, when an error occurs in writing the target data to a certain database address, an error log may be recorded. Then, a retry or rollback operation may be performed through the error log to ensure the reliability and consistency of data writing.
[0114] According to an embodiment of the present disclosure, during the process of the stream processing framework Flink executing the data distribution task, the running metrics of the stream processing framework Flink may be detected, such as resource utilization, connection status, number of error reports, etc., so as to generate a warning message when the running metrics exceed a predetermined standard. The specific metric values and exception handling situations of the running metrics may also be recorded through the error log.
[0115] Figure 5 The application scenario diagram of the data distribution method based on the stream processing framework according to a specific embodiment of the present disclosure is schematically shown.
[0116] Such as Figure 5As shown, data source 501 can transmit stream data to a stream processing framework, which is Flink. The stream processing framework performs data conversion operations on the stream data through data conversion module 502, such as conversion operation 1, conversion operation 1, ... conversion operation P. The data conversion module 502 can convert the data form of the stream data, and can also perform simple data judgment and data operations on the stream data. For example, the resource slot used to distribute the stream data to the stream processing framework is determined according to the type of the stream data.
[0117] The stream processing framework may include multiple resource slots, for example, resource slot 1 503_1, resource slot 2 503_2, resource slot 3 503_3, resource slot m 503_m, ..., resource slot M 503_M. The mapping relationship between the multiple resource slots and the multiple database addresses is pre-generated in the process of starting the stream processing framework. For example, there is a mapping relationship between database address 1 504_1 and resource slot 1 503_1, there is a mapping relationship between database address 2 504_2 and resource slot 2 503_2, there is a mapping relationship between database address 3 504_3 and resource slot 1 503_3, and there is a mapping relationship between database address N 504_N and resource slot m 503_m, ... resource slot M 503_M.
[0118] Therefore, after distributing the stream data to the target resource slot of the stream processing framework, the stream data transmitted by the data source 501 can be distributed to the target database address based on the mapping relationship between the target resource slot and the target data address.
[0119] For example, the stream data transmitted by the data source 501 is distributed to the resource slot 2 503_2, and based on the mapping relationship between the database address 2 504_2 and the resource slot 2 503_2, the stream data is distributed to the database address 2 504_2.
[0120] In the embodiments of the present disclosure, by extending the JDBC Sink, the connection between the resource slot and the JDBC URL is directly established in the JDBC Sink class, which simplifies the data distribution logic and state management, and provides a more direct and efficient writing method. At the same time, the solution also involves considerations such as data consistency, exception handling, and monitoring, and fully guarantees the data integrity in the JDBC URL writing scenario.
[0121] Figure 6 A structural block diagram of a data distribution device based on a stream processing framework according to an embodiment of the present disclosure is schematically shown; and
[0122] like Figure 6 As shown, the data distribution device 600 based on the stream processing framework of this embodiment includes a generation module 610 , a connection module 620 and a distribution module 630 .
[0123] The generation module 610 is used to generate a mapping relationship between at least one resource slot in the stream processing framework and at least one database address in response to receiving a start request for the stream processing framework, wherein the resource slot is used to execute at least one data distribution task of the stream processing framework.
[0124] The connection module 620 is used to establish a connection between at least one resource slot and at least one database address based on the mapping relationship.
[0125] The distribution module 630 is configured to distribute the target data to a target database address in at least one database address based on the connection in response to receiving a request for distributing the target data.
[0126] According to an embodiment of the present disclosure, the generation module 610 includes an acquisition submodule, a first determination submodule and a first generation submodule.
[0127] The acquisition submodule is used to obtain target configuration parameters from the data interaction class of the stream processing framework in response to receiving a start request for the stream processing framework, wherein the target configuration parameters are newly added parameters customized by the user.
[0128] The first determination submodule is used to determine the number and database addresses of databases related to the stream processing framework according to the target configuration parameters, and obtain at least one database address.
[0129] The first generating submodule is used to generate a mapping relationship between at least one resource slot of the stream processing framework and at least one database address.
[0130] According to an embodiment of the present disclosure, the data distribution device 600 based on the stream processing framework further includes an update module for updating target configuration parameters by restarting the stream processing framework when the number of databases and / or the database addresses change.
[0131] According to an embodiment of the present disclosure, the generation module 610 further includes a calculation submodule and a second generation submodule.
[0132] The calculation submodule is used to calculate the calculation result of at least one resource slot respectively based on the mapping strategy and the first number of at least one resource slot respectively.
[0133] The second generating submodule is used to generate a mapping relationship between at least one resource slot of the stream processing framework and at least one database address according to the calculation result of each of the at least one resource slots and the second number of the at least one database address.
[0134] According to an embodiment of the present disclosure, the connection module 620 includes a connection sub-module for establishing M connections between resource slots and database addresses with a mapping relationship, where among the M connections, there is a primary connection and (M - 1) secondary connections, and M is a positive integer greater than 1.
[0135] According to an embodiment of the present disclosure, the data distribution device 600 based on the stream processing framework further includes a connection pool management module for managing the M connections through a connection pool.
[0136] According to an embodiment of the present disclosure, the connection pool management module further includes a detection sub-module and an update sub-module.
[0137] The detection sub-module is used to detect the connection status of the primary connection when the stream processing framework is in a running state.
[0138] The update sub-module is used to, in response to detecting that the connection status of the primary connection is abnormal, use any one of the (M - 1) secondary connections with a normal connection status as the new primary connection.
[0139] According to an embodiment of the present disclosure, the distribution module 630 includes a second determination sub-module, a third determination sub-module, and a call sub-module.
[0140] The second determination sub-module is used to, in response to receiving a request for distributing target data, determine a target resource slot for executing a distribution task related to the target data.
[0141] The third determination sub-module is used to determine a target database address corresponding to the target resource slot according to the mapping relationship.
[0142] The call sub-module is used to call the distribution function of the stream processing framework and distribute the target data to the target database address based on the primary connection between the target resource slot and the target database address.
[0143] According to an embodiment of the present disclosure, any multiple of the generation module 610, the connection module 620, and the distribution module 630 can be combined and implemented in one module, or any one of them can be split into multiple modules. Or, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module.
[0144] According to an embodiment of the present disclosure, at least one of the generation module 610, the connection module 620 and the distribution module 630 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or may be implemented by hardware or firmware such as any other reasonable way of integrating or packaging the circuit, or implemented in any one of the three implementation methods of software, hardware and firmware or in an appropriate combination of any of them. Alternatively, at least one of the generation module 610, the connection module 620 and the distribution module 630 may be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding function may be executed.
[0145] It should be noted that the data distribution device part based on the stream processing framework in the embodiments of the present disclosure corresponds to the data distribution method part based on the stream processing framework in the embodiments of the present disclosure. The description of the data distribution device part based on the stream processing framework specifically refers to the data distribution method part based on the stream processing framework, which will not be repeated here.
[0146] Figure 7 A block diagram of an electronic device suitable for a data distribution method based on a stream processing framework according to an embodiment of the present disclosure is schematically shown.
[0147] like Figure 7 As shown, the electronic device 700 according to an embodiment of the present disclosure includes a processor 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage part 708 into a random access memory (RAM) 703. The processor 701 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 701 may also include an onboard memory for caching purposes. The processor 701 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0148] In RAM 703, various programs and data required for the operation of electronic device 700 are stored. Processor 701, ROM 702 and RAM 703 are connected to each other via bus 704. Processor 701 performs various operations of the method flow according to the embodiment of the present disclosure by executing the program in ROM 702 and / or RAM 703. It should be noted that the program can also be stored in one or more memories other than ROM 702 and RAM 703. Processor 701 can also perform various operations of the method flow according to the embodiment of the present disclosure by executing the program stored in the one or more memories.
[0149] According to an embodiment of the present disclosure, the electronic device 700 may further include an input / output (I / O) interface 705, which is also connected to the bus 704. The electronic device 700 may further include one or more of the following components connected to the input / output I / O interface 705: an input portion 706 including a keyboard, a mouse, etc.; an output portion 707 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage portion 708 including a hard disk, etc.; and a communication portion 709 including a network interface card such as a LAN card, a modem, etc. The communication portion 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 710 as needed, so that a computer program read therefrom is installed into the storage portion 708 as needed.
[0150] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist independently without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiment of the present disclosure is implemented.
[0151] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, may include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, an apparatus or a device. For example, according to an embodiment of the present disclosure, a computer-readable storage medium may include the ROM 702 and / or RAM 703 described above and / or one or more memories other than ROM 702 and RAM 703.
[0152] The embodiment of the present disclosure also includes a computer program product, which includes a computer program, and the computer program contains program code for executing the method shown in the flowchart. When the computer program product is run in a computer system, the program code is used to enable the computer system to implement the above method provided by the embodiment of the present disclosure.
[0153] The above functions defined in the system / device of the embodiment of the present disclosure are performed when the computer program is executed by the processor 701. According to the embodiment of the present disclosure, the system, device, module, unit, etc. described above can be implemented by a computer program module.
[0154] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices, magnetic storage devices, etc. In another embodiment, the computer program may also be transmitted and distributed in the form of signals on a network medium, and downloaded and installed through the communication part 709, and / or installed from the removable medium 711. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0155] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 709, and / or installed from the removable medium 711. When the computer program is executed by the processor 701, the above functions defined in the system of the embodiment of the present disclosure are performed. According to the embodiment of the present disclosure, the system, device, means, module, unit, etc. described above can be implemented by a computer program module.
[0156] According to an embodiment of the present disclosure, the program code for executing the computer program provided by the embodiment of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level process and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, Java, C++, python, "C" language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on the remote computing device, or entirely on the remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect through the Internet).
[0157] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0158] It will be appreciated by those skilled in the art that the features described in the various embodiments and / or claims of the present disclosure may be combined and / or combined in a variety of ways, even if such combinations and / or combinations are not explicitly described in the present disclosure. In particular, the features described in the various embodiments and / or claims of the present disclosure may be combined and / or combined in a variety of ways without departing from the spirit and teachings of the present disclosure. All of these combinations and / or combinations fall within the scope of the present disclosure.
[0159] The specific embodiments described above further illustrate the purpose, technical solutions and beneficial effects of the present disclosure. It should be understood that the above description is only a specific embodiment of the present disclosure and is not intended to limit the present disclosure. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present disclosure should be included in the protection scope of the present disclosure.
Claims
1. A data distribution method based on a stream processing framework, include: In response to receiving a start request for a stream processing framework, generating a mapping relationship between at least one resource slot in the stream processing framework and at least one database address, wherein the resource slot is used to execute at least one data distribution task of the stream processing framework; Based on the mapping relationship, establishing a connection between at least one of the resource slots and at least one of the database addresses; and In response to receiving a request for distributing target data, the target data is distributed to a target database address among at least one of the database addresses based on the connection.
2. The method according to claim 1, in, The step of generating, in response to receiving a start request for a stream processing framework, a mapping relationship between at least one resource slot in the stream processing framework and at least one database address comprises: In response to receiving a start request for a stream processing framework, obtaining a target configuration parameter from a data interaction class of the stream processing framework, wherein the target configuration parameter is a newly added parameter customized by a user; Determine the number and address of databases related to the stream processing framework according to the target configuration parameters, and obtain at least one of the database addresses; and A mapping relationship between at least one resource slot of the stream processing framework and at least one database address is generated.
3. The method according to claim 2, further comprising: include: When the number of databases and / or the address of the databases changes, the target configuration parameters are updated by restarting the stream processing framework.
4. The method according to claim 1 or 2, in, The generating a mapping relationship between at least one resource slot of the stream processing framework and at least one database address includes: Based on the mapping strategy, according to the first number of at least one resource slot, calculate the calculation result of at least one resource slot; A mapping relationship between at least one resource slot of the stream processing framework and the at least one database address is generated according to the calculation result of each of the at least one resource slots and the second number of the at least one database address.
5. The method according to claim 1, in, The establishing a connection between at least one of the resource slots and at least one of the database addresses based on the mapping relationship includes: M connections are established between the resource slot and the database address that have a mapping relationship, wherein the M connections include a master connection and (M-1) slave connections, and M is a positive integer greater than 1.
6. The method according to claim 5, further comprising: include: Managing the M connections through a connection pool; The managing the M connections through the connection pool further includes: When the stream processing framework is in a running state, detecting a connection state of the primary connection; as well as In response to detecting that the connection state of the master connection is abnormal, any one of the (M-1) slave connections whose connection state is normal is used as a new master connection.
7. The method according to claim 5 or 6, in, In response to receiving the request for distributing the target data, distributing the target data to a target database address in at least one of the database addresses based on the connection, comprises: In response to receiving a request for distributing target data, determining a target resource slot for executing a distribution task associated with the target data; Determining a target database address corresponding to the target resource slot according to the mapping relationship; and The distribution function of the stream processing framework is called to distribute the target data to the target database address based on the primary connection between the target resource slot and the target database address.
8. A data distribution device based on a stream processing framework, include: a generating module, configured to generate, in response to receiving a start request for a stream processing framework, a mapping relationship between at least one resource slot in the stream processing framework and at least one database address, wherein the resource slot is used to execute at least one data distribution task of the stream processing framework; a connection module, configured to establish a connection between at least one of the resource slots and at least one of the database addresses based on the mapping relationship; and The distribution module is used for distributing the target data to a target database address of at least one of the database addresses based on the connection in response to receiving a request for distributing the target data.
9. An electronic device, include: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors are enabled to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, causes the processor to execute the method according to any one of claims 1 to 7.
11. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.