Streaming data processing method and related equipment

By determining the scheduling nodes and transmission nodes in the distributed log stream processing system and dynamically allocating sub-partition flow groups, the problem of insufficient load balancing in existing systems during load fluctuations and node failures is solved, and more efficient load balancing and system performance improvement is achieved.

CN120017590APending Publication Date: 2025-05-16TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311537763.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-16
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

When the existing distributed log stream processing system handles load fluctuations, dynamic throughput requirements and node failures, the load balancing strategy is simple and fixed, and it is difficult to effectively deal with problems, resulting in limited system stability and performance improvement.

Method used

By determining the scheduling node and transmission node in the preset server cluster, receiving the load index information of the transmission node, dividing the to-processed stream data into sub-partition flows and combining them into sub-partition flow groups, and dynamically allocating the sub-partition flow groups to the transmission node according to the load index information.

Benefits of technology

It realizes more flexible and efficient load balancing and dynamic adjustment, can effectively deal with load fluctuations and node failures, and improves system stability and performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017590A_ABST
    Figure CN120017590A_ABST
Patent Text Reader

Abstract

The invention discloses a streaming data processing method and related equipment. The embodiment of the invention can be applied to the technical field of computers. The method comprises the following steps: determining a scheduling node for load adjustment and a transmission node for processing streaming data from a preset server cluster; the scheduling node receives the load index information uploaded by the transmission node; stream data to be processed is divided into a plurality of sub-partition streams, and the sub-partition streams are combined to obtain at least one sub-partition stream group; and the scheduling node allocates each sub-partition flow group to the transmission node according to the load index information. According to the invention, the sub-partition streams are combined, the plurality of sub-partition streams are distributed into the same group, the sub-partition stream group is used as the minimum unit of load balancing scheduling, the data scale for load calculation is reduced, the sub-partition streams with larger quantity and scale can be supported, and the scheduling node is used for detecting the load condition of the transmission node, so that the load balancing scheduling efficiency is improved. Therefore, the load of the transmission node can be dynamically balanced and adjusted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of computers, and in particular to a stream data processing method and related equipment. Background Art

[0002] In existing distributed log stream processing system solutions, when creating a cluster log stream transmission channel, it is usually necessary to rely on system preset values ​​or manual estimates to predict the future throughput of the log stream. In order to process the log stream, the existing technical solution divides it into multiple sub-partition streams, and then evenly distributes and mounts these sub-partition streams to each transmission node in the cluster. In this way, each transmission node is responsible for the collection, processing, and forwarding of the sub-partition stream data mounted under its name.

[0003] For transmission nodes that fail or crash, the leader node in the cluster will be responsible for redistributing the sub-partition flows under the crashed node and assigning them to other normally running nodes to achieve failover. This existing method provides certain solutions in basic load balancing and failover. When processing distributed log stream data, it can meet the needs of general scenarios, but it still has deficiencies in handling load fluctuations, dynamic throughput requirements, and node failures. The specific shortcomings are as follows:

[0004] When the load of each sub-partition flow fluctuates over time during system operation, the load balancing strategy of the existing solution is relatively simple and fixed, and it is difficult to effectively deal with throughput changes and load imbalance problems.

[0005] When a large throughput log stream appears in the cluster, the existing solution may cause traffic peaks on local nodes and affect the normal operation and stability of the entire system.

[0006] During cluster expansion, existing solutions fail to fully utilize the resources of newly added nodes, resulting in the potential of newly added nodes not being fully utilized, limiting the improvement of overall system performance. Summary of the invention

[0007] The embodiments of the present application provide a stream data processing method and related equipment. The related equipment may include a stream data processing device, an electronic device, a computer-readable storage medium and a computer program product, which can flexibly, comprehensively and efficiently realize load balancing and dynamic adjustment.

[0008] The present application provides a method for processing stream data, including:

[0009] Determine a scheduling node for load adjustment and a transmission node for processing stream data from a preset server cluster;

[0010] The scheduling node receives the load indicator information uploaded by the transmission node;

[0011] Dividing the stream data to be processed into a plurality of sub-partition streams, and combining the sub-partition streams to obtain at least one sub-partition stream group;

[0012] The scheduling node allocates each of the sub-partition flow groups to the transmission node according to the load indicator information.

[0013] Furthermore, the stream data to be processed is divided into a plurality of sub-partition streams, and the sub-partition streams are combined to obtain at least one sub-partition stream group, including:

[0014] Extracting a keyword of the stream data to be processed, and calculating a hash value of the keyword using a hash function;

[0015] Determine the subpartition corresponding to the keyword in a preset hash table according to the hash value, and assign the keyword to the corresponding subpartition to obtain multiple subpartition streams;

[0016] The sub-partition flows are combined into a plurality of sub-partition flow groups according to the numbering intervals of the sub-partitions.

[0017] Further, combining the sub-partition flows into a plurality of sub-partition flow groups according to the numbering intervals of the sub-partitions includes:

[0018] Determining the initial number of groups according to the number of sub-partition flow groups;

[0019] Based on the numbering intervals of the subpartitions, the multiple subpartition flows are grouped into an initial number of subpartition flow groups.

[0020] Further, the scheduling node allocates each of the sub-partition flow groups to the transmission node according to the load indicator information, including:

[0021] The scheduling node determines a low-load node and an idle node from the transmission nodes according to the load indicator information;

[0022] The scheduling node allocates the sub-partition flow group to the low-load node or the idle node.

[0023] Further, after the scheduling node allocates each of the sub-partition flow groups to the transmission node according to the load indicator information, the method further includes:

[0024] According to the load indicator information, determining high-load nodes and abnormal nodes from the transmission nodes;

[0025] The scheduling node unloads part of the sub-partition flow groups in the high-load node, and unloads all the sub-partition flow groups in the abnormal node;

[0026] The sub-partition flow groups obtained by unloading are allocated to the low-load nodes or the idle nodes.

[0027] Furthermore, before determining the high-load node and the abnormal node from the transmission node according to the load indicator information, the method further includes:

[0028] Based on a preset time interval, the load indicator information uploaded by the transmission node is obtained.

[0029] Further, after allocating the unloaded sub-partition flow group to the low-load node or the idle node, the method further includes:

[0030] After the low-load node or the idle node receives the unloaded sub-partition flow group, updating the load indicator information;

[0031] Upload updated load indicator information to the scheduling node.

[0032] Further, after the scheduling node allocates each of the sub-partition flow groups to the transmission node according to the load indicator information, the method further includes:

[0033] The low-load node or the idle node receives the sub-partition flow group and obtains the data throughput corresponding to the sub-partition flow group;

[0034] The scheduling node receives the data throughput corresponding to each of the sub-partition flow groups uploaded by the low-load node or the idle node;

[0035] If the data throughput is higher than a preset threshold, the scheduling node divides the sub-partition flow group into at least two sub-partition flow sub-groups;

[0036] The scheduling node allocates each of the sub-partition flow subgroups to a low-load node or an idle node according to the load indicator information.

[0037] Further, the scheduling node divides the sub-partition flow group into at least two sub-partition flow sub-groups, including:

[0038] The scheduling node obtains the sub-partition numbering interval corresponding to the sub-partition flow group;

[0039] Based on the numbering intervals of the sub-partitions, the sub-partition flow group is evenly divided into at least two sub-partition flow sub-groups.

[0040] Furthermore, after the scheduling node allocates the sub-partition flow subgroup to a low-load node or an idle node according to the load indicator information, the method further includes:

[0041] The low-load node or the idle node receives the sub-partition flow sub-group and obtains the data throughput corresponding to the sub-partition flow sub-group;

[0042] The scheduling node receives the data throughput corresponding to each of the sub-partition flow subgroups uploaded by the low-load node or the idle node.

[0043] Further, the scheduling node receives the load indicator information uploaded by the transmission node, including:

[0044] The transmission node obtains the data throughput corresponding to each of the sub-partition flow groups mounted on the transmission node and the basic performance index information of the transmission node as the load index information;

[0045] Based on a preset time interval, the transmission node uploads the load indicator information to the scheduling node.

[0046] Accordingly, an embodiment of the present application provides a stream data processing device, including:

[0047] A determination unit, used to determine a scheduling node for load adjustment and a transmission node for processing stream data from a preset server cluster;

[0048] An acquisition unit, configured for the scheduling node to receive the load indicator information uploaded by the transmission node;

[0049] A combining unit, used for dividing the stream data to be processed into a plurality of sub-partition streams, and combining the sub-partition streams to obtain at least one sub-partition stream group;

[0050] An allocating unit is used for the scheduling node to allocate each of the sub-partition flow groups to the transmission node according to the load indicator information.

[0051] An electronic device provided in an embodiment of the present application includes a processor and a memory, wherein the memory stores a plurality of instructions, and the processor loads the instructions to execute the steps in the stream data processing method provided in the embodiment of the present application.

[0052] An embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the steps in the stream data processing method provided in the embodiment of the present application are implemented.

[0053] In addition, an embodiment of the present application also provides a computer program product, including a computer program or instructions, which, when executed by a processor, implements the steps in the stream data processing method provided in the embodiment of the present application.

[0054] The embodiment of the present application provides a method for processing stream data and related equipment, which can determine a scheduling node for load adjustment and a transmission node for processing stream data from a preset server cluster; the scheduling node receives load index information uploaded by the transmission node; the stream data to be processed is divided into multiple sub-partitioned streams, and the sub-partitioned streams are combined to obtain at least one sub-partitioned stream group; the scheduling node assigns each of the sub-partitioned stream groups to the transmission node according to the load index information. The present application combines sub-partitioned streams, assigns multiple sub-partitioned streams to the same group, and uses sub-partitioned stream groups as the minimum unit of load balancing scheduling, which reduces the data scale used for load calculation, can support a larger number of sub-partitioned streams, and uses scheduling nodes to detect the load conditions of transmission nodes, so that the load of transmission nodes can be dynamically balanced and adjusted. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0056] Figure 1 It is a scenario diagram of the stream data processing method provided in an embodiment of the present application;

[0057] Figure 2 is a first flow chart of the stream data processing method provided by an embodiment of the present application;

[0058] Figure 3 is a second flow chart of the stream data processing method provided in an embodiment of the present application;

[0059] Figure 4 is a structural diagram of a stream data processing device provided in an embodiment of the present application;

[0060] Figure 5 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application;

[0061] Figure 6a It is a schematic diagram of the relationship between sub-partition flows and groups provided in an embodiment of the present application;

[0062] Figure 6b This is a schematic diagram of scheduling node election and load indicator information collection provided by this application;

[0063] Figure 6c It is a schematic diagram of triggering load balancing adjustment when the node load provided by this application is too high;

[0064] Figure 6dThis is a schematic diagram of the grouping dichotomy provided by this application. DETAILED DESCRIPTION

[0065] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.

[0066] The embodiment of the present application provides a stream data processing method and related equipment, and the related equipment may include a stream data processing device, an electronic device, a computer readable storage medium and a computer program product. The stream data processing device may be integrated in an electronic device, and the electronic device may be a terminal or a server.

[0067] It is understandable that the stream data processing method of this embodiment can be executed on a terminal, can be executed on a server, or can be executed by both a terminal and a server. The above examples should not be construed as limiting the present application.

[0068] like Figure 1 As shown, a terminal and a server jointly execute a stream data processing method as an example. The stream data processing system provided in the embodiment of the present application includes a terminal 10 and a server 11, etc. The terminal 10 and the server 11 are connected via a network, such as a wired or wireless network connection, etc., wherein the stream data processing device can be integrated in the server.

[0069] Among them, the server 11 can be used to: determine the scheduling node for load adjustment and the transmission node for processing stream data from the preset server cluster; the scheduling node receives the load index information uploaded by the transmission node; divide the stream data to be processed into multiple sub-partition flows, and combine the sub-partition flows to obtain at least one sub-partition flow group; the scheduling node assigns each sub-partition flow group to the transmission node according to the load index information. Among them, the server 11 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.

[0070] In an embodiment of the present application, a server cluster is a server cluster or a distributed system composed of multiple servers. Among them, a distributed system is a software system built on a network. It processes various assisted tasks and then integrates the results. A distributed system can solve the problem that the amount of computing is too large and the memory and CPU of a single node cannot handle it. The storage and computing tasks are shared on ordinary machines, and the growth of data volume is coped with by dynamically adding nodes, but the disadvantage is that the management of multiple nodes and the scheduling of tasks are more troublesome, which is also a problem that distributed systems study and solve.

[0071] In the distributed system of this application, a hash algorithm is used to divide the sub-partition streams into several groups based on the characteristics of log stream data. This allows each group to be mounted on a transmission node separately, responsible for forwarding log data of all sub-partition streams under the group. This partitioning strategy optimizes data distribution, reduces the pressure on a single node, and thus enhances the scalability of the system.

[0072] Select the Leader node from the distributed cluster nodes to perform real-time detection and data collection on key performance indicators such as the system's CPU, memory, and network load. The Leader node is responsible for collecting information from each cluster node and analyzing it to guide load balancing adjustments.

[0073] When a transmission node is detected to be overloaded or down, the Leader node will start the load balancing mechanism, unload some or all groups mounted on the transmission node, and redistribute them to other nodes with lower loads. This dynamic adjustment strategy helps ensure the stability and high availability of the system when facing performance pressure.

[0074] When a group's throughput is detected to be overloaded, the leader node is responsible for unloading the group from the original node and then dividing it into two new groups. Then, the leader node will redistribute and mount the two new groups to nodes with lower loads. This innovation enables the system to respond quickly and balance the load when processing high-throughput requests, thereby improving the transmission throughput of the entire cluster.

[0075] In an embodiment of the present application, a server cluster is a server cluster or a distributed system composed of multiple servers. Among them, a distributed system is a software system built on a network. It processes various assisted tasks and then integrates the results. A distributed system can solve the problem that the amount of computing is too large and the memory and CPU of a single node cannot handle it. The storage and computing tasks are shared on ordinary machines, and the growth of data volume is coped with by dynamically adding nodes, but the disadvantage is that the management of multiple nodes and the scheduling of tasks are more troublesome, which is also a problem that distributed systems study and solve.

[0076] In the distributed system of this application, a hash algorithm is used to divide the sub-partition streams into several groups based on the characteristics of log stream data. This allows each group to be mounted on a transmission node separately, responsible for forwarding log data of all sub-partition streams under the group. This partitioning strategy optimizes data distribution, reduces the pressure on a single node, and thus enhances the scalability of the system.

[0077] Select the Leader node from the distributed cluster nodes to perform real-time detection and data collection on key performance indicators such as the system's CPU, memory, and network load. The Leader node is responsible for collecting information from each cluster node and analyzing it to guide load balancing adjustments.

[0078] When a transmission node is detected to be overloaded or down, the Leader node will start the load balancing mechanism, unload some or all groups mounted on the transmission node, and redistribute them to other nodes with lower loads. This dynamic adjustment strategy helps ensure the stability and high availability of the system when facing performance pressure.

[0079] When a group's throughput is detected to be overloaded, the leader node is responsible for unloading the group from the original node and then dividing it into two new groups. Then, the leader node will redistribute and mount the two new groups to nodes with lower loads. This innovation enables the system to respond quickly and balance the load when processing high-throughput requests, thereby improving the transmission throughput of the entire cluster.

[0080] The terminal 10 can be used for developers to interact with the server 11. For example, developers can set grouping algorithm rules and scheduling node election rules through the terminal 10. The terminal 10 can include a mobile phone, an intelligent voice interaction device, a smart home appliance, a vehicle terminal, an aircraft, a tablet computer, a laptop computer, or a personal computer (PC). A client can also be set on the terminal 10, and the client can be an application client or a browser client, etc.

[0081] It should be noted that the description order of the following embodiments is not intended to limit the preferred order of the embodiments.

[0082] This embodiment will be described from the perspective of a stream data processing device, which can be integrated into an electronic device, such as a server or a terminal. This embodiment can be applied to a scenario where distributed log stream data is automatically load balanced according to the load of cluster nodes during transmission and forwarding.

[0083] Among them, the log stream is the basic unit of log reading and writing. Log streams can be created in log groups to facilitate further classification and management of logs. Log reading and writing are based on log streams. You can specify log streams when writing, classify and store different types of logs, and after collecting logs, package multiple log data and send them to the cloud log service in log streams. The log stream reading and writing method can minimize the number of reads and writes and improve business efficiency.

[0084] Among them, stream data is a set of data sequences that arrive in an orderly, large, fast and continuous manner. Generally speaking, stream data can be regarded as a dynamic data set that grows infinitely over time.

[0085] The English name of load balancing is Load Balance, which means balancing the load (work tasks) and distributing them to multiple operation units for operation, such as FTP servers, Web servers, enterprise core application servers and other main task servers, so as to complete the work tasks in a collaborative manner.

[0086] The technical solution of the present invention is suitable for the transmission, collection and storage of sequential, continuous and large amounts of log data. Typical application scenarios include message queues, etc. The specific applications are as follows:

[0087] Real-time stream processing and analysis: When used in conjunction with a streaming computing engine, the technical solution of the present invention can be used as a data source and data result storage for the computing engine. Through real-time processing and computing analysis, real-time data analysis, anomaly detection and early warning functions can be realized to meet the real-time data needs of various industries.

[0088] Application log and detection data collection: The technical solution of the present invention can be used to collect application operation logs and detection data, helping developers and operation and maintenance teams to discover and locate potential problems, thereby improving the maintainability, reliability and performance of the application.

[0089] AI model and machine learning training data set storage: By collecting and organizing large amounts of data, the technical solution of the present invention can be used as a storage tool for AI models and machine learning training data sets, improving the training effect and performance of the model, and is suitable for various AI and data science applications.

[0090] IoT device data reception and storage: As a tool for receiving and storing large amounts of real-time data generated by IoT devices, the technical solution of the present invention can meet the diverse data access needs such as sensor data and detection information, and provide a data foundation for building intelligent IoT applications.

[0091] Distributed transaction log management: In a distributed system, the technical solution of the present invention can be used as a storage and management tool for transaction logs. The efficient load balancing strategy provided by the present invention helps maintain the consistency and reliability of distributed transactions and provides support for scenarios such as distributed databases and microservice architectures.

[0092] like Figure 2 As shown, the specific process of the stream data processing method can be as follows:

[0093] 201. Determine a scheduling node for load adjustment and a transmission node for processing stream data from a preset server cluster.

[0094] Distributed systems or components generally include a leader election process, such as the leader node election of ZooKeeper, the leader node election of Redis Sentinel, and the master node election in Redis Cluster.

[0095] In a distributed server cluster, all nodes have three states: Leader, Follower, and Candidate.

[0096] When the system enters the voting state, each node will vote. If a node gets more than half of the votes, it will be elected as the Leader node. In this application, it serves as a scheduling node for load adjustment, while other nodes serve as transmission nodes.

[0097] 202. The scheduling node receives load indicator information uploaded by the transmission node.

[0098] refer to Figure 6b In the server cluster, the scheduling node and the transmission node are connected wirelessly or wired to send and receive information. The transmission node will upload the load indicator information to the scheduling node so that the scheduling node can detect the load status of the transmission node based on the load indicator information.

[0099] It is understandable that in some embodiments, in order to promptly process high-load nodes or abnormal nodes, the transmission node needs to detect the load status of the transmission node in real time, and the transmission node can send load indicator information to the scheduling node in real time and continuously.

[0100] In other embodiments, in order to reduce the amount of calculation of the scheduling node, a time interval may be preset, and the transmission node sends the load index information to the scheduling node once every time interval. The scheduling node may obtain the load index information uploaded by the transmission node based on the preset time interval.

[0101] The load indicator information may include key performance indicators such as CPU, memory, and network load of the transmission node.

[0102] In some embodiments, the scheduling node receives the load indicator information uploaded by the transmission node, including the following steps:

[0103] The transmission node obtains the data throughput corresponding to each of the sub-partition flow groups mounted on the transmission node and the basic performance index information of the transmission node as the load index information;

[0104] Based on a preset time interval, the transmission node uploads the load indicator information to the scheduling node.

[0105] Among them, the basic performance indicator information includes the CPU, memory and other information of the transmission node, and the data throughput can be used to determine the network load of the transmission node.

[0106] 203. Divide the stream data to be processed into a plurality of sub-partition streams, and combine the sub-partition streams to obtain at least one sub-partition stream group.

[0107] Partitioning means that messages will be distributed on all nodes in the server cluster according to partitions. The stream data to be processed is divided into multiple sub-partition streams, and each sub-partition stream is distributed and stored on multiple nodes.

[0108] The load balancing of the present invention is based on groups. In order to support massive sub-partition flow transmission, if it is scheduled according to a single sub-partition flow, it will bring very large computing pressure during load balancing scheduling. Therefore, the present application adopts a group method to manage sub-partition flows. The minimum unit of load balancing is designed as a group. Multiple sub-partition flows are allocated to the same group, which reduces the data scale used for load calculation, thereby increasing the number of supported sub-partition flows.

[0109] In some embodiments, dividing the stream data to be processed into a plurality of sub-partition streams, and combining the sub-partition streams to obtain at least one sub-partition stream group may include the following steps:

[0110] Extracting a keyword of the stream data to be processed, and calculating a hash value of the keyword using a hash function;

[0111] Determine the subpartition corresponding to the keyword in a preset hash table according to the hash value, and assign the keyword to the corresponding subpartition to obtain multiple subpartition streams;

[0112] The sub-partition flows are combined into a plurality of sub-partition flow groups according to the numbering intervals of the sub-partitions.

[0113] Among them, Redis cluster solves the problem of uniform distribution by adding another layer between data and nodes, which is called hash slot, to manage the relationship between data and nodes. Now it is equivalent to putting slots on nodes and data in slots. The hash slot is actually an array space, and the array [0,2^14-1] forms a hash solt space. The hash slot is the sub-partition in this application. The data in the hash slot is the sub-partition stream.

[0114] Redis cluster has 16384 built-in hash slots, and Redis will map hash slots to different transmission nodes roughly equally according to the number of nodes. When a key-value needs to be placed in the Redis cluster, Redis first uses the crc16 algorithm to calculate a result for the key (keyword), and then calculates the remainder of the result to 16384, so that each key (keyword) will correspond to a hash slot numbered between 0 and 16383, that is, it is mapped to a transmission node.

[0115] In some embodiments, combining the sub-partition streams into a plurality of sub-partition stream groups according to the numbering intervals of the sub-partitions comprises the following steps:

[0116] Determining the initial number of groups according to the number of sub-partition flow groups;

[0117] Based on the numbering intervals of the subpartitions, the multiple subpartition flows are grouped into an initial number of subpartition flow groups.

[0118] In some embodiments of the present application, the initial number of groups can be determined based on the number of sub-partition flow groups. For example, the sub-partition flow groups can be evenly combined into several groups, and the number of sub-partition flows in each group is the same. For another example, the maximum number of sub-partition flows included in each sub-partition flow group can be set, and the sub-partition flows can be combined into several groups based on the maximum number.

[0119] 204. Allocate the sub-partition flow group to the transmission node based on the load indicator information.

[0120] The scheduling node can evenly distribute the sub-partition flow groups to the transmission nodes according to the load status of the transmission nodes.

[0121] In some embodiments, the scheduling node allocates each of the sub-partition flow groups to the transmission node according to the load indicator information, including the following steps:

[0122] The scheduling node determines a low-load node and an idle node from the transmission nodes according to the load indicator information;

[0123] The scheduling node allocates the sub-partition flow group to the low-load node or the idle node.

[0124] The developer can pre-set the load threshold and determine the load threshold based on the load indicator information. When the load threshold is lower than the preset value, the transmission node is considered to be a low-load node. When the transmission node is not assigned a sub-partition flow group, the transmission node is considered to be an idle node.

[0125] In some embodiments, the scheduling node may be configured to preferentially allocate sub-partition flow groups to idle nodes. When there are no idle nodes in the server cluster, the sub-partition flow groups are allocated to low-load nodes.

[0126] After the allocation is completed, the scheduling node can also dynamically adjust the allocation of sub-partition flow groups according to the load situation. When the transmission node is overloaded or the transmission node is abnormal, some or all sub-partition flow groups can be unloaded from the transmission node and reallocated to the transmission node with lower load.

[0127] In some embodiments, after the scheduling node allocates each of the sub-partition flow groups to the transmission node according to the load indicator information, the following steps are further included:

[0128] According to the load indicator information, determining high-load nodes and abnormal nodes from the transmission nodes;

[0129] The scheduling node unloads part of the sub-partition flow groups in the high-load node, and unloads all the sub-partition flow groups in the abnormal node;

[0130] The sub-partition flow groups obtained by unloading are allocated to the low-load nodes or the idle nodes.

[0131] Among them, the scheduling node can determine whether the transmission node is down based on the load indicator information. If any of the CPU, memory or network load of the transmission node is abnormal, the transmission node is considered to be an abnormal node.

[0132] This dynamic adjustment strategy helps ensure the stability and high availability of the system when facing performance pressure.

[0133] After the load adjustment is completed, the scheduling node will continue to detect the load status of each transmission node. If the transmission node regularly uploads load index information, in order to facilitate the scheduling node to check the load adjustment status in a timely manner, after the load adjustment, the transmission node that receives the obtained sub-partition flow group can immediately update and re-upload the load index information. Specifically, after the sub-partition flow group obtained by unloading is assigned to the low-load node or the idle node, the following steps are also included:

[0134] After the low-load node or the idle node receives the unloaded sub-partition flow group, updating the load indicator information;

[0135] Upload updated load indicator information to the scheduling node.

[0136] In order to cope with the situation of high throughput load, in some embodiments, after the scheduling node allocates each of the sub-partition flow groups to the transmission node according to the load indicator information, the following steps are also included:

[0137] The low-load node or the idle node receives the sub-partition flow group and obtains the data throughput corresponding to the sub-partition flow group;

[0138] The scheduling node receives the data throughput corresponding to each of the sub-partition flow groups uploaded by the low-load node or the idle node;

[0139] If the data throughput is higher than a preset threshold, the scheduling node divides the sub-partition flow group into at least two sub-partition flow sub-groups;

[0140] The scheduling node allocates each of the sub-partition flow subgroups to a low-load node or an idle node according to the load indicator information.

[0141] The throughput refers to the amount of data successfully transmitted by the sub-partition flow group per unit time (measured in bits, bytes, packets, etc.).

[0142] The transmission node can obtain the throughput of the sub-partition flow group assigned to itself, and upload the throughput of each sub-partition flow group to the scheduling node. The scheduling node continuously detects the throughput of each sub-partition flow group. When a high throughput load is detected, the above-mentioned group dichotomy is executed.

[0143] In some embodiments, each time the grouping binary division is performed, the sub-partition flow group is evenly divided into two sub-partition flow sub-groups, and the scheduling node divides the sub-partition flow group into at least two sub-partition flow sub-groups, including the following steps:

[0144] The scheduling node obtains the sub-partition numbering interval corresponding to the sub-partition flow group;

[0145] Based on the numbering intervals of the sub-partitions, the sub-partition flow group is evenly divided into at least two sub-partition flow sub-groups.

[0146] In some embodiments, in order to facilitate the scheduling node to timely monitor the throughput load after the group is split into two, after the transmission node receives the sub-partition flow sub-group, the throughput of the newly allocated sub-partition flow sub-group is uploaded to the scheduling node so that the scheduling node can adjust the load throughput in time.

[0147] In some embodiments, after the group bisection is performed and the throughput of the sub-partition flow subgroup is under normal conditions, the transmission node sends throughput information to the scheduling node at a preset time interval so that the scheduling node continues to pay attention to the load situation of the sub-partition flow subgroup, thereby ensuring that the system can efficiently cope with throughput fluctuations.

[0148] As can be seen from the above, the embodiments of the present application can determine the scheduling node for load adjustment and the transmission node for processing stream data from the preset server cluster; the scheduling node receives the load index information uploaded by the transmission node; the stream data to be processed is divided into multiple sub-partitioned streams, and the sub-partitioned streams are combined to obtain at least one sub-partitioned stream group; the scheduling node assigns each sub-partitioned stream group to the transmission node according to the load index information. The present application combines the sub-partitioned streams, assigns multiple sub-partitioned streams to the same group, and uses the sub-partitioned stream group as the minimum unit of load balancing scheduling, which reduces the data scale used for load calculation, can support a larger number of sub-partitioned streams, and uses the scheduling node to detect the load of the transmission node, so that the load of the transmission node can be dynamically balanced.

[0149] According to the method described in the previous embodiment, the following will be further described in detail by taking the example of the stream data processing device being specifically integrated into an electronic device. The present application embodiment provides a stream data processing method, such as Figure 3 As shown, the specific process of the stream data processing method can be as follows:

[0150] 301. A hash algorithm is used to divide the log stream into multiple sub-partition streams, and these sub-partition streams are combined into several groups, which are distributed and mounted on each transmission node in the cluster in groups.

[0151] The present invention uses a hash algorithm to divide the stream data into multiple sub-partition streams, and then combines these sub-partition streams into several groups, which are distributed and mounted on each transmission node in the cluster in groups.

[0152] refer to Figure 6a , set the encoding interval range of a sub-partition in a cluster to 0x0000~0xFFFF, and initially divide the cluster interval into N groups. If N=4, the grouping is as follows:

[0153] Sub-partition stream group 1: 0x0000~0x3FFF

[0154] Sub-partition stream group 2: 0x4000~0x7FFF

[0155] Subpartition stream group 3: 0x8000~0xBFFF

[0156] Subpartition stream group 4: 0xC000~0xFFFF

[0157] Through the hash algorithm, the sub-partition streams will be dispersed within the preset encoding range. Each group contains a part of the sub-partition streams, which are allocated and mounted to each transmission node in the cluster in groups. The corresponding transmission node is responsible for the data transmission of the sub-partition streams under the group.

[0158] 302. Select a scheduling node in the server cluster, and periodically collect load indicator information of each transmission node based on the scheduling node.

[0159] The technical solution of the present application is to elect a Leader node (or scheduling node) for a distributed cluster. The Leader node is responsible for coordinating and managing the load balancing of the entire cluster, mastering the global load situation, and thus dynamically optimizing resource utilization.

[0160] The scheduling node periodically collects load indicator information of each transmission node, such as CPU usage, memory usage, and network load. This data enables the scheduling node to accurately grasp the load status of each transmission node in the cluster.

[0161] 303. The scheduling node allocates a sub-partition flow group to each transmission node based on load indicator information.

[0162] Based on the collected load information, the scheduling node assigns sub-partition flow groups to each transmission node. This includes:

[0163] Assign subpartition stream groups to free nodes.

[0164] When a node is found to be overloaded, some sub-partition flow groups are unloaded from the node and redistributed to transmission nodes with lower load.

[0165] Dynamically adjust the division and allocation strategy of sub-partition flow groups according to system load conditions.

[0166] 304. The scheduling node detects the load status of the transmission node based on the load indicator information and dynamically adjusts the load of the transmission node.

[0167] refer to Figure 6c In order to adapt to problems such as load fluctuations and node failures during the operation of the distributed system, the technical solution of this application introduces a dynamic load balancing adjustment mechanism. When it is detected that a transmission node is overloaded or down, the scheduling node is responsible for initiating load balancing adjustment. The specific process is as follows:

[0168] Load anomaly detection: The scheduling node continuously monitors the status of each transmission node and detects load anomalies based on preset thresholds.

[0169] Unloading sub-partition flow groups: When a transmission node is found to be overloaded or down, the scheduling node will unload all or part of the sub-partition flow groups on the node. The conditions for activating the adjustment process can be variable, for example, based on the load indicator exceeding the threshold or the transmission node status being faulty.

[0170] Reassign sub-partition flow groups: The scheduling node reallocates and mounts the sub-partition flow groups unloaded in the previous step to low-load or normally operating transmission nodes. When reallocating, the performance indicator data of the node can be referred to to achieve a more reasonable group allocation and maximize the balanced utilization of resources on each node.

[0171] Load re-detection: After the load adjustment is completed, the scheduling node will continue to monitor the load of each node in the cluster. If necessary, the load balancing adjustment process can be triggered again to ensure that the system continues to operate in the optimal state.

[0172] 305. When a high throughput load is detected, the scheduling node divides the sub-partition flow group whose throughput exceeds a preset threshold into two, and reallocates the two sub-partition flow sub-groups.

[0173] In order to cope with high throughput load, this application introduces a grouping dichotomy. The specific process is as follows:

[0174] Detecting throughput anomalies: The scheduling node continuously monitors the throughput of the sub-partition flow groups. When the throughput of a sub-partition flow group continuously exceeds the preset threshold, it can be regarded as a high throughput load situation.

[0175] Grouping binary division: Under high throughput load, the scheduling node performs grouping binary division to divide the sub-partition flow group exceeding the threshold into two and form two new sub-partition flow sub-groups, such as Figure 6d This helps to spread out high load situations and speed up data transfer.

[0176] Reassign sub-partition flow sub-groups: The scheduling node assigns and mounts the two sub-partition flow sub-groups divided in the previous step to the transmission node with low load. In the reallocation process, the load status and resource usage status of each node can be fully considered to achieve a more reasonable load balancing effect.

[0177] Detect the load status of the new group: After the group division is completed, the scheduling node still needs to continue to pay attention to the load status of the sub-partition flow sub-group. The scheduling node can make further load balancing adjustments to ensure that the system can efficiently cope with throughput fluctuations.

[0178] The solution of the present application brings an efficient, stable and scalable load balancing strategy for distributed log stream data processing. In practical applications, the technical solution of the present invention can overcome the challenges faced by the prior art when processing distributed log stream data and improve the performance and stability of the entire system. In practical applications, the technical solution of the present application can overcome the challenges faced by the prior art when processing distributed log stream data and improve the performance and stability of the entire system.

[0179] As can be seen from the above, the embodiments of the present application can determine the scheduling node for load adjustment and the transmission node for processing stream data from the preset server cluster; the scheduling node receives the load index information uploaded by the transmission node; the stream data to be processed is divided into multiple sub-partitioned streams, and the sub-partitioned streams are combined to obtain at least one sub-partitioned stream group; the scheduling node assigns each of the sub-partitioned stream groups to the transmission node according to the load index information. The present application combines the sub-partitioned streams, assigns multiple sub-partitioned streams to the same group, and uses the sub-partitioned stream group as the minimum unit of load balancing scheduling, which reduces the data scale used for load calculation, can support a larger number of sub-partitioned streams, and uses the scheduling node to detect the load of the transmission node, so that the load of the transmission node can be dynamically balanced.

[0180] In order to better implement the above method, the embodiment of the present application also provides a stream data processing device, such as Figure 4 As shown, the stream data processing device may include a determination unit 401, an acquisition unit 402, a combination unit 403, and an allocation unit 404, as follows:

[0181] The determining unit 401 is used to determine a scheduling node for load adjustment and a transmission node for processing stream data from a preset server cluster;

[0182] An acquisition unit 402 is used for the scheduling node to receive the load indicator information uploaded by the transmission node;

[0183] A combining unit 403 is used to divide the stream data to be processed into a plurality of sub-partition streams, and combine the sub-partition streams to obtain at least one sub-partition stream group;

[0184] The allocating unit 404 is configured to allocate, by the scheduling node, each of the sub-partition flow groups to the transmission node according to the load indicator information.

[0185] Optionally, in some embodiments of the present application, the combining unit 403 further includes a calculating subunit, a partitioning subunit and a combining subunit, as follows:

[0186] A calculation subunit, used for extracting a keyword of the stream data to be processed, and calculating a hash value of the keyword by using a hash function;

[0187] A partitioning subunit, configured to determine the subpartition corresponding to the keyword in a preset hash table according to the hash value, and assign the keyword to the corresponding subpartition to obtain a plurality of subpartition streams;

[0188] The combining subunit is used to combine the sub-partition flows into a plurality of sub-partition flow groups according to the numbering intervals of the sub-partitions.

[0189] Optionally, in some embodiments of the present application, the allocating unit 404 further includes a first determining bullet element and a first allocating subunit, as follows:

[0190] A first determining subunit, configured for the scheduling node to determine a low-load node and an idle node from the transmission node according to the load indicator information;

[0191] The first allocation subunit is used for the scheduling node to allocate the sub-partition flow group to the low-load node or the idle node.

[0192] Optionally, in some embodiments of the present application, the stream data processing device may further include an adjustment unit, configured to:

[0193] According to the load indicator information, determining high-load nodes and abnormal nodes from the transmission nodes;

[0194] The scheduling node unloads part of the sub-partition flow groups in the high-load node, and unloads all the sub-partition flow groups in the abnormal node;

[0195] The sub-partition flow groups obtained by unloading are allocated to the low-load nodes or the idle nodes.

[0196] Optionally, in some embodiments of the present application, the adjustment unit may also be used to:

[0197] After the low-load node or the idle node receives the unloaded sub-partition flow group, updating the load indicator information;

[0198] Upload updated load indicator information to the scheduling node.

[0199] Optionally, in some embodiments of the present application, the acquisition unit 402 may also be configured to acquire the load indicator information uploaded by the transmission node based on a preset time interval.

[0200] Optionally, in some embodiments of the present application, the stream data processing device of the present application further includes a throughput unit, a throughput upload unit, a packet binary unit and a packet redistribution unit, as follows:

[0201] A throughput acquisition unit, configured for the low-load node or the idle node to receive the sub-partition flow group and acquire data throughput corresponding to the sub-partition flow group;

[0202] A throughput uploading unit, configured for the scheduling node to receive the data throughput corresponding to each of the sub-partition flow groups uploaded by the low-load node or the idle node;

[0203] a grouping bisection unit, configured to divide the sub-partitioned flow group into two sub-partitioned flow sub-groups if the data throughput is higher than a preset threshold;

[0204] A grouping redistribution unit is used for the scheduling node to distribute each of the sub-partition flow subgroups to a low-load node or an idle node according to the load indicator information.

[0205] As can be seen from the above, in this embodiment, the determination unit 401 can be used to determine the scheduling node for load adjustment and the transmission node for processing stream data from the preset server cluster; the acquisition unit 402 is used for the scheduling node to receive the load index information uploaded by the transmission node; the combination unit 403 is used to divide the stream data to be processed into multiple sub-partition streams, and the sub-partition streams are combined to obtain at least one sub-partition stream group; the allocation unit 404 is used for the scheduling node to allocate each of the sub-partition stream groups to the transmission node according to the load index information. The embodiment of the present application provides a stream data processing method and related equipment, which can determine the scheduling node for load adjustment and the transmission node for processing stream data from the preset server cluster; the present application combines the sub-partition streams, allocates multiple sub-partition streams to the same group, and uses the sub-partition stream group as the minimum unit of load balancing scheduling, which reduces the data scale used for load calculation, can support a larger number of sub-partition streams, and uses the scheduling node to detect the load of the transmission node, so that the load of the transmission node can be dynamically balanced.

[0206] The present application also provides an electronic device, such as Figure 5 As shown, it shows a schematic diagram of the structure of an electronic device involved in an embodiment of the present application, and the electronic device may be a terminal or a server, etc. Specifically:

[0207] The electronic device may include components such as a processor 501 with one or more processing cores, a memory 502 with one or more computer-readable storage media, a power supply 503, and an input unit 504. Those skilled in the art will appreciate that Figure 5 The electronic device structure shown in the figure does not constitute a limitation on the electronic device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0208] The processor 501 is the control center of the electronic device, and uses various interfaces and lines to connect various parts of the entire electronic device. By running or executing software programs and / or modules stored in the memory 502, and calling data stored in the memory 502, the processor 501 performs various functions of the electronic device and processes data. Optionally, the processor 501 may include one or more processing cores; preferably, the processor 501 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the processor 501.

[0209] The memory 502 can be used to store software programs and modules. The processor 501 executes various functional applications and data processing by running the software programs and modules stored in the memory 502. The memory 502 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 502 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices. Accordingly, the memory 502 may also include a memory controller to provide the processor 501 with access to the memory 502.

[0210] The electronic device also includes a power supply 503 for supplying power to each component. Preferably, the power supply 503 can be logically connected to the processor 501 through a power management system, so as to manage charging, discharging, power consumption and other functions through the power management system. The power supply 503 can also include one or more DC or AC power supplies, recharging systems, power failure detection circuits, power converters or inverters, power status indicators and other arbitrary components.

[0211] The electronic device may further include an input unit 504, which may be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.

[0212] Although not shown, the electronic device may further include a display unit, etc., which will not be described in detail herein. Specifically in this embodiment, the processor 501 in the electronic device will load the executable files corresponding to the processes of one or more application programs into the memory 502 according to the following instructions, and the processor 501 will run the application programs stored in the memory 502, thereby realizing various functions, as follows:

[0213] Determine a scheduling node for load adjustment and a transmission node for processing stream data from a preset server cluster;

[0214] The scheduling node receives the load indicator information uploaded by the transmission node;

[0215] Dividing the stream data to be processed into a plurality of sub-partition streams, and combining the sub-partition streams to obtain at least one sub-partition stream group;

[0216] The scheduling node allocates each of the sub-partition flow groups to the transmission node according to the load indicator information.

[0217] The specific implementation of the above operations can be found in the previous embodiments, which will not be described in detail here.

[0218] As can be seen from the above, the embodiment of the present application provides a method for processing stream data and related equipment, which can determine the scheduling node for load adjustment and the transmission node for processing stream data from the preset server cluster; the scheduling node receives the load index information uploaded by the transmission node; the stream data to be processed is divided into multiple sub-partitioned streams, and the sub-partitioned streams are combined to obtain at least one sub-partitioned stream group; the scheduling node assigns each of the sub-partitioned stream groups to the transmission node according to the load index information. The present application combines sub-partitioned streams, assigns multiple sub-partitioned streams to the same group, and uses sub-partitioned stream groups as the minimum unit of load balancing scheduling, which reduces the data scale used for load calculation, can support a larger number of sub-partitioned streams, and uses scheduling nodes to detect the load conditions of transmission nodes, so that the load of transmission nodes can be dynamically balanced and adjusted.

[0219] A person of ordinary skill in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be completed by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.

[0220] To this end, an embodiment of the present application provides a computer-readable storage medium, in which a plurality of instructions are stored, and the instructions can be loaded by a processor to execute the steps in any of the stream data processing methods provided in the embodiments of the present application. For example, the instructions can execute the following steps:

[0221] Determine a scheduling node for load adjustment and a transmission node for processing stream data from a preset server cluster;

[0222] The scheduling node receives the load indicator information uploaded by the transmission node;

[0223] Dividing the stream data to be processed into a plurality of sub-partition streams, and combining the sub-partition streams to obtain at least one sub-partition stream group;

[0224] The scheduling node allocates each of the sub-partition flow groups to the transmission node according to the load indicator information.

[0225] The specific implementation of the above operations can be found in the previous embodiments, which will not be described in detail here.

[0226] The computer-readable storage medium may include: a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0227] Since the instructions stored in the computer-readable storage medium can execute the steps in any of the stream data processing methods provided in the embodiments of the present application, the beneficial effects that can be achieved by any of the stream data processing methods provided in the embodiments of the present application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.

[0228] According to one aspect of the present application, a computer program product or a computer program is provided, the computer program product or the computer program comprising computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method provided in various optional implementations of the above-mentioned stream data processing.

[0229] The above is a detailed introduction to a stream data processing method and related equipment provided in an embodiment of the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for technical personnel in this field, according to the idea of ​​the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A method for processing stream data, characterized in that: include: Determine a scheduling node for load adjustment and a transmission node for processing stream data from a preset server cluster; The scheduling node receives the load indicator information uploaded by the transmission node; Dividing the stream data to be processed into a plurality of sub-partition streams, and combining the sub-partition streams to obtain at least one sub-partition stream group; The scheduling node allocates each of the sub-partition flow groups to the transmission node according to the load indicator information.

2. The stream data processing method according to claim 1, characterized in that: Dividing the stream data to be processed into a plurality of sub-partition streams, and combining the sub-partition streams to obtain at least one sub-partition stream group, including: Extracting a keyword of the stream data to be processed, and calculating a hash value of the keyword using a hash function; Determine the subpartition corresponding to the keyword in a preset hash table according to the hash value, and assign the keyword to the corresponding subpartition to obtain multiple subpartition streams; The sub-partition flows are combined into a plurality of sub-partition flow groups according to the numbering intervals of the sub-partitions.

3. The stream data processing method according to claim 2, characterized in that: The combining the sub-partition streams into a plurality of sub-partition stream groups according to the numbering intervals of the sub-partitions comprises: Determining the initial number of groups according to the number of sub-partition flow groups; Based on the numbering intervals of the subpartitions, the multiple subpartition flows are grouped into an initial number of subpartition flow groups.

4. The stream data processing method according to claim 1, characterized in that: The scheduling node allocates each of the sub-partition flow groups to the transmission node according to the load indicator information, including: The scheduling node determines a low-load node and an idle node from the transmission nodes according to the load indicator information; The scheduling node allocates the sub-partition flow group to the low-load node or the idle node.

5. The stream data processing method according to claim 4, characterized in that: After the scheduling node allocates each of the sub-partition flow groups to the transmission node according to the load indicator information, the method further includes: According to the load indicator information, determining high-load nodes and abnormal nodes from the transmission nodes; The scheduling node unloads part of the sub-partition flow groups in the high-load node, and unloads all the sub-partition flow groups in the abnormal node; The sub-partition flow groups obtained by unloading are allocated to the low-load nodes or the idle nodes.

6. The stream data processing method according to claim 3, characterized in that: Before determining the high-load node and the abnormal node from the transmission node according to the load indicator information, the method further includes: Based on a preset time interval, the load indicator information uploaded by the transmission node is obtained.

7. The stream data processing method according to claim 5, characterized in that: After allocating the unloaded sub-partition flow group to the low-load node or the idle node, the method further includes: After the low-load node or the idle node receives the unloaded sub-partition flow group, updating the load indicator information; Upload updated load indicator information to the scheduling node.

8. The stream data processing method according to claim 4, characterized in that: After the scheduling node allocates each of the sub-partition flow groups to the transmission node according to the load indicator information, the method further includes: The low-load node or the idle node receives the sub-partition flow group and obtains the data throughput corresponding to the sub-partition flow group; The scheduling node receives the data throughput corresponding to each of the sub-partition flow groups uploaded by the low-load node or the idle node; If the data throughput is higher than a preset threshold, the scheduling node divides the sub-partition flow group into at least two sub-partition flow sub-groups; The scheduling node allocates each of the sub-partition flow subgroups to a low-load node or an idle node according to the load indicator information.

9. The stream data processing method according to claim 8, characterized in that: The scheduling node divides the sub-partition flow group into at least two sub-partition flow sub-groups, including: The scheduling node obtains the sub-partition numbering interval corresponding to the sub-partition flow group; Based on the numbering intervals of the sub-partitions, the sub-partition flow group is evenly divided into at least two sub-partition flow sub-groups.

10. The stream data processing method according to claim 8, characterized in that: After the scheduling node allocates the sub-partition flow subgroup to a low-load node or an idle node according to the load indicator information, the method further includes: The low-load node or the idle node receives the sub-partition flow sub-group and obtains the data throughput corresponding to the sub-partition flow sub-group; The scheduling node receives the data throughput corresponding to each of the sub-partition flow subgroups uploaded by the low-load node or the idle node.

11. The method for processing stream data according to claim 1, characterized in that: The scheduling node receives the load indicator information uploaded by the transmission node, including: The transmission node obtains the data throughput corresponding to each of the sub-partition flow groups mounted on the transmission node and the basic performance indicator information of the transmission node as the load indicator information; Based on a preset time interval, the transmission node uploads the load indicator information to the scheduling node.

12. A stream data processing device, characterized in that: include: A determination unit, used to determine a scheduling node for load adjustment and a transmission node for processing stream data from a preset server cluster; An acquisition unit, used for the scheduling node to receive the load indicator information uploaded by the transmission node; A combining unit, used for dividing the stream data to be processed into a plurality of sub-partition streams, and combining the sub-partition streams to obtain at least one sub-partition stream group; An allocating unit is used for the scheduling node to allocate each of the sub-partition flow groups to the transmission node according to the load indicator information.

13. An electronic device, characterized in that: It comprises a memory and a processor; the memory stores an application program, and the processor is used to run the application program in the memory to execute the operations in the stream data processing method according to any one of claims 1 to 11.

14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the steps in the stream data processing method according to any one of claims 1 to 11.

15. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps in the stream data processing method according to any one of claims 1 to 11 are implemented.