Data processing method and device, computer equipment and readable storage medium
By obtaining and processing target ETCD data when it is detected in the cloud computing resource management system, the problem of K8S storage capacity bottleneck is solved, the efficiency of cloud computing resource processing is improved, and the operation of large-scale clusters is supported.
Patent Information
- Application Number
- CN202510360021.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-20
AI Technical Summary
Kubernetes (K8S) has storage capacity bottlenecks in distributed storage systems based on ETCD, which cannot support large-scale clusters, thereby reducing the efficiency of cloud computing resource processing.
By detecting the update of ETCD data in the main node API management layer of the cloud computing resource management system, the target ETCD data is obtained, and the target controller partition is determined from the candidate controller partition based on the data information, and data is sent to it for processing. The method includes data slicing processing and load balancing partition selection to improve processing efficiency.
It effectively alleviates the problem of K8S storage capacity bottleneck, improves the efficiency of cloud computing resource processing, and supports the operation of large-scale clusters.
Smart Images

Figure CN120179733A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of cloud computing technology, and in particular to a data processing method, apparatus, computer device, and readable storage medium. Background Art
[0002] With the continuous development of cloud computing infrastructure service IaaS, in order to reasonably schedule the resources of a cloud computing cluster, an open-source container orchestration and management system (Kubernetes, K8S) can be used to reasonably schedule cloud computing resources.
[0003] However, since K8S is built based on the distributed storage system ETCD, due to the storage capacity limitation of ETCD itself, K8S also has a storage capacity bottleneck, which will cause K8S not to support large-scale clusters, thereby reducing the efficiency of cloud computing resource processing. Summary of the Invention
[0004] Based on this, in view of the above technical problems, it is necessary to provide a data processing method, apparatus, computer device, and readable storage medium that can improve the efficiency of cloud computing resource processing.
[0005] In a first aspect, this application provides a data processing method, which is applied to the application programming interface API management layer of the master node in a cloud computing resource management system, and the API management layer is connected to each candidate controller partition in the master node; the method includes:
[0006] When it is detected that there is data update in the distributed storage system ETCD, obtain the target ETCD data;
[0007] According to the data information of the target ETCD data, determine the target controller partition corresponding to the target ETCD data from each candidate controller partition;
[0008] Send the target ETCD data to the target controller partition to process the target ETCD data within the target controller partition.
[0009] In one of the embodiments, the data information includes data type and data volume; according to the data information of the target ETCD data, determining the target controller partition corresponding to the target ETCD data from each candidate controller partition includes:
[0010] According to the data type of the target ETCD data, determine each alternative controller partition adapted to the target ETCD data from each candidate controller partition;
[0011] According to the magnitude relationship between the data volume of the target ETCD data and the data volume threshold, and the current working information of each alternative controller partition, select the target controller partition from each alternative controller partition.
[0012] In one embodiment, selecting a target controller partition from each alternative controller partition according to the magnitude relationship between the data volume of the target ETCD data and a data volume threshold, and the current working information of each alternative controller partition, includes:
[0013] When the data volume of the target ETCD data is greater than the data volume threshold, performing a splitting process on the target ETCD data to obtain at least two target ETCD sub - data;
[0014] Selecting a target number of target controller partitions from each alternative controller partition according to the current working information of each alternative controller partition; wherein, the target number is the number of target ETCD sub - data;
[0015] Correspondingly, sending the target ETCD data to the target controller partition to process the target ETCD data within the target controller partition, includes:
[0016] Sending one target ETCD sub - data to each target controller partition respectively to process the corresponding target ETCD sub - data within the target controller partition; wherein, the target ETCD sub - data corresponding to different target controller partitions are different.
[0017] In one embodiment, the method further includes:
[0018] When detecting a data query request, selecting at least one query - to - be - processed controller partition from each candidate controller partition according to the data type of the data to be queried in the data query request;
[0019] Issuing the data query request to at least one query - to - be - processed controller partition, and obtaining data query sub - results from at least one query - to - be - processed controller partition;
[0020] Processing the obtained data query sub - results to obtain a data query result corresponding to the data query request.
[0021] In one embodiment, the method further includes:
[0022] Sending a heartbeat acquisition request to each working node in the cloud computing resource management system to instruct each working node to split dynamic heartbeat information from standard heartbeat information and feedback the dynamic heartbeat information to the API management layer; wherein, the standard heartbeat information includes static node information and dynamic heartbeat information; the static node information is generated based on the configuration information of the working node; the dynamic heartbeat information is generated based on a preset data packet;
[0023] Determining the running state of the cloud computing resource management system according to the dynamic heartbeat information fed back by each working node.
[0024] In a second aspect, the present application provides a data processing method, which is applied to a target controller partition having a connection relationship with the API management layer in the master node. The method includes:
[0025] Obtaining target ETCD data sent by the API management layer; wherein, the target controller partition is determined from each candidate controller partition based on the data information of the target ETCD data;
[0026] According to the event change type included in the target ETCD data, calling a corresponding event processing function to process the target ETCD data; and,
[0027] Adding the target ETCD data to the data storage tree in the target controller partition; wherein, the data storage tree is obtained by processing the ETCD data stored in the target controller partition using a tree - like structure storage method.
[0028] In a third aspect, the present application further provides a data processing apparatus, which is applied to the application programming interface (API) management layer of the master node in a cloud computing resource management system. The apparatus includes:
[0029] A first acquisition module, configured to obtain target ETCD data when detecting data updates in the distributed storage system ETCD;
[0030] A controller determination module, configured to determine a target controller partition corresponding to the target ETCD data from each candidate controller partition according to the data information of the target ETCD data;
[0031] A data sending module, configured to send the target ETCD data to the target controller partition for processing the target ETCD data within the target controller partition.
[0032] In a fourth aspect, the present application further provides a data processing apparatus, which is applied to a target controller partition having a connection relationship with the API management layer in the master node. The apparatus includes:
[0033] A second acquisition module, configured to obtain the target ETCD data sent by the API management layer; wherein, the target controller partition is determined from each candidate controller partition based on the data information of the target ETCD data;
[0034] A data processing module, configured to call a corresponding event processing function according to the event change type included in the target ETCD data to process the target ETCD data; and,
[0035] A data storage module for adding target ETCD data to a data storage tree in a target controller partition; wherein, the data storage tree is obtained by processing the ETCD data stored in the target controller partition using a tree - like structure storage method.
[0036] In a fifth aspect, the present application further provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0037] When it is detected that there is data update in the distributed storage system ETCD, obtain the target ETCD data;
[0038] According to the data information of the target ETCD data, determine the target controller partition corresponding to the target ETCD data from each candidate controller partition;
[0039] Send the target ETCD data to the target controller partition to process the target ETCD data within the target controller partition.
[0040] In a sixth aspect, the present application further provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0041] Obtain the target ETCD data sent by the API management layer; wherein, the target controller partition is determined from each candidate controller partition based on the data information of the target ETCD data;
[0042] According to the event change type included in the target ETCD data, call the corresponding event processing function to process the target ETCD data; and,
[0043] Add the target ETCD data to the data storage tree in the target controller partition; wherein, the data storage tree is obtained by processing the ETCD data stored in the target controller partition using a tree - like structure storage method.
[0044] In a seventh aspect, the present application further provides a computer - readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0045] When it is detected that there is data update in the distributed storage system ETCD, obtain the target ETCD data;
[0046] According to the data information of the target ETCD data, determine the target controller partition corresponding to the target ETCD data from each candidate controller partition;
[0047] Send the target ETCD data to the target controller partition to process the target ETCD data within the target controller partition.
[0048] In a eighth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0049] Obtain the target ETCD data sent by the API management layer; wherein, the target controller partition is determined from each candidate controller partition based on the data information of the target ETCD data;
[0050] According to the event change type included in the target ETCD data, call the corresponding event processing function to process the target ETCD data; and,
[0051] Add the target ETCD data to the data storage tree in the target controller partition; wherein, the data storage tree is obtained by processing the ETCD data stored in the target controller partition using a tree-like structure storage method.
[0052] In a ninth aspect, the present application further provides a computer program product, including a computer program. When the computer program is executed by a processor, the following steps are implemented:
[0053] When it is detected that there is data update in the distributed storage system ETCD, obtain the target ETCD data;
[0054] According to the data information of the target ETCD data, determine the target controller partition corresponding to the target ETCD data from each candidate controller partition;
[0055] Send the target ETCD data to the target controller partition to process the target ETCD data within the target controller partition.
[0056] In a tenth aspect, the present application further provides a computer program product, including a computer program. When the computer program is executed by a processor, the following steps are implemented:
[0057] Obtain the target ETCD data sent by the API management layer; wherein, the target controller partition is determined from each candidate controller partition based on the data information of the target ETCD data;
[0058] According to the event change type included in the target ETCD data, call the corresponding event processing function to process the target ETCD data; and,
[0059] Add the target ETCD data to the data storage tree in the target controller partition; wherein, the data storage tree is obtained by processing the ETCD data stored in the target controller partition using a tree-like structure storage method.
[0060] In the above data processing method, device, computer device and readable storage medium, the API management layer connected to each candidate controller partition in the master node obtains the target ETCD data when detecting data updates in the distributed storage system ETCD, and determines the target controller partition corresponding to the target ETCD data from each candidate controller partition according to the data information of the target ETCD data, and then sends the target ETCD data to the target controller partition to process the target ETCD data within the target controller partition. By adopting the above method, through the distributed deployment of the controller, the problem of the storage capacity bottleneck existing in K8S can be effectively alleviated, thereby improving the efficiency of cloud computing resource processing. Brief Description of the Drawings
[0061] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0062] Figure 1 It is an application environment diagram of the data processing method in an embodiment;
[0063] Figure 2 It is a flowchart of the data processing method in an embodiment;
[0064] Figure 3 It is the intention of the controller partition in an embodiment;
[0065] Figure 4 It is a flowchart of determining the target controller partition in an embodiment;
[0066] Figure 5 It is a flowchart of determining the target controller partition in another embodiment;
[0067] Figure 6 It is a flowchart of data query in an embodiment;
[0068] Figure 7 It is a schematic diagram of query result processing in an embodiment;
[0069] Figure 8 It is a flowchart of determining the system operation state in an embodiment;
[0070] Figure 9 It is a schematic flow chart of a data processing method in another embodiment;
[0071] Figure 10 It is a schematic diagram of a data record storage structure in an embodiment;
[0072] Figure 11 It is a schematic diagram of an index data storage structure in an embodiment;
[0073] Figure 12 It is a schematic flow chart of a data processing method in another embodiment;
[0074] Figure 13 It is a structural block diagram of a data processing device in an embodiment;
[0075] Figure 14 It is a structural block diagram of a data processing device in another embodiment;
[0076] Figure 15 It is an internal structure diagram of a computer device in an embodiment. Detailed implementation manners
[0077] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.
[0078] With the continuous development of cloud computing infrastructure service IaaS, in order to reasonably schedule the resources of a cloud computing cluster, an open-source container orchestration and management system K8S can be adopted to reasonably schedule cloud computing resources.
[0079] However, since K8S is built based on the distributed storage system ETCD, and due to the limitation of the storage capacity upper limit of 8G of ETCD itself, K8S will also have a storage capacity bottleneck, which will cause K8S not to support large-scale clusters, thereby reducing the efficiency of cloud computing resource processing.
[0080] Based on this, in an exemplary embodiment, a data processing method is provided, and this method is applied to the application programming interface API management layer of the master node in a cloud computing resource management system. Among them, reference can be made to Figure 1The framework of the cloud computing resource management system shown. The cloud computing resource management system includes a master node Master, each worker node Node, and a distributed storage system ETCD. Further, the master node Master includes an API management layer (API Server), a scheduler Scheduler, and each candidate controller partition Controller_Partition connected to the API management layer. The controller Controller is a distributed controller.
[0081] Further, as Figure 2 shown, the data processing method specifically includes the following steps:
[0082] S201, when it is detected that there is data update in the distributed storage system ETCD, obtain the target ETCD data.
[0083] Among them, the so-called target ETCD data is the updated data in ETCD. Further, ETCD is an important basic component in the cloud native architecture. In the K8S cluster, it can not only be used for service registration and discovery, but also as a middleware for key-value storage.
[0084] Optionally, a preset interface can be used to monitor the data update situation in ETCD in real time; then, when it is detected that there is data update in ETCD, obtain the updated ETCD data from ETCD.
[0085] S202, according to the data information of the target ETCD data, determine the target controller partition corresponding to the target ETCD data from each candidate controller partition.
[0086] Among them, the distributed storage system (ETC Distributed, ETCD) is a distributed, consistent key-value storage system; the so-called data information is the data attribute of the ETCD data, which may include but is not limited to information such as data type and data volume. The so-called candidate storage controller is any controller connected to the API management layer in the master node. The so-called target controller partition is the partition in each candidate controller partition that is adapted to the target ETCD data.
[0087] It can be understood that in this embodiment, based on the design goal of the cloud computing resource management system's bearing capacity upper limit of 10w Host, and at the same time considering the List & Watch performance bottleneck existing in K8S, the controller Controller can be designed as a distributed architecture, and the ETCD data can be partitioned and cached at the controller Controller. That is, based on the data volume borne by the cloud computing resource management system, the controller Controller is split into multiple Partitions, and each Partition is distributedly deployed.
[0088] Among them, the K8s Host mode means that the smallest scheduling unit, the Pod, is directly scheduled to run on the host machine instead of on the Kubernetes worker node, the Node. In this mode, the Pod does not allocate any resources but directly uses the CPU, memory, and storage resources of the host machine. The List & Watch mechanism is a resource monitoring mechanism provided by the Kubernetes API management layer. It allows the client to obtain a list (List) of resource objects through the API management layer and to monitor the changes of resource objects in real time by establishing a long connection (Watch).
[0089] Optionally, after obtaining the target ETCD data, a target controller partition adapted to the target ETCD data can be selected from each candidate controller partition based on data information such as the substring, data volume, and data type of the target ETCD data.
[0090] Exemplarily, the target ETCD data can be hashed, and the character information obtained from the hashing can be used as index information to perform a consistency query in the correspondence between each candidate controller partition and the partition identifier to obtain the target controller partition.
[0091] Alternatively, the target controller partition adapted to the target ETCD data can be selected from each candidate control partition according to the data type of the target ETCD data. Or, according to the data volume of the target ETCD data, a target controller partition can be selected from each candidate control partition whose currently survivable data volume is greater than the data volume of the target ETCD data.
[0092] S203, Send the target ETCD data to the target controller partition to process the target ETCD data within the target controller partition.
[0093] Optionally, after the API management layer sends the target ETCD data to the target controller partition, the target ETCD data can be processed within the target controller partition.
[0094] Exemplarily, such as Figure 3As shown, after the Informer in the target controller partition monitors the target ETCD data using the List & Watch mechanism, it will use the Refector to put the target ETCD data into the Delta FIFO queue for processing. Among them, Refector is the Reflector class defined in the cache package, which is used to monitor the resource changes in K8S, and its function is implemented by the List & Watch function. The Delta FIFO Queue for changed resources is a first-in-first-out queue used to cache the change events and resource objects pulled by the Refector. Among them, Delta represents the storage of changed resource objects, including the type and data of the resource objects.
[0095] When processing the target ETCD data in the Delta FIFO Queue, on the one hand, the Indexer module can be directly used to locally cache the target ETCD data. On the other hand, according to the data type of the target ETCD data (for example, add, update, and delete), the corresponding processing functions (for example, the add function AddFunc, the update function UpdateFunc, and the delete function DeleteFunc) can be called to process the target ETCD data in the Work Queue, and the processing results will be sent to the Control Loop for processing after processing.
[0096] Among them, the Work Queue is used to balance the processing speeds of the Informer and the Control Loop and decouple them at the same time; the Control Loop is the core implementation of the controller, realizing task-driven oriented to the final state. Further, the internal driving logic of the ControlLoop is driven by a state machine, and the state machines corresponding to each step in the life cycle of the cloud host can be defined, and the state machine is driven by the Loop until the entire business logic is completed. The driving logic implicitly includes an automatic retry mechanism for failed business logics until success.
[0097] The implementation of the Control Loop is divided into three steps: 1) Start the main Control Loop (MainControl Loop) of the current controller; 2) Define the state machine of the controller; 3) Implement the state-based handler Handler. Among them, Handler represents a program that receives requests to create containers and also corresponds to the running of a container.
[0098] The interface design of the master node needs to ensure idempotency, and the underlying virtualization also needs to ensure the idempotency of operations. For scenarios that require asynchronous callbacks, they are also implemented in the controller, providing callback interfaces, and the Loop continues to drive after the asynchronous callback.
[0099] In the above data processing method, when the API management layer connected to each candidate controller partition in the master node detects data updates in the distributed storage system ETCD, it obtains the target ETCD data, and based on the data information of the target ETCD data, determines the target controller partition corresponding to the target ETCD data from each candidate controller partition, and then sends the target ETCD data to the target controller partition to process the target ETCD data within the target controller partition. By adopting the above method, through the distributed deployment of controllers, the problem of storage capacity bottleneck existing in K8S can be effectively alleviated, thereby improving the efficiency of cloud computing resource processing.
[0100] To ensure the accuracy of determining the target controller partition, based on the above embodiments, in the embodiments of the present application, the data information includes the data type and the data volume. Based on this, an optional way to determine the target controller partition is provided, as Figure 4 shown, which specifically includes the following steps:
[0101] S401, according to the data type of the target ETCD data, determine each alternative controller partition adapted to the target ETCD data from each candidate controller partition.
[0102] Among them, the so-called alternative controller partition is the controller partition whose stored data type is the same as the data type of the target ETCD data.
[0103] Among them, the core function of ETCD is to store key-value pair data, and both the key and the value are of string type. The key is usually used to identify data, and the value can be a string in any format, such as the lightweight data interchange format JSON, the extensible markup language XML, or plain text. In addition, other types of data can also be included in ETCD.
[0104] Exemplarily, it can include structured data. Although ETCD itself does not directly support complex data structures (such as lists, dictionaries, etc.), structured data can be stored by serializing the data into a string (such as JSON). For example, a JSON object can be stored as a value. It can also include configuration data. ETCD is often used to store configuration information of distributed systems, such as service addresses, port numbers, environment variables, etc. These data are usually stored in the form of key-value pairs, which is convenient for dynamic update and query. It can also include metadata. In a distributed system, ETCD can store metadata of the cluster, such as node information, service status, version number, etc. These data are crucial for system coordination and consistency.
[0105] Optionally, to ensure the adaptability between the target controller partition and the target ETCD data, at least one alternative controller partition that is adapted to the target ETCD data can be screened out from each candidate controller partition according to the data type of the target ETCD data first.
[0106] Exemplarily, the data type of the target ETCD data can be hashed, and the character information obtained by the hashing process can be used as index information to perform a consistency query in the correspondence between each candidate controller partition and the data type identifier, and the candidate controller partition with a consistent query result is used as the alternative controller partition.
[0107] S402. Select a target controller partition from each alternative controller partition according to the size relationship between the data volume of the target ETCD data and the data volume threshold, and the current working information of each alternative controller partition.
[0108] Among them, the so-called data volume threshold is a value used to measure the size of the ETCD data volume. The so-called current working information is the data processing volume within the controller partition at the current moment.
[0109] Optionally, to improve the data processing speed of the target ETCD data, it can be determined whether the target ETCD data needs to be split according to the size relationship between the data volume of the target ETCD data and the data volume threshold; then, the number of controller partitions required can be determined according to the splitting situation of the target ETCD data.
[0110] Exemplarily, if the data volume of the target ETCD data is greater than the data volume threshold, the target ETCD data needs to be split; if the data volume of the target ETCD data is less than or equal to the data volume threshold, the target ETCD data does not need to be split.
[0111] Furthermore, a corresponding number of controller partitions with higher working efficiency can be selected from each alternative controller partition according to the current working information of each alternative controller partition as the target controller partition.
[0112] In the embodiments of the present application, by selecting the target controller partition from each candidate controller partition according to the data type and data volume of the target ETCD data, the rationality of the selection of the target controller partition is improved, and further the accuracy of the determined target controller partition is ensured.
[0113] To further ensure the accuracy of determining the target controller partition, on the basis of the above embodiments, in the embodiments of the present application, another optional method for determining the target controller partition is provided, as Figure 5 shown, which specifically includes the following steps:
[0114] S501. When the data volume of the target ETCD data is greater than the data volume threshold, perform splitting processing on the target ETCD data to obtain at least two target ETCD sub - data.
[0115] Among them, the so - called target ETCD sub - data is the data split from the target ETCD data.
[0116] Optionally, when the data volume of the target ETCD data is greater than the data volume threshold, based on a preset number of strings, perform splitting processing on the target ETCD data to obtain multiple target ETCD sub - data.
[0117] S502. According to the current working information of each alternative controller partition, select a target number of target controller partitions from each alternative controller partition.
[0118] Among them, the target number is the number of target ETCD sub - data.
[0119] Optionally, according to the current working information of each alternative controller partition, sort the current working efficiency of each alternative controller partition in descending order; then, the first target number of alternative controller partitions in the arrangement are used as the target controller partitions.
[0120] S503. Send a target ETCD sub - data to each target controller partition respectively to process the corresponding target ETCD sub - data within the target controller partition.
[0121] Among them, the target ETCD sub - data corresponding to different target controller partitions are different.
[0122] Optionally, a target ETCD sub - data can be sent to each target controller partition respectively. At this time, for each target controller partition, the corresponding target ETCD sub - data can be processed simultaneously within the target controller partition, thereby improving the efficiency of target ETCD data processing.
[0123] In addition, when the data volume of the target ETCD data is less than or equal to the data volume threshold, there is no need to split the target ETCD data. Instead, according to the current working information of each alternative controller partition, select a controller partition with the highest working efficiency from each alternative controller partition as the target controller partition, and send the target ETCD data to the target controller partition.
[0124] In the embodiment of the present application, when the data volume of the target ETCD data is greater than the data volume threshold, the number of target ETCD sub-data obtained after splitting the target ETCD data is used as the target number of the target controller partition, which can ensure the rationality of the determined target number, thereby ensuring the accuracy of the target controller partition.
[0125] To ensure the rationality of data processing, based on the above embodiment, in the embodiment of the present application, an optional method for data query is provided, as Figure 6 shown, which specifically includes the following steps:
[0126] S601, when a data query request is detected, at least one query target controller partition is selected from each candidate controller partition according to the data type of the data to be queried in the data query request.
[0127] Among them, the so-called data query request is a request to query the ETCD data in the controller; the so-called data to be queried is the ETCD data to be queried; the so-called query target controller partition is the controller partition that needs to perform data query based on the data query request.
[0128] It can be understood that since the ETCD data may be stored in any controller partition, after a data query request is detected, it is necessary to perform data query processing in each controller partition based on the data query request.
[0129] Based on this, to ensure the efficiency of data query, at least one query target controller partition that may store the data to be queried can be determined from each candidate controller partition according to the data type of the data to be queried in the data query request.
[0130] Optionally, the data type of the data to be queried can be obtained from the fixed string position in the data query request; subsequently, the data type can be hashed, and the character information obtained by the hashing process can be used as index information to perform a consistency query in the correspondence between each candidate controller partition and the data type identifier, and the candidate controller partition with a consistent query result is used as the query target controller partition.
[0131] S602, the data query request is sent to at least one query target controller partition, and data query sub-results are obtained from at least one query target controller partition.
[0132] Among them, the so-called data query sub-result is the query result obtained from the query target controller partition based on the data query request. Further, the query result can be ETCD data.
[0133] Optionally, the data query request can be sent to the controller partition to be queried, and the data query request can be processed in the data stored in the controller partition to be queried, so as to obtain a data query sub-result; then, the data query sub-result is fed back to the API management layer.
[0134] S603. Process the obtained data query sub-result to obtain the data query result corresponding to the data query request.
[0135] Optionally, in the case of obtaining only one, the query sub-result can be directly used as the data query result corresponding to the data query request; in the case of obtaining multiple data query sub-results, the data query sub-results can be combined and processed to obtain the data query result.
[0136] Exemplarily, in the case of obtaining multiple data query sub-results, the data query sub-results can be directly concatenated to obtain the data query result; or, a trained combination model can be used to process the data query sub-results to output the data query result.
[0137] Optionally, referring to Figure 7 As shown in the schematic diagram of query result processing, when the API management layer obtains multiple data query sub-results, the Merge layer in the fusion layer can be used to sort the data query sub-results fed back by each controller partition, and merge the data query sub-results after sorting. Among them, the Partition Store in the partition cache layer mainly realizes the local cache of the mapping relationship between the partition and the instance, and realizes data synchronization through the List & Watch mechanism.
[0138] In the embodiment of the present application, by processing the data query sub-result obtained from the controller partition to be queried to obtain the data query result, both the efficiency of data query and the accuracy of data query can be ensured.
[0139] Based on the above embodiments, in the embodiment of the present application, an optional method for determining the system operation state is provided. As Figure 8 shown, it specifically includes the following steps:
[0140] S801. Send a heartbeat acquisition request to each working node in the cloud computing resource management system to instruct each working node to split the dynamic heartbeat information from the standard heartbeat information and feed the dynamic heartbeat information back to the API management layer.
[0141] Among them, the so-called heartbeat acquisition request is a request to obtain heartbeat information from each working node; the standard heartbeat information includes static node information and dynamic heartbeat information; the static node information is generated based on the configuration information of the working node and is static and unchangeable, and may include the identification information, name, address information, etc. of the working node; the dynamic heartbeat information is generated based on a preset data packet and is dynamically changeable.
[0142] It can be understood that when the scale of the cloud computing resource management system reaches 100,000, the Agent in each working node will report the standard heartbeat information every 10 seconds, and each piece of data is 10k. Therefore, 1G of data volume update will be generated every 10 seconds. At this time, the API management layer will generate a huge CPU overhead due to the process of data serialization and deserialization, thus affecting the performance. Among them, Agent is an entity that can perceive the environment and make autonomous decisions.
[0143] Based on this, the standard heartbeat information reported by the Agent can be split into static node information and dynamic heartbeat information. Since the static node information is unchanged, when the working node receives the heartbeat acquisition request, it can split the dynamic heartbeat information from the standard heartbeat information and only feedback the dynamic heartbeat information to the API management layer.
[0144] Among them, the dynamic heartbeat information is stored using a lease, and the data size is about several hundred B. Originally, the Host was updated every 10 seconds, and now only the Lease needs to be updated every 10 seconds, greatly reducing the amount of updated data.
[0145] S802, determine the running state of the cloud computing resource management system according to the dynamic heartbeat information fed back by each working node.
[0146] Optionally, after obtaining the dynamic heartbeat information fed back by each working node, the dynamic heartbeat information can be subjected to security verification, and according to the security verification result, the running state of the cloud computing resource management system can be determined. Exemplarily, when the security verification of each working node passes, it is determined that the cloud computing resource management system is running stably; when the security verification of any working node fails, it is determined that there is an abnormality in the cloud computing resource management system. In addition, if there is a situation where the dynamic heartbeat information fed back by any working node is not obtained, it directly proves that there is an abnormality in the cloud computing resource management system.
[0147] It can be understood that in order to ensure the reliability of data processing, when it is determined that the running state of the cloud computing resource management system is stable, and then when it is detected that there is data update in the distributed storage system ETCD, the target ETCD data can be obtained and the subsequent data processing logic can be executed.
[0148] In an embodiment of the present application, by instructing each working node to split the dynamic heartbeat information from the standard heartbeat information and feedback the dynamic heartbeat information to the API management layer, the operating state of the cloud computing resource management system is determined, greatly reducing the amount of data to be updated, thereby improving the operating performance of the cloud computing resource management system.
[0149] In an exemplary embodiment, a data processing method is provided, and this method is applied to a target controller partition having a connection relationship with the API management layer in the master node, as Figure 9 shown, specifically including the following steps:
[0150] S901, Obtain the target ETCD data sent by the API management layer.
[0151] Among them, the target controller partition is determined from each candidate controller partition based on the data information of the target ETCD data.
[0152] Optionally, when the target controller partition detects that the API management layer has sent the target ETCD data, it obtains the target ETCD data.
[0153] S902, According to the event change type included in the target ETCD data, call the corresponding event processing function to process the target ETCD data.
[0154] Among them, the event change type is the change type of the event represented by the target ETCD data, and can include types such as addition, deletion, and update.
[0155] Optionally, the event change type can be obtained from the target ETCD data first, and then, the event processing function corresponding to the event change type of the target ETCD data is called from the preset event processing functions to process the target ETCD data.
[0156] S903, Add the target ETCD data to the data storage tree in the target controller partition.
[0157] Among them, the data storage tree is obtained by processing the ETCD data stored in the target controller partition using a tree structure storage method. Further, the tree structure storage method can be a balanced search tree B+Tree.
[0158] It can be understood that since the ETCD data mainly appears in the form of key-value pairs, in the IaaS business scenario, due to the existence of relatively complex retrieval scenarios such as querying by time range, simply using the key-value pair storage form cannot meet the requirements. Therefore, when storing the ETCD data in the controller partition, a tree structure storage method can be considered to store the ETCD data.
[0159] Optionally, after obtaining the target ETCD data, a tree-like structure storage method can be adopted to add the target ETCD data to the target controller partition for storage.
[0160] Exemplarily, since the B + Tree is a commonly used data storage structure in databases and has good performance in addition, deletion, and modification. Therefore, the B + Tree structure can be used to store the ETCD data, and at the same time, the index and data records are stored separately. That is, the B + Tree structure can be used to store the indexes and data records of each ETCD data already stored in the controller partition to obtain a data storage tree; subsequently, after obtaining new ETCD data, the indexes and data records corresponding to the new ETCD data can continue to be added to the data storage tree based on the B + Tree structure.
[0161] Among them, when using the B + Tree structure to store data, consider Figure 10 The schematic diagram of the data record storage structure shown. Among them, ID1-6 are the identification information of the data records; V1-6 are the values Value. Refer to Figure 11 The schematic diagram of the index data storage structure shown. Among them, K1-6 are the keys Key.
[0162] In the embodiment of the present application, by adopting the tree-like structure storage method to store the target ETCD data, the flexibility of data storage can be ensured, and the reliability of data storage is improved.
[0163] Figure 12 It is a schematic flowchart of a data processing method in another embodiment. On the basis of the above embodiment, this embodiment provides another optional example of a data processing method. Combining Figure 12 , the specific implementation process is as follows:
[0164] S1201, send a heartbeat acquisition request to each working node in the cloud computing resource management system to instruct each working node to split the dynamic heartbeat information from the standard heartbeat information and feedback the dynamic heartbeat information to the API management layer.
[0165] Among them, the standard heartbeat information includes static node information and dynamic heartbeat information; the static node information is generated based on the configuration information of the working node; the dynamic heartbeat information is generated based on a preset data packet.
[0166] S1202, determine the running state of the cloud computing resource management system according to the dynamic heartbeat information fed back by each working node.
[0167] S1203, when the running state is stable operation and it is detected that there is data update in the distributed storage system ETCD, obtain the target ETCD data.
[0168] S1204. Determine each alternative controller partition adapted to the target ETCD data from each candidate controller partition according to the data type of the target ETCD data.
[0169] S1205. In the case where the data volume of the target ETCD data is greater than the data volume threshold, perform splitting processing on the target ETCD data to obtain at least two target ETCD sub-data.
[0170] In the case where the data volume of the target ETCD data is less than or equal to the data volume threshold, determine the target controller partition corresponding to the target ETCD data from each candidate controller partition according to the data information of the target ETCD data, and send the target ETCD data to the target controller partition to process the target ETCD data within the target controller partition.
[0171] S1206. Select a target number of target controller partitions from each alternative controller partition according to the current working information of each alternative controller partition.
[0172] Wherein, the target number is the number of the target ETCD sub-data.
[0173] S1207. Send one target ETCD sub-data to each target controller partition respectively, so as to call the corresponding event processing function within the target controller partition according to the event change type included in the target ETCD data to process the target ETCD sub-data; and add the target ETCD sub-data to the data storage tree in the target controller partition.
[0174] Wherein, the data storage tree is obtained by processing the ETCD data stored in the target controller partition in a tree-like structure storage mode.
[0175] S1208. In the case of detecting a data query request, select at least one query controller partition to be queried from each candidate controller partition according to the data type of the data to be queried in the data query request.
[0176] S1209. Send the data query request to at least one query controller partition to be queried, and obtain data query sub-results from at least one query controller partition to be queried.
[0177] S1210. Process the obtained data query sub-results to obtain the data query result corresponding to the data query request.
[0178] The specific processes of the above S1201 - S1210 can refer to the description of the above method embodiments, and their implementation principles and technical effects are similar, which will not be elaborated here.
[0179] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this document, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0180] Based on the same inventive concept, an embodiment of the present application further provides a data processing device for implementing the data processing method described above. The solution provided by this device for solving problems is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the data processing device provided below can refer to the limitations on the data processing method in the foregoing, and will not be repeated here.
[0181] In an exemplary embodiment, as Figure 13 shown, a data processing device 1 is provided, which is applied to the API management layer of the master node in the cloud computing resource management system, and includes: a first acquisition module 10, a controller determination module 20, and a data sending module 30, where:
[0182] The first acquisition module 10 is configured to acquire target ETCD data when it is detected that there is data update in the distributed storage system ETCD;
[0183] The controller determination module 20 is configured to determine a target controller partition corresponding to the target ETCD data from each candidate controller partition according to the data information of the target ETCD data;
[0184] The data sending module 30 is configured to send the target ETCD data to the target controller partition to process the target ETCD data within the target controller partition.
[0185] In an exemplary embodiment, the data information includes data type and data volume; the controller determination module 20 includes:
[0186] The first determination unit is configured to determine each alternative controller partition adapted to the target ETCD data from each candidate controller partition according to the data type of the target ETCD data;
[0187] A second determination unit, configured to select a target controller partition from each alternative controller partition according to the magnitude relationship between the data volume of the target ETCD data and the data volume threshold, and the current working information of each alternative controller partition.
[0188] In an exemplary embodiment, the second determination unit is specifically configured to:
[0189] In the case where the data volume of the target ETCD data is greater than the data volume threshold, perform splitting processing on the target ETCD data to obtain at least two target ETCD sub-data;
[0190] According to the current working information of each alternative controller partition, select a target number of target controller partitions from each alternative controller partition; wherein, the target number is the number of target ETCD sub-data;
[0191] Correspondingly, the data sending module 30 is configured to:
[0192] Send one target ETCD sub-data to each target controller partition respectively, so as to process the corresponding target ETCD sub-data within the target controller partition; wherein, the target ETCD sub-data corresponding to different target controller partitions are different.
[0193] In an exemplary embodiment, the data processing device 1 further includes a data query module, wherein the data query module is specifically configured to:
[0194] In the case of detecting a data query request, select at least one query controller partition to be queried from each candidate controller partition according to the data type of the data to be queried in the data query request;
[0195] Send the data query request to at least one query controller partition to be queried, and obtain data query sub-results from at least one query controller partition to be queried;
[0196] Process the obtained data query sub-results to obtain a data query result corresponding to the data query request.
[0197] In an exemplary embodiment, the data processing device 1 further includes a status determination module, wherein the status determination module is specifically configured to:
[0198] Send a heartbeat acquisition request to each working node in the cloud computing resource management system, so as to instruct each working node to split out dynamic heartbeat information from the standard heartbeat information and feedback the dynamic heartbeat information to the API management layer; wherein, the standard heartbeat information includes static node information and dynamic heartbeat information; the static node information is generated based on the configuration information of the working node; the dynamic heartbeat information is generated based on a preset data packet;
[0199] Determine the operating status of the cloud computing resource management system according to the dynamic heartbeat information fed back by each working node.
[0200] In an exemplary embodiment, as Figure 14 shown, another data processing device 2 is provided, which is applied to a target controller partition having a connection relationship with the API management layer in the master node, and includes: a second acquisition module 40, a data processing module 50, and a data storage module 60, including:
[0201] The second acquisition module 40 is configured to obtain target ETCD data sent by the API management layer; wherein, the target controller partition is determined from each candidate controller partition based on the data information of the target ETCD data;
[0202] The data processing module 50 is configured to call a corresponding event processing function according to the event change type included in the target ETCD data to process the target ETCD data; and,
[0203] The data storage module 60 is configured to add the target ETCD data to the data storage tree in the target controller partition; wherein, the data storage tree is obtained by processing the ETCD data stored in the target controller partition using a tree structure storage method.
[0204] Each module in the above data processing device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above modules.
[0205] In an exemplary embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 15 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store ETCD data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, a data processing method is implemented.
[0206] Those skilled in the art can understand that Figure 15 the structure shown in [the figure] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.
[0207] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0208] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0209] In one embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0210] It should be noted that the data involved in this application (including but not limited to ETCD data, etc.) are all data authorized by users or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0211] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.
[0212] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope recorded in the present application.
[0213] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A data processing method, characterized in that: An application programming interface API management layer applied to a master node in a cloud computing resource management system, wherein the API management layer is connected to each candidate controller partition in the master node; the method comprises: When it is detected that there is data update in the distributed storage system ETCD, the target ETCD data is obtained; According to the data information of the target ETCD data, determining the target controller partition corresponding to the target ETCD data from each candidate controller partition; The target ETCD data is sent to the target controller partition so as to process the target ETCD data in the target controller partition.
2. The method according to claim 1, characterized in that The data information includes a data type and a data amount; and determining the target controller partition corresponding to the target ETCD data from each candidate controller partition according to the data information of the target ETCD data includes: According to the data type of the target ETCD data, determining from the candidate controller partitions each candidate controller partition adapted to the target ETCD data; According to the size relationship between the data volume of the target ETCD data and the data volume threshold, and the current working information of each candidate controller partition, a target controller partition is selected from each candidate controller partition.
3. The method according to claim 2, characterized in that The selecting a target controller partition from each candidate controller partition according to the size relationship between the data volume of the target ETCD data and the data volume threshold, and the current working information of each candidate controller partition, includes: When the data volume of the target ETCD data is greater than the data volume threshold, the target ETCD data is segmented to obtain at least two target ETCD sub-data; According to the current working information of each candidate controller partition, a target number of target controller partitions are selected from each candidate controller partition; wherein the target number is the number of the target ETCD sub-data; Accordingly, sending the target ETCD data to the target controller partition so as to process the target ETCD data in the target controller partition includes: Send one target ETCD sub-data to each target controller partition respectively, so as to process the corresponding target ETCD sub-data in the target controller partition; wherein the target ETCD sub-data corresponding to different target controller partitions are different.
4. The method according to claim 1, characterized in that: The method further comprises: In the case where a data query request is detected, selecting at least one controller partition to be queried from each candidate controller partition according to the data type of the data to be queried in the data query request; Sending the data query request to the at least one controller partition to be queried, and obtaining a data query sub-result from the at least one controller partition to be queried; The acquired data query sub-results are processed to obtain a data query result corresponding to the data query request.
5. The method according to claim 1, characterized in that The method further comprises: Sending a heartbeat acquisition request to each working node in the cloud computing resource management system to instruct each working node to split dynamic heartbeat information from standard heartbeat information and feed back the dynamic heartbeat information to the API management layer; wherein the standard heartbeat information includes static node information and dynamic heartbeat information; the static node information is generated based on the configuration information of the working node; and the dynamic heartbeat information is generated based on a preset data packet; The operating status of the cloud computing resource management system is determined according to the dynamic heartbeat information fed back by each working node.
6. A data processing method, characterized in that: Applied to a target controller partition that is connected to an API management layer in a master node, the method includes: Acquire the target ETCD data sent by the API management layer; wherein the target controller partition is determined from each candidate controller partition based on the data information of the target ETCD data; According to the event change type contained in the target ETCD data, calling the corresponding event processing function to process the target ETCD data; and, The target ETCD data is added to the data storage tree in the target controller partition; wherein the data storage tree is obtained by processing the ETCD data stored in the target controller partition in a tree structure storage manner.
7. A data processing device, characterized in that: An application programming interface API management layer applied to a master node in a cloud computing resource management system, the device comprising: The first acquisition module is used to acquire target ETCD data when detecting that there is data update in the distributed storage system ETCD; A controller determination module, configured to determine a target controller partition corresponding to the target ETCD data from each candidate controller partition according to data information of the target ETCD data; The data sending module is used to send the target ETCD data to the target controller partition so as to process the target ETCD data in the target controller partition.
8. A data processing device, characterized in that: Applicable to a target controller partition that is connected to an API management layer in a master node, the device comprises: A second acquisition module is used for the target ETCD data sent by the API management layer; wherein the target controller partition is determined from each candidate controller partition based on the data information of the target ETCD data; A data processing module, used to call a corresponding event processing function according to the event change type contained in the target ETCD data to process the target ETCD data; and A data storage module is used to add the target ETCD data to the data storage tree in the target controller partition; wherein the data storage tree is obtained by processing the ETCD data stored in the target controller partition in a tree structure storage manner.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Cited By
Intelligent electric control system and control circuit board with same
CN120610480A