Data processing method, apparatus, device, medium and program product
By diversion of data call requests in the computing node cluster and storage node cluster, the high bandwidth pressure problem faced by databases in large-scale computing node clusters is solved, and more efficient data transmission and model training is achieved.
Patent Information
- Application Number
- CN202110796724.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-14
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-07-14
AI Technical Summary
In a large-scale computing node cluster, the database faces extremely large bandwidth pressure and data transmission pressure, resulting in information blocking and unable to complete the model training task in time.
By obtaining the data request, the calling address of the data unit is determined, and the data units are called separately in the computing node cluster and the storage node cluster, and combined into the target data to divert the data call request of the database.
It reduces the data transmission pressure of the database when concurrent data calls are called in large-scale clusters, improves data call efficiency, and avoids information blockage.
Smart Images

Figure CN113535429B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer applications, and particularly to a data processing method, apparatus, device, medium, and program product. Background Art
[0002] With the continuous development of big data and deep learning technologies, a corresponding learning and training of a model is carried out by using a large amount of labeled or unlabeled data through deep learning methods, and finally a relatively accurate cognitive model is obtained. The trained deep learning model can reveal the complex and rich information carried in big data and make more accurate predictions for future or unknown events.
[0003] However, the training of the model requires a large amount of computing resources. In a distributed system, each computing node undertakes its own computing tasks, and when these tasks are executed, they will encounter calling pre-stored data in the database. In a large-scale computing node cluster, when a large number of computing nodes concurrently apply to the database to call data, it will cause great bandwidth pressure and data transmission pressure on the database, and in severe cases, information blockage will occur, and the training task of the model cannot be completed in time.
[0004] Therefore, how to perform data shunting on data call requests of the database has become a technical problem to be solved urgently. Summary of the Invention
[0005] This application provides a data processing method, apparatus, device, medium, and program product to solve the technical problem of how to perform data shunting on data call requests of the database.
[0006] In a first aspect, this application provides a data processing method, including:
[0007] Obtain a data request, where the data request is used to apply for calling target data for at least one first computing node in a computing node cluster, and the target data is composed of multiple data units;
[0008] Determine the call address of each data unit according to the data request, where the call address includes: a first address, and / or a second address, the first address is the address of at least one second computing node in the computing node cluster, and the second address is the storage address in the database of at least one target storage node in a storage node cluster;
[0009] Call all the data units according to the call address and combine all the data units into the target data.
[0010] In a possible design, when the method is applied to a management node in a computing node cluster or a storage node cluster, the obtaining the data request includes:
[0011] Receive, by a management node in the computing node cluster, the data request sent by at least one of the first computing nodes;
[0012] Correspondingly, determining the call addresses of the respective data units according to the data request includes:
[0013] Send the data request by the management node to at least one of the target storage nodes, so that the target storage node determines the call address according to the target data;
[0014] Receive, by the management node, the call address fed back by the target storage node;
[0015] Correspondingly, calling all the data units according to the call address and combining all the data units into the target data includes:
[0016] Through the management node, combine all the data units into the target data according to the call address, and send the target data to the first computing node.
[0017] In a possible design, when the method is applied to a computing node in a computing node cluster, obtaining the data request includes:
[0018] In the first computing node, in response to a trigger instruction of a preset task, determine the data request, where the data request is used to cause the first computing node to call the target data to execute the preset task;
[0019] Correspondingly, determining the call addresses of the respective data units according to the data request includes:
[0020] Through the first computing node, according to the target data and a preset segmentation method, send a second data request to at least one other computing node in the computing node cluster, where the second data request is used to obtain the data units from the respective data nodes;
[0021] Receive the response results returned by the respective other computing nodes, and determine whether all the data units have been received according to the response results;
[0022] If so, combine all the data units into the target data;
[0023] If not, send a third data request to at least one of the target storage units, where the third data request is used to obtain the remaining data units from the target storage unit.
[0024] In a possible design, before sending the data request to at least one of the target storage nodes, it further includes:
[0025] Obtain the working status information of each storage node in the storage node cluster;
[0026] According to the working status information, screen out at least one of the target storage nodes that meet the preset requirements from each of the storage nodes.
[0027] In a possible design, when this method is applied to a storage node in a storage node cluster, the obtaining of the data request includes:
[0028] Receive the data request through the target storage node;
[0029] Correspondingly, determining the call address of each data unit according to the data request includes:
[0030] Through the target storage node, determine each data unit according to the target data in the data request;
[0031] Through the target storage node, determine the first address corresponding to some or all of the data units in each computing node of the computing node cluster;
[0032] If all the data units are not included in the second computing node, determine the second address corresponding to the remaining data units in the database.
[0033] In a possible design, the target data includes a Docker image, and the Docker image is used to complete the construction of a target virtual environment on a host, and the target virtual environment corresponds to a preset user.
[0034] In a possible design, the storage node cluster includes multiple storage nodes, and each storage node includes: a Docker Registry component and an interface component, and various types of Docker images are stored in the Registry image library in the Docker Registry component.
[0035] Optionally, the interface component includes a URL unified resource locator interface based on the Nginx service platform.
[0036] In a possible design, the functions of the interface component include: caching the Docker image and authenticating the identity information of the user.
[0037] In a second aspect, the present application provides a data processing device, including:
[0038] An acquisition module, configured to acquire a data request for requesting to invoke target data for at least one first computing node in a computing node cluster, where the target data consists of multiple data units;
[0039] A processing module, configured to determine a call address for each of the data units according to the data request, where the call address includes: a first address, and / or a second address, the first address is the address of at least one second computing node in the computing node cluster, and the second address is a storage address in a database of at least one target storage node in a storage node cluster;
[0040] The processing module is further configured to call all the data units according to the call address and combine all the data units into the target data.
[0041] In a possible design, when the device is disposed on a management node in a computing node cluster or a storage node cluster, the acquisition module is configured to receive the data request sent by at least one of the first computing nodes through the management node in the computing node cluster;
[0042] The processing module is configured to:
[0043] Send the data request to at least one of the target storage nodes through the management node, so that the target storage node determines the call address according to the target data; receive the call address fed back by the target storage node through the management node;
[0044] The processing module is further configured to, through the management node, combine all the data units into the target data according to the call address and send the target data to the first computing node.
[0045] In a possible design, when the device is disposed on a computing node in a computing node cluster, the acquisition module is configured to determine the data request in response to a trigger instruction of a preset task in the first computing node, where the data request is used to enable the first computing node to invoke the target data to execute the preset task;
[0046] The processing module is configured to, through the first computing node, send a second data request to at least one other computing node in the computing node cluster according to the target data and a preset splitting method, where the second data request is used to obtain the data units from each of the data nodes;
[0047] The acquisition module is configured to receive response results returned by each of the other computing nodes;
[0048] The processing module is configured to determine whether all of the data units have been received according to the response result; if so, combine all of the data units into the target data; if not, send a third data request to at least one of the target storage units, where the third data request is used to obtain the remaining data units from the target storage unit.
[0049] In a possible design, the obtaining module is further configured to obtain the working status information of each storage node in the storage node cluster;
[0050] The processing module is further configured to screen out at least one of the target storage nodes that meet the preset requirements from each of the storage nodes according to the working status information.
[0051] In a possible design, the working status information includes: workload and available status.
[0052] In a possible design, when the device is disposed on a storage node in the storage node cluster, the obtaining module is configured to receive the data request through the target storage node;
[0053] The processing module is configured to determine each of the data units according to the target data in the data request through the target storage node; determine the first address corresponding to some or all of the data units in each computing node of the computing node cluster through the target storage node; if the second computing node does not include all of the data units, determine the second address corresponding to the remaining data units in the database.
[0054] In a possible design, the target data includes a Docker image, and the Docker image is used to complete the construction of a target virtual environment on a host machine, and the target virtual environment corresponds to a preset user.
[0055] In a possible design, the storage node cluster includes a plurality of storage nodes, and each storage node includes: a Docker Registry component and an interface component, and various types of Docker images are stored in the Registry image library in the Docker Registry component.
[0056] Optionally, the interface component includes a URL uniform resource locator interface based on the Nginx service platform.
[0057] In a possible design, the functions of the interface component include: caching the Docker image and authenticating the identity information of the user.
[0058] In a third aspect, the present application provides a cluster system, including: a computing node cluster and a storage node cluster based on a preset application container engine; wherein,
[0059] The computing node cluster includes a plurality of computing nodes and at least one management node. The computing nodes are used to execute preset tasks, and the management node is used to handle data interaction between the computing node cluster and the storage node cluster;
[0060] The storage node cluster includes a plurality of storage nodes. Each storage node includes: an image library component and an interface component. In the image library of the image library component, various types of image files are stored. The interface component includes: a uniform resource locator interface based on a preset service platform. The interface component is configured to: cache the image files and authenticate the user's identity information;
[0061] The cluster system is used to implement the data processing method described in any one of claims 1-5.
[0062] Optionally, the preset application container engine includes: a Docker application container engine, the image file includes: a Docker image, and the image library component includes: a Docker Registry component.
[0063] In a possible design, the preset task includes: a logical calculus task for a deep learning model.
[0064] In a possible design, the interface component includes: a URL uniform resource locator interface based on the Nginx service platform.
[0065] In a possible design, the storage node cluster further includes at least one storage management node.
[0066] In a fourth aspect, the present application provides an electronic device, including:
[0067] a processor; and,
[0068] a memory for storing executable instructions of the processor;
[0069] Wherein, the processor is configured to execute any possible data processing method provided in the first aspect by executing the executable instructions.
[0070] In a fifth aspect, the present application provides an air conditioner, including: a display element, an energy storage element, and any possible electronic device provided in the third aspect.
[0071] In a sixth aspect, the present application further provides a storage medium, in which a computer program is stored, and the computer program is used to execute any one of the possible data processing methods provided in the first aspect.
[0072] In a seventh aspect, the present application further provides a computer program product, including a computer program, which when executed by a processor, implements any one of the possible data processing methods provided in the first aspect.
[0073] The present application provides a data processing method, device, equipment, medium and program product. By obtaining a data request initiated for invoking target data for at least one first computing node in a computing node cluster, where the target data is composed of multiple data units, then according to the data request, the call addresses of the respective corresponding data units are found in the computing node cluster and the storage node cluster respectively, and then all the data units are obtained according to the call addresses, so as to combine and obtain the target data. The technical problem of how to perform data shunting on the data call request of the database is solved. The technical effect of obtaining the target data by scattering it to multiple locations instead of relying only on the database in the storage node cluster for transmission, thereby reducing the data transmission pressure of the database during large-scale cluster concurrent data calls is achieved. Description of the Drawings
[0074] In order to more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0075] Figure 1 It is a schematic structural diagram of a cluster system provided by the present application;
[0076] Figure 2 It is a schematic flowchart of the first data processing method provided by an embodiment of the present application;
[0077] Figure 3 It is a schematic flowchart of the second data processing method provided by an embodiment of the present application;
[0078] Figure 4 It is a schematic flowchart of the third data processing method provided by an embodiment of the present application;
[0079] Figure 5 It is a schematic flowchart of the fourth data processing method provided by an embodiment of the present application;
[0080] Figure 6 It is a schematic structural diagram of a data processing device provided by the present application;
[0081] Figure 7 A structural schematic diagram of an electronic device provided for this application. Specific implementation manners
[0082] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only some of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts, including but not limited to combinations of multiple embodiments, fall within the scope of protection of this application.
[0083] Terms such as "first", "second", "third", "fourth", etc. (if any) in the specification, claims, and above-mentioned accompanying drawings of this application are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of this application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these process, method, product, or device.
[0084] The following first defines and explains the professional terms involved in this application:
[0085] Docker is an open-source application container engine that allows developers to package an application and its dependencies (i.e., the application's running environment) into a portable container and then publish it to any popular Linux machine. It can also achieve virtualization. Containers use a sandbox mechanism completely and there will be no interfaces between them. It should be noted that Docker itself is not a container, but a tool for creating containers.
[0086] The emergence of the container concept enables various types or different versions of application programs built on different platforms or with different dependent running environments to run simultaneously in the same operating system without conflicts. If the operating system is compared to the sea, then the container is a ship on the sea, and the image is the container on the ship. The container is the medium between various applications and the operating system.
[0087] The three core concepts of Docker technology are: Image, Container, and Repository.
[0088] A Docker image can be understood as a replica of a developed application and its dependencies.
[0089] A Docker repository is a database used to store Docker images.
[0090] A Docker container is used to host multiple Docker images, enabling applications that are isolated from each other and unable to communicate with each other to run properly on the same operating system.
[0091] In the Docker Application Container Engine, the Docker Registry service is responsible for managing Docker images, and its role is similar to that of a warehouse administrator.
[0092] There are three roles in the Docker Registry service, namely index, registry, and registryclient.
[0093] The index is responsible for and maintains information about user accounts, image verification, and public namespaces. It uses the following components to maintain this information: Web UI, metadata storage, authentication service, and symbolization.
[0094] The registry is a repository for images and charts. However, it does not have a local database and does not provide user authentication. Database support is provided by S3, cloud files, and the local file system. In addition, authentication is performed through the Token method of the Index Auth service. Registries can have different types, such as:
[0095] (1) Sponsor Registry: A third-party registry image repository for customers and the Docker community to use.
[0096] (2) Mirror Registry: A third-party registry image repository only for customers to use.
[0097] (3) Vendor Registry: A registry image repository provided by the vendor that publishes Docker images.
[0098] (4) Private Registry: A registry image repository provided by a private entity with a firewall and an additional security layer.
[0099] Registry Client: Docker acts as a registry client to be responsible for maintaining push and pull tasks, as well as client authorization.
[0100] In the prior art, when a user wants to obtain and download an image, the working process of the Docker Registry service is as follows:
[0101] First, the user sends a request to the index to download the image;
[0102] Then the index issues a response, returning three relevant pieces of information: the registry where the image is located, the checksum of all layers included in the image, and a Token for authorization purposes.
[0103] It should be noted that the Token is only returned when there is an X-Docker-Token in the request header. Basic authentication is required for private repositories, which is not mandatory for public repositories.
[0104] Next, the user communicates with the registry through the Token returned in the response. The registry is fully responsible for the image and is used to store the basic image and inherited layers;
[0105] Then, the registry now needs to confirm with the index that the token is authorized;
[0106] Finally, the index will send a "true" or "false" flag to the registry image repository to determine whether to allow the user to download the required image.
[0107] However, when the above process of the user obtaining the image is applied in a large-scale cluster system, problems will occur because each computing node in the large-scale cluster system is equivalent to each user client. When multiple computing nodes simultaneously request to obtain images from the registry image repository, it will cause a rapid increase in the bandwidth requirements and data transmission pressure of the registry image repository.
[0108] The inventive concept of this application is introduced below:
[0109] Currently, in a large-scale cluster system, images are generally deployed through the Docker method. Each computing node in the computing node cluster pulls the corresponding image by requesting the storage node where the registry image repository is located. In a large-scale concurrent scenario, it greatly increases the bandwidth and data transmission pressure of the storage node.
[0110] Although the bandwidth and data transmission pressure of a single storage node can be alleviated by deploying multiple storage nodes with complete registry image repositories, due to cost limitations, it is impossible to increase the storage nodes indefinitely. Therefore, how to perform data shunting on the data call requests of the registry image library has become a technical problem to be solved urgently.
[0111] The inventors of this application found that for a large-scale cluster system, the images requested by multiple computing nodes may be the same, or images with the same constituent units or the same data units. For example, if multiple applications are based on the Java platform environment, then the data units of the Java platform environment in multiple images are the same. Therefore, the computing nodes that have obtained the images or have the same data units can be used as data sources for calling. This makes it possible to fully utilize the resources of each computing node in the large-scale cluster system as dynamic data sources, instead of only a limited number of storage nodes with complete registry image repositories as data sources as before. The call of the target data will not be congested on the storage nodes of the registry image repository, thus solving the data shunting problem of data call requests.
[0112] Next, several embodiments will be combined to introduce the specific steps of the data processing method provided by this application in detail.
[0113] Figure 1 It is a schematic structural diagram of a cluster system provided by this application. As Figure 1 shown, the cluster system 100 includes: a computing node cluster 110 and a storage node cluster 120. Both the computing node cluster 110 and the storage node cluster 120 are set based on the Docker application container engine.
[0114] Among them, the computing node cluster 110 includes: multiple computing nodes 111 and at least one management node 112. The computing nodes 111 are used to execute preset tasks, including logical calculus tasks for deep learning models;
[0115] The management node 112 is used to handle data interaction between the computing node cluster 110 and the storage node cluster 120;
[0116] The storage node cluster 120 includes: a plurality of storage nodes 121 and / or at least one management node (not shown in the figure). Each storage node 121 includes: a Docker Registry component 1211 and an interface component 1212. In the Registry image library of the Docker Registry component 1211, various types of Docker images are stored. The interface component 1212 includes: a URL (Uniform Resource Locator) unified resource location interface based on the Nginx service platform. The interface component 1212 is configured to: cache the Docker images and authenticate the identity information of users.
[0117] It should be noted that for caching Docker images, the native Dokcer Registry image library is based on the file system and has poor concurrency. To reduce the load pressure on the backend storage, a caching policy needs to be added. This application uses Nginx to implement the caching of URL (Uniform Resource Locator) interface data. To avoid excessive cache, the cache expiration time can be configured to only cache the hot data read within a preset recent time. By configuring the caching policy, the performance of concurrent downloading of the same image is greatly improved.
[0118] It should also be noted that the data processing method provided by this application can be applied to Figure 1 the management node of the computing node cluster 110 or the storage node cluster 120 of the cluster system shown, or can also be applied to the computing node 111 or the storage node 121.
[0119] First, through Figure 2 the embodiments shown below, an overall introduction to the data processing method provided by this application will be given, and then through Figure 3 、 Figure 4 and Figure 5 the embodiments shown, the specific implementation methods when applied to different nodes will be introduced.
[0120] Figure 2 This is a schematic flowchart of the first data processing method provided by the embodiments of this application. As Figure 2 shown, the specific steps of this data processing method include:
[0121] S201. Obtain a data request.
[0122] In this step, the data request is used to apply for invoking target data for at least one first computing node in the computing node cluster. The target data consists of multiple data units.
[0123] It should be noted that the data unit can be segmented according to a preset segmentation form or according to the original file organization form of the target data. For example, when a large document is transmitted, it is divided into multiple data packets, and the data size of each data packet has a limited requirement. For example, each data packet is 4M in size, then this data packet can be used as a data unit for management.
[0124] In this embodiment, the target data includes various types of Docker images.
[0125] It should also be noted that the data request is for the computing node to apply for and call a Docker image in order to complete or execute a preset task, such as the logical calculation task of a deep learning algorithm.
[0126] Specifically, at least one computing node in the computing node cluster triggers a data request when executing a preset task, and this computing node directly intercepts this data request;
[0127] Or,
[0128] The management node in the computing node cluster or the storage node cluster, such as the super node, intercepts the data request initiated by the computing node;
[0129] Or, the storage node in the storage node cluster receives the data request initiated by the computing node.
[0130] For the specific situation, reference can be made to other embodiments below.
[0131] S202. Determine the call addresses of each data unit according to the data request.
[0132] In this step, the call address includes: the first address, and / or the second address. The first address is the address of at least one second computing node in the computing node cluster, and the second address is the storage address in the database of at least one target storage node in the storage node cluster.
[0133] Specifically, the identifier of the target data is extracted from the data request, such as the number of the Docker image, and then the identifiers of each data unit that makes up the Docker image are determined according to the preset segmentation method or the organization form of the Docker image.
[0134] Then, first, according to the identifiers of each data unit, each computing node in the computing node cluster is traversed to check whether the corresponding data unit is stored in the computing node. If it exists, this computing node is the second computing node, and the address where it stores the corresponding data unit, such as the IP address, is the first address.
[0135] If all data units are found by traversing each computing node, then return each first address to the first computing node that initiated the data request.
[0136] If all available computing nodes have been traversed and not all data units have been found yet, then send the identifiers of the unfound data units to the database, such as the Registry image library. The Registry image library sends the storage addresses of these data units that were not found in the computing nodes, i.e., the second addresses, to the first computing node.
[0137] S203. Call all data units according to the call address and combine all data units into the target data.
[0138] In this step, call the data units from the call addresses corresponding to each data unit. Since each data unit can come from multiple second computing nodes in the computing node cluster, this makes it so that instead of the database needing to transmit the entire target data, it only needs to transmit partial data units, achieving the shunting of data calls to the database. Even more, if all data units can be obtained from the second computing nodes, then the data transmission pressure on the database is even smaller.
[0139] After obtaining all data units, just combine them into the target data according to the preset splitting method or the data connection identifiers between each data unit, and then the first computing node can achieve the call of the target data.
[0140] The data processing method provided in this embodiment obtains a data request initiated for at least one first computing node in the computing node cluster to request the call of the target data, where the target data is composed of multiple data units. Then, according to the data request, find the call addresses of the corresponding respective data units in the computing node cluster and the storage node cluster respectively, and then obtain all data units according to the call address, thereby combining to obtain the target data. This solves the technical problem of how to perform data shunting for data call requests to the database. It achieves the technical effect of obtaining the target data by dispersing it to multiple locations instead of relying only on the database in the storage node cluster for transmission, reducing the data transmission pressure on the database during large-scale cluster concurrent data calls.
[0141] When the data processing method provided in this application is applied to the management node in the computing node cluster or the storage node cluster, its specific implementation steps are as Figure 3 shown.
[0142] Figure 3 is a schematic flowchart of the second data processing method provided in the embodiment of this application. As Figure 3 shown, the specific steps of this data processing method include:
[0143] S301. Receive, by a management node in a computing node cluster, data requests sent by at least one first computing node.
[0144] In this step, when at least one first computing node in the computing node cluster executes a preset task, it triggers a call requirement for at least one target data, so that the first computing node sends a data request to obtain the target data.
[0145] The computing node sends the data request to the management node, or the computing node originally intended to send the data request to the storage node where the database is located, but in this embodiment, the management node intercepts the data request.
[0146] S302. Send, by the management node, a data request to at least one target storage node, so that the target storage node determines a call address according to the target data.
[0147] In this step, there are multiple storage nodes in the storage node cluster that contain complete database data, and the management node is responsible for the data interaction between the computing node cluster and the storage node cluster.
[0148] Specifically, the management node sends a connection request to at least one storage node, i.e., the target storage node, that is in an available state and has a load rate lower than a preset load threshold.
[0149] Then, the management node can establish a data connection, such as an HTTP connection, with the first target storage node that feedbacks the connection request, and send a data request to the target storage node through the HTTP protocol.
[0150] Optionally, the management node can also split the target data according to a preset splitting method to determine multiple data units. Then the management node establishes data connections with multiple target storage nodes, and then through the data connections, applies to the multiple target storage nodes to call the corresponding data units respectively.
[0151] In a possible design, before sending a data request to at least one target storage node, it further includes:
[0152] Obtain the working state information of each storage node in the storage node cluster;
[0153] According to the working state information, screen out at least one of the target storage nodes that meet the preset requirements from each of the storage nodes.
[0154] The working state information includes: workload and available state.
[0155] Specifically, such as Figure 1As shown in the figure, when the Docker Daemon server on the first computing node resolves the domain name address corresponding to the DockerRegistry image repository, it first filters out the storage nodes in the available state from each storage node containing the Docker Registry image repository; then obtains its workload, such as obtaining the load values of each storage node according to a preset load calculation model; then sorts the load values from low to high, and selects the top N storage nodes as the target storage nodes, or selects the storage nodes with load values lower than the preset load threshold as the target storage nodes.
[0156] Optionally, select the storage node with the lowest load value as the target storage node.
[0157] Next, resolve the domain name address corresponding to the target storage node through DNS (Domain Name System).
[0158] S303. Receive the call address fed back by the target storage node through the management node.
[0159] In this step, the target storage node traverses each computing node in the computing node cluster to determine whether it contains some or all of the data units corresponding to the target data, and feeds back the first address of the second computing node storing the corresponding data unit to the management node.
[0160] If all the data units are not included in the computing nodes of the computing node cluster, then the target storage node feeds back the storage address of the remaining data units in the database, that is, the second address, to the management node.
[0161] In this way, data shunting of the database is realized, reducing the bandwidth and data transmission pressure of the database.
[0162] In this embodiment, the target data includes Docker images. Through the P2P mode, the Docker images are divided into image chunks (i.e., data units) of a preset size (such as 4M), and the data requests for the Docker images are redirected, so that only a small part of the data call requests will actually reach the image repository (i.e., the database on the target storage node) to pull image chunks, and most requests will obtain image chunks from other computing nodes in the computing node cluster.
[0163] S304. Through the management node, combine all data units into the target data according to the call address.
[0164] S305. Send the target data to the first computing node.
[0165] The data processing method provided in this embodiment obtains a data request initiated for requesting the invocation of target data for at least one first computing node in a computing node cluster. The target data consists of multiple data units. Then, according to the data request, the invocation addresses of the respective data units are found in the computing node cluster and the storage node cluster respectively, and all the data units are obtained according to the invocation addresses, so as to combine and obtain the target data. This solves the technical problem of how to perform data shunting for data invocation requests in a database. It achieves the technical effect of obtaining the target data by scattering it to multiple locations instead of relying only on the database in the storage node cluster for transmission, thus reducing the data transmission pressure on the database during large-scale cluster concurrent data invocation.
[0166] When the data processing method provided in this application is applied to a computing node in a computing node cluster, its specific implementation steps are as Figure 4 shown.
[0167] Figure 4 It is a schematic flowchart of the third data processing method provided in an embodiment of this application. As Figure 4 shown, the specific steps of this data processing method include:
[0168] S401. In a first computing node, in response to a trigger instruction of a preset task, determine a data request.
[0169] In this step, the data request is used to enable the first computing node to invoke target data to execute a preset task.
[0170] In this embodiment, the target data includes Docker images. When the first computing node executes a preset task, such as a logical calculation task of a deep learning model or a virtual environment creation task, due to task requirements, the corresponding Docker images need to be used.
[0171] In the prior art, the first computing node directly pulls the corresponding Docker image from the storage node where the Docker Registry image library is located. As a result, when a large number of concurrent Docker image pulls occur, it causes bandwidth congestion in the Docker Registry image library, restricting the computing node from completing the preset task.
[0172] In this embodiment, after the computing node triggers the Docker image invocation requirement, it determines the data request, but does not directly send a pull request to the storage node.
[0173] S402. Through the first computing node, according to the target data and a preset segmentation method, send a second data request to at least one other computing node in the computing node cluster.
[0174] In this step, the second data request is used to obtain data units from each data node, and the target data in the data request in S401 can be divided into multiple data units according to a preset segmentation method.
[0175] The preset segmentation method includes: a data packet segmentation method corresponding to the data transmission protocol.
[0176] Specifically, the first computing node traverses other computing nodes in the computing node cluster to determine whether any of them contain some or all of the data units corresponding to the target data. If so, it obtains the first address of the second computing node storing the corresponding data units.
[0177] S403. Receive the response results returned by each other computing node, and determine whether all data units have been received according to the response results.
[0178] In this step, if all data units corresponding to the target data can be obtained from other computing nodes and transferred to the first computing node, then execute step S404; otherwise, execute step S405.
[0179] S404. Combine all data units into the target data.
[0180] S405. Send a third data request to at least one target storage unit.
[0181] In this step, the third data request is used to obtain the remaining data units from the target storage unit.
[0182] In a possible design, before sending the third data request to at least one target storage node, it further includes:
[0183] Obtaining the working status information of each storage node in the storage node cluster;
[0184] According to the working status information, screen out at least one target storage node that meets the preset requirements from each storage node.
[0185] The working status information includes: workload and available status.
[0186] Specifically, such as Figure 1As shown in the figure, when the Docker Daemon server on the first computing node resolves the domain name address corresponding to the DockerRegistry image repository, it first filters out the storage nodes in the available state from each storage node containing the Docker Registry image repository; then obtains its workload, such as obtaining the load values of each storage node according to the preset load calculation model; then sorts the load values from low to high, and selects the top N storage nodes as the target storage nodes, or selects the storage nodes with load values lower than the preset load threshold as the target storage nodes.
[0187] Optionally, select the storage node with the lowest load value as the target storage node.
[0188] Next, resolve the domain name address corresponding to the target storage node through DNS (Domain Name System).
[0189] S406. Combine all data units into the target data.
[0190] The data processing method provided in this embodiment obtains a data request for applying to call target data by at least one first computing node in a computing node cluster, where the target data is composed of multiple data units, and then according to the data request, finds the call addresses of the corresponding data units in the computing node cluster and the storage node cluster respectively, and then obtains all the data units according to the call addresses, so as to combine and obtain the target data. It solves the technical problem of how to perform data shunting on the data call request of the database. It achieves the technical effect of obtaining the target data by dispersing it to multiple locations, rather than relying solely on the database in the storage node cluster for transmission, thereby reducing the data transmission pressure of the database during large-scale cluster concurrent data calls.
[0191] When the data processing method provided in this application is applied to a storage node in a storage node cluster, its specific implementation steps are as Figure 5 shown.
[0192] Figure 5 It is a schematic flowchart of the fourth data processing method provided in an embodiment of this application. As Figure 5 shown, the specific steps of this data processing method include:
[0193] S501. Receive a data request through the target storage node.
[0194] In this step, the data request is used to enable the first computing node to call the target data to execute a preset task.
[0195] In this embodiment, the target data includes Docker images. When a preset task is executed on the first computing node, such as a logical calculus task of a deep learning model or a task of creating a virtual environment, the corresponding Docker image needs to be used due to task requirements.
[0196] The target storage node receives a data request sent by the first computing node.
[0197] S502. Through the target storage node, determine each data unit according to the target data in the data request.
[0198] In this step, the target storage node determines each data unit that can form the complete target data by parsing the target data, such as through a preset segmentation method or the document composition form of the target data.
[0199] S503. Through the target storage node, determine the first addresses corresponding to some or all of the data units among the computing nodes in the computing node cluster.
[0200] In this step, the target storage node traverses each computing node in the computing node cluster to determine whether it contains some or all of the data units corresponding to the target data. If so, obtain the first address of the second computing node storing the corresponding data unit. The first address is the storage address of the data unit in the second computing node.
[0201] S504. If the second computing node does not contain all the data units, determine the second addresses corresponding to the remaining data units in the database.
[0202] S505. Send the first address and the second address to the first computing node so that the first computing node can call each data unit to form the target data.
[0203] The data processing method provided in this embodiment obtains a data request initiated for at least one first computing node in the computing node cluster to request the target data. The target data is composed of multiple data units. Then, according to the data request, the call addresses of the corresponding data units are found in the computing node cluster and the storage node cluster respectively, and all the data units are obtained according to the call addresses, so as to combine the target data. It solves the technical problem of how to perform data diversion for data call requests in the database. It achieves the technical effect of obtaining the target data by scattering it to multiple locations instead of relying only on the database in the storage node cluster for transmission, reducing the data transmission pressure on the database during large-scale cluster concurrent data calls.
[0204] For the above Figures 2 - 5For any of the embodiments, in a possible design, the target data includes Docker images, which are used to complete the construction of a target virtual environment on a host machine, and the target virtual environment corresponds to a preset user.
[0205] In a possible design, the storage node cluster includes multiple storage nodes, and each storage node includes: a Docker Registry component and an interface component. Various types of the Docker images are stored in the Registry image library in the Docker Registry component.
[0206] Optionally, the interface component includes a URL (Uniform Resource Locator) unified resource location interface based on the Nginx service platform.
[0207] In a possible design, the functions of the interface component include: caching the Docker images and authenticating the identity information of users.
[0208] In any of the above embodiments, the authentication method for the identity information of each computing node in this application includes:
[0209] User authentication is performed using the base 64 encoding method of basic auth provided by the official Docker Registry, and then the OpenRestry extension of nginx is used to perform authorization for different URLs (Uniform Resource Locators). Through the container orchestration tool, the container is created with the uid user identity and gid group identity of the real user in Linux. The container orchestration tool needs to start docker with root privileges in Linux, and mount the home directory of this user and the private key of the root user when starting. The home directory of the user provides the files required by the user, and mounting the private key of the root user is for container scheduling and passwordless ssh access.
[0210] This user authentication method also supports user authentication after container solidification. By recording the uid user identity and gid group identity of the current user when the container orchestration tool starts, the previous user record will be cleared and a new user uid user identity and gid group identity will be created during the next access, thus ensuring that there is only one user in a container and preventing conflicts of user uid user identity or gid group identity. Therefore, the problem of having only one user in a single container is solved, and this user has sudo privileges, which is equivalent to having root ownership of the container, thus realizing user permission isolation.
[0211] Figure 6 The structural schematic diagram of a data processing device provided for this application. The data processing device can be implemented through software, hardware, or a combination of both.
[0212] As Figure 6 shown, the data processing device 600 provided in this embodiment includes:
[0213] An acquisition module 601, configured to acquire a data request, where the data request is used to apply for invoking target data for at least one first computing node in the computing node cluster, and the target data is composed of multiple data units;
[0214] A processing module 602, configured to determine the invocation address of each data unit according to the data request, where the invocation address includes: a first address, and / or a second address, the first address is the address of at least one second computing node in the computing node cluster, and the second address is the storage address in the database of at least one target storage node in the storage node cluster;
[0215] The processing module 602 is further configured to invoke all data units according to the invocation address and combine all data units into target data.
[0216] In a possible design, when the device is set on the management node in the computing node cluster or the storage node cluster, the acquisition module 601 is configured to receive the data request sent by at least one first computing node through the management node in the computing node cluster;
[0217] The processing module 602 is configured to:
[0218] Send a data request to at least one target storage node through the management node so that the target storage node determines the invocation address according to the target data; receive the invocation address fed back by the target storage node through the management node;
[0219] The processing module 602 is further configured to, through the management node, combine all data units into target data according to the invocation address and send the target data to the first computing node.
[0220] In a possible design, when the device is set on a computing node in the computing node cluster, the acquisition module 601 is configured to determine a data request in response to a trigger instruction of a preset task in the first computing node, where the data request is used to enable the first computing node to invoke target data to execute a preset task;
[0221] The processing module 602 is configured to send a second data request to at least one other computing node in the computing node cluster through the first computing node according to the target data and a preset segmentation method, where the second data request is used to obtain data units from each data node;
[0222] An acquisition module 601, configured to receive response results returned by each other computing node;
[0223] A processing module 602, configured to determine whether all data units have been received according to the response results; if so, combine all data units into target data; if not, send a third data request to at least one target storage unit, where the third data request is used to obtain the remaining data units from the target storage unit.
[0224] In a possible design, the acquisition module 601 is further configured to obtain the working status information of each storage node in the storage node cluster;
[0225] The processing module 602 is further configured to screen out at least one target storage node that meets the preset requirements from each storage node according to the working status information.
[0226] In a possible design, the working status information includes: workload and available status.
[0227] In a possible design, when the device is disposed on a storage node in the storage node cluster, the acquisition module 601 is configured to receive a data request through the target storage node;
[0228] The processing module 602 is configured to determine each data unit according to the target data in the data request through the target storage node; determine the first addresses corresponding to some or all of the data units in each computing node in the computing node cluster through the target storage node; if the second computing node does not contain all the data units, determine the second addresses corresponding to the remaining data units in the database.
[0229] In a possible design, the target data includes a Docker image, and the Docker image is used to complete the construction of a target virtual environment on a host, and the target virtual environment corresponds to a preset user.
[0230] In a possible design, the storage node cluster includes multiple storage nodes, and each storage node includes: a Docker Registry component and an interface component, and various types of Docker images are stored in the Registry image library in the Docker Registry component.
[0231] Optionally, the interface component includes a URL uniform resource locator interface based on the Nginx service platform.
[0232] In a possible design, the functions of the interface component include: caching Docker images and authenticating the identity information of users.
[0233] It should be noted thatFigure 6 The data processing device provided by the illustrated embodiment can execute the methods provided in any of the above method embodiments. Its specific implementation principles, technical features, explanations of professional terms, and technical effects are similar and will not be elaborated here.
[0234] Figure 7 This is a schematic structural diagram of an electronic device provided by the present application. As Figure 7 shown, the electronic device 700 may include: at least one processor 701 and a memory 702. Figure 7 The illustrated electronic device takes one processor as an example.
[0235] The memory 702 is used to store programs. Specifically, the program may include program code, and the program code includes computer operation instructions.
[0236] The memory 702 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk memory.
[0237] The processor 701 is used to execute the computer execution instructions stored in the memory 702 to implement the methods described in the above method embodiments.
[0238] Among them, the processor 701 may be a central processing unit (CPU for short), or an application specific integrated circuit (ASIC for short), or one or more integrated circuits configured to implement the embodiments of the present application.
[0239] Optionally, the memory 702 may be either independent or integrated with the processor 701. When the memory 702 is a device independent of the processor 701, the electronic device 700 may further include:
[0240] A bus 703 for connecting the processor 701 and the memory 702. The bus may be an industry standard architecture (ISA) bus, a peripheral component interconnect (PCI) bus, or an extended industry standard architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc., but it does not mean that there is only one bus or one type of bus.
[0241] Optionally, in a specific implementation, if the memory 702 and the processor 701 are integrated on a single chip, the memory 702 and the processor 701 can communicate through an internal interface.
[0242] The present application also provides a computer-readable storage medium, which may include: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes. Specifically, the computer-readable storage medium stores program instructions for the methods in the above-described embodiments.
[0243] The present application also provides a computer program product, including a computer program that, when executed by a processor, implements the methods in the above-described embodiments.
[0244] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A data processing method, characterized in that, Including: Obtaining a data request for requesting to invoke target data for at least one first computing node in a computing node cluster, where the target data consists of multiple data units; Determining the invocation addresses of all the data units according to the data request, and invoking all the data units according to the invocation addresses to combine all the data units into the target data; The invocation address includes: a first address and a second address, where the first address is the address of at least one second computing node in the computing node cluster, and the second address is the storage address in a database of at least one target storage node in a storage node cluster.
2. The data processing method according to claim 1, wherein When the method is applied to a management node in a computing node cluster or a storage node cluster, the obtaining the data request includes: Receiving, by the management node, the data request sent by at least one of the first computing nodes; Correspondingly, the determining the invocation addresses of all the data units according to the data request includes: Sending, by the management node, the data request to at least one of the target storage nodes, so that the target storage node determines the invocation address according to the target data; Receiving, by the management node, the invocation address fed back by the target storage node; Correspondingly, the invoking all the data units according to the invocation addresses to combine all the data units into the target data includes: Combining, by the management node, all the data units into the target data according to the invocation address, and sending the target data to the first computing node.
3. The data processing method according to claim 1, wherein When the method is applied to a computing node in a computing node cluster, the obtaining the data request includes: In the first computing node, in response to a trigger instruction of a preset task, determining the data request, where the data request is used to enable the first computing node to invoke the target data to execute the preset task; Correspondingly, the determining the invocation addresses of all the data units according to the data request, and invoking all the data units according to the invocation addresses to combine all the data units into the target data includes: Sending, by the first computing node, according to the target data and a preset segmentation method, a second data request to at least one other computing node in the computing node cluster, where the second data request is used to obtain the data units from each of the data nodes; Receiving the response results returned by each of the other computing nodes, and determining whether all the data units have been received according to the response results; If so, combining all the data units into the target data; If not, sending a third data request to at least one of the target storage nodes, where the third data request is used to obtain the remaining data units from the target storage node.
4. The data processing method according to claim 2, wherein Before sending a data request to at least one of the target storage nodes, it further includes: Obtaining the working state information of each storage node in the storage node cluster; According to the working state information, screening out at least one of the target storage nodes that meet the preset requirements from each of the storage nodes.
5. The data processing method according to claim 1, characterized in that, When the method is applied to a storage node in a storage node cluster, the obtaining of the data request includes: Receiving the data request through the target storage node; Correspondingly, determining the call addresses of the respective data units according to the data request includes: Determining the respective data units according to the target data in the data request through the target storage node; Determining, through the target storage node, the first addresses corresponding to some or all of the data units among the respective computing nodes in the computing node cluster; If all of the data units are not included in the second computing node, determining the second addresses corresponding to the remaining data units in the database; Correspondingly, calling all of the data units according to the call addresses to combine all of the data units into the target data includes: Sending the first address and the second address to the first computing node so that the first computing node calls the respective data units to combine into the target data.
6. A cluster system, characterized in that, Including: A computing node cluster and a storage node cluster based on a preset application container engine; wherein, The computing node cluster includes a plurality of computing nodes and at least one management node, the computing nodes are used to execute preset tasks, and the management node is used to handle data interaction between the computing node cluster and the storage node cluster; The storage node cluster includes a plurality of storage nodes, and each storage node includes: an image library component and an interface component. In the image library component, various types of image files are stored in the image library. The interface component includes: a uniform resource locator interface based on a preset service platform. The interface component is configured to: cache the image files and authenticate the user's identity information; The cluster system is used to implement the data processing method according to any one of claims 1-5.
7. A data processing device, characterized in that, Including: An obtaining module, configured to obtain a data request, where the data request is used to apply for calling target data for at least one first computing node in a computing node cluster, and the target data is composed of a plurality of data units; A processing module, configured to determine the call addresses of the respective data units according to the data request, and call all of the data units according to the call addresses to combine all of the data units into the target data; The call address includes: a first address and a second address. The first address is the address of at least one second computing node in the computing node cluster, and the second address is the storage address in the database of at least one target storage node in the storage node cluster.
8. An electronic device, characterized in that, Including: A processor; And, A memory, configured to store an executable computer program of the processor; Wherein, the processor is configured to execute the data processing method according to any one of claims 1 to 5 by executing the executable computer program.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the data processing method according to any one of claims 1 to 5.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the data processing method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Method and system for arranging server cluster in cloud computing cluster environment
CN103037002A
Distributed storage based Docker image downloading method
CN106506587A