Request scheduling method and device, electronic equipment and storage medium
By obtaining the number of accesses of data nodes in a distributed system and scheduling, the problem of unbalanced load of data nodes under large pressure requests is solved, and the system stability is improved.
Patent Information
- Application Number
- CN202311484611.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-08
- Publication Date
- 2025-05-09
AI Technical Summary
In the scenario of high pressure request, there is a time lag in request scheduling using periodically reported load information, resulting in unbalanced load of data nodes and poor stability of the entire distributed system.
By obtaining the number of accesses of each data node within a specified cycle, the node to be accessed is determined from multiple data nodes according to the access type and number of accesses, and the data access request is scheduled to the node to achieve load balancing.
It effectively avoids the load imbalance of data nodes, improves the stability of the entire distributed system, and can better respond to large pressure requests.
Smart Images

Figure CN119960924A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a request scheduling method, device, electronic device and storage medium. Background Art
[0002] Request scheduling refers to the process of allocating received data access requests to multiple data nodes for processing in a distributed system to ensure load balancing of each data node. Currently, requests can be scheduled based on the load information periodically reported by each data node.
[0003] In the scenario of high-pressure requests, thousands of data access requests can be obtained in a short period of time. There is a time lag in using periodically reported load information for scheduling. If load information is still used for request scheduling, it is easy to cause load imbalance on data nodes, resulting in poor stability of the entire distributed system. Summary of the invention
[0004] The embodiments of the present application provide a request scheduling method, device, electronic device and storage medium, which can reasonably schedule data access requests to avoid load imbalance and improve the stability of the entire distributed system.
[0005] The present application provides a request scheduling method, the method comprising:
[0006] Obtaining a data access request of a file system to be accessed and an access type corresponding to the data access request, wherein the file system to be accessed includes a plurality of data nodes;
[0007] Obtaining the number of accesses to each of the data nodes within a specified period, where the specified period is consistent with a period for obtaining load information of each of the data nodes;
[0008] Determine at least one node to be accessed from the plurality of data nodes according to the access type and the number of accesses corresponding to each of the data nodes;
[0009] The data access request is dispatched to the at least one node to be accessed so that the at least one node to be accessed processes the data access request.
[0010] The embodiment of the present application also provides a request scheduling device, which includes:
[0011] A request acquisition unit, configured to acquire a data access request of a file system to be accessed and an access type corresponding to the data access request, wherein the file system to be accessed includes a plurality of data nodes;
[0012] An access count acquisition unit, used to acquire the access count of each of the data nodes within a specified period, where the specified period is consistent with a period for acquiring the load information of each of the data nodes;
[0013] A determination unit, configured to determine at least one node to be accessed from the plurality of data nodes according to the access type and the number of accesses corresponding to each of the data nodes;
[0014] The scheduling unit is used to schedule the data access request to the at least one node to be accessed, so that the at least one node to be accessed processes the data access request.
[0015] In some embodiments, the access type includes a read data type, and the determining unit further includes:
[0016] The parsing subunit is used to parse the data access request to obtain the data information to be read;
[0017] a positioning subunit, configured to locate at least one candidate data node from the plurality of data nodes according to the data information to be read;
[0018] The first determining subunit is used to determine a node to be visited from the candidate data nodes according to the number of visits corresponding to the candidate data nodes.
[0019] In some embodiments, the first determining subunit is further configured to:
[0020] Calculating the network distance between the requesting end and each of the candidate data nodes according to the connection relationship between the multiple data nodes and the requesting end, where the requesting end is a terminal that issues the data access request;
[0021] Sorting the candidate data nodes according to the network distance between the request end and each of the candidate data nodes to obtain a candidate data node sequence;
[0022] Using the access times corresponding to the candidate data nodes, the candidate data node sequence is sorted to obtain a target candidate data node sequence;
[0023] Based on the target candidate data node sequence, a node to be visited is determined from the candidate data nodes.
[0024] In some embodiments, the access type includes a write data type, and the determining unit further includes:
[0025] A calculation subunit, configured to calculate an average number of processing threads of the file system to be accessed in the specified period according to the number of processing threads of each data node in the specified period;
[0026] A random subunit, used to determine a data node to be used from the multiple data nodes;
[0027] A second determining subunit is used to determine the node to be accessed according to the average number of processing threads, the number of processing threads of the data node to be used, and the number of accesses;
[0028] The loop subunit is used to return to the step of determining a data node to be used from the multiple data nodes and subsequent steps if the number of the nodes to be visited is less than the preset number, until the number of the nodes to be visited reaches the preset number.
[0029] In some embodiments, the second determining subunit is further configured to:
[0030] Performing weighted processing on the average number of processing threads according to a preset weight to obtain a target average number of threads;
[0031] If the number of processing threads of the data node to be used is less than or equal to the target average number of threads, determining the node to be accessed according to the number of accesses to the data node to be used and the average number of processing threads;
[0032] If the number of processing threads of the data node to be used is greater than the target average number of threads, return to the step of determining a data node to be used from the multiple data nodes and subsequent steps until the number of nodes to be accessed reaches the preset number.
[0033] In some embodiments, the second determining subunit is further configured to:
[0034] Obtain a mapping relationship between a preset thread number range and an access threshold;
[0035] Determining a target threshold from the access thresholds according to a preset thread number range in which the average number of processing threads is located;
[0036] If the number of accesses to the data node to be used is less than or equal to the target threshold, determining the data node to be used as a node to be accessed;
[0037] If the number of visits to the data node to be used is greater than the target threshold, return to the step of determining a data node to be used from the multiple data nodes and subsequent steps until the number of nodes to be visited reaches the preset number.
[0038] In some embodiments, the request scheduling device further includes a data updating unit, which is used to: after determining at least one node to be accessed from the plurality of data nodes according to the access type and the number of accesses corresponding to each of the data nodes, the data updating unit is used to:
[0039] If the specified period has not ended, updating the number of visits corresponding to the at least one node to be visited;
[0040] If the specified period has ended, the number of accesses corresponding to each of the data nodes is reset.
[0041] An embodiment of the present application also provides an electronic device, including a memory storing multiple instructions; the processor loads instructions from the memory to execute the steps in any one of the request scheduling methods provided in the embodiments of the present application.
[0042] An embodiment of the present application also provides a computer-readable storage medium, which stores a plurality of instructions, and the instructions are suitable for a processor to load to execute the steps in any one of the request scheduling methods provided in the embodiments of the present application.
[0043] An embodiment of the present application also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps in any one of the request scheduling methods provided in the embodiments of the present application.
[0044] The embodiment of the present application can obtain a data access request of a file system to be accessed, and an access type corresponding to the data access request; then obtain the number of accesses to each data node within a specified period; determine a node to be accessed from multiple data nodes based on the access type and the number of accesses to each data node, and then schedule the data access request to the node to be accessed so that the node to be accessed can process the data access request.
[0045] The specified period is consistent with the period for obtaining the load information of the data node. By obtaining the number of accesses within the specified period, the instantaneous load situation of each data node can be known. When a large number of data access requests are received in a short period of time, the instantaneous load scheduling requests of each data node can be fully considered to avoid the load imbalance of the data nodes, thereby improving the stability of the file system to be accessed, so as to better cope with the scenario of high-pressure requests. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0047] Figure 1a It is a schematic diagram of an application scenario of the request scheduling method provided in an embodiment of the present application;
[0048] Figure 1bIt is a flowchart of the request scheduling method provided in an embodiment of the present application;
[0049] Figure 1c is a schematic diagram of calculating network distance provided in an embodiment of the present application;
[0050] Figure 1d is a schematic diagram of a multi-level sorting provided in an embodiment of the present application;
[0051] Figure 1e It is a schematic diagram of a process for determining a preset number of nodes to be visited provided in an embodiment of the present application;
[0052] Figure 1f is a schematic diagram of a mapping relationship between a preset thread number range and an access threshold provided in an embodiment of the present application;
[0053] Figure 1g It is a schematic diagram of updating the number of accesses provided in an embodiment of the present application;
[0054] Figure 1h It is a schematic diagram of resetting the number of accesses provided in an embodiment of the present application;
[0055] Figure 2a is a flowchart of a request scheduling method provided by another embodiment of the present application;
[0056] Figure 2b This is a schematic diagram of the overall architecture of the HDFS provided in the embodiment of the present application;
[0057] Figure 3 is a schematic diagram of the structure of a request scheduling device provided in an embodiment of the present application;
[0058] Figure 4 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0059] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.
[0060] Embodiments of the present application provide a request scheduling method, device, electronic device, and storage medium.
[0061] The request scheduling device can be integrated in an electronic device, which can be a terminal, a server, or other equipment. The terminal can be a mobile phone, a tablet computer, a smart Bluetooth device, a laptop, or a personal computer (PC), an intelligent voice interaction device, a smart home appliance, a vehicle terminal, an aircraft, or other equipment; the server can be a single server or a server cluster composed of multiple servers.
[0062] In some embodiments, the request scheduling device may also be integrated into multiple electronic devices. For example, the request scheduling device may be integrated into multiple servers, and the request scheduling method of the present application may be implemented by multiple servers.
[0063] In some embodiments, the server may also be implemented in the form of a terminal.
[0064] For example, refer to Figure 1a , shows a schematic diagram of an application scenario of the request scheduling method, wherein the application scenario may include a terminal 101 and a file system to be accessed 102. A client may be installed on the terminal 101, and a user may initiate a data access request to the file system to be accessed 102 through the client.
[0065] The file system 102 to be accessed may be a distributed file system, and may include a metadata node and multiple data nodes, wherein the metadata node is used to manage the namespace and data block mapping information of the file system 102 to be accessed, and maintain the metadata and data block replication status of the file system 102 to be accessed, etc. The data node is used to store and manage data blocks, and perform read and write operations on data blocks.
[0066] The metadata node may include multiple functional modules, such as read scheduling, write scheduling, instantaneous indicator module, load collection, etc. Read scheduling can be used to schedule data access requests of read data types, and write scheduling can be used to schedule data access requests of write data types; the instantaneous indicator module can be used to count the number of accesses to each data node within a specified period; and load collection can be used to obtain the load information reported by each data node according to a specified period.
[0067] That is, the file system 102 to be accessed can obtain the data access request sent by the terminal 101, and the access type corresponding to the data access request; obtain the number of accesses to each data node within a specified period, and the specified period is consistent with the period for obtaining the load information of each data node; determine at least one node to be accessed from multiple data nodes according to the access type and the number of accesses corresponding to each data node; and schedule the data access request to the node to be accessed so that the node to be accessed can process the data access request.
[0068] The following are detailed descriptions. It should be noted that the order of the following embodiments is not intended to limit the preferred order of the embodiments. It is understood that in the specific implementation of the present application, when the above embodiments of the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.
[0069] Cloud technology refers to a hosting technology that unifies hardware, software, network and other resources within a wide area network or local area network to achieve data computing, storage, processing and sharing.
[0070] Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. It can form a resource pool, which can be used on demand and is flexible and convenient. Cloud computing technology will become an important support. The backend services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites and more portal websites. With the rapid development and application of the Internet industry, in the future, each item may have its own identification mark, which needs to be transmitted to the backend system for logical processing. Data of different levels will be processed separately. All kinds of industry data need strong system backing support, which can only be achieved through cloud computing.
[0071] Cloud storage is a new concept that extends and develops from the concept of cloud computing. A distributed cloud storage system (hereinafter referred to as storage system) refers to a storage system that uses cluster applications, grid technology, and distributed storage file systems to bring together a large number of different types of storage devices (storage devices are also called storage nodes) in the network through application software or application interfaces to work together and provide external data storage and business access functions.
[0072] At present, the storage method of the storage system is: create a logical volume, and when creating a logical volume, allocate physical storage space for each logical volume. The physical storage space may be composed of disks of a storage device or several storage devices. The client stores data on a logical volume, that is, stores the data on the file system. The file system divides the data into many parts, each of which is an object. The object contains not only data but also additional information such as data identification (ID, ID entity). The file system writes each object into the physical storage space of the logical volume, and the file system records the storage location information of each object, so that when the client requests to access the data, the file system can allow the client to access the data according to the storage location information of each object.
[0073] The process of the storage system allocating physical storage space to a logical volume is as follows: based on the estimated capacity of the objects stored in the logical volume (this estimate often has a large margin relative to the actual capacity of the objects to be stored) and the grouping of independent redundant disk arrays (RAID, Redundant Array of Independent Disks), the physical storage space is pre-divided into stripes. A logical volume can be understood as a stripe, thereby allocating physical storage space to the logical volume.
[0074] In short, a database can be seen as an electronic filing cabinet - a place where electronic files are stored, and users can add, query, update, delete, and other operations on the data in the files. The so-called "database" is a collection of data that is stored together in a certain way, can be shared with multiple users, has as little redundancy as possible, and is independent of the application.
[0075] Big data refers to a collection of data that cannot be captured, managed, and processed by conventional software tools within a certain time frame. It is a massive, high-growth, and diverse information asset that requires new processing models to have stronger decision-making power, insight discovery, and process optimization capabilities. With the advent of the cloud era, big data has also attracted more and more attention. Big data requires special technologies to effectively process large amounts of data within a tolerable time frame. Technologies applicable to big data include large-scale parallel processing databases, data mining, distributed file systems, distributed databases, cloud computing platforms, the Internet, and scalable storage systems.
[0076] The file system to be accessed in the embodiment of the present application is a distributed system, which can provide file storage capabilities for big data services. Since big data services involve a large amount of calculations, the file system to be accessed needs to process a large number of data access requests. The embodiment of the present application schedules data access requests according to the number of accesses of the data node within a specified period, so that the load of each data node in the file system to be accessed is balanced, which can adapt to the scenario of big data services and provide guarantees for big data services.
[0077] In this embodiment, a request scheduling method is provided, such as Figure 1b As shown, the specific process of the request scheduling method can be as follows:
[0078] 110. Obtain a data access request for a file system to be accessed and an access type corresponding to the data access request.
[0079] The file system to be accessed refers to a data storage system to be accessed, which may refer to a system that provides file storage capabilities for big data services. The file system to be accessed may be a distributed system composed of multiple servers, each of which may be referred to as a data node, and each data node is a node that actually stores data blocks in the file system to be accessed.
[0080] A data access request refers to a request for accessing data, and an access type refers to the type of a data access request. Generally, the type of a data access request may include a read data type and a write data type. Among them, a data access request of a read data type refers to a request to read a corresponding data block from a data node, and a data access request of a write data type refers to a request to write a data block to a data node of a file system to be accessed, so as to store the data block in the file system to be accessed.
[0081] In some embodiments, the data access request may be initiated by the object, for example, when the object needs to access the file system to be accessed, the data access request may be initiated to the file system to be accessed through the client on the terminal, so that the file system to be accessed may receive the data access request. In other embodiments, the data access request may be automatically generated after the preset conditions are met, for example, in a scheduled task, a data access request is automatically generated at a specified interval.
[0082] 120. Obtain the number of accesses to each of the data nodes within a specified period.
[0083] The specified period is consistent with the period for obtaining the load information of each data node. The specified period can be set according to actual needs, for example, it can be 3 seconds, 5 seconds, etc., that is, every time the specified period is passed, each data node can report its own load information to the metadata node.
[0084] Since load information is periodic, if a large number of data access requests are received within a specified period, and the load information is still used for request scheduling, load imbalance is likely to occur, resulting in poor stability of the entire distributed system. Therefore, the number of accesses to the data node within the specified period can be obtained to capture the instantaneous load of the data node within the specified period, where the greater the number of accesses, the higher the instantaneous load.
[0085] The number of accesses to a data node in a specified period refers to the number of data access requests allocated to the data node in the specified period. The number of data access requests allocated to each data node may be counted each time a specified period is entered, that is, each time a data access request is allocated to each data node, the number of accesses to the data node increases by one.
[0086] 130. Determine at least one node to be accessed from the multiple data nodes according to the access type and the number of accesses corresponding to each of the data nodes.
[0087] The node to be accessed refers to the data node that processes the data access request. When the data access request is obtained, the access type corresponding to the data access request can be obtained at the same time, and at least one node to be accessed can be determined from multiple data nodes based on the access type and the number of accesses corresponding to each data node.
[0088] Data access requests of different access types need to use different parameters in addition to the number of accesses to determine the node to be accessed. As an implementation method, a mapping relationship between a preset access type and a preset parameter may be preset, and the preset parameter is a parameter other than the number of accesses required to determine the node to be accessed. For example, when the preset access type is a read data type, the corresponding preset parameter includes the network distance between the terminal that issues the access request and the data node; when the preset access type is a write data type, the corresponding preset parameter includes the load information reported by each data node. Among them, the load information reported by each data node may be the number of processing threads.
[0089] It can be seen that when the access type is a read data type, the parameters required to determine the node to be accessed include the network distance between the terminal that issues the access request and the data node and the number of accesses. When the access type is a write data type, the parameters required to determine the node to be accessed include the load information of each data node and the number of accesses. The following will explain them in detail.
[0090] In some embodiments, if the access type is a read data type, when determining the node to be accessed, the data access request can be parsed and processed to obtain the data information to be read; based on the data information to be read, at least one candidate data node is located from the multiple data nodes; and based on the number of accesses corresponding to the candidate data nodes, the node to be accessed is determined from the candidate data nodes.
[0091] When the access type of the data access request is a read data type, it indicates that a certain data needs to be read from the file system to be accessed. By parsing the data access request, the data information corresponding to the data to be read carried therein can be obtained. It should be noted that in order to improve the reliability and fault tolerance of the data in the file system to be accessed, when writing the data block to the file system to be accessed, the data block needs to be copied multiple times to obtain multiple data copies, and these data copies are usually distributed on different data nodes.
[0092] Therefore, after obtaining the data information to be read, the data information to be read can be used to locate at least one candidate data node from multiple data nodes, and the candidate data node is the data node storing the data to be read.
[0093] Since there is at least one candidate data node, and when reading data, it is only necessary to read from one candidate data node, the node to be accessed can be determined from the candidate data nodes based on the number of accesses to the candidate data node. In actual use, the file system to be accessed often has the characteristics of storage and computing integration, that is, the client and the data node may run on the same physical server. If the data node happens to be a candidate data node, reading data directly from the data node can effectively improve the access speed and save network bandwidth. In other words, in this case, the device running the client and the data node do not need to go through an intermediate device, which is equivalent to the network distance between the client and the data node being 0. Therefore, the node to be accessed can be determined from the candidate data nodes in combination with the network distance and the number of accesses.
[0094] Optionally, the network distance between the requesting end and each of the candidate data nodes may be calculated based on the connection relationship between the multiple data nodes and the requesting end, where the requesting end is a terminal that issues the data access request; the candidate data nodes may be sorted according to the network distance between the requesting end and each of the candidate data nodes to obtain a candidate data node sequence; the candidate data node sequence may be sorted using the number of accesses corresponding to the candidate data nodes to obtain a target candidate data node sequence; and the node to be accessed may be determined from the candidate data nodes based on the target candidate data node sequence.
[0095] The requesting end refers to the terminal that issues the data access request. In the network topology of the file system to be accessed, it may include data nodes, requesting ends, and multiple intermediate devices. The data nodes and requesting ends may be connected via intermediate devices, which may be routers or switches, etc. Based on the connection relationship between multiple data nodes and the requesting end, the network distance between the requesting end and each candidate data node may be calculated.
[0096] As an implementation method, for each candidate data node, the common ancestor node of the requesting end and the candidate data node can be calculated; the distance between the requesting end and the common ancestor node and the distance between the candidate data node and the common ancestor node are summed to obtain the network distance between the requesting end and the candidate data node.
[0097] For example, see Figure 1c, shows a schematic diagram of calculating the network distance, where the common ancestor node of data node 1 and the request end is switch 1, and the distance between the request end and switch 1 is 1, the distance between data node 1 and switch 1 is 1, and the network distance between the request end and data node 1 is 2. For another example, the common ancestor node between the request end and data node 3 is switch 2, and the distance between the request end and switch 2 is 2, the distance between data node 3 and switch 2 is 2, and the network distance between the request end and data node 3 is 4.
[0098] As an implementation method, when determining the common ancestor node, a node can be selected from the network topology structure as the target node, and usually an intermediate device can be used as the target node; then, starting from the request end and the candidate data node, a connection path to the target node can be found. At this time, two paths can be generated, and then the intersection of the two paths is used as the common ancestor node of the request end and the candidate data node.
[0099] The smaller the network distance between the requester and the candidate data node, the faster the data access request is processed and the less network bandwidth is consumed. The candidate data nodes can be sorted in multiple levels using the network distance and the number of accesses to obtain a target candidate data node sequence, which can be used to determine the node to be accessed. For example, see Figure 1d , showing a schematic diagram of multi-level sorting. After calculating the network distance between the request end and the candidate data node, the candidate data nodes can be sorted based on the network distance. In some embodiments, the first-level sorting refers to sorting the candidate data nodes according to the first rule and the network distance, wherein the first rule can be set according to actual needs, for example, it can be a sorting rule based on the network distance from large to small, or it can be a sorting rule based on the network distance from small to large. It should be noted that in the candidate data node sequence, there may be two adjacent candidate data nodes with the same network distance. For example, Figure 1c The obtained candidate data node sequence is data node 1, data node 2, and data node 3, where the network distances from data node 2 and data node 3 to the requesting end are both 4.
[0100] The second-level sorting re-sorts the candidate data node sequence according to the access times of the candidate data nodes to obtain the target candidate data node sequence. In some embodiments, the candidate data node sequence may be sorted according to the second rule and the access times of the candidate data nodes, where the second rule is consistent with the first rule. That is, if the first rule is sorting in ascending order, the second rule is also sorting in ascending order. Optionally, adjacent candidate data nodes with equal network distances may be determined from the candidate data node sequence, and the adjacent candidate data nodes with equal network distances are re-sorted using the access times.
[0101] The following only takes sorting in ascending order as an example for illustration. For example, the candidate data nodes include data node 1, data node 2, and data node 3. The network distances between the requesting end and each candidate data node are denoted as L1, L2, and L3, and L1 < L2 = L3. When sorting based on the network distance, L1 and L2 may be compared first. Since L1 < L2, the order remains unchanged. Then L2 and L3 are compared. Since L2 = L3, it remains unchanged until, among adjacent data nodes, the network distance of the previous data node is not greater than that of the subsequent data node. The finally obtained candidate data node sequence is data node 1, data node 2, data node 3.
[0102] Then, the re-sorting using the access times continues. Since L2 = L3, only data node 2 and data node 3 in the candidate data node sequence need to be sorted. If the access times of data node 2 is 3 and the access times of data node 3 is 1, then the order between data node 2 and data node 3 is swapped until, among adjacent data nodes, the access times of the previous data node are not greater than those of the subsequent data node. The finally obtained target candidate data node sequence is data node 1, data node 3, data node 2.
[0103] Then, the node to be accessed may be determined from the candidate data nodes according to the target candidate data node sequence. If both the first rule and the second rule are sorting in ascending order, the candidate data node ranked most靠前 in the target candidate data node sequence is determined as the node to be accessed. This candidate data node ranked most靠前 has the smallest network distance and the smallest access times. Determining it as the node to be accessed can effectively improve the data processing speed. In the aforementioned example, data node 1 may be directly determined as the node to be accessed.
[0104] In some implementations, for each candidate data node, the network distance between the candidate data node and the requesting end and the number of accesses to the candidate data node may be used to calculate the parameters to be used of the candidate data node. For example, a first weight may be set for the network distance, and a second weight may be set for the number of accesses; the network distance may be weighted based on the first weight to obtain a weighted distance; the number of accesses may be weighted based on the second weight to obtain a weighted number; and finally, the weighted distance and the weighted number may be summed to obtain the parameters to be used. The first weight and the second weight may be set according to actual needs, and are not specifically limited here. In this way, the parameters to be used corresponding to each candidate data node may be calculated, and then the candidate data node corresponding to the smallest parameter to be used may be determined as the node to be accessed.
[0105] In some embodiments, if the access type is a read data type, when determining the node to be accessed, the average number of processing threads of the file system to be accessed in the specified period can be calculated based on the number of processing threads of each of the data nodes in the specified period; a data node to be used is determined from the multiple data nodes; the node to be accessed is determined based on the average number of processing threads, the number of processing threads of the data node to be used, and the number of accesses; if the number of nodes to be accessed is less than a preset number, return to execute the step of determining a data node to be used from the multiple data nodes and subsequent steps until the number of nodes to be accessed reaches the preset number.
[0106] At the beginning of a specified period, each data node can report its load information, such as the number of processing threads, to the metadata node, so that at the beginning of the specified period, the number of processing threads corresponding to each data node can be obtained. The number of processing threads can be understood as the number of threads that are processing data on the data node. The larger the number of processing threads, the higher the load.
[0107] Based on the number of processing threads of each data node in a specified period, the average number of processing threads of the entire file system to be accessed in a specified period can be calculated. For example, the number of processing threads corresponding to all data nodes can be summed to obtain the total number of processing threads; then the total number of processing threads can be divided by the total number of data nodes to obtain the average number of processing threads.
[0108] When the access type of the data access request is a write data type, in order to improve the reliability and fault tolerance of the data in the file system to be accessed, the data block to be written can be copied a preset number of copies and written to a preset number of data nodes respectively. Therefore, it is necessary to determine a preset number of nodes to be accessed from multiple data nodes.
[0109] Optionally, a data node to be used may be determined from a plurality of data nodes, and the data node to be used may be any one of the plurality of data nodes. Then, the node to be accessed may be determined according to the average number of processing threads, the number of processing threads of the data node to be used, and the number of accesses; if the number of nodes to be accessed is less than a preset number, the step of determining a data node to be used from the plurality of data nodes and subsequent steps are returned to be executed until the number of nodes to be accessed reaches the preset number. For example, see Figure 1e , showing a schematic diagram of a process for determining a preset number of nodes to be visited.
[0110] In some embodiments, when determining the node to be accessed based on the average number of processing threads, the number of processing threads of the data node to be used, and the number of accesses, the average number of processing threads can be weighted according to a preset weight to obtain a target average number of threads; if the number of processing threads of the data node to be used is less than or equal to the target average number of threads, the node to be accessed is determined based on the number of accesses to the data node to be used and the average number of processing threads; if the number of processing threads of the data node to be used is greater than the target average number of threads, return to execute the step of determining a data node to be used from the multiple data nodes and subsequent steps until the number of nodes to be accessed is the preset number.
[0111] The preset weight is a preset weight of the average number of processing threads, and the preset weight can be set according to actual needs. For example, the preset weight can be 2 or 3.
[0112] After weighting the average number of processing threads according to the preset weight, the target average number of threads can be obtained, and the target average number of threads can be used to determine whether the data node to be used can be determined as the node to be accessed. For example, it can be determined whether the number of processing threads of the data node to be used is less than or equal to the target average number of threads; if the number of processing threads of the data node to be used is less than or equal to the target average number of threads, it indicates that the load of the data node to be used does not exceed the predetermined load, and the node to be accessed can continue to be determined based on the number of accesses to the data node to be used and the average number of processing threads.
[0113] In some embodiments, when determining the node to be accessed based on the number of accesses to the data node to be used and the average number of processing threads, a mapping relationship between a preset thread number range and an access threshold may be obtained; a target threshold is determined from the access threshold based on the preset thread number range in which the average number of processing threads is located; if the number of accesses to the data node to be used is less than or equal to the target threshold, the data node to be used is determined as the node to be accessed; if the number of accesses to the data node to be used is greater than the target threshold, return to execute the step of determining a data node to be used from the multiple data nodes and subsequent steps until the number of nodes to be accessed is the preset number.
[0114] The mapping relationship between the preset thread number range and the access threshold can be preset. For example, see Figure 1f , showing a schematic diagram of the mapping relationship between the preset thread number range and the access threshold. Figure 1f In the figure, the horizontal axis is the preset number of threads, and the vertical axis is the access threshold. When the preset number of threads is in the range of (0-500], the corresponding access threshold is 30; when the preset number of threads is in the range of (500-1000], the corresponding access threshold is 10; when the preset number of threads is in the range of (1000, +∞], the corresponding access threshold is 5.
[0115] According to the preset thread number range in which the average number of processing threads is located, the target threshold corresponding to the average number of processing threads can be determined. For example, when the average number of processing threads is in (500-1000], the target threshold is 10.
[0116] Then, the number of accesses to the data node to be used and the target threshold value can be compared. If the number of accesses to the data node to be used is less than or equal to the target threshold value, it indicates that the instantaneous load of the current data node to be used is small and can be determined as the node to be accessed. If the number of accesses to the data node to be used is greater than the target threshold value, it indicates that the current data node to be used has been assigned a large number of data access requests within the specified period, and the instantaneous load is high, so it cannot be determined as the node to be accessed. Then, the step of determining a data node to be used from multiple data nodes and subsequent steps can be returned to execute.
[0117] It should be noted that each time a data node to be used is determined, the data node to be processed is marked. When returning to execute to determine a data node to be used from multiple data nodes, a new data node to be used can be directly determined from the unmarked data nodes to avoid repeatedly determining the same data node as the data node to be used, thereby improving the efficiency of determining the data node to be accessed. For example, there are 10 data nodes in total. Assuming that the first data node is determined as the data node to be used, a preset mark is added to the first data node. When the data node to be used is subsequently returned to be determined, it is determined from the remaining 9 data nodes that do not have the preset mark. In the above manner, after determining the preset number of nodes to be accessed, the loop can end.
[0118] 140. Schedule the data access request to the at least one node to be accessed, so that the at least one node to be accessed processes the data access request.
[0119] After determining at least one node to be accessed, the address corresponding to the node to be accessed can be sent to the requesting end, so that the requesting end can send a data access request to the node to be accessed based on the address, and the node to be accessed can respond to the data access request.
[0120] If the access type of the data access request is a read data type, the node to be accessed may feed back the requested data to the requesting end. If the access type of the data access request is a write data type, the node to be accessed may write the data in the request so that the data is stored in the node to be accessed.
[0121] It should be noted that after determining the node to be visited, the number of visits to the data node within the specified period can also be updated. For example, if the specified period has not ended, the number of visits corresponding to the at least one node to be visited is updated; if the specified period has ended, the number of visits corresponding to each of the data nodes is reset.
[0122] Before the specified period ends, the number of visits to at least one node to be visited can be updated. For example, see Figure 1g , showing a schematic diagram of updating the access count, wherein the access count corresponding to the node to be accessed may be increased by a preset value, for example, the access count is added by 1 to obtain an updated access count, so that the updated access count can be used for scheduling other data access requests within a specified period.
[0123] If the specified cycle ends, a new specified cycle is entered, and the access count corresponding to each data node can be directly reset, that is, the access count is reset to the initial value, which can be 0.
[0124] The end of a specified period can be determined based on the load information reported by the data node. For example, see Figure 1h , showing a schematic diagram of resetting the number of accesses. Each data node in the file system to be accessed can report load information according to a specified period, and the load information can be obtained by the load collection module in the metadata node. When the load collection module obtains the reported load information of each data node, it can report notification information to the instantaneous indicator module. When the instantaneous indicator module receives the notification information, it can be considered that the specified period has ended, and the number of accesses corresponding to each data node can be reset.
[0125] The request scheduling scheme provided in the embodiment of the present application can be applied in various scenarios for scheduling data requests. For example, taking the scheduling of data access requests to a distributed file system as an example, the scheme provided in the embodiment of the present application can schedule data access requests to data nodes with less load, and load balancing can be achieved under high-pressure requests, thereby improving the stability of the distributed file system.
[0126] The method provided by the embodiment of the present application can obtain the data access request and the corresponding access type of the file system to be accessed; obtain the number of accesses of each data node within a specified period. According to the access type and the number of accesses of each data node, the node to be accessed is determined, and the data access request is scheduled to the node to be accessed so that the node to be accessed can process the data access request. The specified period can be the period in which the data node reports the number of processing threads. During this period, by counting the number of accesses to each data node, the instantaneous load of the data node can be captured, and then the instantaneous load is used for request scheduling, which can achieve load balancing under high-pressure requests. In this way, in the scenario of high-pressure requests, the stability of the file system to be accessed can be effectively improved.
[0127] The method described in the above embodiment will be further described in detail below.
[0128] In this embodiment, the method of the embodiment of the present application is described in detail by taking the file system to be accessed as the Hadoop Distributed File System (HDFS) as an example.
[0129] like Figure 2a As shown, the specific process of a request scheduling method is as follows:
[0130] 210. Obtain a data access request of HDFS, where the data access request includes a read data request and a write data request.
[0131] 220. Obtain load parameters of each data node within a specified period, where the load parameters include the number of processing threads and the number of accesses.
[0132] 230. For a data read request, a node to be accessed is determined using the network distance between the requesting end and each data node and the number of accesses.
[0133] 240. For a data write request, determine a node to be accessed using the number of processing threads and the number of accesses of each data node.
[0134] 250. Schedule the data access request to the node to be accessed so that the node to be accessed processes the data access request.
[0135] The contents of steps 210 to 250 can refer to the corresponding parts of the above-mentioned embodiment, and will not be repeated here. Figure 2b , showing the overall architecture of HDFS. Figure 2b The request scheduling method is described in detail. It should be noted that the request scheduling method may be executed by a metadata node of the HDFS system.
[0136] HDFS can be used to store and manage large-scale data sets. HDFS can include metadata nodes and data nodes. Metadata nodes can be used to manage file system naming controls and data block mapping relationships, dimension file system metadata and data block replication status, etc. Data nodes can be responsible for storing and managing data blocks and performing read and write operations on data blocks.
[0137] HDFS can provide file storage capabilities for big data. Big data services usually involve a lot of computing, such as MapReduce, Spark, HBase, etc. When using big data services, high-pressure scenarios often occur, such as receiving tens of thousands of data access requests from thousands of clients in an instant. These data access requests are finally located to the data node where the data block is located through the metadata node. However, due to the uneven distribution of data on the data node or the uneven request access volume, the data node load will be unbalanced. Each data node can report its number of processing threads to the load collection module of the metadata node according to a specified period. Since the number of processing threads is reported periodically, the specified period is usually 3 seconds. If in a scenario of instantaneous high pressure, the number of processing threads used for scheduling data access requests may not have been updated, and it is difficult to cope with the scenario of high-pressure requests. For this reason, the present application adds an instantaneous indicator module in the metadata node, which can store the number of times each data node is selected within a specified period, that is, the number of accesses, and use the number of accesses as an instantaneous indicator.
[0138] When a node to be accessed is determined according to each data access request, the number of accesses to the node to be accessed may be increased by one, and when the next designated cycle arrives, the number of accesses to the node to be accessed may be reset to an initial value.
[0139] Optionally, for a read data request, the read scheduling module can obtain the data node when the corresponding data is written as a candidate data node. Then, it can be optimized through multi-level sorting, which mainly includes two levels. The first level is to sort the candidate data nodes in ascending order based on the network distance between the request end and the candidate data node to obtain a candidate data node sequence. The second level is to sort the candidate data nodes with the same network distance again in ascending order based on the number of accesses on the basis of the candidate data node sequence to obtain a target candidate data node sequence, and finally sort the candidate data node at the front as the node to be accessed.
[0140] For write data requests, the write scheduling module can calculate the average number of processing threads for the entire distributed file system based on the number of processing threads reported by each data node. Since data is usually written to three data nodes, it is necessary to select three nodes to be accessed, and one can be selected from multiple data nodes as the data node to be used. Then, if the number of processing threads of the data node to be used is not greater than twice the average number of processing threads, and the number of accesses to the data node to be used is less than the target threshold corresponding to the average number of processing threads, it can be determined as a node to be accessed, thereby determining three nodes to be accessed.
[0141] After determining the node to be accessed, the metadata node can send the address of the node to be accessed to the requesting end, so that the requesting end can communicate directly with the node to be accessed based on the address. When reading data, the node to be accessed can feed back the data block to the requesting end, and when writing data, the node to be accessed can write the data block. Therefore, in the scenario of high-pressure requests, the load of the data node in a short period of time can be known based on the number of accesses, and request scheduling can be performed based on this, which can achieve load balancing under high-pressure requests and effectively improve the stability of the distributed file system.
[0142] From the above, it can be seen that the request scheduling method provided in the embodiment of the present application can obtain the number of accesses to each data node within a specified period, which can be the period in which the data node reports the number of processing threads. The instantaneous load of the data node within the specified period can be captured, and the instantaneous load is used for request scheduling, which can achieve load balancing under high-pressure requests. In this way, in the scenario of high-pressure requests, the stability of the file system to be accessed can be effectively improved.
[0143] In order to better implement the above method, the embodiment of the present application also provides a request scheduling device, which can be integrated in an electronic device, and the electronic device can be a terminal, a server, etc. Among them, the terminal can be a mobile phone, a tablet computer, a smart Bluetooth device, a laptop, a personal computer, an intelligent voice interaction device, a smart home appliance, a vehicle terminal, an aircraft, etc. The server can be a single server or a server cluster composed of multiple servers.
[0144] For example, in this embodiment, the method of the embodiment of the present application will be described in detail by taking the specific integration of the request scheduling device in the server as an example.
[0145] For example, Figure 3 As shown, the request scheduling device may include a request acquisition unit 310, an access number acquisition unit 320, a determination unit 330, and a scheduling unit 340, as follows:
[0146] (I) Request acquisition unit 310
[0147] The method is used to obtain a data access request of a file system to be accessed and an access type corresponding to the data access request, wherein the file system to be accessed includes multiple data nodes.
[0148] (II) Access Count Acquisition Unit 320
[0149] Used to obtain the number of accesses to each of the data nodes within a specified period, where the specified period is consistent with the period for obtaining the load information of each of the data nodes.
[0150] (III) Determination Unit 330
[0151] Used to determine at least one to-be-accessed node from the multiple data nodes according to the access type and the number of accesses corresponding to each of the data nodes.
[0152] In some embodiments, the access type includes a read data type, and the determining unit 330 further includes:
[0153] The parsing subunit is used to parse the data access request to obtain the data information to be read;
[0154] a positioning subunit, configured to locate at least one candidate data node from the plurality of data nodes according to the data information to be read;
[0155] The first determining subunit is used to determine a node to be visited from the candidate data nodes according to the number of visits corresponding to the candidate data nodes.
[0156] In some embodiments, the first determining subunit is further configured to:
[0157] Calculating the network distance between the requesting end and each of the candidate data nodes according to the connection relationship between the multiple data nodes and the requesting end, where the requesting end is a terminal that issues the data access request;
[0158] Sorting the candidate data nodes according to the network distance between the request end and each of the candidate data nodes to obtain a candidate data node sequence;
[0159] Using the access times corresponding to the candidate data nodes, the candidate data node sequence is sorted to obtain a target candidate data node sequence;
[0160] Based on the target candidate data node sequence, a node to be visited is determined from the candidate data nodes.
[0161] In some embodiments, the access type includes a write data type, and the determining unit 330 further includes:
[0162] A calculation subunit, configured to calculate an average number of processing threads of the file system to be accessed in the specified period according to the number of processing threads of each data node in the specified period;
[0163] A random subunit, used to determine a data node to be used from the multiple data nodes;
[0164] A second determining subunit is used to determine the node to be accessed according to the average number of processing threads, the number of processing threads of the data node to be used, and the number of accesses;
[0165] The loop subunit is used to return to the step of determining a data node to be used from the multiple data nodes and subsequent steps if the number of the nodes to be visited is less than the preset number, until the number of the nodes to be visited reaches the preset number.
[0166] In some embodiments, the second determining subunit is further configured to:
[0167] Performing weighted processing on the average number of processing threads according to a preset weight to obtain a target average number of threads;
[0168] If the number of processing threads of the data node to be used is less than or equal to the target average number of threads, determining the node to be accessed according to the number of accesses to the data node to be used and the average number of processing threads;
[0169] If the number of processing threads of the data node to be used is greater than the target average number of threads, return to the step of determining a data node to be used from the multiple data nodes and subsequent steps until the number of nodes to be accessed reaches the preset number.
[0170] In some embodiments, the second determining subunit is further configured to:
[0171] Obtain a mapping relationship between a preset thread number range and an access threshold;
[0172] Determining a target threshold from the access thresholds according to a preset thread number range in which the average number of processing threads is located;
[0173] If the number of accesses to the data node to be used is less than or equal to the target threshold, determining the data node to be used as a node to be accessed;
[0174] If the number of visits to the data node to be used is greater than the target threshold, return to the step of determining a data node to be used from the multiple data nodes and subsequent steps until the number of nodes to be visited reaches the preset number.
[0175] (IV) Scheduling unit 340
[0176] Used to schedule the data access request to the at least one node to be accessed, so that the at least one node to be accessed processes the data access request.
[0177] In some embodiments, the request scheduling device further includes a data updating unit, which is used to: after determining at least one node to be accessed from the plurality of data nodes according to the access type and the number of accesses corresponding to each of the data nodes, the data updating unit is used to:
[0178] If the specified period has not ended, updating the number of visits corresponding to the at least one node to be visited;
[0179] If the specified period has ended, the number of accesses corresponding to each of the data nodes is reset.
[0180] In specific implementation, the above units can be implemented as independent entities, or can be arbitrarily combined to be implemented as the same or several entities. The specific implementation of the above units can refer to the previous method embodiments, which will not be repeated here.
[0181] From the above, it can be seen that the request scheduling device of this embodiment can obtain the number of accesses to each data node within a specified period, and thus can know the instantaneous load situation of each data node. When a large number of data access requests are received in a short period of time, the instantaneous load scheduling requests of each data node can be fully considered, which can avoid the load imbalance of data nodes, thereby improving the stability of the file system to be accessed, so as to better cope with the scenario of high-pressure requests.
[0182] The embodiment of the present application also provides an electronic device, which can be a terminal, a server, etc. Among them, the terminal can be a mobile phone, a tablet computer, a smart Bluetooth device, a laptop, a personal computer, an intelligent voice interaction device, a smart home appliance, a vehicle terminal, an aircraft, etc. The server can be a single server or a server cluster composed of multiple servers, etc.
[0183] In some embodiments, the request scheduling device may also be integrated into multiple electronic devices. For example, the request scheduling device may be integrated into multiple servers, and the request scheduling method of the present application may be implemented by multiple servers.
[0184] In this embodiment, the electronic device of this embodiment is a server as an example for detailed description, for example, Figure 4 As shown, it shows a schematic diagram of the structure of the electronic device involved in the embodiment of the present application, specifically:
[0185] The electronic device may include components such as a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, an input module 404, and a communication module 405. Those skilled in the art will appreciate that Figure 4 The electronic device structure shown in the figure does not constitute a limitation on the electronic device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0186] The processor 401 is the control center of the electronic device. It uses various interfaces and lines to connect various parts of the entire electronic device. By running or executing software programs and / or modules stored in the memory 402, and calling data stored in the memory 402, it performs various functions of the electronic device and processes data, thereby performing overall detection of the electronic device. In some embodiments, the processor 401 may include one or more processing cores; in some embodiments, the processor 401 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the processor 401.
[0187] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 402 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices. Accordingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.
[0188] The electronic device also includes a power supply 403 for supplying power to various components. In some embodiments, the power supply 403 can be logically connected to the processor 401 through a power management system, so as to manage charging, discharging, and power consumption through the power management system. The power supply 403 can also include any components such as one or more DC or AC power supplies, recharging systems, power failure detection circuits, power converters or inverters, and power status indicators.
[0189] The electronic device may further include an input module 404, which may be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.
[0190] The electronic device may further include a communication module 405. In some embodiments, the communication module 405 may include a wireless module. The electronic device may perform short-range wireless transmission through the wireless module of the communication module 405, thereby providing the user with wireless broadband Internet access. For example, the communication module 405 may be used to help the user send and receive emails, browse web pages, and access streaming media.
[0191] Although not shown, the electronic device may further include a display unit, etc., which will not be described in detail herein. Specifically in this embodiment, the processor 401 in the electronic device will load the executable files corresponding to the processes of one or more application programs into the memory 402 according to the following instructions, and the processor 401 will run the application programs stored in the memory 402, thereby realizing various functions, as follows:
[0192] Obtaining a data access request of a file system to be accessed and an access type corresponding to the data access request, wherein the file system to be accessed includes a plurality of data nodes;
[0193] Obtaining the number of accesses to each of the data nodes within a specified period, where the specified period is consistent with a period for obtaining load information of each of the data nodes;
[0194] Determine at least one node to be accessed from the plurality of data nodes according to the access type and the number of accesses corresponding to each of the data nodes;
[0195] The data access request is dispatched to the at least one node to be accessed so that the at least one node to be accessed processes the data access request.
[0196] The specific implementation of the above operations can be found in the previous embodiments, which will not be described in detail here.
[0197] From the above, it can be seen that the embodiment of the present application can obtain the number of accesses to each data node within a specified period, and thus the instantaneous load situation of each data node can be known. When a large number of data access requests are received in a short period of time, the instantaneous load scheduling requests of each data node can be fully considered, which can avoid the load imbalance of the data nodes, thereby improving the stability of the file system to be accessed, so as to better cope with the scenario of high-pressure requests.
[0198] A person of ordinary skill in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be completed by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.
[0199] To this end, an embodiment of the present application provides a computer-readable storage medium, in which a plurality of instructions are stored, and the instructions can be loaded by a processor to execute the steps in any one of the request scheduling methods provided in the embodiments of the present application. For example, the instructions can execute the following steps:
[0200] Obtaining a data access request of a file system to be accessed and an access type corresponding to the data access request, wherein the file system to be accessed includes a plurality of data nodes;
[0201] Obtaining the number of accesses to each of the data nodes within a specified period, where the specified period is consistent with a period for obtaining load information of each of the data nodes;
[0202] Determine at least one node to be accessed from the plurality of data nodes according to the access type and the number of accesses corresponding to each of the data nodes;
[0203] The data access request is dispatched to the at least one node to be accessed so that the at least one node to be accessed processes the data access request.
[0204] The storage medium may include: a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0205] According to one aspect of the present application, a computer program product or a computer program is provided, the computer program product or the computer program including computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the method provided in various optional implementations of the request scheduling aspect provided in the above embodiments.
[0206] Since the instructions stored in the storage medium can execute the steps in any one of the request scheduling methods provided in the embodiments of the present application, the beneficial effects that can be achieved by any one of the request scheduling methods provided in the embodiments of the present application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.
[0207] The above is a detailed introduction to a request scheduling method, device, electronic device and storage medium provided in an embodiment of the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for technical personnel in this field, according to the idea of the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A request scheduling method, characterized in that: The method comprises: Obtaining a data access request of a file system to be accessed and an access type corresponding to the data access request, wherein the file system to be accessed includes a plurality of data nodes; Obtaining the number of accesses to each of the data nodes within a specified period, where the specified period is consistent with a period for obtaining load information of each of the data nodes; Determine at least one node to be accessed from the plurality of data nodes according to the access type and the number of accesses corresponding to each of the data nodes; The data access request is dispatched to the at least one node to be accessed so that the at least one node to be accessed processes the data access request.
2. The method according to claim 1, characterized in that The access type includes a read data type, and determining at least one to-be-accessed node from the plurality of data nodes according to the access type and the number of accesses corresponding to each of the data nodes includes: Parsing the data access request to obtain the data information to be read; Locating at least one candidate data node from the multiple data nodes according to the data information to be read; According to the number of visits corresponding to the candidate data nodes, a node to be visited is determined from the candidate data nodes.
3. The method according to claim 2, characterized in that The determining the node to be visited from the candidate data nodes according to the number of visits corresponding to the candidate data nodes includes: Calculating the network distance between the requesting end and each of the candidate data nodes according to the connection relationship between the multiple data nodes and the requesting end, where the requesting end is a terminal that issues the data access request; Sorting the candidate data nodes according to the network distance between the request end and each of the candidate data nodes to obtain a candidate data node sequence; Using the access times corresponding to the candidate data nodes, the candidate data node sequence is sorted to obtain a target candidate data node sequence; Based on the target candidate data node sequence, a node to be visited is determined from the candidate data nodes.
4. The method according to claim 1, characterized in that The access type includes a write data type, and determining at least one to-be-accessed node from the plurality of data nodes according to the access type and the number of accesses corresponding to each of the data nodes includes: Calculate the average number of processing threads of the file system to be accessed in the specified period according to the number of processing threads of each data node in the specified period; Determine a data node to be used from the plurality of data nodes; Determine the node to be accessed according to the average number of processing threads, the number of processing threads of the data node to be used, and the number of accesses; If the number of the nodes to be accessed is less than the preset number, return to the step of determining a data node to be used from the multiple data nodes and subsequent steps until the number of the nodes to be accessed reaches the preset number.
5. The method according to claim 4, characterized in that: The determining the node to be accessed according to the average number of processing threads, the number of processing threads of the data node to be used, and the number of accesses includes: Performing weighted processing on the average number of processing threads according to a preset weight to obtain a target average number of threads; If the number of processing threads of the data node to be used is less than or equal to the target average number of threads, determining the node to be accessed according to the number of accesses to the data node to be used and the average number of processing threads; If the number of processing threads of the data node to be used is greater than the target average number of threads, return to the step of determining a data node to be used from the multiple data nodes and subsequent steps until the number of nodes to be accessed reaches the preset number.
6. The method according to claim 5, characterized in that The step of determining the node to be accessed according to the number of accesses to the data node to be used and the average number of processing threads includes: Obtain a mapping relationship between a preset thread number range and an access threshold; Determining a target threshold from the access thresholds according to a preset thread number range in which the average number of processing threads is located; If the number of accesses to the data node to be used is less than or equal to the target threshold, determining the data node to be used as a node to be accessed; If the number of visits to the data node to be used is greater than the target threshold, return to the step of determining a data node to be used from the multiple data nodes and subsequent steps until the number of nodes to be visited reaches the preset number.
7. The method according to any one of claims 1 to 6, characterized in that: After determining at least one node to be accessed from the plurality of data nodes according to the access type and the number of accesses corresponding to each of the data nodes, the method further includes: If the specified period has not ended, updating the number of visits corresponding to the at least one node to be visited; If the specified period has ended, the number of accesses corresponding to each of the data nodes is reset.
8. A request scheduling device, characterized in that: The device comprises: A request acquisition unit, configured to acquire a data access request of a file system to be accessed and an access type corresponding to the data access request, wherein the file system to be accessed includes a plurality of data nodes; An access count acquisition unit, used to acquire the access count of each of the data nodes within a specified period, where the specified period is consistent with a period for acquiring the load information of each of the data nodes; A determination unit, configured to determine at least one node to be accessed from the plurality of data nodes according to the access type and the number of accesses corresponding to each of the data nodes; The scheduling unit is used to schedule the data access request to the at least one node to be accessed, so that the at least one node to be accessed processes the data access request.
9. An electronic device, characterized in that: It comprises a processor and a memory, wherein the memory stores a plurality of instructions; the processor loads instructions from the memory to execute the steps in the request scheduling method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the steps in the request scheduling method according to any one of claims 1 to 7.