A computing power scheduling method and system

By collecting resource information based on the entire network's computing power resources under the computing power scheduling task, using a distributed architecture and CFN protocol, and combining virtual machine and container scheduling mechanisms, the problem of insufficient computing resources in computing tasks is solved, and efficient and reliable computing power scheduling is achieved.

CN119127473BActive Publication Date: 2026-04-07NANJING SUYI IND +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-19
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

In existing technologies, computing tasks cannot be executed due to insufficient resources caused by the allocation of unsuitable computing resources.

Method used

By determining the target scheduling path of computing power service requirements under different preset dimensions based on the computing power resource information of the entire network under the computing power scheduling task, and collecting computing power node resource information by adopting a distributed architecture and CFN protocol, combined with the scheduling mechanism of virtual machines and containers, efficient and reliable computing power scheduling is achieved.

Benefits of technology

It solves the problem of computing tasks failing to execute due to insufficient resources, achieves efficient and reliable execution of computing power scheduling, and avoids computing task failures caused by insufficient resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119127473B_ABST
    Figure CN119127473B_ABST
Patent Text Reader

Abstract

A computing power scheduling method and system. The method comprises: under a computing power scheduling task, determining computing power service demand based on configuration information selected by a user, task resources, and workflow definition; determining a target scheduling path of the computing power service demand in different preset dimensions based on preset scheduling mechanism and network-wide computing power resource information, and completing the computing power scheduling task based on the target scheduling path, wherein the preset dimensions include one of a cost dimension and a distance dimension. The scheme can realize scheduling of computing power resources.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of computing power scheduling, and particularly relates to a computing power scheduling method and system. BACKGROUND

[0002] An artificial intelligence platform needs to process a large amount of computing tasks, including model training, reasoning, etc. However, different computing tasks have different demands for computing power. If computing tasks are allocated to inappropriate computing resources, the computing tasks may not be executed due to insufficient resources.

[0003] The above content is only used to assist in understanding the technical solutions of the present application and does not represent the acknowledgement of the above content as prior art. SUMMARY

[0004] In order to solve the problems in the prior art, the application provides a computing power scheduling method and system to solve the technical problem that inappropriate computing resources are allocated to computing tasks, resulting in the computing tasks being unable to be executed due to insufficient resources.

[0005] To solve the above technical problems, the application adopts the following technical solutions.

[0006] The application first discloses a computing power scheduling method, which comprises the following steps:

[0007] Under the computing power scheduling task, the computing power business demand is determined according to the configuration information selected by the user, the task resource and the workflow definition.

[0008] Based on the full-network computing power resource information, the target scheduling path of the computing power business demand under different preset dimensions is determined according to a preset scheduling mechanism, and the computing power scheduling task is completed based on the target scheduling path, wherein the preset dimensions include one of a cost dimension and a distance dimension.

[0009] The application further comprises the following preferred schemes:

[0010] Before the step of determining the target scheduling path of the computing power business demand under different preset dimensions based on the full-network computing power resource information and according to a preset scheduling mechanism, the application further comprises the following steps:

[0011] Receiving computing power characteristic information, computing power service information and network index information fed back by a computing power node;

[0012] Determining the time delay information and the path information between different computing power nodes;

[0013] Inputting the computing power characteristic information, the computing power service information, the network index information, the time delay information and the path information into a target prediction model to obtain resource information of the computing power node;

[0014] Based on the resource information of the computing power nodes, the computing power resource information of the entire network is determined.

[0015] The step of determining the network-wide computing power resource information based on the resource information of the computing power nodes further includes:

[0016] The CFN protocol, based on a distributed architecture, receives resource information of computing nodes sent by the CFN router to determine the computing resource information of the entire network.

[0017] The distributed architecture is an intelligent distributed deployment composite component, which integrates PXE service, DHCP service, TFTP service, and MySQL service at the central deployment node; wherein...

[0018] The PXE service is responsible for network bootstrapping during the initial stages of node startup.

[0019] The DHCP service matches MAC addresses for deployed nodes and assigns IP addresses;

[0020] The TFTP service is responsible for transferring boot files, which include kernel files and an initialization root file system.

[0021] The MySQL service is responsible for recording node deployment information to facilitate system management and maintenance.

[0022] The distributed architecture is used to automatically generate network and storage configurations for nodes based on the cluster's hardware information, and generate Kolla configuration for OpenStack deployment parameters based on the optimized placement of the registry. After obtaining components from the official OpenStack GitHub, the installation packages are deployed. Finally, the nodes are prepared and configured, and Kolla is used to deploy an OpenStack cluster with one click.

[0023] The CFN protocol based on a distributed architecture receives resource information of computing nodes sent by the CFN router to determine the computing resource information of the entire network, and further includes:

[0024] Based on a preset collection method, the resource information of the computing power node is collected through the CFN router, and then the resource information of the computing power node sent by the CFN router is received. The preset collection method includes a first collection method and a second collection method. The first collection method is that the computing power node registers its resource information with the CFN router, and the second collection method is that the CFN router periodically collects information.

[0025] The preset scheduling mechanism includes a first scheduling mechanism and a second scheduling mechanism, the first scheduling mechanism is a virtual machine-based scheduling mechanism, and the second scheduling mechanism is a container-based scheduling mechanism; wherein

[0026] The full-network computing power resource information is used to determine a target scheduling path of the computing power service demand in different preset dimensions according to the preset scheduling mechanism, and the computing power scheduling task is completed based on the target scheduling path, and the method further includes the following steps:

[0027] The full-network computing power resource information is used to determine a target scheduling path of the computing power service demand in different preset dimensions according to the first scheduling mechanism, and the computing power scheduling task is completed based on the target scheduling path.

[0028] The full-network computing power resource information is used to determine a target scheduling path of the computing power service demand in different preset dimensions according to the second scheduling mechanism, and the computing power scheduling task is completed based on the target scheduling path.

[0029] The application also discloses a computing power scheduling system using the computing power scheduling method.

[0030] A determining module is configured to determine a computing power service demand under a computing power scheduling task according to configuration information, task resources and a work flow definition selected by a user.

[0031] A completing module is configured to determine a target scheduling path of the computing power service demand in different preset dimensions according to a preset scheduling mechanism based on full-network computing power resource information, and complete the computing power scheduling task based on the target scheduling path, wherein the preset dimensions include one of a cost dimension and a distance dimension.

[0032] Correspondingly, the application also discloses a device including a processor and a storage medium.

[0033] The storage medium is used to store instructions.

[0034] The processor is used to operate according to the instructions to perform the steps of the computing power scheduling method.

[0035] Correspondingly, the application also discloses a computer readable storage medium having a computer program stored thereon, and the program is executed by a processor to implement the steps of the computing power scheduling method.

[0036] The beneficial effect of the present application is that, compared with the prior art, the present application provides a computing power scheduling method and system, which determines computing power service demand according to user-selected configuration information, task resources and workflow definition; determines the target scheduling path of the computing power service demand in different preset dimensions based on the preset scheduling mechanism based on the whole network computing power resource information, and completes the computing power scheduling task based on the target scheduling path, solves the technical problem that the computing task cannot be executed due to insufficient resources caused by the allocation of inappropriate computing resources to the computing task, makes the execution of computing power scheduling more efficient and reliable, and further avoids the execution of some computing tasks due to insufficient resources. BRIEF DESCRIPTION OF DRAWINGS

[0037] Figure 1 A flowchart is provided for the first embodiment of the computing power scheduling method of the present application.

[0038] Figure 2 A flowchart is provided for the second embodiment of the computing power scheduling method of the present application.

[0039] Figure 3 A computing environment migration schematic diagram of virtual machines and containers is provided for the second embodiment of the present application.

[0040] Figure 4 A module structure schematic diagram of the computing power scheduling system of the present application is provided.

[0041] Figure 5 A device structure schematic diagram of the hardware running environment involved in the computing power scheduling method in the embodiment of the present application is provided. DETAILED DESCRIPTION

[0042] In order to make the purpose, technical scheme and advantages of the present application clearer, the technical scheme of the present application will be described clearly and completely below in combination with the drawings in the embodiments of the present application.

[0043] The embodiments described in the present application are only embodiments of part of the present application, not all embodiments. Based on the spirit of the present application, other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present application.

[0044] In view of the deficiencies of the prior art, the present application provides a computing power scheduling method and system, which determines computing power service demand according to user-selected configuration information, task resources and workflow definition under the computing power scheduling task; determines the target scheduling path of the computing power service demand in different preset dimensions based on the preset scheduling mechanism based on the whole network computing power resource information, and completes the computing power scheduling task based on the target scheduling path. The preset dimensions include one of the cost dimension and the distance dimension.

[0045] Reference Figure 1As shown, the computing power scheduling method disclosed by the application comprises the following steps:

[0046] Step S10, under the computing power scheduling task, the computing power service demand is determined according to the configuration information, task resources and workflow definition selected by the user.

[0047] Specifically, the scheduling center classifies and displays the platform-wide computing power, counts the number of data centers actually connected, the number of successfully run scheduling tasks, the number of scheduling requests, the number of resources, the number of workflow models, etc., and provides a one-glance resource overview for the user, facilitating the user to judge and select the required computing power according to the actual demand.

[0048] Step S20, based on the network-wide computing power resource information, the target scheduling path of the computing power service demand in different preset dimensions is determined according to the preset scheduling mechanism, and the computing power scheduling task is completed based on the target scheduling path, wherein the preset dimensions include one of cost dimension and distance dimension.

[0049] Among them, the network-wide computing power resource information refers to the resource information of all computing power nodes in the network, and the preset scheduling mechanism includes a virtual machine-based scheduling mechanism and a container-based scheduling mechanism.

[0050] In specific implementation, the determination of the target scheduling path is as follows:

[0051] Step S21, by selecting the corresponding data center, the scheduling task creation process is entered, the scheduling task is created, and the connection between the data centers is displayed in the form of visual dynamic effect.

[0052] Step S22, after entering the task process, the task scene and computing power configuration are selected according to the actual demand of the user, and the project selection or project creation is performed, the scheduling task will run in the project created or selected by the user, if no resource configuration is specified, the platform will flexibly schedule according to the idle condition of all resources.

[0053] Step S23, according to the prompt, upload the task resources required to run, the platform provides two resource modes of file resources and data sources, after successful uploading, the resources will be managed uniformly in the resource center.

[0054] Step S24, the platform supports flexible creation mode with drag and drop, selects the workflow model, and provides model creation for various big data workflows such as package SHELL, SQL and FLINK, and application scripts such as python and java.

[0055] Step S25, according to the configuration information, resource size and workflow definition selected by the user, the target scheduling path is recommended through cost, distance, comprehensive conditions and other dimensions.

[0056] In a specific implementation, the unified interface definition and implementation of data scheduling and task scheduling can be completed based on a general AI training / inference (Tensorflow / Pytorch), big data analysis (Spark / Flink), scientific calculation (Vasp), and the like. Specifically, the general AI training (Tensorflow / Pytorch) has two implementation forms in the computing power operation platform: encapsulating parameters and data required by a model and transferring them to the model, selecting a container image and a region where computing power resources are located, copying a model file to the container, and then performing execution, obtaining an execution result, and monitoring; or accepting a container image to run the model training, provided that the user has encapsulated the model into the container image in a predetermined form. The general AI inference (Tensorflow / Pytorch) deploys a model trained and parameterized to the computing power operation platform in the form of a k8s service, and exposes a corresponding port to accept external service requests. The big data analysis (Spark / Flink) and the scientific calculation (Vasp) are similar to the general AI training. The above are all completed by writing programs in the java computer language to define and implement the unified interface of data scheduling and task scheduling.

[0057] In a specific implementation, the Paas layer cross-domain resource management (monitoring and management of storage resources) realizes unified monitoring and unified management of functions such as alarm, permission, report, and linkage control by deploying a distributed collection unit and a monitoring system. Specifically, the monitoring data collection uses the open-source solution categraf, and the monitoring system uses the cloud monitoring module integrated in the computing power operation platform, which can realize functions such as chart form viewing of monitoring data, alarm and notification, report, and simple automatic fault processing.

[0058] The computing power scheduling method provided in the application solves the technical problem that a computing task cannot be executed due to insufficient resources caused by the allocation of the computing task to inappropriate computing resources, determines computing power business requirements based on configuration information, task resources, and workflow definitions selected by a user, and then designs a more appropriate target scheduling path for the user according to the computing power business requirements, so that the execution of the computing power scheduling is more efficient and reliable, thereby avoiding that some computing tasks cannot be executed due to insufficient resources.

[0059] In the second embodiment of the application, with reference to Figure 2 , the step S20 is further provided with steps S201-S204 before the step S20:

[0060] In step S201, the computing power feature information, the computing power service information, and the network index information fed back by the computing power node are received.

[0061] The computing power feature information can include processing capability of a central processing unit (CPU) / graphics processing unit (GPU), load information, deployment position, and the like; the computing power service information can include service type and service duration, and the like; and the network index information can include time delay, traffic, packet loss, and the like.

[0062] In step S202, time delay information and path information between different computing power nodes are determined.

[0063] Specifically, since the computing power resource is dynamically changed, the computing power awareness center also needs to measure the time delay information and path information between different computing power nodes.

[0064] In step S203, the computing power feature information, the computing power service information, the network index information, the time delay information, and the path information are input into a target prediction model to obtain resource information of the computing power node.

[0065] Preferably, the target prediction model is a prediction model learned by combining an AI traffic prediction model with an AI deep neural algorithm, and the target prediction model can timely predict the resource information of the computing power node, thereby improving the resource configuration speed and utilization efficiency of the whole network.

[0066] In step S204, the resource information of the computing power node is used to determine whole-network computing power resource information.

[0067] In the preferred embodiment, the determination of the whole-network computing power resource information based on the resource information of the computing power node further includes receiving, based on a CFN protocol of a distributed architecture, resource information of the computing power node sent by a CFN router to determine the whole-network computing power resource information.

[0068] Specifically, the ultimate purpose of the computing power network is to provide computing resources in the form of application services for users, and the computing resources are located in the infrastructure layer, and the physical location is generally different from that of the user, which requires the help of network functions to schedule tasks needed to be processed by the user to the computing resources. In the initial network, the location of the computing resource and the amount of resource it has are unknown to the user and the whole computing power network, and a communication message needs to be used as a carrier to interact in the network according to a specific protocol to complete the sharing of the computing resource information.

[0069] Specifically, in the TCP / IP architecture, as long as the IP of the authorized computing resource is reachable, the computing resource can be considered as available, and the communication protocol carrying the computing resource information can be located at any layer above the network layer (including the network layer), and the computing resource information is forwarded based on the IP message based on the network layer protocol.

[0070] Specifically, the computing power network controls and distributes computing resource information based on a CFN (Computing First Network) protocol, can solve the problems of complex MEC deployment, low efficiency, and low resource reuse rate, and enables the network to have the capability of dynamic routing of built-in computing services.

[0071] The CFN protocol publishes the computing resource status and network status as routing information to the network, and routes computing task messages to the most suitable computing node based on a virtual service ID, so as to achieve the purposes of optimal user experience, optimal computing resource utilization, and optimal network efficiency. The CFN protocol inherits the design idea of the traditional label forwarding protocol, is carried on the IP network, establishes a session between adjacent CFN protocol-enabled routers, publishes the obtained computing resource information to adjacent CFN routers by means of a routing protocol, and realizes the global diffusion of computing resource information. Meanwhile, the CFN router constructs a service routing information table according to different services, guides the forwarding of service messages with the service ID as the destination address, and thus realizes the utilization of distributed computing resources in a service manner.

[0072] Specifically, in view of the difficulty of fast replication and deployment of IAAS virtualization resource pool, PAAS cloud service platform, and upper-layer application in multiple data centers, a large-scale OpenStack deployment architecture with automatic placement of distributed repositories is adopted, and containerized deployment of multiple repositories is used, so that the deployment is more lightweight and easy to operate. In a large-scale heterogeneous environment of multiple data centers, an OpenStack deployment architecture with intelligent distributed registries placement (i.e., a distributed architecture) is adopted.

[0073] In a preferred embodiment, the distributed architecture is an intelligent distributed deployment synthetic component, such as an intelligent distributed registry deployment (IDRD) synthetic component. The intelligent distributed deployment synthetic component integrates a Pre-boot Execution Environment (PXE) service, a Dynamic Host Configuration Protocol (DHCP) service, a Trivial File Transfer Protocol (TFTP) service, and a MySql service in a central deployment node; wherein the PXE service is responsible for network booting in the initial stage of node startup; the DHCP service matches the MAC address of the deployed node and issues an IP address; the TFTP service is responsible for transferring a boot file, wherein the boot file includes a kernel file (vmlinuz) and an initialization root file system (initrd); and the MySql service is responsible for recording information such as node deployment status, so as to facilitate system management and maintenance.

[0074] Further, a plurality of registries are arranged in the distributed architecture, respectively placed on different nodes in the multi-data center cluster, wherein the registries are copies of the registry in the central server, and the registries schedule the requests from the cluster nodes to pull the images. The requests are distributed to different registries to achieve load balancing to alleviate the pressure on the central server.

[0075] All nodes in the multi-data center cluster are linked together through a high-speed deployment network and a Baseboard Management Controller (BMC) control network. An IDRD agent component is placed on each node to intelligently generate a node configuration file and a deployment file.

[0076] In the preferred embodiment, the distributed architecture is used to automatically generate network configuration and storage configuration of the nodes according to hardware information of the cluster, generate configuration of deployment parameter Kolla of OpenStack according to optimized placement of the registry, obtain components from the official GitHub of OpenStack to package a one-click deployment installation package, and finally complete preparation and configuration of the nodes, and deploy the nodes into an OpenStack cluster using Kolla.

[0077] Specifically, the distributed architecture includes three basic modules: a node preparation module, an automatically generated configuration file module, and a cluster deployment module. Network configuration and storage configuration of the nodes are automatically generated according to hardware information of the cluster, and configuration of deployment parameter Kolla of OpenStack is generated according to optimized placement of the registry. Bootstrap is a live system that does not affect data of a server when starting. Therefore, in the node preparation module, components are obtained from the official GitHub of OpenStack and packaged into a one-click deployment installation package to simplify operations in the later deployment process and install an OS for a bare-metal server. Finally, the nodes are prepared and configured, and the nodes are deployed into an OpenStack cluster using Kolla.

[0078] In the preferred embodiment, the CFN protocol based on the distributed architecture receives resource information of the computing power nodes sent by a CFN router, determines resource information of the entire network, and further includes:

[0079] The resource information of the computing power nodes is received after the resource information of the computing power nodes is collected by the CFN router based on a preset collection method, wherein the preset collection method includes a first collection method and a second collection method. The first collection method is a method in which the computing power nodes register the resource information to the CFN router, and the second collection method is a method in which the CFN router periodically collects information.

[0080] In the computing power network, to realize the integration of computing resource information and the use of computing resource information anytime and anywhere, the network-wide synchronization of information must be completed. The CFN router is responsible for collecting local computing resource information, and diffuses the information through IP packets or routing protocol packets. All CFN routers generate service routing information tables locally according to the complete computing resource information obtained and in combination with the network topology information, which are used to guide the forwarding of service packets. The specific implementation process is as follows:

[0081] ①CFN routers A and D complete the collection of local computing resource information. The collection process can adopt the way that the local computing resource nodes register the computing resource information to the CFN routers, or the way that the CFN routers periodically collect information;

[0082] ②CFN routers A and D carry the computing resource information in IP protocol or routing protocol, and publish it to other CFN routers in the network to realize the network-wide sharing of information;

[0083] ③CFN routers generate service routing information tables locally according to the network-wide information obtained and in combination with the network topology learned through the routing protocol, to guide the forwarding of service packets.

[0084] Routers B and C, as transit routers, can not necessarily support the CFN protocol, because the computing resource information is carried in IP protocol or routing protocol. B and C only need to forward the IP packets or routing protocol packets carrying the computing resource information, and do not analyze the CFN-related information in the packets.

[0085] Different applications in the computing power network have different focuses on the demand for computing resources. For example, the processing of two-dimensional pictures requires higher CPU, the processing of videos and AI requires higher GPU, and the processing of network packets requires higher NPU. According to different application services and different required computing resources, different service routing information entries will be generated on the computing power network router. Each service routing information entry on each computing power network router will guide the forwarding according to the different computing resource requirements.

[0086] When the number of application services is huge and the network scale is large, each router needs to obtain the information of the whole network and then independently calculate the path for each application service, the whole network maintenance workload is very huge, and the details of the interaction between the CFN protocol, the convergence, the IGP and the BGP and the interaction between the AS have not been maturely studied, so in order to the feasibility of the computing power network operation, the computing power network needs to be uniformly managed, the synchronization of information and the calculation of path are centralized, the service routing information table item is calculated and then distributed to the router, and the router is only responsible for the data layer service message forwarding, which is consistent with the idea of SDN.

[0087] In a specific implementation, the CFN protocol can also be based on a centralized architecture, which is different from the distributed architecture in that the routers do not need to directly communicate with each other and do not need to generate service routing information tables through local calculation, but only need to generate table items locally according to the distribution table items of the computing power network controller to guide forwarding.

[0088] In the design of the centralized architecture, in order to determine whether the computing resource information is directly sent to the computing power network controller for unified calculation, or the idea in the distributed architecture is followed to send the computing resource information to the router and then send it to the computing power network controller. Compared with the router, considering the large number of computing resource nodes, if each computing resource node needs to communicate with the computing power network controller, the pressure on the computing power network controller is too large, so the router can continue to bear the responsibility of collecting computing resource information, and the specific implementation process of the centralized control architecture is as follows:

[0089] ①Routers A and D complete the collection of local computing resource information, and the collection process can adopt the way that the local computing resource nodes register the computing resource information to the router, or the way that the router periodically collects information;

[0090] ②Routers A and D carry the computing resource information in the IP protocol or the routing protocol and publish it to the computing power network controller;

[0091] ③The computing power network controller generates a service information flow table according to the complete computing resource information and the completed network topology calculation;

[0092] ④The computing power network controller distributes the service information flow table to routers A and D;

[0093] ⑤Routers A and D generate a service information flow table locally according to the received computing power network controller information to guide the forwarding of service messages.

[0094] In a specific implementation, a data skew theory and a quantitative evaluation model are proposed for the difficulty in describing internal data skew, and it is found that the non-uniformity of data internal key values and the characteristics of long tail effect are the main reasons for low efficiency of data-intensive computing of supercomputing systems; and based on Pareto / Zipf power law distribution, a quantitative evaluation model, a benchmark test set and a test method are proposed, so that it is possible to accurately calculate the impact of data skew on node load and improve scheduling efficiency.

[0095] In a preferred embodiment, the preset scheduling mechanism includes a first scheduling mechanism and a second scheduling mechanism, the first scheduling mechanism is a virtual machine-based scheduling mechanism, and the second scheduling mechanism is a container-based scheduling mechanism; wherein, based on the full-network computing resource information, the target scheduling path of the computing service demand in different preset dimensions is determined according to the preset scheduling mechanism, and the computing scheduling task is completed based on the target scheduling path, further comprising:

[0096] Based on the full-network computing resource information, the target scheduling path of the computing service demand in different preset dimensions is determined according to the first scheduling mechanism, and the computing scheduling task is completed based on the target scheduling path; or, based on the full-network computing resource information, the target scheduling path of the computing service demand in different preset dimensions is determined according to the second scheduling mechanism, and the computing scheduling task is completed based on the target scheduling path.

[0097] Specifically, as shown in Figure 3 , in view of the explosive increase of user access concurrency and the explosive increase of data computing caused by emergency response, the standardized computing environment is quickly replicated and deployed through virtual machine and container standardized environment migration technology and distributed architecture, and unified elastic computing, storage and network resource services are provided to the outside.

[0098] Specifically, the virtual machine-based scheduling is realized through OpenStack-Train and cloud management, and the container-based scheduling is realized through the computing power operation platform based on k8s.

[0099] The application also provides a computing power scheduling system, please refer to Figure 4 , the computing power scheduling system comprises:

[0100] The determining module 10 is configured to determine the computing service demand under the computing power scheduling task according to the configuration information, task resources and workflow definition selected by the user.

[0101] The completing module 20 is configured to determine the target scheduling path of the computing service demand in different preset dimensions according to the preset scheduling mechanism based on the full-network computing resource information, and complete the computing scheduling task based on the target scheduling path, wherein the preset dimensions include one of cost dimension and distance dimension.

[0102] The application has the beneficial effect that, compared with the prior art, the application provides a computing power scheduling method and system, determines computing power service demand for user-selected configuration information, task resources and workflow definition, determines target scheduling paths of the computing power service demand in different preset dimensions based on preset scheduling mechanisms based on network-wide computing power resource information, and completes the computing power scheduling task based on the target scheduling paths, solves the technical problem that computing tasks cannot be executed due to insufficient resources caused by the allocation of computing tasks to inappropriate computing resources, makes the execution of computing power scheduling more efficient and reliable, and thus avoids the execution of some computing tasks due to insufficient resources.

[0103] Reference Figure 5 , which shows a structural schematic diagram of a computing power scheduling device suitable for being used to implement the embodiments of the present application. The computing power scheduling device in the embodiments of the present application can include but is not limited to mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (Personal Digital Assistant), PADs (Portable Application Description), PMPs (Portable Media Player), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), and the like, and fixed terminals such as digital TVs, desktop computers, and the like. Figure 5 The illustrated computing power scheduling device is only an example and should not bring any limitation to the functions and use range of the embodiments of the present application.

[0104] As Figure 5As shown, the computing power scheduling device can include a processing apparatus 1001 (e.g., a central processor, a graphics processor, etc.) that can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 1002 or loaded from a storage apparatus 1003 into a random access memory (RAM) 1004. In the RAM 1004, various programs and data required for operation of the xxx device are also stored. The processing apparatus 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 1006: input apparatuses 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; output apparatuses 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; the storage apparatus 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication apparatus 1009. The communication apparatus 1009 can allow the computing power scheduling device to communicate wirelessly or by wire with other devices to exchange data.

[0105] Based on the spirit of the present application, those skilled in the art can easily think of a computer program product based on the foregoing computing power scheduling method. The computer program product can include a computer readable storage medium on which computer readable program instructions for causing a processor to implement various aspects of the present disclosure are loaded. That is, the present application also includes a device including a processor and a storage medium; the storage medium is used to store instructions; the processor is used to operate according to the instructions to perform the steps according to the foregoing computing power scheduling method.

[0106] Computer readable storage media can be tangible storage media which can retain and store instructions for use by an instruction execution device. Computer readable storage media can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer readable storage media include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.

[0107] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.

[0108] Computer readable program instructions for carrying out operations of the present disclosure can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state setting data, or any combination of source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.

[0109] Finally, it should be noted that the above-mentioned embodiments are merely used to illustrate the technical solutions of the present application, rather than limiting the present application. Although the present application has been described in detail with reference to the above-mentioned embodiments, those skilled in the art should understand that the specific embodiments of the present application can be modified or replaced equivalently without departing from the spirit and scope of the present application, and any modification or equivalent replacement without departing from the spirit and scope of the present application should be covered in the protection scope of the claims of the present application.

Claims

1. A computing power scheduling method, characterized in that, Includes the following steps: Under the computing power scheduling task, the computing power service requirements are determined based on the configuration information, task resources and workflow definition selected by the user; Based on the network's computing power resource information, the target scheduling path for the computing power service demand under different preset dimensions is determined according to the preset scheduling mechanism, and the computing power scheduling task is completed based on the target scheduling path. The preset dimensions include either cost dimension or distance dimension. Before the step of determining the target scheduling path of the computing power service demand under different preset dimensions based on the computing power resource information of the entire network and according to the preset scheduling mechanism, the method further includes: Receive computing power characteristic information, computing power service information, and network indicator information fed back by computing power nodes; Determine the latency and path information between different computing power nodes; The computing power characteristic information, computing power service information, network indicator information, latency information, and path information are input into a prediction model that combines AI traffic prediction model with AI deep neural algorithm learning to obtain the resource information of computing power nodes. Based on the resource information of the computing power nodes, determine the computing power resource information of the entire network; The step of determining the network-wide computing power resource information based on the resource information of the computing power nodes further includes: The CFN protocol, based on a distributed architecture, receives resource information of computing nodes sent by the CFN router to determine the computing resource information of the entire network. The CFN protocol publishes the computing resource status and network status as routing information to the network, and routes computing task messages to the most suitable computing node based on the virtual service ID.

2. The method according to claim 1, characterized in that, The distributed architecture is an intelligent distributed deployment composite component, which integrates PXE service, DHCP service, TFTP service, and MySQL service at the central deployment node; wherein... The PXE service is responsible for network bootstrapping during the initial stages of node startup. The DHCP service matches MAC addresses for deployed nodes and assigns IP addresses; The TFTP service is responsible for transferring boot files, which include kernel files and an initialization root file system. The MySQL service is responsible for recording node deployment information to facilitate system management and maintenance.

3. The method according to claim 1, characterized in that, The distributed architecture is used to automatically generate network and storage configurations for nodes based on the cluster's hardware information, and generate Kolla configuration for OpenStack deployment parameters based on the optimized placement of the registry. After obtaining components from the official OpenStack GitHub, the installation packages are deployed. Finally, the nodes are prepared and configured, and Kolla is used to deploy an OpenStack cluster with one click.

4. The method according to claim 1, characterized in that, The CFN protocol based on a distributed architecture receives resource information of computing nodes sent by the CFN router to determine the computing resource information of the entire network, and further includes: Based on a preset collection method, the resource information of the computing power node is collected through the CFN router, and then the resource information of the computing power node sent by the CFN router is received. The preset collection method includes a first collection method and a second collection method. The first collection method is that the computing power node registers its resource information with the CFN router, and the second collection method is that the CFN router periodically collects information.

5. The method according to claim 1, characterized in that, The preset scheduling mechanism includes a first scheduling mechanism and a second scheduling mechanism. The first scheduling mechanism is a virtual machine-based scheduling mechanism, and the second scheduling mechanism is a container-based scheduling mechanism. The step of determining the target scheduling path for the computing power service demand under different preset dimensions based on the network-wide computing power resource information and according to a preset scheduling mechanism, and completing the computing power scheduling task based on the target scheduling path, further includes: Based on the network-wide computing power resource information, the target scheduling path for the computing power service demand under different preset dimensions is determined according to the first scheduling mechanism, and the computing power scheduling task is completed based on the target scheduling path; or... Based on the network's computing power resource information, the target scheduling path for the computing power service demand under different preset dimensions is determined according to the second scheduling mechanism, and the computing power scheduling task is completed based on the target scheduling path.

6. A computing power scheduling system, characterized in that, The system includes: The determination module is used to determine computing power service requirements based on the configuration information, task resources, and workflow definition selected by the user under the computing power scheduling task. The completion module is used to determine the target scheduling path of the computing power service demand under different preset dimensions according to the full network computing power resource information and a preset scheduling mechanism, and to complete the computing power scheduling task based on the target scheduling path. The preset dimensions include one of the cost dimension and the distance dimension. Before the step of determining the target scheduling path of the computing power service demand under different preset dimensions based on the computing power resource information of the entire network and according to the preset scheduling mechanism, the method further includes: Receive computing power characteristic information, computing power service information, and network indicator information fed back by computing power nodes; Determine the latency and path information between different computing power nodes; The computing power characteristic information, computing power service information, network indicator information, latency information, and path information are input into a prediction model that combines AI traffic prediction model with AI deep neural algorithm learning to obtain the resource information of computing power nodes. Based on the resource information of the computing power nodes, determine the computing power resource information of the entire network; The step of determining the network-wide computing power resource information based on the resource information of the computing power nodes further includes: The CFN protocol, based on a distributed architecture, receives resource information of computing nodes sent by the CFN router to determine the computing resource information of the entire network. The CFN protocol publishes the computing resource status and network status as routing information to the network, and routes computing task messages to the most suitable computing node based on the virtual service ID.

7. A device comprising a processor and a storage medium; characterized in that: The storage medium is used to store instructions; The processor is configured to operate according to the instructions to execute the steps of the computing power scheduling method according to any one of claims 1-5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the computing power scheduling method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Computing power resource allocation method and device, storage medium and electronic equipment

    CN116828538A