A computing power network resource scheduling method, device, equipment and medium

By iteratively optimizing the neural network model and comprehensively considering multiple key factors of the computing power network, the problem of resource scheduling separation in the computing power network is solved, and efficient integrated scheduling and intelligent management of resources are realized.

CN118827504BActive Publication Date: 2026-01-20CHINA MOBILE GROUP DESIGN INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311016573.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-11
Publication Date
2026-01-20
Estimated Expiration
2043-08-11

AI Technical Summary

Technical Problem

In a computing power network, the scheduling of the network and computing power are separated, which makes it impossible to form a complete resource scheduling system, affecting the overall operating efficiency and making it impossible to achieve the global optimal solution for computing network resources.

Method used

By adopting a self-learning approach and iteratively optimizing through a neural network model, the optimal computing network path is determined by comprehensively considering multiple key factors in the computing network, thereby achieving integrated resource scheduling.

Benefits of technology

It improves the utilization efficiency and intelligence of computing network resources, avoids the separation of computing power and network, and realizes integrated orchestration and fusion scheduling of resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118827504B_ABST
    Figure CN118827504B_ABST
Patent Text Reader

Abstract

A computing power network resource scheduling method and device, equipment and medium are disclosed. According to a preset service requirement, a plurality of computing network key factors in a computing power network are obtained, a preset neural network model is used to calculate a computing network path and a total reward value of the computing network path. The computing network path is formed by all computing network nodes. The weight parameters of the neural network model are iteratively optimized until the objective function of the neural network model meets the requirements, and a trained neural network model is obtained. The objective function is defined according to the total reward value of the computing network path. The optimal computing network path is obtained according to the trained neural network model, and the computing power network resource scheduling is performed according to the optimal computing network path. The present application can consider multiple computing network key factors in the computing power network, iteratively optimize the best computing network path through autonomous learning, and realize the fusion scheduling of computing power network resources.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of communication technology, and in particular to a computing power network resource scheduling method and device, equipment and medium. BACKGROUND

[0002] The computing power network is a new type of information infrastructure oriented to computing network integration, which is based on computing and network and realizes the on-demand allocation and flexible scheduling of computing, storage and network resources among cloud, edge and end. The computing power network is a brand-new network system, and the computing network brain belongs to the orchestration management layer of the computing power network architecture, which is the core link for unified management, orchestration, scheduling and operation and maintenance of computing network resources and capability elements, and is the management center of the entire computing power network.

[0003] The computing power network contains a large number of resource categories and resource quantities, and the computing network business is rich in variety, so the computing power network presents high complexity. As the central decision system of the computing power network, the decision-making ability of the computing network brain will directly affect the operation efficiency of the entire computing power network. At present, the scheduling of the network and the scheduling of the computing power in the computing power network are separated, and each manages its own affairs, which cannot form a complete computing network resource scheduling system, so it cannot achieve the global optimal solution of the computing network resources, which will affect the overall operation efficiency of the computing power network in the future. How to realize the scheduling of the multi-element fusion of the computing power network is a major technical problem faced by the computing power network. SUMMARY

[0004] The purpose of the embodiments of the present application is to provide a computing power network resource scheduling method, device, equipment and medium, which can consider multiple computing network key factors in the computing power network, and iteratively optimize the best computing network path through autonomous learning, so as to realize the fusion scheduling of the computing power network resources and improve the intelligentization of the computing power network resource scheduling.

[0005] To achieve the above purpose, the embodiments of the present application provide a computing power network resource scheduling method, comprising:

[0006] According to the preset business demand, a plurality of computing network key factors in the computing power network are obtained;

[0007] According to the plurality of computing network key factors, a preset neural network model is used for calculation to determine a computing network path and a total reward value of the computing network path; wherein the computing network path is a path formed by all computing network nodes;

[0008] The weight parameters of the neural network model are iteratively optimized until the objective function of the neural network model meets the requirements, and a trained neural network model is obtained; wherein the objective function is defined according to the total reward value of the computing network path;

[0009] An optimal computing network path is obtained according to the trained neural network model, and computing network resource scheduling is performed according to the optimal computing network path.

[0010] As an improvement of the above scheme, the weight parameters of the neural network model are iteratively optimized until the objective function of the neural network model meets the requirements, and a trained neural network model is obtained, comprising:

[0011] The gradient ascent method is used to iteratively optimize the weight parameters of the neural network model until the objective function of the neural network model reaches a maximum value, and a trained neural network model is obtained; wherein the objective function is to maximize the mathematical expectation of the total reward value.

[0012] As an improvement of the above scheme, the preset neural network model is used to calculate according to the plurality of computing network key factors to determine the computing network path, comprising:

[0013] Any one computing network node is selected as an access node;

[0014] Each computing network node that has passed at the current time is taken as a current state, a multi-factor matrix is constructed according to the value of each computing network key factor under each computing network node of the current state, and input to the preset neural network model for calculation to determine the next passed computing network node to enter the next state;

[0015] When all the computing network nodes are passed, the computing network path is obtained.

[0016] As an improvement of the above scheme, the total reward value of the computing network path is determined, comprising:

[0017] According to the value of each computing network key factor under each computing network node of each state, the total key factor value of each computing network key factor under each state is calculated;

[0018] According to the total key factor value of each computing network key factor under each state, the reward value of each state is calculated;

[0019] According to the reward value of each state, the total reward value of the computing network path is calculated.

[0020] As an improvement of the above scheme, according to the value of each computing network key factor under each computing network node of each state, the total key factor value of each computing network key factor under each state is calculated, comprising:

[0021] Each metric factor to which each computing network key factor belongs is determined;

[0022] According to the value of each said algorithm network key factor under each said state of each said algorithm network node, a corresponding preset total key factor value calculation formula of the metric factor is used to calculate the total key factor value of each said algorithm network key factor under each said state.

[0023] As an improvement of the above scheme, the metric factor includes an additive metric factor, a multiplicative metric factor and a concave metric factor.

[0024] The algorithm network key factor belonging to the additive metric factor includes a calculation factor, a storage factor, a time delay factor, a distance factor and an energy consumption factor, the algorithm network key factor belonging to the multiplicative metric factor includes a reliability factor, and the algorithm network key factor belonging to the concave metric factor includes a bandwidth factor.

[0025] As an improvement of the above scheme, the calculation of the reward value of each said state according to the total key factor value of each said algorithm network key factor under each said state includes:

[0026] According to the business requirement, a weight coefficient of each said algorithm network key factor is determined;

[0027] According to the weight coefficient of each said algorithm network key factor and the total key factor value of each said algorithm network key factor under each said state, a unified metric value of each said state is calculated;

[0028] According to the unified metric value of each said state, a reward value of each said state is calculated.

[0029] As an improvement of the above scheme, the calculation of the unified metric value of each said state according to the weight coefficient of each said algorithm network key factor and the total key factor value of each said algorithm network key factor under each said state includes:

[0030] According to the weight coefficient of each said algorithm network key factor and the total key factor value of each said algorithm network key factor under each said state, a unified metric value of each said state is calculated;

[0031]

[0032] The calculation of the reward value of each said state according to the unified metric value of each said state includes:

[0033] According to the unified metric value of each said state, a reward value of each said state is calculated;

[0034] r k =1 / M k ;

[0035] wherein, M kr is a uniform degree value of the kth state k δ is a reward value of the kth state, the kth state including k passed algorithm network nodes i wi is a weight coefficient of the ith algorithm network key factor ri is a total key factor value of the ith algorithm network key factor in the kth state, n is the number of the algorithm network key factors, and

[0036] As an improvement of the above scheme, the total reward value of the algorithm network path is calculated according to the reward value in each state, and the total reward value of the algorithm network path is calculated according to the reward value in each state.

[0037] The reward value of each state is de-meaned to obtain a processed reward value.

[0038] According to the processed reward value in each state, the total reward value of the algorithm network path is obtained by summation.

[0039]

[0040] Wherein, R is the total reward value, r k δ is a reward value of the kth state, K is the number of algorithm network nodes of the algorithm network path.

[0041] As an improvement of the above scheme, after the plurality of algorithm network key factors in the algorithm network are obtained, the method further comprises:

[0042] The values of the plurality of algorithm network key factors are normalized.

[0043] The embodiment of the application also provides a scheduling device for algorithm network resources, comprising:

[0044] An algorithm network key factor acquisition module is configured to acquire a plurality of algorithm network key factors in an algorithm network according to a preset service requirement;

[0045] An algorithm network path determination module is configured to calculate an algorithm network path and a total reward value of the algorithm network path by using a preset neural network model according to the plurality of algorithm network key factors; wherein the algorithm network path is a path formed by all algorithm network nodes.

[0046] A neural network model learning module is configured to iteratively optimize weight parameters of the neural network model until a target function of the neural network model meets a requirement, so as to obtain a trained neural network model; wherein the target function is defined according to the total reward value of the algorithm network path.

[0047] The computing power network resource scheduling module is configured to obtain an optimal computing power network path according to the trained neural network model, and perform computing power network resource scheduling based on the optimal computing power network path.

[0048] The present application also provides a computing power network resource scheduling device, which comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, and the processor implements the computing power network resource scheduling method according to any one of the above embodiments when executing the computer program.

[0049] The present application also provides a computer readable storage medium comprising a stored computer program, wherein the computer readable storage medium controls the device where the computer readable storage medium is located to execute the computing power network resource scheduling method according to any one of the above embodiments when the computer program is running.

[0050] Compared with the prior art, the computing power network resource scheduling method, device, equipment and medium disclosed by the present application obtain a plurality of algorithm network key factors in the computing power network according to a preset business requirement, calculate the algorithm network path and the total reward value of the algorithm network path by using a preset neural network model according to the plurality of algorithm network key factors, perform iterative optimization on the weight parameters of the neural network model until the objective function of the neural network model meets the requirements to obtain a trained neural network model, wherein the objective function is defined according to the total reward value of the algorithm network path, obtain an optimal algorithm network path according to the trained neural network model, and perform computing power network resource scheduling based on the optimal algorithm network path. The present application comprehensively considers various indexes such as computing power, network and benefit factors in the computing power network, uses the self-learning ability of an intelligent agent, adopts a machine learning method, iteratively optimizes the best algorithm network path, realizes integrated arrangement and fusion scheduling of algorithm network resources, can effectively avoid the separation of computing power and network, improves the utilization efficiency of algorithm network resources, and improves the intelligence of computing power network resource scheduling. BRIEF DESCRIPTION OF DRAWINGS

[0051] Figure 1 is a flowchart of a computing power network resource scheduling method provided by the present application embodiment;

[0052] Figure 2 is a principle diagram of determining an algorithm network path in the present application embodiment;

[0053] Figure 3 is a principle diagram of a neural network model in the present application embodiment;

[0054] Figure 4is a structural schematic diagram of a computing power network resource scheduling device provided by an embodiment of the present application.

[0055] Figure 5 is a structural schematic diagram of a computing power network resource scheduling device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0056] The technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative labor fall within the protection scope of the present application.

[0057] Referring to Figure 1 is a flow schematic diagram of a computing power network resource scheduling method provided by an embodiment of the present application. The present application provides a computing power network resource scheduling method, which includes the following steps S11 to S14.

[0058] S11, according to a preset service requirement, obtaining a plurality of algorithm network key factors in a computing power network;

[0059] S12, according to the plurality of algorithm network key factors, using a preset neural network model to perform calculation, determining an algorithm network path and a total reward value of the algorithm network path; wherein the algorithm network path is a path formed by all algorithm network nodes;

[0060] S13, performing iterative optimization on weight parameters of the neural network model until a target function of the neural network model meets a requirement, obtaining a trained neural network model; wherein the target function is defined according to the total reward value of the algorithm network path;

[0061] S14, obtaining an optimal algorithm network path according to the trained neural network model, and performing computing power network resource scheduling according to the optimal algorithm network path.

[0062] In the embodiments of the present application, according to different services and SLA (Service Level Agreement, service level agreement) requirements, considering various factors in the computing power network, including but not limited to computing power factors, network factors and benefit factors, a plurality of algorithm network key factors are identified to form an algorithm network key factor set.

[0063] As an example, the algorithm network key factors relied on in the computing power network decision are shown in Table 1 as follows:

[0064] Table 1

[0065]

[0066] It can be understood that the types of the above-mentioned calculation network key factors are only optional embodiments, and in actual application, corresponding calculation network key factors can be obtained according to different business requirements, and do not constitute a limitation on the present application.

[0067] Preferably, after obtaining the plurality of calculation network key factors in the step S11, the method further comprises:

[0068] S11', normalizing the values of the plurality of calculation network key factors.

[0069] In the embodiment of the present application, in order to facilitate optimization training, each calculation network key factor is normalized. Optionally, the Z-score method is used for normalization, and the normalization formula is as follows:

[0070]

[0071] Wherein, x i is the value of the calculation network key factor, X i is the normalized value of the calculation network key factor, μ is the average value of all calculation network key factors x i , and σ is the standard deviation of all calculation network key factors x i .

[0072] It should be noted that since the calculation network path passes through a plurality of calculation network nodes, the normalized calculation network key factor value under each calculation network node is as shown in Table 2:

[0073] Table 2

[0074]

[0075] Further, referring to Figure 2 , the principle diagram for determining the calculation network path in the embodiment of the present application, the neural network model is pre-set in the embodiment of the present application, which is used to generate the calculation network path from the access node to the target node under the condition of comprehensively considering the plurality of calculation network key factors.

[0076] Specifically, a neural network model is initialized, the several algorithm network key factors are calculated by using the neural network model to obtain an algorithm network path, and a total reward value of the algorithm network path is calculated. A target function of the neural network model is calculated according to the total reward value, if the target function does not meet the requirement, the weight parameters of the neural network model are updated to obtain an updated neural network model, the several algorithm network key factors are calculated by using the updated neural network model to obtain a new algorithm network path, and a total reward value of the algorithm network path is calculated, the target function of the neural network model is calculated according to the total reward value, and whether the target function meets the requirement is judged, so as to realize iterative optimization of the weight parameters of the neural network model until the target function of the neural network model meets the requirement, and a trained neural network model is obtained, that is, an optimal algorithm network path is determined. The optimal algorithm network path is the best route from an access node to a target node in the case of considering the several algorithm network key factors corresponding to the current business demand. Further, the optimal algorithm network path is used for algorithm network resource scheduling.

[0077] By using the technical means of the embodiment of the present application, various indexes such as computing power, network and benefit factors in the computing power network are comprehensively considered, the self-learning ability of an agent is used, a machine learning method is used, the best algorithm network path is iteratively optimized, the integrated arrangement and fusion scheduling of algorithm network resources are realized, the separation of computing power and network can be effectively avoided, the utilization efficiency of algorithm network resources is improved, and the intelligentization of computing power network resource scheduling is improved.

[0078] As a preferred embodiment, the embodiment of the present application is further implemented on the basis of the above-mentioned embodiment, in step S12, the several algorithm network key factors are calculated by using a preset neural network model to determine an algorithm network path, comprising:

[0079] S121, an algorithm network node is randomly selected as an access node;

[0080] S122, each algorithm network node that has passed at the current time is taken as a current state, a multi-factor matrix is constructed according to the value of each algorithm network key factor under each algorithm network node of the current state, and is input to the preset neural network model for calculation to determine a next passed algorithm network node to enter a next state;

[0081] S123, when all the algorithm network nodes are passed, the algorithm network path is obtained.

[0082] Preferably, in step S12, the total reward value of the algorithm network path is determined, comprising:

[0083] S124, calculating a total key factor value of each of the algorithm network key factors in each of the states according to the values of each of the algorithm network key factors in each of the algorithm network nodes in each of the states;

[0084] S125, calculating a reward value of each of the states according to the total key factor values of each of the algorithm network key factors in each of the states;

[0085] S126, calculating a total reward value of the algorithm network path according to the reward values of each of the states.

[0086] Referring to Figure 3 , which is a principle schematic diagram of a neural network model in an embodiment of the present application, the embodiment of the present application takes n algorithm network key factors and K algorithm network nodes as an example, in which an algorithm network node is arbitrarily selected as an access node, the node corresponds to an initial state S1, and a multi-factor matrix in the state S1 is constructed as follows:

[0087]

[0088] An initial neural network model π θ is initialized, where θ is a weight parameter of the neural network model, the multi-factor matrix corresponding to the initial state S1 is input into the neural network model, an action a1(θ) is obtained, and a reward value r1(θ) is obtained at the same time, the reward value r1(θ) is calculated according to the values of the n algorithm network key factors in an algorithm network node corresponding to the state S1 (that is, the access node). Then, path selection is performed according to the action a1(θ), the next algorithm network node is determined, the state S2 is entered, and a multi-factor matrix in the state S2 is constructed as follows:

[0089]

[0090] The multi-factor matrix corresponding to the initial state S2 is input into the neural network model, an action a2(θ) and a reward value r2(θ) are obtained, the reward value r2(θ) is calculated by calculating n total key factor values of the n algorithm network key factors in two algorithm network nodes corresponding to the state S2, and the n total key factor values are calculated according to the n algorithm network key factors. Then, path selection is performed according to the action a2(θ), the next algorithm network node is determined, and the state S3 is entered. By analogy, a multi-factor matrix in a final state S K is obtained as follows:

[0091]

[0092] The multi-factor matrix corresponding to the final state S K is continuously input into the neural network model π θ , a reward r K(θ), while the action a K (θ) stop working, and the process ends, so as to obtain the algorithm network path, and according to the reward values in the above K states, the total reward value of the algorithm network path is calculated.

[0093] Preferably, the reward value r k of each state and the calculation process of the total reward value R of the algorithm network path are explained.

[0094] Step S124, that is, calculating the total key factor value of each algorithm network key factor in each state according to the value of each algorithm network key factor at each algorithm network node in each state, includes:

[0095] S1241, determining the measurement factor to which each algorithm network key factor belongs;

[0096] S1242, calculating the total key factor value of each algorithm network key factor in each state according to the value of each algorithm network key factor at each algorithm network node in each state, using the calculation formula of the corresponding preset total key factor value of the measurement factor.

[0097] Specifically, the several algorithm network key factors are attribute-labeled for different measurement factors in the algorithm network. The measurement factors are divided into the following three types of different attributes, including additive measurement factors, multiplicative measurement factors and concave measurement factors.

[0098] Among them, the algorithm network key factors belonging to the additive measurement factors include the calculation factor, the storage factor, the time delay factor, the distance factor and the energy consumption factor, the algorithm network key factors belonging to the multiplicative measurement factors include the reliability factor, and the algorithm network key factors belonging to the concave measurement factors include the bandwidth factor.

[0099] For the additive measurement factor, the calculation formula of the corresponding key factor value is:

[0100] For the multiplicative measurement factor, the calculation formula of the corresponding key factor value is:

[0101] For the concave measurement factor, the calculation formula of the corresponding key factor value is:

[0102] As an example, taking the seven algorithm network key factors mentioned in Table 1 above as an example, for the algorithm network key factors corresponding to the additive measurement factor:

[0103] The total calculation factor value of the calculation factor x1 in the kth state is:

[0104] The total storage factor value of the storage factor x2 in the kth state is:

[0105] The total latency factor value of the latency factor x4 in the kth state is:

[0106] The total distance factor value of the distance factor x5 in the kth state is:

[0107] The total energy consumption factor value of the energy consumption factor x7 in the kth state is:

[0108] For the multiplicative metric factor, the corresponding algorithm network key factor is:

[0109] The total reliability factor value of the reliability factor x6 in the kth state is:

[0110] For the concave metric factor, the corresponding algorithm network key factor is:

[0111] The total wideband factor value of the wideband factor x3 in the kth state is:

[0112] Further, the step S125, that is, calculating the reward value of each state according to the total key factor value of each algorithm network key factor in each state, comprises:

[0113] S1251, determining the weight coefficient of each algorithm network key factor according to the service requirement;

[0114] S1252, calculating the unified metric value of each state according to the weight coefficient of each algorithm network key factor and the total key factor value of each algorithm network key factor in each state;

[0115] S1253, calculating the reward value of each state according to the unified metric value of each state.

[0116] Specifically, in the embodiment of the present application, the unified metric value is designed to be used for representing the performance of the route. According to the current service requirement, the importance of each algorithm network key factor is determined, so as to set the weight coefficient of each algorithm network key factor, respectively δ1, δ2, …, δn, and the formula of the unified metric value is as follows:

[0117]

[0118] According to the unified metric value, the reward value of the state is calculated, and the formula is as follows:

[0119] rk = 1 / M k ;

[0120] wherein, M k is the unified dimension value of the kth state, r k is the reward value of the kth state, the kth state includes k passed algorithm network nodes, δ i is the weight coefficient of the ith algorithm network key factor, is the total key factor value of the ith algorithm network key factor in the kth state, n is the number of the algorithm network key factors, and

[0121] Further, step S126, that is, calculating the total reward value of the algorithm network path according to the reward value in each state, comprises:

[0122] S126, performing mean removal processing on the reward value of each state to obtain a processed reward value;

[0123] S127, obtaining the total reward value of the algorithm network path in a summation manner according to the processed reward value in each state.

[0124] In the embodiment of the present application, in order to avoid the reward value in each state being possibly always positive, and ensure that the punishment mechanism is effectively implemented, mean removal operation is performed on the reward value of each state. Further, the total reward value of the algorithm network path is:

[0125]

[0126] wherein, R is the total reward value, r k is the reward value of the kth state, and K is the number of algorithm network nodes of the algorithm network path.

[0127] By using the technical means of the embodiment of the present application, the computing power, network and benefit factors and other various indexes in the computing power network are comprehensively considered, the unified dimension value parameter is designed, the neural network model of state, action and reward value is constructed through the mapping relationship between the unified dimension value and the reward value by using the method of machine learning, the self-learning ability of the agent is used for iterative optimization, and the optimal algorithm network path is obtained. The present application uses the mapping relationship between the unified dimension value and the reward value, applies the unified dimension value to the whole optimization strategy, so as to ensure that the target is finally achieved, and the mean removal operation is proposed for the reward value, so as to ensure that the reward and punishment jointly act, and the efficiency of data training is improved. The present application can effectively realize the integrated arrangement and fusion scheduling of the algorithm network resources, avoids the fragmentation of computing power and network, comprehensively considers the influence of multiple factors, reduces the system complexity, improves the utilization efficiency of the algorithm network resources, improves the intelligentization of the computing power network resource scheduling, and has strong implementability.

[0128] As a preferred embodiment, the embodiment of the present application is further implemented on the basis of any of the above embodiments, and step S13, that is, the weight parameters of the neural network model are iteratively optimized until the objective function of the neural network model meets the requirement, to obtain a trained neural network model, comprising:

[0129] The weight parameters of the neural network model are iteratively optimized by using the gradient ascent method until the objective function of the neural network model reaches the maximum value, to obtain a trained neural network model; wherein the objective function is to maximize the mathematical expectation of the total reward value.

[0130] In the embodiment of the present application, for a neural network model π θ The reward value in each state of the algorithm network path is represented as r k (θ), and the total reward value of the algorithm network path is represented as R(θ):

[0131]

[0132] The objective function of the neural network model is determined according to the total reward value, and the weight parameters of the neural network model are iteratively optimized by using the gradient ascent method to obtain the optimal algorithm network path.

[0133] In an optional embodiment, the objective function J(θ) is defined The weight parameters θ are obtained to maximize the mathematical expectation of the total reward value.

[0134] According to the law of large numbers, it is obtained that: Wherein, N represents the number of random trials.

[0135] Based on the gradient ascent strategy, the weight parameter optimization is performed according to the gradient formula to obtain the best path. The gradient formula based on the law of large numbers in the embodiment of the present application is:

[0136]

[0137] The weight parameters θ are updated The iterative optimization is performed until J(θ) is maximized, the weight parameters θ at this time are determined, a trained neural network model is obtained, and thus the optimal algorithm network path is determined, and the optimization operation on the algorithm network path is completed.

[0138] By using the technical means of the embodiment of the present application, the computing power, network and benefit factors and various indexes in the computing power network are comprehensively considered, and a unified dimension value parameter is designed. Through the mapping relationship between the unified dimension value and the reward value, the neural network model of the state, action and reward value is constructed, and the objective function J(θ) is introduced, and the gradient The calculation method utilizes a gradient ascent algorithm to realize iterative optimization of the link weight parameters of the algorithm network, obtain a global solution of the algorithm network routing, and obtain an optimal algorithm network path of the algorithm network.

[0139] It can be understood that the specific calculation process involved in the iteration of the weight parameters of the neural network model to obtain the optimal solution is only a preferred embodiment, and in actual application, other gradient ascent methods with similar functions can be used to realize the iteration of the weight parameters to obtain the optimal solution, and this does not constitute a limitation on the present application.

[0140] Referring to Figure 4 The embodiment of the present application provides a structure schematic diagram of a scheduling device for algorithm network resources.

[0141] The algorithm network key factor acquisition module 21 is configured to acquire a plurality of algorithm network key factors in the algorithm network according to a preset service demand.

[0142] The algorithm network path determination module 22 is configured to calculate an algorithm network path and a total reward value of the algorithm network path by using a preset neural network model according to the plurality of algorithm network key factors.

[0143] The neural network model learning module 23 is configured to iteratively optimize the weight parameters of the neural network model until a target function of the neural network model meets a requirement, so as to obtain a trained neural network model.

[0144] The algorithm network resource scheduling module 24 is configured to obtain an optimal algorithm network path according to the trained neural network model, and schedule algorithm network resources by using the optimal algorithm network path.

[0145] The technical means of the embodiment of the present application comprehensively considers various indexes such as the computing power, the network and the benefit factors in the algorithm network, iteratively optimizes the best algorithm network path by using the self-learning ability of the agent and the method of machine learning, realizes the integrated arrangement and the integrated scheduling of the algorithm network resources, can effectively avoid the separation of the computing power and the network, improves the utilization efficiency of the algorithm network resources, and improves the intelligentization of the algorithm network resource scheduling.

[0146] Preferably, the device further comprises:

[0147] The normalization processing module is configured to normalize the values of the plurality of algorithm network key factors.

[0148] As a preferred implementation, the algorithm network path determination module 22 is specifically configured to:

[0149] arbitrarily select one algorithm network node as an access node; take each algorithm network node that has elapsed at the current time as a current state, construct a multi-factor matrix according to the value of each algorithm network key factor under each algorithm network node of the current state, and input the multi-factor matrix to the preset neural network model for calculation to determine a next elapsed algorithm network node to enter a next state; and when all the algorithm network nodes are elapsed, obtain the algorithm network path.

[0150] According to the value of each algorithm network key factor under each algorithm network node of each state, calculate a total key factor value of each algorithm network key factor under each state; according to the total key factor value of each algorithm network key factor under each state, calculate a reward value of each state; and according to the reward value of each state, calculate a total reward value of the algorithm network path.

[0151] Preferably, the calculation of the total key factor value of each algorithm network key factor under each state according to the value of each algorithm network key factor under each algorithm network node of each state comprises:

[0152] determining a metric factor to which each algorithm network key factor belongs;

[0153] According to the value of each algorithm network key factor under each algorithm network node of each state, using the calculation formula of the corresponding preset total key factor value of the metric factor to which the algorithm network key factor belongs, calculate the total key factor value of each algorithm network key factor under each state.

[0154] Preferably, the metric factor comprises an additive metric factor, a multiplicative metric factor and a concave metric factor; the algorithm network key factor belonging to the additive metric factor comprises a calculation factor, a storage factor, a time delay factor, a distance factor and an energy consumption factor, the algorithm network key factor belonging to the multiplicative metric factor comprises a reliability factor, and the algorithm network key factor belonging to the concave metric factor comprises a bandwidth factor.

[0155] Preferably, the calculation of the reward value of each state according to the total key factor value of each algorithm network key factor under each state comprises:

[0156] According to the business requirement, determine a weight coefficient of each algorithm network key factor;

[0157] According to the weight coefficient of each of the algorithm network key factors and the total key factor value of each of the algorithm network key factors in each of the states, a unified measure value of each of the states is calculated;

[0158] According to the unified measure value of each of the states, a reward value of each of the states is calculated.

[0159] Preferably, the calculation of the unified measure value of each of the states according to the weight coefficient of each of the algorithm network key factors and the total key factor value of each of the algorithm network key factors in the state comprises:

[0160] According to the weight coefficient of each of the algorithm network key factors and the total key factor value of each of the algorithm network key factors in each of the states, the unified measure value of each of the states is calculated by using the following calculation formula:

[0161]

[0162] The calculation of the reward value of each of the states according to the unified measure value of each of the states comprises:

[0163] According to the unified measure value of each of the states, the reward value of each of the states is calculated by using the following calculation formula:

[0164] r k =1 / M k ;

[0165] Wherein, M k is the unified measure value of the kth state, r k is the reward value of the kth state, the kth state includes k algorithm network nodes that have been passed, δ i is the weight coefficient of the ith algorithm network key factor, and is the total key factor value of the ith algorithm network key factor in the kth state, n is the number of the algorithm network key factors, and

[0166] Preferably, the calculation of the total reward value of the algorithm network path according to the reward value in each of the states comprises:

[0167] The reward value of each of the states is de-meaned to obtain a processed reward value;

[0168] According to the processed reward value in each of the states, the total reward value of the algorithm network path is obtained by using summation:

[0169]

[0170] Wherein, R is the total reward value, r kThe reward value of the kth state, K is the number of the computing network nodes of the computing network path.

[0171] As a preferred implementation, the neural network model learning module 23 is specifically configured to:

[0172] The weight parameters of the neural network model are iteratively optimized by using the gradient ascent method until the objective function of the neural network model reaches a maximum value, thereby obtaining a trained neural network model; wherein the objective function is to maximize the mathematical expectation of the total reward value.

[0173] It should be noted that the computing network resource scheduling device provided by the embodiments of the present application is used to execute all process steps of the computing network resource scheduling method of the above-mentioned embodiments, and the working principles and beneficial effects of the two are one-to-one correspondence, so they will not be repeated.

[0174] Referring to Figure 5 is a structural schematic diagram of a computing network resource scheduling device provided by the embodiments of the present application, the embodiments of the present application also provide a computing network resource scheduling device 30, which comprises a processor 31, a memory 32, and a computer program stored in the memory and configured to be executed by the processor, and the processor implements the computing network resource scheduling method of any one of the above-mentioned embodiments when executing the computer program.

[0175] The embodiments of the present application also provide a computer readable storage medium, which comprises a stored computer program, wherein the computer readable storage medium controls the device where the computer readable storage medium is located to execute the computing network resource scheduling method of any one of the above-mentioned embodiments when the computer program is running.

[0176] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by a computer program instructing related hardware, and the program can be stored in a computer readable storage medium, and the program can include the processes of the above-mentioned embodiments when executed. The storage medium can be a magnetic disc, an optical disc, a read-only memory (ROM) or a random access memory (RAM) and the like.

[0177] The above-mentioned is the preferred implementation of the present application, and it should be noted that for those of ordinary skill in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, and these improvements and refinements are also considered to be within the protection scope of the present application.

Claims

1. A method for scheduling computing power network resources, characterized in that, The application relates to a method for determining an optimal computing network path based on a preset service demand. According to the preset service demand, a plurality of computing network key factors in a computing network are acquired; According to the plurality of computing network key factors, a preset neural network model is used for calculation to determine a computing network path and a total reward value of the computing network path; wherein the computing network path is a path formed by all computing network nodes; The weight parameters of the neural network model are iteratively optimized until the target function of the neural network model meets the requirement, and a trained neural network model is obtained; wherein the target function is defined according to the total reward value of the computing network path; According to the trained neural network model, an optimal computing network path is obtained, and computing network resource scheduling is performed based on the optimal computing network path. The computing network key factors include a calculation factor, a storage factor, a time delay factor, a distance factor, an energy consumption factor, a reliability factor and a bandwidth factor; wherein the calculation factor, the storage factor, the time delay factor, the distance factor and the energy consumption factor belong to the computing network key factors of additive measurement factors, the reliability factor belongs to the computing network key factor of multiplicative measurement factors, and the bandwidth factor belongs to the computing network key factor of concave measurement factors. 2.The method of claim 1, wherein, The weight parameters of the neural network model are iteratively optimized until the target function of the neural network model meets the requirement, and a trained neural network model is obtained; wherein the target function is defined according to the total reward value of the computing network path; The weight parameters of the neural network model are iteratively optimized until the target function of the neural network model meets the requirement, and a trained neural network model is obtained; wherein the target function is defined according to the total reward value of the computing network path; 3.The method of claim 1, wherein, Any one computing network node is selected as an access node; Each computing network node that has passed at the current time is taken as a current state, a multi-factor matrix is constructed according to the value of each computing network key factor under each computing network node of the current state, and the multi-factor matrix is input into the preset neural network model for calculation to determine a next passed computing network node to enter a next state; When all the computing network nodes are passed, the computing network path is obtained. The total reward value of the computing network path is determined by the following steps: 4.The method of Claim 3, wherein, According to the value of each computing network key factor under each computing network node of each state, the total key factor value of each computing network key factor under each state is calculated; According to the total key factor value of each computing network key factor under each state, the reward value of each state is calculated; According to the reward value of each state, the total reward value of the computing network path is calculated. The total key factor value of each computing network key factor under each state is calculated according to the value of each computing network key factor under each computing network node of each state, and the total key factor value of each computing network key factor under each state is calculated according to the value of each computing network key factor under each computing network node of each state. 5.The method of claim 4, wherein, ​ ​ According to the value of each said algorithm network key factor under each said state of each said algorithm network node, a corresponding preset total key factor value calculation formula of the metric factor is used to calculate the total key factor value of each said algorithm network key factor under each said state. 6.The method of claim 4, wherein, The calculation of the reward value of each said state according to the total key factor value of each said algorithm network key factor under each said state comprises: According to the business requirement, the weight coefficient of each said algorithm network key factor is determined; According to the weight coefficient of each said algorithm network key factor and the total key factor value of each said algorithm network key factor under each said state, the unified metric value of each said state is calculated; According to the unified metric value of each said state, the reward value of each said state is calculated.

7. The method of Claim 6, wherein, The calculation of the unified metric value of each said state according to the weight coefficient of each said algorithm network key factor and the total key factor value of each said algorithm network key factor under each said state comprises: According to the weight coefficient of each said algorithm network key factor and the total key factor value of each said algorithm network key factor under each said state, the unified metric value of each said state is calculated using the following calculation formula: ; The calculation of the reward value of each said state according to the unified metric value of each said state comprises: According to the unified metric value of each said state, the reward value of each said state is calculated using the following calculation formula: ; wherein, is a unit measure of a state, k is a reward value of a state, k k k is a weight coefficient of a key factor of a network, i is a total key factor value of a key factor of a network in a state, i k n is a number of the key factors of the network, and .​​​​​​ 8.The method of claim 4, wherein, The calculation of the total reward value of the algorithm network path according to the reward value under each said state comprises: The reward value of each said state is de-meaned to obtain a processed reward value; According to the processed reward value under each said state, the total reward value of the algorithm network path is obtained by summation: ; wherein R is a total reward value, is the reward value for the k st state, and K is the number of MCTS nodes of the MCTS path. 9.The method of Claim 1, wherein, After the several algorithm network key factors in the computing power network are obtained, the method further comprises: The values of the several algorithm network key factors are normalized.

10. A computing power network resource scheduling apparatus characterized in that, Comprise: An algorithm network key factor acquisition module is configured to acquire several algorithm network key factors in a computing power network according to a preset business requirement; An algorithm network path determination module is configured to determine an algorithm network path and a total reward value of the algorithm network path by calculating using a preset neural network model according to the several algorithm network key factors; wherein the algorithm network path is a path formed by all algorithm network nodes; A neural network model learning module is configured to iteratively optimize the weight parameters of the neural network model until the objective function of the neural network model meets the requirements to obtain a trained neural network model; wherein the objective function is defined according to the total reward value of the algorithm network path; A computing power network resource scheduling module is configured to obtain an optimal algorithm network path according to the trained neural network model, and perform computing power network resource scheduling according to the optimal algorithm network path. The key factors of the computing network include computing factor, storage factor, latency factor, distance factor, energy consumption factor, reliability factor, and bandwidth factor; among them, computing factor, storage factor, latency factor, distance factor, and energy consumption factor belong to the additive measurement factors of computing network key factors, reliability factor belongs to the multiplicative measurement factors of computing network key factors, and bandwidth factor belongs to the concave measurement factors of computing network key factors.

11. A computing power network resource scheduling device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor, when executing the computer program, implements the method for scheduling computing network resources as described in any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to perform the scheduling method for computing network resources as described in any one of claims 1 to 9.