Resource scheduling method and device based on large electric power model and computer equipment
By constructing a hardware-model matching evaluation function and a pruning and compression strategy, the resource scheduling of the large-scale power model is optimized, which solves the problems of low computing resource utilization and high energy consumption of the large-scale power model on different hardware platforms, and realizes the collaborative adaptation between hardware platforms and model algorithms.
Patent Information
- Application Number
- CN202511559817.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-29
- Publication Date
- 2026-01-09
AI Technical Summary
In practical applications, large-scale power models have low computational resource utilization and high energy consumption. Software optimization methods and hardware optimization methods are not well compatible, and it is difficult to coordinate and adapt with hardware platforms of different architectures.
By constructing a hardware capability matrix and a model requirement matrix, a hardware-model matching evaluation function is generated, a pruning strategy is determined, and model compression is performed. Combined with computation graph decomposition and hardware instruction scheduling, resource scheduling is optimized to improve adaptability.
It improves the utilization rate of computing resources in the large-scale power model, reduces overall energy consumption, enables it to run efficiently on different hardware platforms, and solves the problem of insufficient adaptability.
Smart Images

Figure CN121301019A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of power control, and in particular to a resource scheduling method, apparatus, and computer equipment based on a large power model. Background Technology
[0002] With the development of artificial intelligence applications in the field of power control, large-scale power models have become a key technology for the intelligent upgrading of power systems. They can deeply mine multi-source heterogeneous data such as grid operation and equipment monitoring, and have core functions such as operation monitoring and fault early warning. However, large-scale power models are highly dependent on high-performance hardware (such as GPU clusters), resulting in problems such as high consumption of computing resources.
[0003] In related technologies, there is a lack of compatibility between software and hardware optimization methods for large-scale power modeling. At the software level, optimization techniques (such as knowledge distillation and model pruning) generally focus only on the number of parameters and computational complexity, neglecting hardware storage structures and underlying parameters. At the hardware level, different hardware platforms struggle to achieve seamless integration with software model algorithms, resulting in slow inference speeds and memory overflows. This situation leads to low computational resource utilization and high energy consumption in practical applications of large-scale power models, hindering their ability to fully leverage their performance advantages. Summary of the Invention
[0004] Based on this, embodiments of this application provide a resource scheduling method, apparatus, and computer equipment based on a large power model, which can improve the compatibility problem between software optimization methods and hardware optimization methods of the large power model, effectively reduce the overall energy consumption of the large power model in practical applications and improve the utilization rate of computing resources, enabling the large power model to achieve collaborative adaptation with hardware platforms of different architectures.
[0005] To achieve the above objectives, in a first aspect, some embodiments of this application provide a resource scheduling method based on a large-scale power model. This resource scheduling method includes the following steps.
[0006] Obtain the storage structure characteristics, power large model parameter quantity, and computation graph topology of the target hardware platform;
[0007] Based on the storage structure characteristics and the number of parameters of the power large model, a hardware capability matrix and a model requirement matrix are constructed respectively.
[0008] Based on the hardware capability matrix and the model requirement matrix, a hardware-model matching degree evaluation function is constructed;
[0009] The hardware-model matching evaluation result is determined based on the hardware-model matching evaluation function.
[0010] Based on the hardware-model matching evaluation results, a pruning strategy is determined;
[0011] The storage capacity of the model parameters is determined based on the pruning strategy and hardware storage granularity.
[0012] Based on the storage capacity of the model parameters, the large power model is pruned and compressed to obtain the target compressed model.
[0013] Based on the computation graph topology, the target compression model is decomposed into multiple subgraphs.
[0014] Based on the hardware capability matrix, a computing unit is allocated to each subgraph to generate a hardware instruction scheduling sequence;
[0015] Resource scheduling is performed based on the hardware instruction scheduling sequence.
[0016] In some embodiments, constructing a hardware-model matching evaluation function based on the hardware capability matrix and the model requirement matrix includes the following steps.
[0017] The hardware capability matrix and the model requirement matrix are respectively divided into hierarchical parts to obtain the hardware hierarchical structure and the model hierarchical structure.
[0018] Within the hardware hierarchy and the model hierarchy, the correlation between each hardware index and each model index is determined based on cosine similarity, resulting in a hardware correlation matrix and a model correlation matrix, respectively.
[0019] Based on the hardware correlation matrix and the model correlation matrix, and according to the hardware hierarchy and the model hierarchy, the correlation of the lower-level indicators in the hierarchy is passed to the upper level to obtain the hardware correlation target matrix and the model correlation target matrix respectively.
[0020] The hardware correlation target matrix and the model correlation target matrix are integrated to construct a bipartite network; the bipartite network uses the hardware indicators as hardware nodes and the model indicators as model nodes; the weight of the edge of the bipartite network is the correlation between the corresponding hardware indicator and the model indicator.
[0021] The binary network is analyzed to generate hardware centrality vectors and model centrality vectors;
[0022] Based on the hardware centrality vector and the model centrality vector, construct the hardware-model matching evaluation function.
[0023] In some embodiments, the step of passing the correlation of lower-level indicators in the hierarchy to the upper level according to the hardware hierarchy and the model hierarchy includes the following steps.
[0024] The matching strength between the hardware node and the model node is calculated based on the centrality of the hardware node, the centrality of the model node, and the correlation between the hardware node and the model node.
[0025] In some embodiments, constructing the hardware-model matching evaluation function based on the hardware centrality vector and the model centrality vector includes the following steps.
[0026] Based on the hardware centrality vector and the model centrality vector, the matching strength between the hardware node and the model node is determined, and a matching strength matrix is constructed.
[0027] The rows and columns of the matching strength matrix are summed to obtain the hardware-side matching strength vector sum and the model-side matching strength vector sum;
[0028] The hardware-side matching strength vector and the model-side matching strength vector are respectively mean-processed, and the mean-processed result is determined as the hardware-model matching degree evaluation function.
[0029] In some embodiments, determining the hardware-model matching evaluation result based on the hardware-model matching evaluation function includes the following steps.
[0030] The average hardware-side matching strength is determined based on the number of hardware nodes involved in the calculation and the sum of the hardware-side matching strength vectors.
[0031] The mean matching strength of the model side is determined based on the number of model nodes involved in the calculation and the sum of the matching strength vectors of the model side.
[0032] The hardware-model matching degree evaluation result is determined based on the average of the mean matching strength on the hardware side and the mean matching strength on the model side.
[0033] In some embodiments, determining the pruning strategy based on the hardware-model matching evaluation result includes the following steps.
[0034] If the matching degree value in the hardware-model matching degree evaluation result is greater than or equal to the first threshold, a structured pruning strategy is adopted.
[0035] If the matching degree value in the hardware-model matching degree evaluation result is less than the first threshold and greater than or equal to the second threshold, a strategy combining structured pruning and unstructured pruning is adopted; wherein, the first threshold is greater than the second threshold.
[0036] If the matching degree value in the hardware-model matching degree evaluation result is less than the second threshold, an unstructured pruning strategy is adopted.
[0037] Secondly, this application also provides a resource scheduling device based on a large power model according to some embodiments; the resource scheduling device includes: a first acquisition module, a first construction module, a second construction module, a first determination module, a second determination module, a third determination module, a model processing module, a model analysis module, a first generation module, and an optimization execution module. The first acquisition module is used to acquire the storage structure characteristics, power large model parameter quantity, and computation graph topology of the target hardware platform; the first construction module is used to construct a hardware capability matrix and a model requirement matrix based on the storage structure characteristics and the power large model parameter quantity, respectively; the second construction module is used to construct a hardware-model matching degree evaluation function based on the hardware capability matrix and the model requirement matrix; the first determination module is used to determine the hardware-model matching degree evaluation result based on the hardware-model matching degree evaluation function; the second determination module is used to determine a pruning strategy based on the hardware-model matching degree evaluation result; the third determination module is used to determine the model parameter storage capacity based on the pruning strategy and hardware storage granularity; the model processing module is used to prune and compress the power large model based on the model parameter storage capacity to obtain a target compressed model; the model analysis module is used to decompose the computation graph of the target compressed model based on the computation graph topology to obtain multiple subgraphs; the first generation module is used to allocate computing units to each subgraph according to the hardware capability matrix and generate a hardware instruction scheduling sequence; the optimization execution module is used to perform resource scheduling based on the hardware instruction scheduling sequence.
[0038] Thirdly, according to some embodiments, this application also provides a computer device including a memory and a processor, wherein the memory stores a computer program; and the processor executes the computer program to implement the steps of the method described in the first aspect of this application.
[0039] Fourthly, according to some embodiments, this application also provides a computer-readable storage medium having a computer program stored thereon; when executed by a processor, the computer program implements the steps of the method described in the first aspect of this application.
[0040] Fifthly, according to some embodiments, this application also provides a computer program product, including a computer program; when executed by a processor, the computer program implements the steps of the method described in the first aspect of this application.
[0041] The embodiments of this application may have, or at least have, the following advantages:
[0042] In this embodiment, a hardware capability matrix and a model requirement matrix are constructed based on storage structure characteristics and the number of parameters in a large-scale power model. A hardware-model matching evaluation function is then built based on these two matrices, accurately quantifying the compatibility between hardware and the model to determine the hardware-model matching evaluation result. This addresses the issues of software optimization methods focusing solely on model parameters and computational complexity, and hardware optimization methods being detached from model algorithms, thereby effectively improving the compatibility between software and hardware optimization methods for large-scale power models. By using a pruning strategy determined based on the hardware-model matching evaluation result, combined with hardware storage granularity and the pruning strategy, the large-scale power model is pruned and compressed to obtain a target compressed model. This removes redundant parameters and connections from the large-scale power model, effectively reducing its demand for computational and storage resources, improving computational resource utilization, and thus reducing overall energy consumption. This effectively solves the problems of low computational resource utilization and high energy consumption in practical applications of large-scale power models. By decomposing the target compressed model into multiple subgraphs based on the computational graph topology, and allocating computational units to each subgraph according to the hardware capability matrix, this approach enables collaborative work between the hardware platform and the model algorithm. This solves the problem of incompatibility between hardware platforms with different architectures and the model algorithm, allowing the large-scale power model to operate efficiently on various hardware platforms. Through the combined effect of these technical features, this application improves the compatibility issue between software and hardware optimization methods for large-scale power models, effectively reducing the overall energy consumption of large-scale power models in practical applications and improving the utilization rate of computational resources, thus enabling the large-scale power model to achieve collaborative adaptation with hardware platforms of different architectures.
[0043] Details of one or more embodiments of this application are set forth in the following drawings and description. Other features, objects, and advantages of this application will become apparent from the specification, drawings, and claims. Attached Figure Description
[0044] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0045] Figure 1 This is an application environment diagram of a resource scheduling method provided in some embodiments;
[0046] Figure 2 This is a flowchart illustrating a resource scheduling method provided in some embodiments;
[0047] Figure 3This is a flowchart illustrating step S300 provided in some embodiments;
[0048] Figure 4 This is a flowchart illustrating step S360 provided in some embodiments;
[0049] Figure 5 This is a flowchart illustrating step S400 provided in some embodiments;
[0050] Figure 6 This is a structural block diagram of a resource scheduling device provided in some embodiments;
[0051] Figure 7 This is a schematic diagram of the internal structure of a computer device provided in some embodiments.
[0052] Explanation of reference numerals in the attached figures:
[0053] 1-First Acquisition Module, 2-First Construction Module, 3-Second Construction Module, 4-First Determination Module, 5-Second Determination Module, 6-Third Determination Module, 7-Model Processing Module, 8-Model Analysis Module, 9-First Generation Module, 10-Optimization Execution Module. Detailed Implementation
[0054] To facilitate understanding of this application, a more complete description will be provided below with reference to the accompanying drawings, which illustrate preferred embodiments of the application. However, this application may be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided so that the disclosure of this application will be thorough and complete.
[0055] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the application.
[0056] This application provides a resource scheduling method, apparatus, and computer equipment based on a large power model, which can improve the compatibility problem between software and hardware optimization methods of the large power model, effectively reduce the overall energy consumption of the large power model in practical applications and improve the utilization rate of computing resources, enabling the large power model to achieve collaborative adaptation with hardware platforms of different architectures.
[0057] The method for processing family hardship levels provided in this application embodiment can be applied to, for example, Figure 1In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or placed on a cloud or other network server. Server 104 acquires the storage structure characteristics, power model parameter quantity, and computation graph topology of the target hardware platform; based on the storage structure characteristics and power model parameter quantity, it constructs a hardware capability matrix and a model requirement matrix; based on the hardware capability matrix and model requirement matrix, it constructs a hardware-model matching degree evaluation function; based on the hardware-model matching degree evaluation function, it determines the hardware-model matching degree evaluation result; based on the hardware-model matching degree evaluation result, it determines a pruning strategy; based on the pruning strategy and hardware storage granularity, it determines the model parameter storage capacity; based on the model parameter storage capacity, it prunes and compresses the power model to obtain a target compressed model; based on the computation graph topology, it decomposes the target compressed model into multiple subgraphs; based on the hardware capability matrix, it allocates computational units to each subgraph and generates a hardware instruction scheduling sequence; based on the hardware instruction scheduling sequence, it performs resource scheduling; and it sends the resource scheduling result to terminal 102. The terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, and smart in-vehicle systems. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted devices. The server 104 can be implemented using a standalone server or a server cluster consisting of multiple servers.
[0058] In some embodiments, please refer to Figure 2 The resource scheduling method based on the power big model includes the following steps S100 to S1000.
[0059] Step S100: Obtain the storage structure characteristics, power large model parameter quantity, and computation graph topology of the target hardware platform.
[0060] In some examples, hardware detection tools are used to directly detect and obtain storage structure characteristics of the target hardware platform.
[0061] For example, storage structure features include, but are not limited to, storage capacity, read / write speed, and storage hierarchy; wherein, the storage hierarchy includes the hierarchical relationship between cache, memory, and hard disk.
[0062] In some examples, the number of parameters of the power large model is obtained from the parameter file obtained after training the power large model.
[0063] For example, the parameters of a large power model include the number of connection weights for neurons in each layer.
[0064] In some examples, the computation graph topology is obtained through visualization tools or parsers of model building frameworks such as TensorFlow or PyTorch.
[0065] For example, computation graph topology is used to characterize the connections and data flow between computation nodes in a large power model.
[0066] Step S200: Based on the storage structure characteristics and the number of parameters in the power large model, construct the hardware capability matrix and the model requirement matrix respectively.
[0067] In some examples, during the process of constructing the hardware capability matrix in step S200, it is necessary to quantify hardware indicators such as storage capacity and read / write speed into matrix elements.
[0068] For example, consider a target hardware platform with a three-tier storage structure consisting of cache, memory, and hard disk; the hardware capability matrix is a three-dimensional matrix; and the matrix elements are H. ijk , represents the hardware performance metrics (e.g., data transfer rate) when the j-th level storage interacts with the k-th level storage under the i-th operation type (e.g., read or write).
[0069] In some examples, during the process of constructing the model demand matrix in step S200, it is necessary to quantify the storage capacity requirements of the power large-scale model based on the number of power large-scale model parameters and the characteristics of model operation.
[0070] For example, the matrix element is M. mn , representing the storage capacity requirement of the nth level for the mth operation in the large power model.
[0071] In this embodiment of the application, by constructing a hardware capability matrix and a model requirement matrix, the complex characteristics of hardware and models can be transformed into a computable and comparable mathematical form, which facilitates subsequent matching degree evaluation and decision optimization.
[0072] For example, the large-scale power model is a power fault diagnosis model; the number of parameters is 5 million; the target hardware platform has three levels of storage: 16MB cache, 8GB RAM, and 1TB hard drive. In step S100, the hardware detection tool acquires information such as storage capacity, read / write speed, and storage hierarchy structure, and constructs a hardware capability matrix H; in step S200, the read and write requirements of the large-scale power fault diagnosis model for data at different storage levels during the operation are analyzed, and a model requirement matrix M is constructed. During the training process of the power fault diagnosis model, some parameter update operations require frequent reading of data from the cache; based on this reading requirement, the required value is filled into the corresponding position in the model requirement matrix M.
[0073] Step S300: Construct a hardware-model matching degree evaluation function based on the hardware capability matrix and the model requirement matrix.
[0074] Step S400: Determine the hardware-model matching degree evaluation result based on the hardware-model matching degree evaluation function.
[0075] For example, the specific value of the hardware-model matching evaluation function obtained in step S300 is the hardware-model matching evaluation result.
[0076] Step S500: Determine the pruning strategy based on the hardware-model matching evaluation results.
[0077] Step S600: Determine the storage capacity of model parameters based on the pruning strategy and hardware storage granularity.
[0078] It should be noted that hardware storage granularity refers to the smallest data unit size that the target hardware platform can process at one time, such as bit, byte, or word.
[0079] For example, if the hardware storage granularity is 32 bits, then the target hardware platform can process 32 bits of data at a time.
[0080] Step S700: Based on the storage capacity of the model parameters, prune and compress the large power model to obtain the target compressed model.
[0081] It should be explained that in steps S600 to S700, the storage capacity of the model parameters is determined based on the hardware storage granularity. For parameter data smaller than the hardware storage granularity, merging or padding can be performed. Then, the power model is pruned and compressed according to the pruning strategy determined in step S500 to remove redundant parameters and connections, thereby obtaining the target compressed model.
[0082] Step S800: Based on the computation graph topology, the target compression model is decomposed into multiple subgraphs.
[0083] It should be noted that each subgraph is relatively independent and is used to represent a part of the computational task.
[0084] For example, the computation graph is G = (V, E); where V is the set of nodes (representing the computation task) and E is the set of edges (representing data dependencies); the computation graph G is decomposed into multiple subgraphs G1, G2, ..., G using graph theory algorithms (such as topological sorting). k .
[0085] Step S900: Allocate computing units to each subgraph according to the hardware capability matrix and generate a hardware instruction scheduling sequence.
[0086] For example, let the set of hardware computing units be C = {c1, c2, ..., c...} n}; Calculate the complexity function Complexity(G) based on the subgraph. i ) and the hardware computing unit processing capacity function Capacity(c j Assign computational units to each subgraph, that is, find the optimal mapping f: G i →c j This is used to generate a hardware instruction scheduling sequence.
[0087] For example, taking the computation graph of the target compression model as an example, which includes tasks such as matrix multiplication and convolution operations; the types of subgraphs obtained by decomposition in step S800 include, but are not limited to, matrix multiplication subgraphs and convolution subgraphs.
[0088] Based on the above example, taking a target hardware platform containing 4 CPU cores and 8 GPU stream processors as an example, in step S900: since matrix multiplication has a large computational load and is suitable for parallel computing, the matrix multiplication subgraph needs to be allocated to the GPU stream processor; since convolution operation is relatively complex and has strong data dependencies, the convolution subgraph needs to be allocated to the CPU core; and the hardware instruction scheduling sequence is generated according to the data dependencies and hardware execution efficiency.
[0089] In this embodiment of the application, in step S900, by analyzing the hardware capability matrix, information such as the computing power and data processing speed of each computing unit (such as CPU cores and GPU stream processors) of the target hardware platform can be obtained. By allocating the most suitable computing unit to each subgraph according to the computational complexity, data dependency relationship and hardware computing unit characteristics of the subgraph, and determining the execution order of each subgraph computing task, a hardware instruction scheduling sequence is generated, which can make full use of the parallel computing capability of the target hardware platform, thereby improving the utilization rate of computing resources.
[0090] Step S1000: Resource scheduling is performed based on the hardware instruction scheduling sequence.
[0091] It should be noted that by dynamically allocating and managing hardware resources (such as computing units, memory, and cache) based on hardware instruction scheduling sequences, the execution progress and resource usage of each computing task can be monitored during execution, and resource allocation can be adjusted in a timely manner. This ensures that computing tasks can be executed efficiently according to the scheduling sequence, avoiding resource conflicts and waste.
[0092] For example, in an example where the target hardware platform contains 4 CPU cores and 8 GPU stream processors, based on the hardware instruction scheduling sequence, the convolutional subgraph task on the CPU can be executed first during resource scheduling, and then the matrix multiplication subgraph task can be executed in parallel using the GPU.
[0093] Based on the above example, taking the real-time monitoring of the CPU core and GPU stream processor load during the execution of hardware instruction scheduling sequence as an example; when a GPU stream processor is idle while other tasks are waiting for computing resources, it is necessary to allocate appropriate subgraph tasks to the idle processor to ensure continuous and efficient utilization of resources.
[0094] In this embodiment, a hardware capability matrix and a model requirement matrix are constructed based on storage structure characteristics and the number of parameters in a large-scale power model. A hardware-model matching evaluation function is then built based on these two matrices, accurately quantifying the compatibility between hardware and the model to determine the hardware-model matching evaluation result. This addresses the issues of software optimization methods focusing solely on model parameters and computational complexity, and hardware optimization methods being detached from model algorithms, thereby effectively improving the compatibility between software and hardware optimization methods for large-scale power models. By using a pruning strategy determined based on the hardware-model matching evaluation result, combined with hardware storage granularity and the pruning strategy, the large-scale power model is pruned and compressed to obtain a target compressed model. This removes redundant parameters and connections from the large-scale power model, effectively reducing its demand for computational and storage resources, improving computational resource utilization, and thus reducing overall energy consumption. This effectively solves the problems of low computational resource utilization and high energy consumption in practical applications of large-scale power models. By decomposing the target compressed model into multiple subgraphs based on the computational graph topology, and allocating computational units to each subgraph according to the hardware capability matrix, this approach enables collaborative work between the hardware platform and the model algorithm. This solves the problem of incompatibility between hardware platforms with different architectures and the model algorithm, allowing the large-scale power model to operate efficiently on various hardware platforms. Through the combined effect of these technical features, this application improves the compatibility issue between software and hardware optimization methods for large-scale power models, effectively reducing the overall energy consumption of large-scale power models in practical applications and improving the utilization rate of computational resources, thus enabling the large-scale power model to achieve collaborative adaptation with hardware platforms of different architectures.
[0095] In some embodiments, please refer to Figure 3 Step S300 includes the following steps S310 to S360.
[0096] Step S310: Perform hierarchical division of the hardware capability matrix and model requirement matrix to obtain the hardware hierarchical structure and model hierarchical structure.
[0097] In some examples, step S310 includes: determining the hardware metrics of the hardware capability matrix and the model metrics of the model requirement matrix.
[0098] For example, hardware metrics include cache capacity, memory bandwidth, number of compute cores, and bus bandwidth.
[0099] For example, model metrics include the number of parameters, computational complexity, and memory access frequency.
[0100] In some examples, the hardware capability matrix is divided into layers according to "storage layer - computing layer - communication layer".
[0101] For example, the hardware capability matrix is H = [h1, h2, h3, h4], where h1 is the cache capacity, h2 is the memory bandwidth, h3 is the number of computing cores, and h4 is the bus bandwidth; then the hardware hierarchy H obtained after hierarchical partitioning is... layer It includes a storage layer [h1], a computing layer [h3], and a communication layer [h2, h4].
[0102] For example, the model requirement matrix is M = [m1, m2, m3, m4], where m1 is the parameter storage amount, m2 is the computation amount, m3 is the number of memory accesses, and m4 is the data transfer amount; then the model hierarchy structure M obtained after hierarchical partitioning is... layer It includes a storage requirement layer [m1], a computing requirement layer [m2], and an interaction requirement layer [m3, m4].
[0103] Step S320: Within the hardware hierarchy and model hierarchy, determine the correlation between each hardware index and each model index based on cosine similarity, and obtain the hardware correlation matrix and the model correlation matrix respectively.
[0104] It should be noted that in step S320, the correlation between hardware metrics and model metrics needs to be calculated in the corresponding hardware and model layers. For example, the correlation between hardware metrics in the hardware storage layer [h1] and model metrics in the model storage requirement layer [m1] is calculated.
[0105] For example, cosine similarity can be determined using the following formula:
[0106]
[0107] Where, r ij For cosine similarity, h i For the i-th metric of the hardware storage layer, m j Store the j-th metric of the requirement layer for the model.
[0108] For example, let's take the cache capacity H of the hardware storage layer as an example. s = [h1] and the parameter storage capacity M of the model storage requirement layer s Taking [|m1], cache capacity h1 = 1024, and parameter storage size m1 = 800 as an example:
[0109] According to the cosine similarity formula, the numerator in the cosine similarity formula is h1·m. j=1024 × 800 = 819200, denominator is Thus, the correlation coefficient r between cache capacity and parameter storage can be calculated. 11 =819200 / (1024×800)=1.
[0110] For example, using hardware communication layer H com = [h2, h4] and model interaction requirement layer M com Taking [m3, m4] as an example, with memory bandwidth h2 = 100GB / s, bus bandwidth h4 = 50GB / s, memory access count m3 = 500 times / second, and data transfer volume m4 = 2000GB; calculate the cosine similarity pairwise using the following formula:
[0111]
[0112] Where, r 23 To represent the correlation between memory bandwidth and the number of memory accesses, r 24 To represent the correlation between memory bandwidth and data transfer volume, r 43 To show the correlation between bus bandwidth and memory access frequency, r 44 This represents the correlation between bus bandwidth and data transfer volume.
[0113] For example, the simplified hardware correlation matrix R h It can be:
[0114] For example, the model correlation matrix R m Similarly, calculations can be performed based on indicators within each level.
[0115] Step S330: Based on the hardware correlation matrix and the model correlation matrix, and according to the hardware hierarchy and the model hierarchy, the correlation of the lower-level indicators in the hierarchy is passed to the upper level to obtain the hardware correlation target matrix and the model correlation target matrix respectively.
[0116] For example, the upper-layer hardware metric H k (For example, overall communication capability) is determined by the lower-level indicator h. k1 and h k2 Taking the components (including memory bandwidth h2 and bus bandwidth h4) as an example, let the weight of the lower-level indicator be w. k1 =0.6, w k2 If the correlation coefficient is 0.4, then the upper-level correlation degree is determined according to the following formula:
[0117] R Hk =W k1 ×r k1i +w k2 ×r k2i ;
[0118] Where, r k1i and r k2i This represents the correlation between lower-level indicators and model indicators.
[0119] For example, suppose the model interaction requirement layer has an indicator m3 (i.e., memory access count), and the correlation degree between the lower-level hardware indicator h2 and m3 is r. 23 =0.8, the correlation degree between h4 and m3 is r 43 =0.6; then the correlation between the overall upper-layer communication capability and m3 is calculated according to the following formula:
[0120] R Hk-m3 =0.6×0.8+0.4×0.6=0.48+0.24=0.72.
[0121] Referring to the examples above, by passing the correlation degree of each lower layer upwards, the final hardware correlation degree target matrix R can be obtained. H The model correlation target matrix R M Among them, the hardware correlation target matrix R H Used to integrate the correlation after passing through all levels.
[0122] Step S340: Integrate the hardware correlation target matrix and the model correlation target matrix to construct a bipartite network. The bipartite network uses hardware indicators as hardware nodes and model indicators as model nodes; the weight of each edge in the bipartite network represents the correlation between the corresponding hardware indicator and the model indicator.
[0123] For example, hardware node H nodes Determined according to the following formula:
[0124] H nodes ={h1,h3,H com}; where h1 is the cache capacity, h3 is the number of computing cores, and H com This is a communication indicator for the upper layer.
[0125] For example, model node M nodes Determined according to the following formula:
[0126] M nodes ={m1, m2, M com}; where m1 is the parameter storage amount, m2 is the computation amount, and M com This is an indicator for upper-level interaction.
[0127] Based on the example above, the weight of an edge is the correlation between hardware metrics and model metrics; for example, the hardware storage layer's weight is the correlation r between the cache capacity h1 and the parameter storage amount m1 of the model storage requirement layer. 11 As the weight of edge (h1, m1); hardware upper-layer communication index Hcom Interaction index M with the upper layer of the model com RH correlation k-mj As an edge (H) com M com The weight of ).
[0128] Step S350: Analyze the bipartite network to generate hardware centrality vectors and model centrality vectors.
[0129] In some embodiments, the PageRank algorithm is used to calculate node centrality, and the calculation formula includes:
[0130]
[0131] Where d is the damping coefficient, in(v) is the node pointed to by v, and out(u) is the node pointed to by u. For example, the damping coefficient d is 0.85.
[0132] In some examples, let's take the number of hardware nodes as h3, representing the number of computational cores; assuming the incoming edges come from the model's computational requirement layer m2, with a correlation degree of r. 32 =0.9, and the PageRank value of m2 is PR(m2) = 0.6, and the sum of the outgoing edge weights of m2 is 0.9 + 0.2 (assuming it is also associated with other hardware nodes), then:
[0133]
[0134] By traversing all hardware nodes, the hardware centrality vector is obtained using the following formula:
[0135] C h =[PR(h1), PR(h3), PR(H)] com )).
[0136] Similarly, the model node centrality is calculated using the following formula to obtain the model centrality vector:
[0137] C m =[PR(m1), PR(m2), PR(M com )).
[0138] Step S360: Construct a hardware-model matching evaluation function based on the hardware centrality vector and the model centrality vector.
[0139] In this embodiment, by using hierarchical partitioning, cosine similarity, and correlation propagation, complex adaptation relationships can be broken down into quantifiable and analyzable modules; it comprehensively covers the full-dimensional evaluation from micro-indicators to macro-levels, and from single correlations to network structures; and it enables the key nodes of the output to be interconnected with the levels, directly guiding subsequent optimization actions such as pruning and scheduling.
[0140] In some embodiments, step S330 involves passing the correlation of lower-level indicators in the hierarchical structure to the upper level based on the hardware hierarchy and the model hierarchy, including the following step S331.
[0141] Step S331: Calculate the matching strength between hardware nodes and model nodes based on the centrality of hardware nodes, the centrality of model nodes, and the correlation between hardware nodes and model nodes.
[0142] For example, the function that passes the correlation between lower-level indicators and higher-level indicators in a hierarchy to the upper level is:
[0143] S ij =C h *C m *r hij ;
[0144] Among them, S ij For hardware node h i With model node m j The matching strength between them, C h For hardware node h i Centrality, C m For model node m j Centrality, r hij Represented as hardware node h i With model node m j The degree of correlation between them.
[0145] In some embodiments, please refer to Figure 4 Step S360 includes the following steps S361 to S363.
[0146] Step S361: Based on the hardware centrality vector and the model centrality vector, determine the matching strength between the hardware node and the model node, and construct the matching strength matrix.
[0147] Optionally, the hardware centrality vector is in, For the centrality of the i-th hardware node, for example, if the hardware nodes include three types of nodes: cache, computing core, and communication bus, then C h =[0.6,0.7,0.8]. The model centrality vector is in, Let C be the centrality of the j-th model node. For example, if the model nodes include three types of nodes: parameter storage, computation tasks, and data interaction nodes, then C... m = [0.5, 0.6, 0.7].
[0148] For example, the association matrix between hardware nodes and model nodes is R = [r ij ]; where rij The degree of association between hardware node i and model node j, for example: the degree of association between cache and parameter storage r. 11 =0.9, the correlation between the computational core and the computational task r 22 =0.8, correlation coefficient r between communication bus and data interaction 33 =0.7, and other weakly correlated r ij For example, caching and computation tasks r 12 =0.3 etc.
[0149] For example, according to the matching strength formula The matching strength of each hardware-model node pair is calculated as follows.
[0150] Hardware Node 1 (Cache) & Model Node 1 (Parameter Storage): S 11 =0.6 × 0.5 × 0.9 = 0.27;
[0151] Hardware Node 1 (Cache) & Model Node 2 (Computation Task): S 12 =0.6 × 0.6 × 0.3 = 0.108;
[0152] Hardware Node 1 (Cache) & Model Node 3 (Data Interaction): S 13 =0.6×0.7×0.2=0.084;
[0153] Hardware Node 2 (Computational Core) & Model Node 1 (Parameter Storage): S 21 =0.7 × 0.5 × 0.2 = 0.07;
[0154] Hardware Node 2 (Computational Core) & Model Node 2 (Computational Task): S 22 =0.7 × 0.6 × 0.8 = 0.336;
[0155] Hardware Node 2 (Computing Core) & Model Node 3 (Data Interaction): S 23 =0.7 × 0.7 × 0.4 = 0.196;
[0156] Hardware Node 3 (Communication Bus) & Model Node 1 (Parameter Storage): S 31 =0.8 × 0.5 × 0.1 = 0.04;
[0157] Hardware Node 3 (Communication Bus) & Model Node 2 (Computation Task): S 32 =0.8 × 0.6 × 0.3 = 0.144;
[0158] Hardware Node 3 (Communication Bus) & Model Node 3 (Data Interaction): S 33 =0.8×0.7×0.7=0.392.
[0159] Based on the example above, an array is constructed according to the matching strength values of each computing hardware-model node pair, resulting in the following matching strength matrix S:
[0160]
[0161] Step S362: Sum the rows and columns of the matching strength matrix to obtain the hardware-side matching strength vector sum and the model-side matching strength vector sum.
[0162] In some examples, based on the obtained hardware-side matching strength vector S h The summation is performed row by row, and each element represents the total matching strength between a hardware node and all model nodes.
[0163] For example, the hardware-side matching strength vector S is obtained by summing the matrix S row by row according to the following formula. h :
[0164] Total match strength of hardware node 1 (cache)
[0165] Total matching strength of hardware node 2 (computing core)
[0166] Total matching strength of hardware node 3 (communication bus)
[0167] From the above, S h = [0.462, 0.602, 0.576].
[0168] In some examples, the model-side matching strength vector S is obtained. m Summing by column, each element is the sum of the matching strengths between a certain model node and all hardware nodes.
[0169] For example, the model-side matching strength vector S is obtained by summing each column of matrix S according to the following formula. m :
[0170] Total matching strength of model node 1 (parameter storage)
[0171] Total matching strength of model node 2 (computation task)
[0172] Total matching strength of model node 3 (data interaction)
[0173] From the above, S m = [0.38, 0.588, 0.672].
[0174] Step S363: The hardware-side matching strength vector sum and the model-side matching strength vector sum are respectively mean-processed, and the mean-processed result is determined as the hardware-model matching degree evaluation function.
[0175] For example, the hardware-model matching evaluation function is:
[0176]
[0177] Among them, S z This represents the hardware-model matching evaluation result; n is the number of nodes involved in the computation on the hardware side. is the matching strength vector element of the i-th node on the hardware side; m is the number of nodes participating in the calculation on the model side; This represents the matching strength vector element of the j-th node on the model side.
[0178] For example, the number of hardware nodes n = 3, and the number of model nodes m = 3; the hardware-side mean is determined according to the hardware-model matching evaluation function as follows: And the mean of the model side is determined to be:
[0179] Based on the above, the overall matching degree is calculated as follows:
[0180] In this embodiment, the complex relationship between hardware and model is simulated throughout the process, from hierarchical structure and correlation to network centrality (such as the correlation between cache, computing core, communication bus and different model requirements); the matching strength integrates centrality and correlation, and the mean processing balances bidirectional adaptation, which can accurately guide the optimization direction of subsequent pruning and scheduling.
[0181] In some embodiments, please refer to Figure 5 Step S400 includes the following steps S410 to S430.
[0182] Step S410: Determine the average hardware-side matching strength based on the number of hardware nodes involved in the calculation and the sum of the hardware-side matching strength vectors.
[0183] Step S420: Determine the mean of the model-side matching strength based on the number of model nodes involved in the calculation and the sum of the model-side matching strength vectors.
[0184] Step S430: Determine the hardware-model matching degree evaluation result based on the average of the hardware-side matching strength mean and the model-side matching strength mean.
[0185] In this embodiment, mean processing can balance the matching contributions of the hardware side and the model side, and avoid the impact of the difference in the number of hardware nodes and model nodes on the results.
[0186] It should be noted that the matching degree value in the model matching degree evaluation results is negatively correlated with the hardware-model matching degree.
[0187] In some embodiments, step S500 includes the following steps S510, S520 and S530.
[0188] Step S510: If the matching degree value in the hardware-model matching degree evaluation result is greater than or equal to the first threshold, a structured pruning strategy is adopted.
[0189] It should be noted that if the matching degree value in the hardware-model matching degree evaluation result is large (e.g., greater than or equal to the first threshold), it indicates that the matching degree between the hardware and the model is low, and the model's demand for hardware resources far exceeds the hardware's capabilities. In this case, the hardware's caching level may not be able to meet the model's frequent data read and write needs, and memory bandwidth will also become a performance bottleneck due to excessive data transfer volume. Therefore, it is necessary to prioritize the use of structured pruning strategies.
[0190] In this embodiment, by pruning the entire convolutional kernel, neuron group, or network layer, the number of model parameters and computational load can be significantly reduced, rapidly lowering the model's hardware resource requirements. For example, for a large power equipment fault diagnosis model containing multiple convolutional layers, if the evaluation results show a severe shortage of memory bandwidth, some redundant convolutional layers can be pruned to directly reduce data processing volume and alleviate memory bandwidth pressure.
[0191] For example, for the cache hierarchy, analyze the model's cache usage patterns. If the model frequently accesses the cache but the cache capacity is insufficient, during structured pruning, focus on pruning those parts that are frequently read from the cache but whose computation results have little impact on the final output; for example, prune some neuron groups that are only used for intermediate computations and whose results do not need to be kept in the cache for a long time, freeing up cache space and improving cache hit rate.
[0192] For example, regarding memory bandwidth, network structures involving a large number of data read and write operations can be pruned to address memory bandwidth bottlenecks; for instance, some fully connected layers with large data input and output scales can be removed to reduce the amount of data transmitted in memory and lower the demand for memory bandwidth.
[0193] Step S520: If the matching degree value in the hardware-model matching degree evaluation result is less than the first threshold and greater than or equal to the second threshold, a strategy combining structured pruning and unstructured pruning is adopted; wherein the first threshold is greater than the second threshold.
[0194] It should be noted that if the matching value in the hardware-model matching evaluation result is moderate (e.g., less than the first threshold and greater than or equal to the second threshold), it indicates that there is a certain adaptation problem between the hardware and the model. In this case, the hardware cache level may experience performance fluctuations in some high-load scenarios, and the memory bandwidth may also approach saturation in certain computing tasks. Therefore, a combination of structured pruning and unstructured pruning is required.
[0195] In this embodiment, structured pruning is used to remove obviously redundant network structures and initially reduce the model size; unstructured pruning addresses the local resource shortage problem that still exists after structured pruning by performing more refined trimming of model parameters, further optimizing resource consumption without affecting model accuracy as much as possible.
[0196] In some examples, structured pruning is used to remove some unimportant convolutional kernels from the large power load forecasting model, and then unstructured pruning is used to prune the smaller weighted connections in the remaining network.
[0197] For example, for cache levels, based on the data access frequency of the model at different cache levels, unstructured pruning is used to optimize the corresponding parts. If a performance degradation is found when the model accesses the second-level cache, unstructured pruning is used to remove parameters that are frequently accessed but contribute little to the cache, adjusting the data storage and access patterns in the cache to improve cache utilization efficiency.
[0198] For example, regarding memory bandwidth, analyze the peak data transfer rates during computational tasks. For computational modules that cause memory bandwidth constraints during peak times, use structured pruning to reduce their size, while using unstructured pruning to optimize internal parameters, balancing computational load and data transfer load, and preventing memory bandwidth from becoming a performance bottleneck.
[0199] Step S530: If the matching degree value in the hardware-model matching degree evaluation result is less than the second threshold, an unstructured pruning strategy is adopted.
[0200] It should be noted that if the matching score in the hardware-model matching evaluation result is small (e.g., less than the second threshold), it indicates that the hardware and model are well matched, but there may still be some local optimization space. In this case, the hardware cache level and memory bandwidth can basically meet the model requirements, but performance can still be further improved through fine-tuning, so unstructured pruning is required.
[0201] In this embodiment, unstructured pruning allows for fine-tuning of model parameters, removing minor redundant connections without disrupting the overall model structure, thus optimizing computational efficiency and resource utilization. For example, in large power system state estimation models, unstructured pruning can be used to remove connections with near-zero weights from the network, further reducing computational load without affecting model accuracy.
[0202] For example, for the cache level, based on the prefetching mechanism of hardware caching and the principle of data locality, the model parameter layout is adjusted through unstructured pruning, so that the storage of data in the cache is more in line with the hardware access habits, thereby improving the cache prefetch hit rate and further improving cache performance.
[0203] For example, regarding memory bandwidth, the continuity and burstiness of data transmission during the analysis model calculation process are considered. Unstructured pruning is used to optimize the parameter storage order, reduce fragmentation during data transmission, improve the effective utilization of memory bandwidth, and make fuller use of hardware resources.
[0204] It should be understood that, although Figures 2-5 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figures 2-5 At least some of the steps in the process may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but may be executed at different times. The execution order of these steps or stages is not necessarily sequential, but may be executed in turn or alternately with other steps or at least some of the steps or stages in other steps.
[0205] This application also provides a resource scheduling device based on a large power model according to some embodiments. This resource scheduling device can be used to execute the resource scheduling method based on a large power model described in the foregoing embodiments. The resource scheduling device also possesses all the technical advantages of the aforementioned resource scheduling methods. It should be noted that the parts that are the same as or corresponding to the above embodiments can be referred to the corresponding descriptions of the foregoing embodiments, and will not be elaborated upon below.
[0206] In some embodiments, please refer to Figure 6The resource scheduling device includes a first acquisition module, a first construction module, a second construction module, a first determination module, a second determination module, a third determination module, a model processing module, a model analysis module, a first generation module, and an optimization execution module. The first acquisition module is used to acquire the storage structure characteristics, power large model parameter quantity, and computation graph topology of the target hardware platform; the first construction module is used to construct a hardware capability matrix and a model requirement matrix based on the storage structure characteristics and the power large model parameter quantity, respectively; the second construction module is used to construct a hardware-model matching degree evaluation function based on the hardware capability matrix and the model requirement matrix; the first determination module is used to determine the hardware-model matching degree evaluation result based on the hardware-model matching degree evaluation function; the second determination module is used to determine a pruning strategy based on the hardware-model matching degree evaluation result; the third determination module is used to determine the model parameter storage capacity based on the pruning strategy and hardware storage granularity; the model processing module is used to prune and compress the power large model based on the model parameter storage capacity to obtain a target compressed model; the model analysis module is used to decompose the computation graph of the target compressed model based on the computation graph topology to obtain multiple subgraphs; the first generation module is used to allocate computing units to each subgraph according to the hardware capability matrix and generate a hardware instruction scheduling sequence; the optimization execution module is used to perform resource scheduling based on the hardware instruction scheduling sequence.
[0207] In some embodiments, the second construction module includes: a first partitioning unit, used to hierarchically partition the hardware capability matrix and the model requirement matrix to obtain a hardware hierarchical structure and a model hierarchical structure; a first acquisition unit, used to determine the correlation between each hardware indicator and each model indicator based on cosine similarity within the hardware hierarchical structure and the model hierarchical structure, to obtain a hardware correlation matrix and a model correlation matrix respectively; a matrix processing unit, used to pass the correlation of lower-level indicators in the hierarchical structure to the upper level based on the hardware correlation matrix and the model correlation matrix, and according to the hardware hierarchical structure and the model hierarchical structure, to obtain a hardware correlation target matrix and a model correlation target matrix respectively; a first construction unit, used to integrate the hardware correlation target matrix and the model correlation target matrix to construct a bipartite network; wherein the bipartite network uses hardware indicators as hardware nodes and model indicators as model nodes; the weight of the edge in the bipartite network is the correlation between the corresponding hardware indicator and the model indicator; a first generation unit, used to analyze the bipartite network to generate a hardware centrality vector and a model centrality vector; and a second construction unit, used to construct a hardware-model matching degree evaluation function based on the hardware centrality vector and the model centrality vector.
[0208] In some embodiments, the matrix processing unit includes: a first calculation subunit, configured to calculate the matching strength between the hardware node and the model node based on the centrality of the hardware node, the centrality of the model node, and the correlation between the hardware node and the model node.
[0209] In some embodiments, the second construction unit includes: a first construction subunit, configured to determine the matching strength between hardware nodes and model nodes based on hardware centrality vectors and model centrality vectors, and construct a matching strength matrix; a second calculation subunit, configured to sum the rows and columns of the matching strength matrix respectively to obtain the sum of hardware-side matching strength vectors and the sum of model-side matching strength vectors; and a mean processing unit, configured to perform mean processing on the sum of hardware-side matching strength vectors and the sum of model-side matching strength vectors respectively, and determine the mean processing result as the hardware-model matching degree evaluation function.
[0210] In some embodiments, the first determining module includes: a first determining unit, configured to determine the average hardware-side matching strength based on the number of hardware nodes participating in the calculation on the hardware side and the sum of hardware-side matching strength vectors; a second determining unit, configured to determine the average model-side matching strength based on the number of model nodes participating in the calculation on the model side and the sum of model-side matching strength vectors; and a third determining unit, configured to determine the hardware-model matching degree evaluation result based on the average of the average hardware-side matching strength and the average model-side matching strength.
[0211] In some embodiments, the second determining module includes a comparison unit and a pruning strategy determining unit. The comparison unit compares the matching degree value, a first threshold, and a second threshold in the hardware-model matching degree evaluation result; the first threshold is greater than the second threshold. The pruning strategy determining unit, based on the comparison result of the comparison unit, employs a structured pruning strategy when the matching degree value in the hardware-model matching degree evaluation result is greater than or equal to the first threshold; employs a strategy combining structured and unstructured pruning when the matching degree value in the hardware-model matching degree evaluation result is less than the first threshold but greater than or equal to the second threshold; and employs an unstructured pruning strategy when the matching degree value in the hardware-model matching degree evaluation result is less than the second threshold.
[0212] Please see Figure 7 This application also provides a computer device according to some embodiments, including a memory and a processor, wherein the memory stores a computer program; the processor executes the computer program to implement the steps of the method described in the foregoing embodiments of this application. This computer device also possesses all the technical advantages of the aforementioned resource scheduling methods. It should be noted that the parts that are the same as or corresponding to the above embodiments can be referred to the corresponding descriptions of the foregoing embodiments, and will not be described in detail below.
[0213] This application also provides a computer-readable storage medium storing a computer program according to some embodiments; when the computer program is executed by a processor, it implements the steps of the methods described in the foregoing embodiments of this application. This computer-readable storage medium also possesses all the technical advantages of the aforementioned resource scheduling methods. It should be noted that the parts that are the same as or corresponding to the above embodiments can be referred to the corresponding descriptions of the foregoing embodiments, and will not be described in detail below.
[0214] This application also provides a computer program product according to some embodiments, including a computer program; when executed by a processor, the computer program implements the steps of the methods described in the foregoing embodiments of this application. The computer program product also possesses all the technical advantages of the aforementioned resource scheduling methods. It should be noted that the parts that are the same as or corresponding to the above embodiments can be referred to the corresponding descriptions of the foregoing embodiments, and will not be described in detail below.
[0215] In the description of this specification, references to terms such as "some embodiments," "some examples," "exemplarily," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative descriptions of the above terms do not necessarily refer to the same embodiments or examples.
[0216] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0217] The embodiments described above are merely examples of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application.
Claims
1. A resource scheduling method based on a large-scale power model, characterized in that, include: Obtain the storage structure characteristics, power large model parameter quantity, and computation graph topology of the target hardware platform; Based on the storage structure characteristics and the number of parameters of the power large model, a hardware capability matrix and a model requirement matrix are constructed respectively. Based on the hardware capability matrix and the model requirement matrix, a hardware-model matching degree evaluation function is constructed; The hardware-model matching evaluation result is determined based on the hardware-model matching evaluation function. Based on the hardware-model matching evaluation results, a pruning strategy is determined; The storage capacity of the model parameters is determined based on the pruning strategy and hardware storage granularity. Based on the storage capacity of the model parameters, the large power model is pruned and compressed to obtain the target compressed model. Based on the computation graph topology, the target compression model is decomposed into multiple subgraphs. Based on the hardware capability matrix, a computing unit is allocated to each subgraph to generate a hardware instruction scheduling sequence; Resource scheduling is performed based on the hardware instruction scheduling sequence.
2. The resource scheduling method according to claim 1, characterized in that, The step of constructing a hardware-model matching evaluation function based on the hardware capability matrix and the model requirement matrix includes: The hardware capability matrix and the model requirement matrix are respectively divided into hierarchical parts to obtain the hardware hierarchical structure and the model hierarchical structure. Within the hardware hierarchy and the model hierarchy, the correlation between each hardware index and each model index is determined based on cosine similarity, resulting in a hardware correlation matrix and a model correlation matrix, respectively. Based on the hardware correlation matrix and the model correlation matrix, and according to the hardware hierarchy and the model hierarchy, the correlation of the lower-level indicators in the hierarchy is passed to the upper level to obtain the hardware correlation target matrix and the model correlation target matrix respectively. The hardware correlation target matrix and the model correlation target matrix are integrated to construct a bipartite network; the bipartite network uses the hardware indicators as hardware nodes and the model indicators as model nodes; the weight of the edge of the bipartite network is the correlation between the corresponding hardware indicator and the model indicator. The binary network is analyzed to generate hardware centrality vectors and model centrality vectors; Based on the hardware centrality vector and the model centrality vector, construct the hardware-model matching evaluation function.
3. The resource scheduling method according to claim 2, characterized in that, The step of passing the correlation between lower-level indicators in the hierarchical structure and the model hierarchy to the upper level includes: The matching strength between the hardware node and the model node is calculated based on the centrality of the hardware node, the centrality of the model node, and the correlation between the hardware node and the model node.
4. The resource scheduling method according to claim 2, characterized in that, The step of constructing the hardware-model matching evaluation function based on the hardware centrality vector and the model centrality vector includes: Based on the hardware centrality vector and the model centrality vector, the matching strength between the hardware node and the model node is determined, and a matching strength matrix is constructed. The rows and columns of the matching strength matrix are summed to obtain the hardware-side matching strength vector sum and the model-side matching strength vector sum; The hardware-side matching strength vector and the model-side matching strength vector are respectively mean-processed, and the mean-processed result is determined as the hardware-model matching degree evaluation function.
5. The resource scheduling method according to claim 4, characterized in that, The step of determining the hardware-model matching degree evaluation result based on the hardware-model matching degree evaluation function includes: The average hardware-side matching strength is determined based on the number of hardware nodes involved in the calculation and the sum of the hardware-side matching strength vectors. The mean value of the model-side matching strength is determined based on the number of model nodes involved in the calculation and the sum of the model-side matching strength vectors. The hardware-model matching degree evaluation result is determined based on the average of the mean matching strength on the hardware side and the mean matching strength on the model side.
6. The resource scheduling method according to claim 1, characterized in that, The step of determining the pruning strategy based on the hardware-model matching evaluation result includes: If the matching degree value in the hardware-model matching degree evaluation result is greater than or equal to the first threshold, a structured pruning strategy is adopted. If the matching degree value in the hardware-model matching degree evaluation result is less than the first threshold and greater than or equal to the second threshold, a strategy combining structured pruning and unstructured pruning is adopted; wherein, the first threshold is greater than the second threshold. If the matching degree value in the hardware-model matching degree evaluation result is less than the second threshold, an unstructured pruning strategy is adopted.
7. A resource scheduling device based on a large power model, characterized in that, include: The first acquisition module is used to acquire the storage structure characteristics, power large model parameter quantity, and computation graph topology of the target hardware platform; The first construction module is used to construct a hardware capability matrix and a model requirement matrix based on the storage structure characteristics and the number of parameters of the power large model, respectively. The second construction module is used to construct a hardware-model matching degree evaluation function based on the hardware capability matrix and the model requirement matrix; The first determining module is used to determine the hardware-model matching degree evaluation result based on the hardware-model matching degree evaluation function. The second determining module is used to determine the pruning strategy based on the hardware-model matching degree evaluation result; The third determining module is used to determine the storage capacity of model parameters based on the pruning strategy and hardware storage granularity. The model processing module is used to prune and compress the large power model according to the storage capacity of the model parameters to obtain the target compressed model; The model analysis module is used to decompose the target compressed model into multiple subgraphs based on the computation graph topology. The first generation module is used to allocate computing units to each of the subgraphs according to the hardware capability matrix and generate a hardware instruction scheduling sequence. An optimized execution module is used for resource scheduling based on the hardware instruction scheduling sequence.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.