Full-link environment pre-check system, method, device, medium and program product
Through the full-link environmental pre-inspection system, the hardware resource status is dynamically monitored and accurate predictions and risk assessments are performed, which solves the problems of insufficient resource utilization and low operation and maintenance efficiency in artificial intelligence model training, and realizes the reliability of job deployment and the improvement of resource utilization efficiency.
Patent Information
- Application Number
- CN202510804069.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-09-09
AI Technical Summary
In existing technologies, artificial intelligence model training faces problems such as insufficient resource utilization, poor training stability, and low operation and maintenance efficiency. Especially when the resource allocation strategy is static and relies on manual operation, it is difficult to achieve dynamic monitoring of resource usage and overall risk assessment.
A full-link environmental pre-inspection system is provided, which receives job parameters through the interface layer, monitors the hardware resource status through the environmental perception layer, makes accurate predictions through the resource demand prediction layer, and uses the risk assessment layer to calculate risk factors for graded alarms, forming a complete link of dynamic perception-accurate prediction-scientific risk assessment to ensure accurate resource matching before job deployment.
It significantly improves the reliability of job deployment and resource utilization efficiency, avoids resource waste, saves hardware costs, improves operation and maintenance efficiency, and provides all-round guarantees for efficient and stable operation of jobs.
Smart Images

Figure CN120610840A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a full-link environment pre-inspection system, method, device, medium and program product. Background Art
[0002] AI model training faces multiple challenges: In terms of resource management, static resource allocation strategies struggle to adapt to dynamically changing training requirements, often leading to underutilization of graphics processor (GPU) memory resources. Regarding training stability, if the demand for data transmission in distributed training exceeds the network bandwidth capacity, it can easily cause job interruptions or failures, impacting the continuity of the training process. At the operations and maintenance level, environmental configuration and status checks rely on manual operations, which is costly and inefficient.
[0003] In related technical solutions, detection systems for artificial intelligence operations usually deploy agent programs to periodically collect processor and memory usage data, and trigger alarms when resource utilization exceeds a preset threshold. However, this method can only achieve passive monitoring of resource usage and can only provide fragmented alarm information, and cannot analyze and judge the potential risks faced by training tasks from a holistic perspective. Summary of the Invention
[0004] The present invention provides a full-link environmental pre-inspection system, method, equipment, medium and program product, which can form a complete link from perception, prediction to evaluation, analyze potential risks from a holistic perspective, and significantly improve the reliability of job deployment and resource utilization efficiency.
[0005] The present invention provides a full-link environment pre-inspection system, comprising: Interface layer, used to receive job parameters; The environment perception layer monitors the hardware resource status in the full-link environment and obtains available video memory capacity, inter-node bandwidth, and available memory capacity in turn. The resource demand prediction layer is used to couple the job parameters with the hardware resource status to build a model, and predict the video memory demand, network demand and memory demand based on the constructed prediction model; The risk assessment layer is used to calculate the risk factor based on the available video memory capacity, inter-node bandwidth and available memory capacity obtained by the environmental perception layer, and the video memory demand, network demand and memory demand predicted by the resource demand prediction layer, and perform risk assessment based on the risk factor and the set graded alarm rules; the risk factor is an indicator that quantifies the degree of resource gap.
[0006] The present invention also provides a full-link environment pre-check method, comprising: Receive job parameters; Monitor the hardware resource status in the full-link environment and obtain available video memory capacity, inter-node bandwidth, and available memory capacity in sequence; Coupling the job parameters with the hardware resource status to model the performance, and predicting the video memory requirement, network requirement, and memory requirement based on the constructed prediction model; Based on the acquired available video memory capacity, inter-node bandwidth and available memory capacity, as well as the predicted video memory demand, network demand and memory demand, a risk factor is calculated, and a risk assessment is performed based on the risk factor and the set graded alarm rules; the risk factor is an indicator that quantifies the degree of resource gap.
[0007] The present invention also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned full-link environment pre-inspection methods when executing the computer program.
[0008] The present invention also provides a computer-readable storage medium, which stores a computer program, wherein when the computer program is executed by a processor, the steps of any of the above-mentioned full-link environment pre-inspection methods are implemented.
[0009] The present invention also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned full-link environment pre-inspection methods when executed by a processor.
[0010] Through the above-mentioned full-link environmental pre-inspection system provided by the present invention, the interface layer is used to receive job parameters, the environmental perception layer is used to dynamically monitor the status of hardware resources in the entire link, and the available video memory capacity, inter-node bandwidth and available memory capacity are obtained in turn. The resource demand prediction layer is used to couple the job parameters with the hardware resource status to model the demand for video memory, network and memory accurately, and the risk assessment layer is used to calculate the risk factor based on the acquired resource data and the predicted resource demand and to issue graded alarms for risk assessment. In this way, a complete link of dynamic perception-accurate prediction-scientific risk assessment is formed, an end-to-end pre-inspection system is realized, and resources are accurately matched before job deployment, which significantly improves the reliability and resource utilization efficiency of job deployment, avoids resource waste caused by static allocation, and can analyze potential risks as a whole, saving hardware costs, improving operation and maintenance efficiency, and providing all-round protection for efficient and stable operation of jobs.
[0011] In addition, the present invention also provides a corresponding pre-inspection method, electronic device, computer-readable storage medium and program product for the full-link environment pre-inspection system, which has the same or corresponding technical features as the above-mentioned full-link environment pre-inspection system and has the same effect as above. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the embodiments of the present invention, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0013] Figure 1 A schematic diagram of the structure of a full-link environment pre-inspection system provided by an embodiment of the present invention; Figure 2 This is a flowchart of the full-link environment pre-check method provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0014] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0015] It should be noted that, in the description of the present invention, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. The terms "first," "second," etc., in the present invention are used to distinguish similar objects, and are not used to describe a particular order or precedence.
[0016] In order to enable those skilled in the art to better understand the solutions of the present invention, the present invention is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0017] In combination with the specific application environment architecture or specific hardware architecture on which the execution of the full-link environment pre-inspection system depends, the specific application environment architecture or specific hardware architecture is described here.
[0018] An embodiment of the present invention provides a full-link environment pre-inspection system. Figure 1 A schematic diagram of the structure of the full-link environment pre-inspection system provided by an embodiment of the present invention is shown as follows: Figure 1 As shown, the system includes: Interface layer 1 is used to receive job parameters; it should be noted that job parameters refer to various parameters and settings used to define task attributes, scopes, conditions or behaviors when completing a task, operation or executing a specific program. In the present invention, job parameters may include model characteristic parameters, hardware requirements, etc. of artificial intelligence jobs; the artificial intelligence jobs here are training or reasoning tasks based on deep learning models, which require computing, storage and network resources; model characteristic parameters may include the number of parameters, batch size, sequence length, number of nodes, etc. of deep learning models; hardware requirements may include the memory capacity range of each node graphics processor. Batch size refers to the number of samples input to the model in a single iteration. The number of nodes refers to the number of physical servers participating in distributed training.
[0019] Environmental Perception Layer 2 monitors the status of hardware resources in the full-link environment, sequentially obtaining available graphics memory capacity, inter-node bandwidth, and available memory capacity. The full-link environment here refers to the operating environment that includes the graphics processor cluster, network topology, storage system, and software dependencies. The full-link environment pre-check system of the present invention is an automated system that completes environmental diagnosis, resource prediction, and risk assessment before job submission. It can achieve collaborative prediction and risk assessment of hardware resources such as graphics memory, network, and memory before job deployment.
[0020] Resource demand prediction layer 3 is used to couple job parameters with hardware resource status and model the demand for video memory, network, and memory based on the constructed prediction models. The prediction models here include video memory demand prediction models, network demand prediction models, and memory demand prediction models.
[0021] Risk assessment layer 4 is used to calculate risk factors based on the available video memory capacity, inter-node bandwidth and available memory capacity obtained by the environmental perception layer 2, as well as the video memory demand, network demand and memory demand predicted by the resource demand prediction layer 3, and perform risk assessment based on the risk factors and set hierarchical alarm rules; the risk factor is an indicator that quantifies the degree of resource gap.
[0022] In the above-mentioned full-link environmental pre-inspection system provided by the embodiment of the present invention, the interface layer 1 is used to receive job parameters, the environmental perception layer 2 is used to dynamically monitor the status of hardware resources in the entire link, and the available video memory capacity, inter-node bandwidth and available memory capacity are obtained in turn. The resource demand prediction layer 3 is used to couple the job parameters with the hardware resource status to model the precise demand for video memory, network and memory, and the risk assessment layer 4 is used to calculate the risk factor based on the acquired resource data and the predicted resource demand and to issue graded alarms for risk assessment. In this way, a complete link of dynamic perception-precise prediction-scientific risk assessment is formed, an end-to-end pre-inspection system is realized, and the reliability and resource utilization efficiency of job deployment are ensured, which avoids the waste of resources caused by static allocation. Potential risks can be analyzed as a whole, saving hardware costs, improving operation and maintenance efficiency, and providing all-round guarantees for efficient and stable operation of jobs.
[0023] Furthermore, in a specific implementation, in the above-mentioned full-link environment pre-inspection system provided in an embodiment of the present invention, the environment perception layer 2 may include: a video memory perception unit, used to obtain the operating data of the graphics processor and obtain the available video memory capacity based on the operating data; a network perception unit, used to measure the effective bandwidth between nodes; and a memory perception unit, used to parse system files about memory usage and obtain the available memory capacity.
[0024] In practice, the environment awareness layer 2 of the present invention can achieve video memory awareness, network awareness, and memory awareness. The video memory awareness unit can obtain real-time operating data of the graphics processor through the NVIDIA Management Library (NVML) interface and obtain available video memory capacity based on this real-time operating data. The network awareness unit can measure the effective bandwidth between nodes using professional tools that can measure network bandwidth performance. The memory awareness unit can parse system files related to memory usage to obtain available memory capacity.
[0025] Furthermore, in a specific implementation, in the above-mentioned full-link environment pre-inspection system provided in an embodiment of the present invention, the video memory perception unit may include: a capacity acquisition subunit, used to obtain the operating data of the graphics processor, and obtain the total video memory capacity and the allocated video memory capacity from the operating data; a first calculation subunit, used to obtain the available video memory capacity based on the difference between the total video memory capacity and the allocated video memory capacity.
[0026] In implementation, the capacity acquisition subunit can use the NVML interface to obtain the real-time operation data of the graphics processor, and obtain the total video memory capacity and the allocated video memory capacity from the operation data. Next, the first calculation subunit uses formula (1) to calculate the current available video memory capacity: ; (1)
[0027] in, is the available video memory capacity (unit: GB), is the total video memory capacity (unit: GB), The allocated video memory capacity (unit: GB).
[0028] This can monitor the usage status of video memory in real time, accurately grasp the remaining space of video memory, and provide data support for system resource scheduling; it can detect abnormal video memory allocation in advance to avoid task interruption or program crash due to insufficient video memory; and help optimize video memory resource allocation strategy, improve computing efficiency, and ensure stable operation of high-load tasks.
[0029] Furthermore, in a specific implementation, in the above-mentioned full-link environment pre-inspection system provided in an embodiment of the present invention, the network perception unit may include: a channel establishment sub-unit, used to establish a data transmission channel between nodes; a data acquisition sub-unit, used to obtain the test start time, test end time and the amount of data transmitted each time after multiple data transmissions on the data transmission channel; a second calculation sub-unit, used to obtain the total test time based on the test start time and test end time; and to obtain the total amount of data transmitted multiple times based on the amount of data transmitted each time and the number of transmissions; a third calculation sub-unit, used to obtain the effective bandwidth between nodes based on the ratio of the total amount of data transmitted multiple times to the total test time.
[0030] In practice, the network perception unit can measure the effective bandwidth between nodes using tools used to measure network bandwidth performance. Specifically, the channel establishment subunit establishes a data transmission channel between nodes; the data acquisition subunit performs multiple data transmission tests, while recording the test start time, test end time, and the amount of data transmitted each time; then, formula (2) is used to obtain the effective bandwidth between nodes: ; (2) in, is the effective bandwidth between nodes; is the amount of data transferred for the i-th time (unit: GB), usually set to To reduce errors; The test end time, is the test start time, and n is the total number of data transmissions.
[0031] For example, if three transmission tests take 0.12s, 0.11s, and 0.13s respectively, the actual bandwidth is calculated as: .
[0032] This approach of eliminating the effects of transient fluctuations through multiple tests allows measurement results to more accurately reflect the true network status, precisely measuring effective network bandwidth and eliminating the impact of transient fluctuations, providing reliable data for network performance evaluation. In scenarios such as distributed computing and cluster deployments, this approach can assist in evaluating inter-node communication capabilities, ensuring the stability of large-scale data transmission and preventing the impact of insufficient network performance on overall system efficiency.
[0033] Furthermore, in a specific implementation, in the above-mentioned full-link environment pre-inspection system provided in an embodiment of the present invention, the memory perception unit may include: a file parsing sub-unit, used to parse the system file about memory usage, and obtain the physical memory capacity that is completely unused in the system, the memory capacity occupied by the system cache, and the memory capacity occupied by the block device buffer; a fourth calculation sub-unit, used to add the physical memory capacity that is completely unused in the system, the memory capacity occupied by the system cache, and the memory capacity occupied by the block device buffer to obtain the available memory capacity.
[0034] In implementation, the memory awareness unit can analyze the (a virtual file used to view memory information) to obtain available memory information. It should be noted that, This file stores detailed information about system memory usage. This file does not exist on physical storage devices but is generated dynamically by the kernel. It records the usage of physical memory, swap space, various caches, and buffers in text format.
[0035] File parsing subunit system By parsing the file, we can obtain the completely unused physical memory capacity in the system, the memory capacity occupied by the system cache (Page Cache), and the memory capacity occupied by the block device buffer (Buffer); the fourth calculation sub-unit uses formula (3) to obtain the available memory capacity: ; (3) in, is the available memory capacity (unit: MB), It is the physical memory capacity that is completely unused in the system (unit: MB). The memory capacity occupied by the system cache, including the file system cache and other quickly released memory (unit: MB). The memory capacity occupied by the block device buffer, used to temporarily store disk input and output operation data (unit: MB).
[0036] The above method can accurately calculate the current actual available memory capacity of the system, including unused physical memory, system cache that can be quickly released, and disk input and output buffer occupancy, providing an important basis for system resource management and scheduling.
[0037] Furthermore, in a specific implementation, in the above-mentioned full-link environment pre-inspection system provided in an embodiment of the present invention, the resource demand prediction layer 3 may include: a video memory demand prediction unit, which is used to input the batch size, sequence length and model parameter quantity in the job parameters into the video memory demand prediction model to predict the video memory demand; a network demand prediction unit, which is used to obtain the gradient data volume and the number of nodes according to the job parameters and input them into the network demand prediction model to predict the network demand; a memory demand prediction unit, which is used to collect the historical memory quantity at multiple time points and input them into the memory demand prediction model to predict the memory demand; the memory demand prediction model is a time series model.
[0038] During implementation, the memory demand prediction unit can input the batch size, sequence length, and model parameter count into the memory demand prediction model to predict memory demand. The network demand prediction unit can calculate the gradient data volume and node count based on job parameters and input these into the network demand prediction model to predict network demand. Gradient data volume refers to the amount of model parameter updates transmitted between nodes during distributed training (unit: GB). The memory demand prediction unit can collect historical memory usage at multiple time points and input this into the memory demand prediction model to predict memory demand. The memory demand prediction model here can be a time series model. The autoregressive integrated moving average (ARIMA) model can be used for time series prediction.
[0039] Furthermore, in a specific implementation, in the above-mentioned full-link environment pre-inspection system provided in an embodiment of the present invention, the video memory demand prediction unit may include: a coefficient value fitting sub-unit, used to fit the video memory occupancy data of historical jobs through the least squares method to obtain the coefficients of the video memory demand prediction model; a first model inference sub-unit, used to extract the batch size, sequence length and model parameter amount from the job parameters, and combine the coefficients of the video memory demand prediction model to calculate the video memory demand.
[0040] In implementation, the memory demand prediction unit takes into account three key factors: batch size, input sequence length, and model parameter quantity. Among them, the batch size directly affects the memory occupancy of the activation value, the input sequence length is in a square relationship with the memory occupancy of the attention mechanism, and the model parameter quantity determines the storage space of the parameters. The coefficient value fitting subunit of the present invention can fit the memory occupancy data of historical jobs by the least squares method to obtain the coefficient value corresponding to the memory demand prediction model. The memory occupancy data of historical jobs includes the actual memory occupancy value under different batch sizes, input sequence lengths, and model parameter quantity combinations. Using the obtained coefficient value combined with the batch size, sequence length, and model parameter quantity extracted from the job parameters, the first model inference subunit calculates the memory demand by formula (4): ; (4) in, 、 、 is the coefficient of the memory demand prediction model, for example: , , ; The batch size directly affects the memory usage of activation values; is the input sequence length, which is squared with the memory usage of the attention mechanism, quantifying the nonlinear growth characteristic of the memory overhead of the attention mechanism; The number of model parameters (unit: million) determines the parameter storage space.
[0041] Formula (4) above can be used as a model for predicting video memory requirements. This formula, through multi-dimensional modeling of batch size, sequence length, and model parameter count, fully considers the nonlinear impact of the attention mechanism on video memory usage, significantly reducing the error rate of video memory requirement prediction and significantly improving accuracy. It can accurately estimate the amount of video memory required during training. This is of great significance for rationally planning computing resources and avoiding training interruptions or performance degradation caused by insufficient video memory. At the same time, by fitting historical data, this method can adapt to the characteristics of different models and training scenarios, improving the accuracy and practicality of predictions.
[0042] Furthermore, in a specific implementation, in the above-mentioned full-link environment pre-inspection system provided in an embodiment of the present invention, the network demand prediction unit may include: a gradient data calculation subunit, used to extract the model parameter quantity and the number of data type bytes from the job parameters; obtain the single iteration gradient data size according to the product of the model parameter quantity and the number of data type bytes; a single node transmission calculation subunit, used to calculate the amount of data transmitted per second by a single node based on the single iteration gradient data size, the number of iterations per second, and the single iteration time; a second model inference subunit, used to extract the number of nodes from the job parameters, and calculate the network demand according to the product of the amount of data transmitted per second by a single node and the number of nodes.
[0043] In implementation, the gradient data calculation subunit can extract the model parameter quantity and the data type byte number from the job parameters; according to the product of the model parameter quantity and the data type byte number, the single iteration gradient data size is obtained by formula (5): ; (5) in, is the size of gradient data for a single iteration (unit: GB).
[0044] The single-node transmission calculation subunit can calculate the amount of data transmitted per second by a single node based on the size of the gradient data in a single iteration, the number of iterations per second, and the single iteration time. The iteration time refers to the time it takes to complete one forward propagation and one backward propagation (unit: seconds). The second model inference subunit can calculate the network demand using formula (6): ; (6) in, is the network demand, is the number of iterations per second, i.e. the iteration frequency, which is determined by the training speed. is the single iteration time (unit: seconds), is the number of nodes.
[0045] The above formula (5) combined with the above formula (6) can be used as a network demand prediction model. The logic of the formula derivation is that the amount of data that a single node needs to transmit per second is: ;Total network requirements: Single node requirements Number of nodes.
[0046] The network demand prediction model of this invention primarily considers four key factors: the size of the gradient data per iteration, the number of iterations per second, the time it takes for a single iteration, and the number of distributed nodes. By introducing a dynamic correlation model between gradient data volume and iteration frequency, it accurately captures the communication peak characteristics of distributed training, significantly reducing the error rate and improving accuracy of network demand prediction. This allows for precise estimation of the network bandwidth required for distributed training, helping users plan cluster configurations and avoid training efficiency losses due to network bottlenecks.
[0047] Furthermore, in a specific implementation, in the above-mentioned full-link environment pre-inspection system provided in an embodiment of the present invention, the memory demand prediction unit may include: a parameter determination subunit, which is used to determine the order of the memory demand prediction model through the autocorrelation graph and partial autocorrelation graph of historical data, and then use the maximum likelihood estimation method to fit the autoregressive coefficient of the memory demand prediction model; a third model inference subunit, which is used to collect the historical memory quantities at multiple time points, and input the historical memory quantities into the memory demand prediction model, use the autoregressive coefficients for weighted summation, and add the random fluctuation term to calculate the memory demand.
[0048] In implementation, the parameter determination subunit can determine the order of the memory demand prediction model through the autocorrelation function (ACF) and partial autocorrelation function (PACF) of historical data, and then use the maximum likelihood estimation (MLE) method to fit the autoregressive coefficients of the memory demand prediction model. The third model inference subunit can collect historical memory quantities at multiple time points and input the historical memory quantities into the memory demand prediction model. The autoregressive coefficients are used for weighted summation, and the random fluctuation term is added to calculate the memory demand using formula (7): ; (7) in, 、 、 is the autoregressive coefficient, e.g. , , . For the Memory requirements at the moment; 、 and yes The amount of historical memory at three time points before the moment; it is a random fluctuation term, subject to .
[0049] The above formula (7) can be used as a memory demand prediction model. The memory demand prediction model can use the ARIMA time series model to estimate future memory demand by analyzing historical memory usage data. The model can be calculated based on the memory usage at the past three time points, combined with the autoregressive coefficient obtained by fitting, while taking into account the influence of random errors. The autoregressive coefficient is obtained by determining the model order through the autocorrelation diagram and partial autocorrelation diagram of historical data, and then solved using the maximum likelihood estimation method. It can effectively identify the periodic fluctuation pattern of memory usage and make a relatively accurate prediction of short-term memory demand. This helps the system to schedule memory resources in advance and avoid service anomalies caused by insufficient memory. It can also provide data support for hardware resource planning and improve resource utilization efficiency. It is especially suitable for scenarios where memory usage has periodic or trend changes.
[0050] Furthermore, in a specific implementation, in the above-mentioned full-link environment pre-inspection system provided in an embodiment of the present invention, the risk assessment layer 4 may include: a video memory risk factor calculation unit, used to calculate the video memory risk factor based on the ratio of video memory demand to available display capacity; a network risk factor calculation unit, used to calculate the network risk factor based on the ratio of network demand to inter-node bandwidth; and a memory risk factor calculation unit, used to calculate the memory risk factor based on the ratio of memory demand to available memory capacity.
[0051] In practice, the risk factor can be calculated by formula (8): ; (8) in, is a risk factor, To predict resource requirements (GPU / network / memory), The current amount of available resources.
[0052] The risk factor calculation method assesses the usage risk of three resource categories: graphics memory, network, and internal memory. This method divides the predicted resource demand by the currently available resource quantity and multiplies the result by 100% to derive the risk factor for each resource type. This method provides a quantitative and intuitive reflection of the potential risk level of each resource. This risk factor allows users to clearly understand the gap between resource demand and available quantity, facilitating proactive resource scheduling and risk warnings, avoiding task interruptions or system failures caused by insufficient resources, and providing a valuable reference for rational resource planning and management.
[0053] It should be noted that if , indicating sufficient resources, marked as risk-free. , indicating that the resource gap exceeds the current total amount and is marked as severely over-limit. The present invention can adopt the worst principle and use the resource with the highest risk as the overall risk level. Table 1 shows the rules for setting graded alarms.
[0054] Table 1 Setting the hierarchical alarm rules
[0055] Furthermore, in specific implementation, the above-mentioned full-link environment pre-inspection system provided in the embodiment of the present invention may also include: a decision engine layer, which is used to arrange resource problems in a set order of risk levels and generate optimization suggestions; the optimization suggestions include: when the memory risk factor is higher than the set memory risk factor threshold, reducing the batch size or enabling model parallelism; when the network risk factor is higher than the set network risk factor threshold, reducing the number of nodes or enabling gradient compression; when the memory risk factor is higher than the set memory risk factor threshold, adjusting data loading or restricting background processes.
[0056] In implementation, the present invention can generate optimization suggestions based on the risk level, such as adjusting the batch size or nodes, etc. Among them, resource problems can be arranged from high to low according to the risk level, such as video memory is greater than the network, and network is greater than memory. When the video memory risk factor is higher than the set video memory risk factor threshold, it means that the video memory is insufficient. When encountering a situation of insufficient video memory, the batch size can be reduced to reduce the amount of data processed by the model during a single training or inference, thereby reducing the video memory occupancy; enabling model parallelism can split the model into multiple parts and distribute them on different devices for calculation, avoiding overloading the video memory of a single device. When the network risk factor is higher than the set network risk factor threshold, it means that the network is insufficient. If the problem of insufficient network occurs, the number of nodes can be reduced to reduce the demand for data transmission between nodes and alleviate network pressure; enabling gradient compression mode, by compressing the gradient data, reducing the amount of data transmitted, improving network transmission efficiency, and balancing the bandwidth load. In the natural language processing scenario, the training speed is accelerated due to network optimization. When the memory risk factor is higher than the set memory risk factor threshold, it means that there is insufficient memory. When memory is insufficient, you can optimize the data loading strategy, such as using asynchronous loading, batch loading, etc., to avoid reading large amounts of data into memory at one time; limit the running of background processes and release memory resources occupied by irrelevant programs, which can effectively ensure the memory requirements of critical tasks.
[0057] The following uses a specific example to illustrate the full-link environment pre-inspection system of the present invention. The interface layer 1 receives the job parameters, which include: the model targeted by the job is a pre-trained language model (Bidirectional EncoderRepresentations from Transformers-large, BERT-large), the model parameter quantity , sequence length , batch size , number of nodes ;Hardware requirements: GPU memory per node .
[0058] The environment perception layer 2 monitors the status of hardware resources such as video memory, network, and memory in the full link environment and obtains the available video memory of a single node. ; Inter-node bandwidth (Average value of 5 tests); Available memory .
[0059] The memory demand predicted by resource demand forecast layer 3 is: .
[0060] Forecasted network demand: First, calculate the amount of gradient data, where the BERT-large parameter size is 340M, using the FP32 data type ( ): ; Then assume 2 iterations per second ( ), single iteration time : .
[0061] Estimated memory requirements: ; The historical memory usage sequence is [70GB, 69GB, 71GB], and the ARIMA(3,0,0) model is used for prediction.
[0062] Next, the risk assessment layer 4 calculates the risk factor.
[0063] Calculate the memory risk factor: (High risk); Calculate the network risk factor: (Serious warning); Calculate the memory risk factor: (High risk).
[0064] Finally, the overall risk level: (Submission prohibited).
[0065] It should be noted that the present invention can effectively save hardware costs through accurate prediction and reasonable allocation of graphics processor resource requirements; the time required for environmental pre-inspection is shortened, efficiency is greatly improved, and the need for manual intervention is greatly reduced, thereby improving operation and maintenance efficiency.
[0066] The present invention can generate risk reports to provide quantitative gap analysis (such as video memory requirements exceeding 64% of available resources) and executable recommendations (such as reducing the batch size to 64), so that operation and maintenance decisions have a clear basis; it can also support dynamic tracking of optimization effects. For example, after the application of gradient compression technology, the system can provide real-time feedback on the reduction in network load (such as bandwidth requirements reduced from 30Gbps to 15Gbps).
[0067] An embodiment of the present invention also provides a full-link environment pre-check method. Figure 2 Flowchart of the full-link environment pre-check method provided by the embodiment of the present invention. Figure 2 As shown, the method includes: S201: Receive operation parameters.
[0068] S202: Monitor the hardware resource status in the full-link environment, and obtain the available video memory capacity, inter-node bandwidth, and available memory capacity in sequence.
[0069] S203: Couple the job parameters with the hardware resource status to create a model, and predict the graphics memory requirement, network requirement, and memory requirement based on the constructed prediction model.
[0070] S204. Calculate risk factors based on the available video memory capacity, inter-node bandwidth, and available memory capacity, as well as the predicted video memory demand, network demand, and memory demand. Perform risk assessment based on the risk factors and the set hierarchical alarm rules. The risk factor is an indicator that quantifies the degree of resource gap.
[0071] In the above-mentioned full-link environment pre-inspection method provided in the embodiment of the present invention, an end-to-end pre-inspection method can be implemented through dynamic perception, accurate prediction and scientific risk assessment to ensure accurate matching of resources before job deployment, significantly improve the reliability of job deployment and resource utilization efficiency, avoid resource waste caused by static allocation, and analyze potential risks as a whole, saving hardware costs, improving operation and maintenance efficiency, and providing all-round guarantees for efficient and stable operation of jobs.
[0072] Since the embodiments of the full-link environment pre-check method correspond to the embodiments of the full-link environment pre-check system, the description of the features in the embodiments corresponding to the full-link environment pre-check method can be found in the relevant description of the embodiments corresponding to the full-link environment pre-check system, and will not be repeated here. And it has the same beneficial effects as the full-link environment pre-check system mentioned above.
[0073] Furthermore, in the specific implementation, in the above-mentioned full-link environment pre-inspection method provided in the embodiment of the present invention, step S202 obtains the available video memory capacity, the bandwidth between nodes and the available memory capacity in sequence, which may specifically include: obtaining the operating data of the graphics processor, and obtaining the available video memory capacity based on the operating data; measuring the effective bandwidth between nodes; parsing the system files about memory usage to obtain the available memory capacity.
[0074] Among them, obtaining the available video memory capacity according to the operating data may specifically include: obtaining the operating data of the graphics processor, obtaining the total video memory capacity and the allocated video memory capacity from the operating data; and obtaining the available video memory capacity according to the difference between the total video memory capacity and the allocated video memory capacity.
[0075] Measuring the effective bandwidth between nodes may specifically include: establishing a data transmission channel between the nodes; after performing multiple data transmissions on the data transmission channel, obtaining the test start time, test end time and the amount of data transmitted each time; obtaining the total test time based on the test start time and test end time; and obtaining the total amount of data transmitted multiple times based on the amount of data transmitted each time and the number of transmissions; and obtaining the effective bandwidth between the nodes based on the ratio of the total amount of data transmitted multiple times to the total test time.
[0076] Parsing system files about memory usage to obtain available memory capacity may specifically include: parsing system files about memory usage to obtain completely unused physical memory capacity in the system, memory capacity occupied by system cache, and memory capacity occupied by block device buffers; adding completely unused physical memory capacity in the system, memory capacity occupied by system cache, and memory capacity occupied by block device buffers to obtain available memory capacity.
[0077] Furthermore, in specific implementation, in the above-mentioned full-link environment pre-inspection method provided in an embodiment of the present invention, step S203 predicts the video memory requirement, network requirement and memory requirement based on the constructed prediction model, which may specifically include: inputting the batch size, sequence length and model parameter quantity in the job parameters into the video memory requirement prediction model to predict the video memory requirement; obtaining the gradient data volume and the number of nodes according to the job parameters and inputting them into the network requirement prediction model to predict the network requirement; collecting the historical memory quantity at multiple time points and inputting them into the memory requirement prediction model to predict the memory requirement; the memory requirement prediction model is a time series model.
[0078] The batch size, sequence length and model parameter amount in the job parameters are input into the video memory demand prediction model to predict the video memory demand. Specifically, the method may include: fitting the video memory occupancy data of historical jobs by the least squares method to obtain the coefficients of the video memory demand prediction model; extracting the batch size, sequence length and model parameter amount from the job parameters, and combining them with the coefficients of the video memory demand prediction model to calculate the video memory demand.
[0079] According to the operation parameters, the gradient data volume and the number of nodes are obtained and input into the network demand prediction model to predict the network demand, which may specifically include: extracting the model parameter volume and the number of data type bytes from the operation parameters; obtaining the single iteration gradient data size according to the product of the model parameter volume and the number of data type bytes; calculating the amount of data transmitted per second by a single node according to the single iteration gradient data size, the number of iterations per second, and the single iteration time; extracting the number of nodes from the operation parameters, and calculating the network demand according to the product of the amount of data transmitted per second by a single node and the number of nodes.
[0080] Collecting historical memory quantities at multiple time points and inputting them into a memory demand prediction model to predict memory demand may specifically include: determining the order of the memory demand prediction model through the autocorrelation graph and partial autocorrelation graph of historical data, and then fitting the autoregressive coefficients of the memory demand prediction model using the maximum likelihood estimation method; collecting historical memory quantities at multiple time points and inputting the historical memory quantities into the memory demand prediction model, performing weighted summation using the autoregressive coefficients, and adding a random fluctuation term to calculate the memory demand.
[0081] Furthermore, in the specific implementation, in the above-mentioned full-link environment pre-inspection method provided in the embodiment of the present invention, step S204 calculates the risk factor, which may specifically include: calculating the video memory risk factor based on the ratio of video memory demand to available display capacity; calculating the network risk factor based on the ratio of network demand to inter-node bandwidth; calculating the memory risk factor based on the ratio of memory demand to available memory capacity.
[0082] Furthermore, in specific implementation, the above-mentioned full-link environment pre-inspection method provided in an embodiment of the present invention may also include: arranging resource issues in a set order of risk levels and generating optimization suggestions; the optimization suggestions include: when the video memory risk factor is higher than the set video memory risk factor threshold, reducing the batch size or enabling model parallelism; when the network risk factor is higher than the set network risk factor threshold, reducing the number of nodes or enabling gradient compression; when the memory risk factor is higher than the set memory risk factor threshold, adjusting data loading or restricting background processes.
[0083] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0084] An embodiment of the present invention also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned full-link environment pre-check method embodiments.
[0085] An embodiment of the present invention also provides a computer-readable storage medium, which stores a computer program, wherein the computer program is configured to execute the steps of any of the above-mentioned full-link environment pre-inspection method embodiments when running.
[0086] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0087] An embodiment of the present invention also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the steps in any of the above-mentioned full-link environment pre-inspection method embodiments.
[0088] An embodiment of the present invention also provides another computer program product, including a non-volatile computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps of any of the above-mentioned full-link environment pre-inspection method embodiments.
[0089] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.
[0090] The above is a detailed introduction to the full-link environment pre-inspection system, method, equipment, medium and program product provided by the present invention. This article uses specific examples to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present invention, the present invention can also be improved and modified, and these improvements and modifications also fall within the scope of protection of the present invention.
Claims
1. A full-link environment pre-inspection system, characterized in that: include: Interface layer, used to receive job parameters; The environment perception layer monitors the hardware resource status in the full-link environment and obtains available video memory capacity, inter-node bandwidth, and available memory capacity in turn. The resource demand prediction layer is used to couple the job parameters with the hardware resource status to build a model, and predict the video memory demand, network demand and memory demand based on the constructed prediction model; The risk assessment layer is used to calculate the risk factor based on the available video memory capacity, inter-node bandwidth and available memory capacity obtained by the environmental perception layer, and the video memory demand, network demand and memory demand predicted by the resource demand prediction layer, and perform risk assessment based on the risk factor and the set graded alarm rules; the risk factor is an indicator that quantifies the degree of resource gap.
2. The full-link environment pre-inspection system according to claim 1, characterized in that: The environmental perception layer includes: A video memory sensing unit, configured to obtain operation data of a graphics processor and obtain available video memory capacity based on the operation data; Network sensing unit, used to measure the effective bandwidth between nodes; The memory awareness unit is used to parse system files about memory usage and obtain available memory capacity.
3. The full-link environment pre-inspection system according to claim 2, characterized in that: The video memory sensing unit includes: a capacity acquisition subunit, configured to acquire operation data of the graphics processor, and acquire the total video memory capacity and the allocated video memory capacity from the operation data; The first calculation subunit is configured to obtain the available video memory capacity according to the difference between the total video memory capacity and the allocated video memory capacity.
4. The full-link environment pre-inspection system according to claim 2, characterized in that: The network sensing unit includes: A channel establishment subunit is used to establish a data transmission channel between nodes; A data acquisition subunit is used to obtain the test start time, test end time and the amount of data transmitted each time after multiple data transmissions on the data transmission channel; The second calculation subunit is configured to obtain a total test time based on the test start time and the test end time; and obtain a total amount of data transmitted multiple times based on the amount of data transmitted each time and the number of transmissions; The third calculation subunit is configured to obtain the effective bandwidth between the nodes according to the ratio of the total amount of data transmitted multiple times to the total test time.
5. The full-link environment pre-inspection system according to claim 2, characterized in that: The memory awareness unit includes: The file parsing subunit is used to parse the system file about memory usage to obtain the completely unused physical memory capacity in the system, the memory capacity occupied by the system cache, and the memory capacity occupied by the block device buffer; The fourth calculation subunit is used to add the completely unused physical memory capacity in the system, the memory capacity occupied by the system cache, and the memory capacity occupied by the block device buffer to obtain the available memory capacity.
6. The full-link environment pre-inspection system according to claim 1, characterized in that: The resource demand prediction layer includes: A video memory demand prediction unit, configured to input the batch size, sequence length, and model parameter quantity in the job parameters into a video memory demand prediction model to predict video memory demand; A network demand prediction unit, configured to obtain the amount of gradient data and the number of nodes according to the operation parameters and input the obtained data into a network demand prediction model to predict the network demand; The memory demand prediction unit is used to collect historical memory quantities at multiple time points and input them into a memory demand prediction model to predict memory demand; the memory demand prediction model is a time series model.
7. The full-link environment pre-inspection system according to claim 6, characterized in that: The video memory demand prediction unit includes: A coefficient value fitting subunit is used to fit the video memory usage data of historical jobs by the least square method to obtain the coefficients of the video memory demand prediction model; The first model inference subunit is used to extract the batch size, sequence length and model parameter quantity from the job parameters, and calculate the video memory requirement in combination with the coefficients of the video memory requirement prediction model.
8. The full-link environment pre-inspection system according to claim 6, characterized in that: The network demand prediction unit includes: The gradient data calculation subunit is used to extract the model parameter quantity and the data type byte number from the operation parameters; and obtain the single iteration gradient data size according to the product of the model parameter quantity and the data type byte number; The single-node transmission calculation subunit is used to calculate the amount of data transmitted per second by a single node based on the size of the gradient data in a single iteration, the number of iterations per second, and the single iteration time; The second model reasoning subunit is used to extract the number of nodes from the operation parameters, and calculate the network demand according to the product of the amount of data transmitted per second by the single node and the number of nodes.
9. The full-link environment pre-inspection system according to claim 6, characterized in that: The memory demand prediction unit includes: a parameter determination subunit, configured to determine the order of the memory demand prediction model through an autocorrelation diagram and a partial autocorrelation diagram of historical data, and then fit the autoregressive coefficient of the memory demand prediction model using a maximum likelihood estimation method; The third model inference sub-unit is used to collect the historical memory quantities at multiple time points, and input the historical memory quantities into the memory demand prediction model, use the autoregressive coefficients for weighted summation, and add the random fluctuation term to calculate the memory demand.
10. The full-link environment pre-inspection system according to claim 1, characterized in that: The risk assessment layer includes: A video memory risk factor calculation unit, configured to calculate a video memory risk factor based on a ratio of video memory demand to available display capacity; A network risk factor calculation unit, configured to calculate a network risk factor based on a ratio of network demand to inter-node bandwidth; The memory risk factor calculation unit is used to calculate the memory risk factor according to the ratio of the memory demand to the available memory capacity.
11. The full-link environment pre-inspection system according to claim 10, characterized in that: Also includes: The decision engine layer is used to sort resource issues according to the set order of risk level and generate optimization suggestions; The optimization suggestion includes: when the memory risk factor is higher than a set memory risk factor threshold, reducing the batch size or enabling model parallelism; When the network risk factor is higher than the set network risk factor threshold, the number of nodes is reduced or gradient compression is enabled; when the memory risk factor is higher than the set memory risk factor threshold, data loading is adjusted or background processes are restricted.
12. A full-link environment pre-check method, characterized in that: include: Receive job parameters; Monitor the hardware resource status in the full-link environment and obtain available video memory capacity, inter-node bandwidth, and available memory capacity in sequence; Coupling the job parameters with the hardware resource status to model the performance, and predicting the video memory requirement, network requirement, and memory requirement based on the constructed prediction model; Calculate risk factors based on the available video memory capacity, inter-node bandwidth, and available memory capacity, as well as the predicted video memory demand, network demand, and memory demand, and perform risk assessment based on the risk factors and the set graded alarm rules; The risk factor is an indicator that quantifies the extent of the resource gap.
13. An electronic device, characterized in that: include: Memory for storing computer programs; A processor is used to implement the steps of the full-link environment pre-check method as claimed in claim 12 when executing the computer program.
14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by the processor, the steps of the full-link environment pre-check method as claimed in claim 12 are implemented.
15. A computer program product comprising a computer program, characterized in that When the computer program is executed by the processor, the steps of the full-link environment pre-check method as claimed in claim 12 are implemented.
Citation Information
Cited By
Efficient structure disease target detection method based on video memory perception
CN121258924A