Calculation power model data analysis and acquisition system based on artificial intelligence
Through the data analysis and acquisition system of computing power model based on artificial intelligence, the acquisition frequency and priority of hardware utilization data is dynamically adjusted, and the resource waste and delay detection problems of traditional acquisition systems are solved, and efficient and accurate hardware monitoring and resource management are achieved.
Patent Information
- Application Number
- CN202510952683.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-09-05
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional acquisition systems have problems such as high-frequency fluctuation data leakage, difficult to detect delay performance abnormalities, increased risk of hardware damage and waste of resources in hardware utilization data collection, and cannot promptly increase the acquisition frequency to position.
The data analysis and acquisition system of computing power model based on artificial intelligence is adopted, and the fluctuations of hardware usage data are analyzed through the digital analysis module to generate usage adjustment signals. The adjustment module dynamically adjusts the acquisition frequency and priority based on historical data and performance indicators to optimize resource allocation.
Accurately judge the fluctuation status of hardware usage, shorten the delay in performance abnormality detection, reduce resource consumption, improve the efficiency and accuracy of the acquisition system, reduce false alarm rates, and ensure timely monitoring of key hardware and reasonable allocation of resources.
Smart Images

Figure CN120596849A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to artificial intelligence technology, specifically to an artificial intelligence-based computing power model data analysis and acquisition system. Background Art
[0002] Traditional data collection systems typically use a fixed frequency to collect CPU, GPU, and other hardware usage data. This can lead to missed data during high-frequency fluctuations during the usage phase, making latency anomalies difficult to detect. During the idle phase, hardware usage remains stable for long periods while data is collected at a high frequency, resulting in a waste of storage resources and network bandwidth.
[0003] Hardware utilization data is often prioritized based on hardware type. When data is collected based on this fixed priority, when unusual fluctuations occur during the execution of different tasks, the collection frequency of the corresponding hardware cannot be increased in time to locate the problem, resulting in delayed fault detection and increased risk of hardware damage.
[0004] In response to the above technical problems, this application proposes a solution. Summary of the Invention
[0005] The purpose of the present invention is to solve the problems raised by the background technology and to propose an artificial intelligence-based computing power model data analysis and collection system.
[0006] The purpose of the present invention can be achieved through the following technical solutions:
[0007] AI-based computing model data analysis and acquisition system, including data acquisition module, data analysis module and adjustment module;
[0008] The data analysis module analyzes the hardware utilization data transmitted by the data acquisition module, determines the fluctuation of the corresponding utilization data, and generates a corresponding utilization adjustment signal based on the fluctuation of the utilization data. The utilization adjustment signal is then transmitted to the adjustment module. The module analyzes the adjustment amount of the five data items contained in the hardware utilization data, and then analyzes the priority of the corresponding data items to arrange the adjustment order.
[0009] The adjustment module obtains the two usage rate data with the highest number of occurrences in the use phase and the idle phase in the historical data based on the adjustment signal and priority generated by the analysis module, and determines the upper and lower limits of the fluctuation range and the normal fluctuation value; after receiving the adjustment signal at different stages, it adjusts the hardware utilization rate to the upper and lower limits of the fluctuation range of the corresponding stage; calculates the difference between the current utilization rate and the target value, multiplies it by the weight coefficient to obtain the single adjustment amount, and completes the hardware utilization rate adjustment in order of priority.
[0010] As a preferred embodiment of the present invention, the steps of data preprocessing by the data analysis module are as follows:
[0011] S1: Calculate the mean A1 and standard deviation B for Z1 data of the same type collected at the same time, set the corresponding data fluctuation range [A1-a1B, A1+a1B] based on the calculated mean A1 and standard deviation B, mark the corresponding data that is not within the fluctuation range at that time point as an outlier, count the number of outliers Y1, and if Y1>k1*Z1, analyze the judgment range of the outliers, where k1 is the preset proportional coefficient;
[0012] S2: Get the time point corresponding to the same time, and calculate the proportion of the same type of data fluctuation range within the time point a b The value of b is the serial number of the corresponding time point; calculate the proportion data a b The mean A2 of the mean is compared with a1. If The fluctuation range setting is determined to be normal, and k2 is the preset proportional coefficient; if Y1>k1*Z1, the acquisition time is determined to be an abnormal time; if a1 is not within the range of A2*[1±k2], the fluctuation range setting is determined to be inaccurate. If Y1>k1*Z1, the proportional data a1 is increased by one, and then the abnormal value is re-judged. If Y1>k1*Z1 still exists, the acquisition time is determined to be an abnormal time; if Y1≤k1*Z1, after eliminating the abnormal value, the mean A3 of the remaining data of the same type is calculated, and the mean A3 is used as the detection data at this time point;
[0013] S3: Mark the time point of the abnormal moment. If the detection time of the corresponding item data after the corresponding time point is all abnormal time, it is determined that the acquisition system has an abnormality, and an acquisition maintenance signal is generated and transmitted to the adjustment module; otherwise, it is determined that the detection data is abnormal.
[0014] As a preferred embodiment of the present invention, the standard processing steps of the hardware usage data by the analysis module are as follows:
[0015] K1: Obtains usage data for each item within a set time period from the current time point, then arranges the usage data in order of collection time and calculates the data difference C1 between adjacent sorted numbers. The data difference C1 represents the change in the usage data for the corresponding item. The data difference C1 is averaged by removing extreme values to obtain the average value A4. This average value A4 is used as the preset usage change threshold for the corresponding item.
[0016] K2: Determine that the change amount of the usage rate data of the corresponding item is a normal fluctuation if it is less than the preset usage rate change threshold of the corresponding item, and count the number of abnormal fluctuations of the corresponding item; sum up the number of abnormal fluctuations of the usage rate data of each item, then the weight coefficient of the usage rate data of the corresponding item is equal to the number of abnormal fluctuations of the usage rate data of the corresponding item divided by the sum of the number of abnormal fluctuations of the usage rate data of each item;
[0017] K3: The hardware usage rate data is equal to the sum of the usage rate data of the corresponding item multiplied by the weight coefficient of the usage rate data of the corresponding item.
[0018] As a preferred embodiment of the present invention, the steps for the data analysis module to generate an adjustment signal are as follows:
[0019] Q1: Retrieve the hardware usage rate data within a set time period from the current time point, and calculate the difference C2 between the hardware usage rate data corresponding to adjacent acquisition time points in the order of acquisition time, and count the number CS of times when the difference C2 within the set time period is greater than the normal fluctuation value of the hardware usage rate data;
[0020] Q2: If the total number Z2*k2 of the hardware usage rate data within the set time period < CS, it is determined that the usage rate fluctuates frequently, and the acquisition frequency should be increased to more accurately capture the change situation, generate a high usage rate adjustment signal, and transmit the high usage rate adjustment signal to the adjustment module; otherwise, it is determined that the usage rate fluctuates stably, and the acquisition frequency should be appropriately reduced to reduce unnecessary data acquisition overhead, generate a low usage rate adjustment signal, and transmit the low usage rate adjustment signal to the adjustment module. [[ID=?]]
[0021] As a preferred embodiment of the present invention, the steps for the data analysis module to divide the importance of the adjustment sequence data are as follows:
[0022] P1: Analyze the historical data of the usage rate data of each item, establish the fluctuation range of the historical usage rate data of the corresponding item according to the mean and standard deviation of the historical data of each item, compare the current detected usage rate data of the corresponding item with the fluctuation range of the historical usage rate data of the corresponding item. If the current detected usage rate data of the corresponding item is within the fluctuation range of the historical usage rate data of the corresponding item, it is determined that the priority of the usage rate data of this corresponding item remains unchanged; if the current detected usage rate data of the corresponding item is not within the fluctuation range of the historical usage rate data of the corresponding item, it is determined that the priority of the usage rate data of this corresponding item is adjusted;
[0023] It should be noted that there seems to be a mistake in the original text where the "?" is marked in the ID number in the translation. The original ID number should be "14" and it is translated as "
[0021] ". This is just for your reference in case there is an error in the original text. If you have any other questions, please feel free to ask.P2: Calculate the difference between the upper and lower limits of the fluctuation range of the historical usage data of the corresponding item requiring priority adjustment. Divide the excess value of the corresponding item's current detected usage data beyond the fluctuation range by the calculated difference between the upper and lower limits of the fluctuation range to obtain the proportion of the excess value to the fluctuation range. Compare the calculated proportion data with the preset proportion data of 50%;
[0024] P3: If the calculated ratio data BL js Less than the preset ratio data BL ys , then the priority adjustment amount ΔYX1=YX dq *(1+BL js ), YX dq Priority of usage data for current corresponding item detection; if the ratio data BL js Greater than the preset ratio data BL ys , then the priority adjustment amount ΔYX1=YX dq *[1+(BL js -BL ys )]+1; then recalculate the priority of each usage data, and the new priority after recalculation is YX X1 =YX dq +ΔYX1, rearrange the order according to priority.
[0025] As a preferred embodiment of the present invention, the steps for dividing the impact of the computing power model performance indicators of the adjustment order by the analysis module are as follows:
[0026] G1: Retrieve the historical usage data for the corresponding items that require priority adjustment, mark the usage data that is outside the fluctuation range of the corresponding historical usage data as outliers, remove the outliers, and then map the remaining historical usage data to the performance indicator XN of the computing power model according to the collection time, marking the time when the historical usage data is missing;
[0027] G2: If the number of consecutive missing time markers is greater than the preset threshold, or the total number of missing time markers is greater than the preset threshold, reselect the target time period until the number of consecutive missing time markers in the selected target time period is less than the preset threshold, and the total number of missing time markers is less than the preset threshold. In the time period where the number of consecutive missing time markers is less than the preset threshold, fill in the missing values.
[0028] G3: Standardize the performance indicator XN of the computing power model and calculate the utilization characteristic SY of each hardware i Pearson correlation coefficient with performance index XN, Pearson correlation coefficient n is the number of samples, SY i,t is the specific value of the i-th hardware usage feature at the t-th sampling time;
[0029] G4: The greater the impact of the corresponding item usage rate data on the performance indicators of the computing power model, the closer the Pearson correlation coefficient of the corresponding item usage rate data is to 1; the preset correlation coefficient threshold r max , if the Pearson correlation coefficient r of the corresponding item usage rate data i ≥r max , then the priority adjustment amount ΔYX2=YX dq *[1+(r i -r max )]+1; otherwise, the priority adjustment amount ΔYX2=YX dq *(1+BL js ); Then recalculate the priority of each usage data, and the new priority YX X2 =YX dq +ΔYX2.
[0030] As a preferred embodiment of the present invention, the steps for determining the priority of the corresponding hardware usage data by the analysis module are as follows:
[0031] L1: Comprehensively consider the impact of data status importance and computing power model performance indicators, and adjust the priority of the corresponding item usage data
[0032] L2: If there are corresponding item usage data of the same priority size, the priority adjustment amount of the corresponding item usage data of the same priority is calculated according to the weight ratios ω1 and ω2 pre-set by the user, and the priority adjustment amount ΔYX = ω1*ΔYX1+ω2*ΔYX2.
[0033] Compared with the prior art, the present invention has the following beneficial effects:
[0034] 1. The analysis module counts the number of times the difference in hardware usage data exceeds the normal fluctuation value within a set time period and compares it with the preset threshold to accurately determine the usage fluctuation status. During the use phase, the acquisition frequency of key hardware such as GPUs is increased, significantly shortening the delay in performance anomaly detection and facilitating timely discovery and resolution of problems. During the idle phase, the reduced acquisition frequency reduces acquisition resource consumption by more than 50%, effectively saving resources. The adjustment module dynamically sets the acquisition target value based on the fluctuation range of hardware during the use and idle phases in historical data, so that the acquisition frequency and accuracy match the corresponding needs at different stages.
[0035] 2. The data analysis module comprehensively prioritizes data based on its importance and its impact on the performance of the computing model. On the one hand, the historical fluctuation range of each utilization data item is analyzed, and the current detection data is compared with the historical range to calculate the excess ratio. On the other hand, the Pearson correlation coefficient between the hardware utilization characteristics and the computing model performance indicators is calculated to quantify the impact of hardware on performance. Based on the results of these two aspects, the priority adjustment amount of each data item is dynamically adjusted. For data that is closely related to performance and exceeds the historical fluctuation range by a large margin, its priority is increased to ensure high-frequency collection. In addition, when the hardware priority is the same, the user is allowed to customize the collection order by pre-setting weights, so that the user can increase the data collection priority of key resources to ensure that key data is collected first. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] To facilitate understanding by those skilled in the art, the present invention is further described below with reference to the accompanying drawings.
[0037] Figure 1 It is a system flow chart of the present invention. DETAILED DESCRIPTION
[0038] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0039] Example:
[0040] See also Figure 1 As shown, the AI-based computing model data analysis and acquisition system includes a data acquisition module, a data analysis module, and an adjustment module;
[0041] The data acquisition module collects hardware usage data and passes the collected data to the data analysis module;
[0042] The data analysis module analyzes the hardware utilization data transmitted by the data acquisition module, determines the fluctuation of the corresponding utilization data, and generates a corresponding utilization adjustment signal based on the fluctuation of the utilization data. The utilization adjustment signal is then transmitted to the adjustment module. The module analyzes the adjustment amount of the five data items contained in the hardware utilization data, and then analyzes the priority of the corresponding data items to arrange the adjustment order.
[0043] Sort the collected data by collection time, calculate the mean A1 and standard deviation B for Z1 data of the same type collected at the same time, set the corresponding data fluctuation range [A1-a1B, A1+a1B] based on the calculated mean A1 and standard deviation B, mark the corresponding data that is not within the fluctuation range at that time point as an outlier, count the number of outliers Y1, and if Y1>k1*Z1, analyze the judgment range of the outliers, where k1 is the preset proportional coefficient;
[0044] By calculating the mean A1 and standard deviation B of the same type of data at the same time point, we construct the fluctuation range [A1-a1B, A1+z1B]. This process is essentially based on the characteristics of the normal distribution. In a normal distribution, approximately 68% of the data lies within the range of ±1 standard deviation of the mean, approximately 95% lies within the range of ±2 standard deviations, and approximately 99.7% lies within the range of ±3 standard deviations. Therefore, when a1 is set to 3, this range covers the vast majority of normal data. Data outside this range is marked as outliers. This setting effectively filters out noise data caused by factors such as transient hardware failures and network transmission interference.
[0045] Get the time point corresponding to the same time, and the proportion data a in the fluctuation range of the same type of data at this time point b The value of b is the serial number of the corresponding time point; calculate the proportion data a b The mean A2 of the mean is compared with a1. If The fluctuation range setting is determined to be normal, and k2 is the preset proportional coefficient; if Y1>k1*Z1, the acquisition time is determined to be an abnormal time; if a1 is not within the range of A2*[1±k2], the fluctuation range setting is determined to be inaccurate. If Y1>k1*Z1, the proportional data a1 is increased by one, and then the abnormal value is re-judged. If Y1>k1*Z1 still exists, the acquisition time is determined to be an abnormal time; if Y1≤k1*Z1, after eliminating the abnormal value, the mean A3 of the remaining data of the same type is calculated, and the mean A3 is used as the detection data at this time point;
[0046] The outlier threshold k1*Z1 reflects the design of the fault-tolerance mechanism. For example, if k1=0.1, when the number of data of the same type Z1=100, if the number of outliers Y1>10, it means that the current fluctuation range setting may not conform to the actual data distribution, and the judgment range needs to be re-analyzed. This design avoids misjudgments caused by single outliers and adapts to scenarios with different data scales through the proportional coefficient k1. In actual applications, if the CPU usage data of a server cluster has 20 outliers at a certain point in time (Z1=200, k1=0.1), the system will automatically adjust the proportional coefficient a1 (for example, from 3 to 3.5), expand the fluctuation range, and re-judge until the number of outliers returns to a reasonable range.
[0047] The marking of abnormal moments and the logic for determining acquisition system anomalies utilize time series continuity analysis. Once a time point is marked as an abnormal moment, the system continuously monitors the status of subsequent detection moments. If multiple consecutive moments (e.g., three or more) are abnormal, it is determined to be an acquisition system anomaly (e.g., sensor hardware failure or data transmission link interruption), and a maintenance signal is generated. If only a single point of abnormality occurs, it is determined to be a test data anomaly (e.g., single-sample distortion caused by instantaneous current fluctuations). This hierarchical determination mechanism significantly reduces the false alarm rate. For example, in practice at one data center, this mechanism reduced the false alarm rate of acquisition system anomalies from 15% to 3%, while also shortening the time to locate the actual fault to within 10 minutes.
[0048] Mark the time point of the abnormal moment. If the detection time of the corresponding item data after the corresponding time point is all abnormal time, it is determined that the acquisition system has an abnormality, and an acquisition maintenance signal is generated and passed to the adjustment module; otherwise, it is determined that the detection data is abnormal.
[0049] The hardware usage data is collected, and the hardware usage data includes CPU usage, GPU usage, memory usage, hard disk usage, and network bandwidth usage data; each usage data item within a set time period from the current time point is obtained, and then the usage data items are arranged in order according to the collection time, and the data difference C1 between adjacent sorting numbers is calculated, where the data difference C1 represents the change in the usage data of the corresponding item; the data difference C1 of the corresponding item is removed from extreme values and averaged to obtain an average value A4, and the average value A4 of the usage data of the corresponding item is used as the preset usage change threshold value of the corresponding item; the usage data change data of the corresponding item whose change in the usage data of the corresponding item is less than the preset usage change threshold value of the corresponding item is determined to be a normal fluctuation, and the number of abnormal fluctuations of the corresponding item is counted; the number of abnormal fluctuations of each usage data item is summed, and the weight coefficient of the usage data of the corresponding item is equal to the number of abnormal fluctuations of the usage data of the corresponding item divided by the sum of the number of abnormal fluctuations of each usage data item; the hardware usage data is equal to the sum of the usage data of the corresponding item multiplied by the weight coefficient of the usage data of the corresponding item;
[0050] By calculating the difference C1 between adjacent data and removing extreme values (such as differences outside 3 times the standard deviation), and then calculating the mean to obtain the preset change amount threshold A4, this step is similar to low-pass filtering and can eliminate the influence of burst pulse noise; for example, when the hard disk usage data suddenly rises from 20% to 100% and then quickly drops back to 25% at adjacent times, this difference will be regarded as an extreme value and removed to avoid the threshold being wrongly raised; the calculation of the weight coefficient is based on the principle of "fluctuation significance" - the more abnormal fluctuations there are, the more unstable the hardware state is, and the higher the importance for system monitoring; specifically, the number of abnormal fluctuations refers to the number of times the usage change amount exceeds the preset threshold A4; assuming the number of abnormal fluctuations of the CPU, GPU, and memory are 15, 8, and 3 times respectively, then their weight coefficients are 15 / (15 + 8 + 3) = 0.5, 8 / 26 ≈ 0.31, 3 / 26 ≈ 0.12; this weight distribution makes key hardware (such as the frequently fluctuating CPU) occupy a larger proportion in the overall hardware usage data and more accurately reflects the core characteristics of the system load;
[0051] The weighted summation formula of the hardware usage data has a clear physical meaning; taking the model training scenario as an example, when the GPU usage is 80% (weight 0.4), the CPU usage is 60% (weight 0.3), and the memory usage is 40% (weight 0.2), the overall hardware usage is 80%×0.4 + 60%×0.3 + 40%×0.2 = 68%, and this value can better reflect the leading role of the GPU in the training task than the simple arithmetic mean (66.7%), providing a more accurate basis for resource scheduling;
[0052] Retrieve the hardware usage data within a set time period from the current time point, and calculate the difference C2 between the corresponding hardware usage data at adjacent acquisition time points in the order of acquisition time. Count the number CS of times when the difference C2 within the set time period is greater than the normal fluctuation value of the hardware usage data. If the total number Z2*k2 of the hardware usage data within the set time period < CS, it is determined that the usage fluctuation is frequent, and the acquisition frequency should be increased to more accurately capture the changes, generate a high usage rate adjustment signal, and transmit the high usage rate adjustment signal to the adjustment module; otherwise, it is determined that the usage fluctuation is stable, and the acquisition frequency should be appropriately reduced to reduce unnecessary data acquisition overhead, generate a low usage rate adjustment signal, and transmit the low usage rate adjustment signal to the adjustment module;
[0053] The essence of the C2 statistic within a set time period is to quantify the intensity of dynamic changes in hardware utilization. For example, within one hour (set time period), the number of times the adjacent difference in a GPU's utilization exceeds the normal fluctuation value is CS = 20 times. If the total number of data in this time period is Z2 = 60 and the preset proportional coefficient k2 = 0.3, then Z2*k2 = 18 < 20, which is considered to be frequent fluctuation. The collection frequency needs to be increased from 30 seconds / time to 10 seconds / time.
[0054] The core goal of adjusting the acquisition frequency is to optimize resource allocation in different scenarios. During the usage phase, high-frequency acquisition (e.g., 10 seconds / time) can capture hardware load peaks in real time, avoiding performance degradation due to untimely heat dissipation. During the idle phase (e.g., during the nighttime task trough), low-frequency acquisition (e.g., 60 seconds / time) can reduce storage and network overhead by more than 50%. After a cloud computing platform applied this mechanism, the recognition rate of abnormal fluctuations in core hardware increased from 80% to 98%, while the bandwidth consumption of the overall acquisition system was reduced by 35%.
[0055] The normal fluctuation value is set based on the statistical characteristics of historical data. Specifically, the system records the two most frequently occurring utilization data points during the active and idle phases (for example, during the active phase, GPU utilization is often 70% and 85%, with a difference of 15% representing the normal fluctuation value). This value reflects the typical fluctuation range of the hardware in different business phases. If the actual difference exceeds this value, it indicates that the hardware status deviates from the normal state and requires increased data collection frequency for closer monitoring.
[0056] The change in hardware usage data in the historical data is obtained, and the two hardware usage data that appear most frequently in the use phase and the idle phase in the historical data are recorded. The difference between the two hardware usage data that appear most frequently in the corresponding phase is used as the normal fluctuation value of the hardware usage data in the corresponding phase, and the two hardware usage data that appear most frequently in the corresponding phase are used as the upper limit and lower limit of the fluctuation range of the hardware usage data in the corresponding phase respectively;
[0057] In the idle stage, when the adjustment module receives a low-usage adjustment signal, the hardware usage data is adjusted to a value corresponding to the lower limit of the fluctuation range of the hardware usage data in the idle stage; in the idle stage, when the adjustment module receives a high-usage adjustment signal, the hardware usage data is adjusted to a value corresponding to the upper limit of the fluctuation range of the hardware usage data in the idle stage; in the use stage, when the adjustment module receives a low-usage adjustment signal, the hardware usage data is adjusted to a value corresponding to the lower limit of the fluctuation range of the hardware usage data in the use stage; in the use stage, when the adjustment module receives a high-usage adjustment signal, the hardware usage data is adjusted to a value corresponding to the upper limit of the fluctuation range of the hardware usage data in the use stage;
[0058] The hardware utilization data includes five items of utilization data. When adjusting the hardware utilization data, the adjustment target value is obtained. The adjustment target value is the value reached after adjusting the hardware utilization data in the corresponding stage; the difference between the current hardware utilization data and the adjustment target value is multiplied by the weight coefficient of the corresponding item utilization data to obtain the adjustment amount of the corresponding item utilization data.
[0059] When adjusting the corresponding item usage data in the hardware usage data, the priority of the corresponding item usage data is divided. The priorities are divided from large to small into: CPU usage, GPU usage, memory usage, hard disk usage and network bandwidth usage data. The difference between adjacent priorities is 1. The order of adjustment of the corresponding item usage data is based on the divided priorities.
[0060] Analyze the historical data of each item of usage rate data, establish the fluctuation range of the historical usage rate data of the corresponding item based on the mean and standard deviation of each item of historical data, compare the current detection usage rate data of the corresponding item with the fluctuation range of the historical usage rate data of the corresponding item, and if the current detection usage rate data of the corresponding item is within the fluctuation range of the historical usage rate data of the corresponding item, determine that the usage rate data priority of the corresponding item remains unchanged; if the current detection usage rate data of the corresponding item is not within the fluctuation range of the historical usage rate data of the corresponding item, determine that the usage rate data priority of the corresponding item is adjusted;
[0061] Calculate the difference between the upper and lower limits of the fluctuation range of the historical usage data of the corresponding item that needs priority adjustment, and divide the excess value of the corresponding item's current detection usage data exceeding the fluctuation range by the difference between the upper and lower limits of the fluctuation range to obtain the proportion of the excess value to the fluctuation range. Compare the calculated proportion data with the preset proportion data of 50%. If the calculated proportion data BL js Less than the preset ratio data BL ys , then the priority adjustment amount ΔYX1=YX dq *(1+BL js ), YX dq Priority of usage data for current corresponding item detection; if the ratio data BL js Greater than the preset ratio data BL ys , then the priority adjustment amount ΔYX1=YX dq *[1+(BL js -BL ys )]+1; then recalculate the priority of each usage data, and the new priority after recalculation is YX X1 =YX dq +ΔYX1, rearrange the order according to the priority;
[0062] Dynamic priority division combines the dual dimensions of historical state deviation and real-time performance impact. js The calculation of the value (such as the difference between the upper and lower limits of the fluctuation range) measures the abnormality of the current state; for example, if the historical fluctuation range of a hardware is 20%-80% (difference 60%), the current detection value is 90%, and the excess value is 10%, then BL js =10% / 60%≈16.7%, if the preset ratio BL js =50%, indicating a low abnormality level and a small priority adjustment amount (ΔYX1 = current priority × (1 + 16.7%));
[0063] The historical usage data of the corresponding items that need priority adjustment are retrieved, and the usage data of the corresponding items that are not within the fluctuation range of the corresponding historical usage data are marked as outliers, and the outliers are eliminated. Then, the remaining historical usage data are matched one-to-one with the performance indicator XN of the computing power model according to the collection time, and the time when the historical usage data is missing is marked; if the number of consecutive missing time tags is greater than the preset number threshold, or the total number of missing time tags is greater than the preset number threshold, the target time period is reselected until the number of consecutive missing time tags in the selected target time period is less than the preset number threshold, and the total number of missing time tags is less than the preset number threshold; in the time period when the number of consecutive missing time tags is less than the preset number threshold, if the missing time tag is c n , then the time stamp is c n-1 、c n-2 、c n+1 、c n+2 The usage rate data of the corresponding items is calculated, and the change in usage rate data of the corresponding items in adjacent time intervals is ΔSY n-1 , ΔSY n+1 Calculation, determination of ΔSY n+1 -ΔSY n-1 =3d, d is the interval value of the change, then c n =c n-1 +d=c n+1 -d, fill in missing values;
[0064] The setting of a consecutive missing time threshold ensures the effectiveness of data repair. If the number of consecutive missing time markers exceeds the preset threshold (e.g., 5), it indicates that the data reliability of that time period is too low, and the system will reselect the target time period to avoid analysis based on erroneous data. In the historical data repair practice of a scientific research institution, this mechanism has increased data integrity from 75% to 98%, providing reliable support for long-term performance trend analysis.
[0065] The collaborative work of missing value filling and outlier removal further improves data quality. The system first removes outliers that fall outside the historical fluctuation range, then performs missing value filling on the remaining data to prevent outliers from contaminating the interpolation results. For example, if a 100% outlier in memory usage data is caused by a hardware failure, removing it and then performing linear interpolation on adjacent missing values ensures that the completed data conforms to the actual fluctuation pattern.
[0066] Standardize the performance index XN of the computing power model and calculate the utilization characteristic SY of each hardware i Pearson correlation coefficient with performance index XN, Pearson correlation coefficient n is the number of samples, SY i,t is the specific value of the i-th hardware utilization feature at the t-th sample time; the greater the impact of the corresponding item utilization data on the performance index of the computing power model, the closer the Pearson correlation coefficient of the corresponding item utilization data is to 1; the preset correlation coefficient threshold r max , if the Pearson correlation coefficient r of the corresponding item usage rate data i ≥r max , then the priority adjustment amount ΔYX2=YX dq *[1+(r i -r max )]+1; otherwise, the priority adjustment amount ΔYX2=YX dq *(1+BL js ); Then recalculate the priority of each usage data, and the new priority YX X2 =YX dq +ΔYX2;
[0067] The computing power model performance index XN is a set of key parameters used to quantitatively evaluate the operating status, efficiency, and reliability of the computing power model. Its definition has dynamic adaptability to scenarios and needs to be flexibly adjusted according to the model application scenario and business needs. This application uses efficiency indicators to focus on the speed at which the computing power model completes tasks and the efficiency of resource utilization. The computing power model performance index XN is obtained by weighted calculation of the time efficiency index, throughput index, and resource efficiency index. The index XN is standardized by Z-Score. Where μ is the mean and σ is the standard deviation), eliminating the dimension effect and ensuring comparability with hardware utilization data;
[0068] Pearson correlation coefficient r i The introduction of quantifies the correlation between hardware utilization and computing model performance. The coefficient ranges from [-1, 1]. The closer the absolute value is to 1, the stronger the correlation is. For example, the r of GPU utilization and model inference time is i=-0.9, indicating that the two are strongly negatively correlated (the more fully loaded the GPU is, the shorter the inference time is), and the hardware has a significant impact on performance; when r i ≥r max (such as 0.7), the priority adjustment formula is ΔYX2=current priority×[1+(r i -r max )]+1 will further increase its priority to ensure high-frequency collection;
[0069] Based on the importance of data status and the impact of computing power model performance indicators, the priority of the corresponding item usage data is adjusted If there are corresponding item usage data with the same priority, the priority adjustment amount is calculated for the corresponding item usage data with the same priority according to the weight ratio ω1 and ω2 preset by the user, and the priority adjustment amount ΔYX = ω1*ΔYX1 + ω2*ΔYX2;
[0070] Comprehensive priority adjustment The design avoids the limitation of a single dimension; for example, the historical fluctuation of a certain hardware is low (BL js = 20%), but the correlation coefficient with performance is high (r i =0.8), it will still receive a higher priority after comprehensive adjustment, ensuring continuous monitoring of performance-critical hardware; in addition, user-defined weights ω1 and ω2 allow flexible adjustment in specific scenarios (such as network bandwidth priority in distributed computing), improving system adaptability.
[0071] The preferred embodiments of the present invention disclosed above are intended only to help illustrate the present invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the present invention to specific embodiments. Obviously, many modifications and variations are possible based on the contents of this specification. These embodiments are selected and described in detail in this specification to better explain the principles and practical applications of the present invention, thereby enabling those skilled in the art to better understand and utilize the present invention. The present invention is limited only by the claims and their full scope and equivalents.
Claims
1. The AI-based computing power model data analysis and acquisition system is characterized by: Including data acquisition module, data analysis module and adjustment module; The data analysis module analyzes the hardware utilization data transmitted by the data acquisition module, determines the fluctuation of the corresponding item utilization data, generates the corresponding utilization adjustment signal according to the fluctuation of the utilization data, and then transmits the utilization adjustment signal to the adjustment module; Analyze the adjustment amounts of the five data items included in the hardware usage data, and then analyze the priorities of the corresponding data items to arrange the adjustment order; The adjustment module obtains the two most frequently occurring usage data in the use phase and idle phase from the historical data based on the adjustment signal and priority generated by the analysis module, and determines the upper and lower limits of the fluctuation range and the normal fluctuation value; After receiving the adjustment signal at different stages, the hardware utilization rate is adjusted to the upper and lower limits of the fluctuation range of the corresponding stage; the difference between the current utilization rate and the target value is calculated and multiplied by the weight coefficient to obtain the single adjustment amount, and the hardware utilization rate adjustment is completed in order of priority.
2. The artificial intelligence-based computing power model data analysis and acquisition system according to claim 1 is characterized in that: The steps for data preprocessing in the data analysis module are as follows: S1: Calculate the mean A1 and standard deviation B for Z1 data of the same type collected at the same time, set the corresponding data fluctuation range [A1-a1B, A1+a1B] based on the calculated mean A1 and standard deviation B, mark the corresponding data that is not within the fluctuation range at that time point as an outlier, count the number of outliers Y1, and if Y1>k1*Z1, analyze the judgment range of the outliers, where k1 is the preset proportional coefficient; S2: Get the time point corresponding to the same time, and calculate the proportion of the same type of data fluctuation range within the time point a b The value of , b is the serial number of the corresponding time point; Calculate the ratio data a b The mean A2 of the mean is compared with a1. If The fluctuation range setting is determined to be normal, and k2 is the preset proportional coefficient; if Y1>k1*Z1, the acquisition time is determined to be an abnormal time; if a1 is not within the range of A2*[1±k2], the fluctuation range setting is determined to be inaccurate. If Y1>k1*Z1, the proportional data a1 is increased by one, and then the abnormal value is re-judged. If Y1>k1*Z1 still exists, the acquisition time is determined to be an abnormal time; if Y1≤k1*Z1, after eliminating the abnormal value, the mean A3 of the remaining data of the same type is calculated, and the mean A3 is used as the detection data at this time point; S3: Mark the time point of the abnormal moment. If the detection time of the corresponding item data after the corresponding time point is all abnormal time, it is determined that the acquisition system has an abnormality, and an acquisition maintenance signal is generated and transmitted to the adjustment module; otherwise, it is determined that the detection data is abnormal.
3. The artificial intelligence-based computing power model data analysis and acquisition system according to claim 1 is characterized in that: The standard processing steps for hardware utilization data in the analysis module are as follows: K1: Obtains usage data for each item within a set time period from the current time point, then arranges the usage data in order of collection time and calculates the data difference C1 between adjacent sorted numbers. The data difference C1 represents the change in the usage data for the corresponding item. The data difference C1 is averaged by removing extreme values to obtain the average value A4. This average value A4 is used as the preset usage change threshold for the corresponding item. K2: Determine that the change in the usage rate data of the corresponding item is less than the preset usage rate change threshold of the corresponding item as normal fluctuation, and count the number of abnormal fluctuations of the corresponding item; The number of abnormal fluctuations of each item of usage rate data is summed up, and the weight coefficient of the corresponding item of usage rate data is equal to the number of abnormal fluctuations of the corresponding item of usage rate data divided by the sum of the number of abnormal fluctuations of each item of usage rate data; K3: The hardware usage data is equal to the sum of the corresponding item usage data multiplied by the weight coefficient of the corresponding item usage data.
4. The artificial intelligence-based computing power model data analysis and acquisition system according to claim 1 is characterized in that: The steps for generating the adjustment signal by the digital analysis module are as follows: Q1: Retrieve the hardware utilization rate data within a set time period from the current time point, calculate the difference C2 between the hardware utilization rate data corresponding to adjacent acquisition time points in the order of acquisition time, and count the number CS of times when the difference C2 within the set time period is greater than the normal fluctuation value of the hardware utilization rate data; Q2: If the total number of hardware utilization rate data within the set time period Z2*k2 < CS, it is determined that the utilization rate fluctuates frequently, and the acquisition frequency should be increased to more accurately capture the changes, generate a high utilization rate adjustment signal, and transmit the high utilization rate adjustment signal to the adjustment module; Otherwise, it is determined that the utilization rate fluctuates stably, and the acquisition frequency should be appropriately reduced to reduce unnecessary data acquisition overhead, generate a low utilization rate adjustment signal, and transmit the low utilization rate adjustment signal to the adjustment module.
5. The artificial intelligence-based computing power model data analysis and acquisition system according to claim 1 is characterized in that: The data analysis module performs the following steps for dividing the importance of data in the adjustment order: P1: Analyze the historical data of each utilization rate data, establish the fluctuation range of the historical utilization rate data of the corresponding item according to the mean and standard deviation of each historical data, compare the current detected utilization rate data of the corresponding item with the fluctuation range of the historical utilization rate data of the corresponding item. If the current detected utilization rate data of the corresponding item is within the fluctuation range of the historical utilization rate data of the corresponding item, it is determined that the priority of the utilization rate data of this corresponding item remains unchanged; If the current detected utilization rate data of the corresponding item is not within the fluctuation range of the historical utilization rate data of the corresponding item, it is determined that the priority of the utilization rate data of this corresponding item is adjusted; P2: Calculate the difference between the upper limit value and the lower limit value of the fluctuation range of the historical utilization rate data of the corresponding item that needs to adjust the priority, and divide the excess value of the current detected utilization rate data of the corresponding item exceeding the fluctuation range of the corresponding item by the difference calculated by the upper and lower limits of the fluctuation range to obtain the proportion of the excess value to the fluctuation range. Compare the calculated proportion data with the preset proportion data of 50%; P3: If the calculated ratio data BL js Less than the preset ratio data BL ys , then the priority adjustment amount ΔYX1=YX dq *(1+BL js ), YX dq Prioritize usage data for current corresponding items; If the ratio data BL is calculated js Greater than the preset ratio data BL ys , then the priority adjustment amount ΔYX1=YX dq *[1+(BL js -BL ys )]+1; then recalculate the priority of each usage data, and the new priority after recalculation is YX X1 =YX dq +ΔYX1, rearrange the order according to priority.
6. The artificial intelligence-based computing power model data analysis and acquisition system according to claim 5 is characterized in that: The data analysis module performs the following steps for dividing the impact on the performance indicators of the computing power model in the adjustment order: G1: Retrieve the historical utilization rate data of the corresponding item that needs to adjust the priority, mark the utilization rate data whose corresponding item data is not within the fluctuation range of the historical utilization rate data of the corresponding item as outliers,剔除 the outliers, and then correspond the remaining historical utilization rate data with the performance indicator XN of the computing power model according to the acquisition time, and mark the time when the historical utilization rate data is missing; G2: If the number of consecutive missing time marks is greater than the preset number threshold, or the total number of missing time marks is greater than the preset number threshold,重新 select the target time period until the number of consecutive missing time marks in the selected target time period is less than the preset number threshold, and the total number of missing time marks is less than the preset number threshold; 补全 the missing values within the time period when the number of consecutive missing time marks is less than the preset number threshold; G3: Standardize the performance indicator XN of the computing power model and calculate the utilization characteristic SY of each hardware i Pearson correlation coefficient with performance index XN, Pearson correlation coefficient n is the number of samples, SY i,t is the specific value of the i-th hardware usage feature at the t-th sampling time; G4: The greater the impact of the utilization rate data of the corresponding item on the performance indicator of the computing power model, the closer the Pearson correlation coefficient of the utilization rate data of the corresponding item is to 1; Preset correlation coefficient threshold r max , if the Pearson correlation coefficient r of the corresponding item usage rate data i ≥r max , then the priority adjustment amount ΔYX2=YX dq *[1+(r i -r max )]+1; otherwise, the priority adjustment amount ΔYX2=YX dq *(1+BL js ); Then recalculate the priority of each usage data, and the new priority YX X2 =YX dq +ΔYX2.
7. The artificial intelligence-based computing power model data analysis and acquisition system according to claim 6 is characterized in that: The data analysis module performs the following steps for determining the priority of the hardware utilization rate data of the corresponding item: L1: Comprehensively consider the impact of data status importance and computing power model performance indicators, and adjust the priority of the corresponding item usage data L2: If there are corresponding item usage data of the same priority size, the priority adjustment amount of the corresponding item usage data of the same priority is calculated according to the weight ratios ω1 and ω2 pre-set by the user, and the priority adjustment amount ΔYX = ω1*ΔYX1+ω2*ΔYX2.