Parallel processor power consumption dynamic optimization method and system

Through the multimodal feature fusion and attention mechanism enhanced high-precision power consumption prediction model, combined with intelligent adaptive scheduling strategy, the problem of difficulty in achieving accurate prediction and dynamic response in parallel processor power consumption management is solved, and the effective reduction of power consumption and stable performance is achieved, which is suitable for high-performance computing and other fields.

CN120029436AActive Publication Date: 2025-05-23SHANDONG INSPUR SCI RES INST CO LTD
View PDF 14 Cites 0 Cited by

Patent Information

Application Number
CN202510510969.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-05-23
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

Existing parallel processor power consumption management technologies are difficult to achieve accurate prediction and dynamic response to future power consumption requirements, and cannot adapt to complex and changeable practical application scenarios, resulting in high power consumption and unstable performance.

Method used

A high-precision power consumption prediction model enhanced by multimodal feature fusion and attention mechanism is adopted, combined with CNN and LSTM, and an attention mechanism is introduced to automatically assign weights to different input features to achieve accurate prediction of power consumption of parallel processors. At the same time, through intelligent adaptive scheduling strategies, voltage, frequency and task allocation are dynamically adjusted to achieve effective reduction of power consumption and stable performance maintenance.

Benefits of technology

Real-time prediction and dynamic optimization of parallel processor power consumption is achieved, energy efficiency ratio is improved, hardware life is extended, and operating costs are reduced. It is suitable for high-performance computing, gaming entertainment, and artificial intelligence fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029436A_ABST
    Figure CN120029436A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic optimization method and system for power consumption of a parallel processor, belongs to the technical field of computer graphic processing and energy efficiency management, and realizes a high-precision power consumption prediction model based on multi-modal feature fusion and attention mechanism enhancement, the multi-modal features comprise code features of an application program and user operation behavior features, and the user operation behavior features comprise the code features of the application program and the attention mechanism. The code characteristics of the application program can reflect the complexity and resource requirements of the application program, and the user operation behavior characteristics can reflect the use habits and performance requirements of the user; the attention mechanism is enhanced, the attention mechanism is introduced on the basis of combination of the CNN and the LSTM, and the attention mechanism can automatically distribute different weights for different input features, so that the model pays more attention to the features having great influence on power consumption prediction. The method can realize effective reduction of power consumption and stable maintenance of performance, and is especially suitable for complex calculation scenes with extremely high requirements on energy consumption and performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer graphics processing and energy efficiency management, and in particular to a method and system for dynamically optimizing power consumption of a parallel processor. Background Art

[0002] With the rapid development of computer graphics processing technology, parallel processors have become the core driving force in high-performance computing, gaming entertainment, artificial intelligence and other fields. However, the high power consumption problem of parallel processors during high-performance operation has become increasingly prominent, which not only greatly increases energy consumption and operating costs, but may also shorten the life of the hardware due to problems such as overheating. The power management technology of parallel processors on the current market is mostly limited to static parameter adjustment or simple feedback control, lacking the ability to accurately predict and dynamically respond to future power consumption requirements, and is difficult to adapt to complex and changing actual application scenarios. Summary of the invention

[0003] The technical task of the present invention is to address the above shortcomings and provide a method and system for dynamic optimization of parallel processor power consumption, which can effectively reduce power consumption and maintain stable performance, and is particularly suitable for complex computing scenarios with extremely high energy consumption and performance requirements.

[0004] The technical solution adopted by the present invention to solve its technical problem is: A parallel processor power consumption dynamic optimization method is proposed, which realizes a high-precision power consumption prediction model based on multimodal feature fusion and attention mechanism enhancement. The multimodal features include application code features and user operation behavior features. The application code features can reflect the complexity and resource requirements of the application, and the user operation behavior features can reflect the user's usage habits and performance requirements. The attention mechanism is enhanced. On the basis of combining CNN and LSTM, an attention mechanism is introduced. The attention mechanism can automatically assign different weights to different input features, so that the model pays more attention to the features that have a great impact on power consumption prediction; The implementation of this method includes the following steps: (1) Data collection and preprocessing; (2) Model training and optimization; (3) Implementation of intelligent adaptive scheduling.

[0005] This method uses advanced algorithms to accurately predict the power consumption requirements of parallel processors, and adjusts their working status in real time through intelligent scheduling, thereby effectively reducing power consumption and maintaining stable performance. It is especially suitable for complex computing scenarios with extremely high energy consumption and performance requirements.

[0006] Furthermore, the data collection and preprocessing specifically include: Data collection: Regularly collect the operation data of the parallel processor. The collection methods include: parallel processor monitoring interface, API provided by the operating system and application log records; the collected information includes: general data, application code features and user operation behavior features; Data cleaning: The collected raw data may contain noise, outliers or missing values. Data cleaning techniques, such as deduplication, missing value filling, and smoothing, are used to ensure the accuracy and completeness of the data. At the same time, corresponding cleaning algorithms are used to remove invalid information based on the newly collected code features and user operation behavior features. Data normalization: In order to eliminate the impact of different dimensions on model training, the cleaned data is normalized and all feature values ​​are scaled to the same scale, usually by mapping the data to the interval [0, 1] or [-1, 1]. For code features and user operation behavior features, a normalization method that adapts to their data characteristics is used.

[0007] Furthermore, the conventional data includes: CPU usage, parallel processor usage, video memory occupancy, temperature, voltage, power consumption, and currently running applications and task types; The code feature data of the application includes code complexity, function call frequency, etc.; The user operation behavior characteristic data includes: operation frequency, operation time interval, etc.

[0008] Furthermore, the model training and optimization are specifically implemented as follows: (2.1) Data division: Divide the preprocessed data into three parts according to a certain ratio (such as 70% training set, 15% validation set, and 15% test set) for model training, validation, and testing; (2.2) Model structure design: The model adopts a structure that combines CNN and LSTM, and introduces an attention mechanism. The CNN part contains multiple convolutional layers and pooling layers to extract spatial features; the LSTM part contains multiple LSTM units to capture long-term dependencies in time series; the attention layer is after the LSTM layer to assign weights to different input features; (2.3) Loss function definition: The mean squared error (MSE) is used as the loss function to evaluate the difference between the model prediction value and the true value; (2.4) Training process: Initialization: Randomly initialize model weights and biases; Forward propagation: input the training set data into the model, perform forward propagation through the CNN and LSTM layers, and calculate the predicted value; Calculate the loss: Use the MSE loss function to calculate the error between the predicted value and the true value; Backpropagation: Use gradient descent methods (such as Adam optimizer) to perform backpropagation, calculate the gradient of the loss function with respect to the model parameters, and update the model parameters to minimize the loss; Batch Normalization: Add a batch normalization layer after each convolutional layer or LSTM layer to speed up the training process and improve model stability; Verification and adjustment: After each training cycle (epoch), the performance of the model is evaluated using the validation set data. If the loss on the validation set begins to increase (i.e., overfitting occurs), the early stopping method is used to stop the training and return the model with the best validation performance. (2.5) Hyperparameter tuning: Use grid search or random search to search for the best hyperparameter combination in the predefined parameter space, including learning rate, batch size, number and parameters of CNN and LSTM layers, parameters of attention layer, etc. Evaluate the performance of the model under different hyperparameter combinations and select the combination with the lowest validation set loss as the final model parameters.

[0009] Furthermore, the intelligent adaptive scheduling implementation includes multi-objective optimization scheduling and real-time risk assessment. The multi-objective optimization scheduling takes into account the three objectives of power consumption, performance and hardware life at the same time. When dynamically adjusting the key parameters of the parallel processor including voltage and frequency and optimizing task allocation, the multi-objective optimization algorithm is used to find the optimal balance point between the three objectives; for example, under the premise of ensuring a certain performance, power consumption is reduced as much as possible, while avoiding excessive use of hardware and extending the hardware life; The real-time risk assessment evaluates the operating risks of the parallel processor in real time, including overheating and overloading. When potential risks are detected, the scheduling strategy immediately takes corresponding measures, including quickly reducing voltage and frequency, and urgently migrating high-risk tasks. At the same time, according to the severity and type of the risk, the parameters of the scheduling strategy are dynamically adjusted to ensure that the parallel processor runs in a safe and stable state.

[0010] Furthermore, the intelligent adaptive scheduling implementation specifically includes: Power consumption prediction module: During the operation of the parallel processor, the current operation status information is collected in real time and input into the trained power consumption prediction model; the model outputs the predicted power consumption value and its confidence interval within the specified time period in the future; at the same time, based on the newly collected code features and user operation behavior features, the prediction results are further refined to improve the accuracy of the prediction; Parameter adjustment module: intelligently adjusts the voltage and frequency of the parallel processor according to the predicted power consumption value; if a high power consumption state is predicted, the voltage and frequency are reduced in advance to reduce power consumption; otherwise, they are increased to improve performance; based on the temperature limit, hardware life and performance requirements of the parallel processor, the adjusted parameters are ensured to be within a safe range; and the amplitude and speed of parameter adjustment are dynamically adjusted according to the real-time risk assessment results; Task allocation module: When a parallel processor reaches the set power consumption limit or performance bottleneck threshold, some high-power consumption tasks are migrated to other idle or low-power consumption parallel processors for execution based on the importance, urgency, code characteristics and user operation behavior characteristics of the current task; this requires the system to have the ability to schedule and communicate tasks across parallel processors, and to be able to quickly and accurately evaluate the migration costs and benefits of tasks; Self-learning mechanism: Continuously optimize and adjust the parameters of the scheduling strategy based on the actual operation results. For example, by recording the power consumption changes, performance, hardware status and user feedback after each adjustment, the system can gradually optimize the scheduling strategy and improve its adaptability and robustness to different application scenarios.

[0011] Furthermore, the method performs software and hardware integrated design, integrates a power consumption optimization acceleration module in the parallel processor hardware, and the module is used to implement parallel processing power consumption prediction and scheduling algorithms, greatly improving the response speed of the system; at the same time, the acceleration module adopts a low-power design, which will not increase excessive power consumption while improving performance; provides an intelligent configuration engine, through which users input their own needs and preferences, such as giving priority to reducing power consumption, giving priority to improving performance, etc. The intelligent configuration engine automatically adjusts the parameters of the prediction model and scheduling algorithm according to the user's input to achieve personalized power consumption optimization; the specific implementation is as follows: Firmware and driver modification: Integrate the prediction model and scheduling algorithm into the firmware and driver of the parallel processor. Modify the underlying control logic of the parallel processor to ensure that the prediction model and scheduling algorithm can obtain the operating status of the parallel processor in real time and perform corresponding optimization operations. At the same time, integrate the hardware acceleration module and intelligent configuration engine into the firmware and driver to achieve the collaborative work of software and hardware. System testing: Conduct system testing in a variety of real-world scenarios, including high-load gaming, large-scale data processing, and deep learning training. Verify the optimization effect of the system by comparing power consumption, performance, hardware temperature, and stability indicators before and after optimization. At the same time, collect user feedback and opinions to further optimize and improve the system. Performance evaluation: In addition to the power consumption optimization effect, it is also necessary to evaluate its impact on the performance of parallel processors. By comparing the performance indicators before and after optimization, including frame rate, response time, task completion time, etc., ensure that the system does not significantly reduce the user experience while reducing power consumption. At the same time, evaluate the impact of the system on the life of the hardware to ensure the stability and reliability of the hardware during long-term operation.

[0012] The present invention also claims protection for a parallel processor power consumption dynamic optimization system, including a data acquisition and preprocessing module, a model training and optimization module, and an intelligent adaptive scheduling module, the system realizes the training and optimization of a high-precision power consumption prediction model based on multimodal feature fusion and attention mechanism enhancement, and realizes the dynamic optimization of the power consumption of the parallel processor based on the high-precision power consumption prediction model; The system specifically implements dynamic optimization of power consumption of parallel processors through the above method.

[0013] The present invention also claims protection for a parallel processor power consumption dynamic optimization device, comprising: at least one memory and at least one processor; The at least one memory is used to store a machine-readable program; The at least one processor is used to call the machine-readable program to implement the above method.

[0014] The present invention also claims protection for a computer-readable medium having computer instructions stored thereon, which implement the above method when executed by a processor.

[0015] Compared with the prior art, the method and system for dynamically optimizing power consumption of parallel processors of the present invention have the following beneficial effects: By introducing deep learning prediction models and intelligent scheduling strategies, real-time prediction and dynamic optimization of parallel processor power consumption are achieved. This improvement improves the energy efficiency of parallel processors, extends hardware life, and reduces operating costs, meeting the urgent needs of high-performance computing, gaming entertainment, and artificial intelligence for high-performance, low-power parallel processors. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a flowchart of a method for dynamically optimizing power consumption of a parallel processor provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0017] The present invention will be further described below in conjunction with specific embodiments.

[0018] The embodiment of the present invention provides a method for dynamically optimizing the power consumption of parallel processors. It uses advanced algorithms to accurately predict the power consumption requirements of parallel processors, and adjusts their working states in real time through intelligent scheduling, thereby effectively reducing power consumption and maintaining stable performance. It is particularly suitable for complex computing scenarios with extremely high energy consumption and performance requirements.

[0019] This method realizes a high-precision power consumption prediction model based on multimodal feature fusion and attention mechanism enhancement, and realizes dynamic optimization of parallel processor power consumption based on the high-precision power consumption prediction model. It mainly includes the following contents: 1. High-precision power consumption prediction model, including: Multimodal feature fusion: In addition to analyzing the historical operation data, current workload, temperature, voltage and other multi-dimensional information of the parallel processor, the code features of the application (such as code complexity, API call mode, and video memory access frequency) and user operation behavior characteristics (operation delay sensitivity and task priority preference) are also introduced. Code features can reflect the complexity and resource requirements of the application, while user operation behavior characteristics can reflect the user's usage habits and performance requirements. Through multimodal feature fusion, the model can more comprehensively capture the factors that affect the power consumption of parallel processors and achieve more accurate predictions of power consumption changes.

[0020] Attention mechanism enhancement: Based on the combination of CNN and LSTM, the attention mechanism is introduced. The attention mechanism can automatically assign different weights to different input features, so that the model pays more attention to the features that have a greater impact on power consumption prediction. For example, in some specific application scenarios, temperature and workload may have a more significant impact on power consumption. The attention mechanism can highlight the role of these features and improve the prediction accuracy of the model.

[0021] 2. Intelligent adaptive scheduling strategy, including: Multi-objective optimization scheduling: Traditional scheduling strategies mainly focus on reducing power consumption, while the scheduling strategy of this method takes into account the three goals of power consumption, performance, and hardware life at the same time. When dynamically adjusting key parameters such as voltage and frequency of parallel processors and optimizing task allocation, a multi-objective optimization algorithm is used to find the optimal balance between these three goals. For example, while ensuring a certain performance, power consumption can be reduced as much as possible, while avoiding excessive use of hardware and extending hardware life.

[0022] Real-time risk assessment and response: The system assesses the operating risks of parallel processors in real time, such as overheating and overload. When potential risks are detected, the scheduling strategy will immediately take corresponding measures, such as quickly reducing voltage and frequency, and urgently migrating high-risk tasks. At the same time, the parameters of the scheduling strategy will be dynamically adjusted according to the severity and type of the risk to ensure that the parallel processors operate in a safe and stable state.

[0023] 3. Software and hardware integrated design: Hardware acceleration module: A dedicated power optimization acceleration module is integrated into the parallel processor hardware. This module can process power prediction and scheduling algorithms in parallel, greatly improving the system's response speed. At the same time, the acceleration module adopts a low-power design, which will not increase excessive power consumption while improving performance.

[0024] Smart configuration engine: Provides a smart configuration engine through which users can input their needs and preferences, such as prioritizing lower power consumption or improving performance. The smart configuration engine automatically adjusts the parameters of the prediction model and scheduling algorithm based on the user's input to achieve personalized power consumption optimization.

[0025] Combined with Figure 1 As shown, the implementation process of this method is as follows: 1. Data collection and preprocessing.

[0026] Data collection: The parallel processor's operating data is collected regularly through the parallel processor's monitoring interface, the API provided by the operating system, and the application's log records. In addition to conventional data such as CPU usage, parallel processor usage, video memory usage, temperature, voltage, power consumption, and currently running applications and task types, the application's code features (such as code complexity, function call frequency, etc.) and user operation behavior features (such as operation frequency, operation time interval, etc.) are also collected.

[0027] Data cleaning: The collected raw data may contain noise, outliers or missing values. Data cleaning techniques such as deduplication, filling missing values, and smoothing are used to ensure the accuracy and integrity of the data. At the same time, special cleaning algorithms are used to remove invalid information based on the newly collected code features and user operation behavior features.

[0028] Data normalization: In order to eliminate the impact of different dimensions on model training, the cleaned data is normalized and all feature values ​​are scaled to the same scale, usually mapping the data to the interval [0, 1] or [-1, 1]. Different normalization methods are used for code features and user operation behavior features to adapt to their data characteristics.

[0029] 2. Model training and optimization.

[0030] (1) Data division: The preprocessed data is divided into three parts according to a certain ratio (such as 70% training set, 15% validation set, and 15% test set) for model training, validation, and testing.

[0031] (2) Model structure design: The model uses a structure that combines CNN and LSTM, and introduces an attention mechanism. The CNN part contains several convolutional layers and pooling layers to extract spatial features; the LSTM part contains multiple LSTM units to capture long-term dependencies in time series; the attention layer is after the LSTM layer to assign weights to different input features. The specific structure may be as follows: CNN part: Conv2D(filters=32, kernel_size=(3, 3), activation='relu') ->MaxPooling2D(pool_size=(2, 2)) ->... (can be repeated for multiple layers) LSTM part: LSTM(units=64, return_sequences=True) ->LSTM(units=32) Attention layer: Attention() Output layer: Dense(units=1, activation='linear') (output predicted power consumption value) (3) Loss function definition: The mean squared error (MSE) is used as the loss function to evaluate the difference between the model prediction value and the true value. The formula of MSE is as follows: ; Where N is the number of samples, y i is the actual power consumption value of the i-th sample, y^ i is the power consumption value of the i-th sample predicted by the model.

[0032] (4) Training process: Initialization: Randomly initialize model weights and biases.

[0033] Forward propagation: The training set data is input into the model, forward propagated through the CNN and LSTM layers, and the predicted value is calculated.

[0034] Calculate the loss: Use the MSE loss function to calculate the error between the predicted value and the true value.

[0035] Backpropagation: Use gradient descent methods (such as the Adam optimizer) to perform backpropagation, calculate the gradient of the loss function with respect to the model parameters, and update the model parameters to minimize the loss.

[0036] Batch Normalization: Add a batch normalization layer after each convolutional layer or LSTM layer to speed up the training process and improve model stability. The formula of the batch normalization layer is as follows: ; Among them, x i is the i-th element in the input batch, m is the batch size, μ B and are the mean and variance of the batch, ϵ is a small constant to avoid division by zero errors, and γ and β are learnable scaling and translation parameters.

[0037] Validation and adjustment: After each training cycle (epoch), the validation set data is used to evaluate the performance of the model. If the loss on the validation set begins to increase (that is, overfitting occurs), the early stopping method is used to stop the training and return the model with the best validation performance.

[0038] (5) Hyperparameter tuning: Use methods such as grid search or random search to search for the optimal hyperparameter combination in the predefined parameter space, such as learning rate, batch size, number and parameters of CNN and LSTM layers, parameters of attention layer, etc.

[0039] Evaluate the performance of the model under different hyperparameter combinations and select the combination with the lowest validation set loss as the final model parameters.

[0040] 3. Implementation of intelligent scheduling strategy.

[0041] Power consumption prediction module: During the operation of the parallel processor, the current operation status information is collected in real time and input into the trained power consumption prediction model. The model outputs the predicted power consumption value and its confidence interval for a period of time in the future. At the same time, based on the newly collected code features and user operation behavior features, the prediction results are further refined to improve the accuracy of the prediction.

[0042] Parameter adjustment module: intelligently adjusts the voltage and frequency of the parallel processor according to the predicted power consumption value. If a high power consumption state is predicted, the voltage and frequency are reduced in advance to reduce power consumption; otherwise, they are appropriately increased to improve performance. At the same time, the temperature limit, hardware life and performance requirements of the parallel processor are considered to ensure that the adjusted parameters are within a safe range. In addition, the amplitude and speed of parameter adjustment are dynamically adjusted according to the real-time risk assessment results.

[0043] Task allocation module: When a parallel processor is about to reach its power consumption limit or performance bottleneck, it intelligently migrates some high-power tasks to other idle or low-power parallel processors for execution based on the importance, urgency, code characteristics, and user operation behavior characteristics of the current task. This requires the system to have the ability to schedule and communicate tasks across parallel processors, and to be able to quickly and accurately evaluate the cost and benefit of task migration.

[0044] Self-learning mechanism: The system also has self-learning capabilities, and can continuously optimize and adjust the parameters of the scheduling strategy according to the actual operation results. For example, by recording the power consumption changes, performance, hardware status and user feedback after each adjustment, the system can gradually optimize the scheduling strategy and improve its adaptability and robustness to different application scenarios.

[0045] 4. Software and hardware integration and testing.

[0046] Firmware and driver modification: Integrate the prediction model and scheduling algorithm into the firmware and driver of the parallel processor. This requires modifying the underlying control logic of the parallel processor to ensure that the prediction model and scheduling algorithm can obtain the operating status of the parallel processor in real time and perform corresponding optimization operations. At the same time, integrate the hardware acceleration module and intelligent configuration engine into the firmware and driver to achieve the collaborative work of software and hardware.

[0047] System testing: Conduct system testing in a variety of real-world scenarios, including high-load gaming, large-scale data processing, deep learning training, etc. Verify the optimization effect of the system by comparing power consumption, performance, hardware temperature, and stability indicators before and after optimization. At the same time, collect user feedback and opinions to further optimize and improve the system.

[0048] Performance evaluation: In addition to the power consumption optimization effect, the system also needs to evaluate its impact on the performance of parallel processors. By comparing performance indicators such as frame rate, response time, and task completion time before and after optimization, ensure that the system does not significantly reduce user experience while reducing power consumption. At the same time, evaluate the impact of the system on the life of the hardware to ensure the stability and reliability of the hardware during long-term operation.

[0049] This method can be widely used in various computer products equipped with parallel processors, such as high-performance servers, workstations, game consoles, and edge computing devices. By optimizing the power consumption of parallel processors, these products will have higher energy efficiency and lower operating costs, thereby improving market competitiveness.

[0050] The embodiment of the present invention also provides a parallel processor power consumption dynamic optimization system, which implements the training and optimization of a high-precision power consumption prediction model based on multimodal feature fusion and attention mechanism enhancement, and implements the dynamic optimization of the power consumption of the parallel processor based on the high-precision power consumption prediction model. The system specifically implements the dynamic optimization of the power consumption of the parallel processor through the parallel processor power consumption dynamic optimization method described in the above embodiment.

[0051] The system includes: 1. Data acquisition and preprocessing module.

[0052] Data collection: The system regularly collects the operation data of the parallel processor through the parallel processor monitoring interface, the API provided by the operating system, and the log records of the application. In addition to conventional data such as CPU usage, parallel processor usage, video memory usage, temperature, voltage, power consumption, and currently running applications and task types, the system also collects application code features (such as code complexity, function call frequency, etc.) and user operation behavior features (such as operation frequency, operation time interval, etc.).

[0053] Data cleaning: The collected raw data may contain noise, outliers or missing values. The system uses data cleaning techniques, such as deduplication, filling missing values, and smoothing, to ensure the accuracy and integrity of the data. At the same time, a special cleaning algorithm is used to remove invalid information based on the newly collected code features and user operation behavior features.

[0054] Data normalization: In order to eliminate the impact of different dimensions on model training, the system normalizes the cleaned data and scales all feature values ​​to the same scale, usually mapping the data to the interval [0, 1] or [-1, 1]. Different normalization methods are used for code features and user operation behavior features to adapt to their data characteristics.

[0055] 2. Model training and optimization module.

[0056] (1) Data division: The preprocessed data is divided into three parts according to a certain ratio (such as 70% training set, 15% validation set, and 15% test set) for model training, validation, and testing.

[0057] (2) Model structure design: The model adopts a structure that combines CNN and LSTM, and introduces an attention mechanism. The CNN part contains several convolutional layers and pooling layers to extract spatial features; the LSTM part contains multiple LSTM units to capture long-term dependencies in time series; the attention layer is after the LSTM layer to assign weights to different input features.

[0058] (3) Loss function definition: The mean squared error (MSE) is used as the loss function to evaluate the difference between the model prediction value and the true value.

[0059] (4) Training process: Initialization: Randomly initialize model weights and biases.

[0060] Forward propagation: The training set data is input into the model, forward propagated through the CNN and LSTM layers, and the predicted value is calculated.

[0061] Calculate the loss: Use the MSE loss function to calculate the error between the predicted value and the true value.

[0062] Backpropagation: Use gradient descent methods (such as the Adam optimizer) to perform backpropagation, calculate the gradient of the loss function with respect to the model parameters, and update the model parameters to minimize the loss.

[0063] Batch Normalization: Add a batch normalization layer after each convolutional layer or LSTM layer to speed up the training process and improve model stability.

[0064] Validation and adjustment: After each training cycle (epoch), the validation set data is used to evaluate the performance of the model. If the loss on the validation set begins to increase (that is, overfitting occurs), the early stopping method is used to stop the training and return the model with the best validation performance.

[0065] (5) Hyperparameter tuning: Use methods such as grid search or random search to search for the optimal hyperparameter combination in the predefined parameter space, such as learning rate, batch size, number and parameters of CNN and LSTM layers, parameters of attention layer, etc.

[0066] Evaluate the performance of the model under different hyperparameter combinations and select the combination with the lowest validation set loss as the final model parameters.

[0067] 3. Intelligent adaptive scheduling module.

[0068] Power consumption prediction submodule: During the operation of the parallel processor, the current operation status information is collected in real time and input into the trained power consumption prediction model. The model outputs the predicted power consumption value and its confidence interval for a period of time in the future. At the same time, based on the newly collected code features and user operation behavior features, the prediction results are further refined to improve the accuracy of the prediction.

[0069] Parameter adjustment submodule: Based on the predicted power consumption value, the system intelligently adjusts the voltage and frequency of the parallel processor. If a high power consumption state is predicted, the voltage and frequency are reduced in advance to reduce power consumption; otherwise, they are appropriately increased to improve performance. At the same time, the temperature limit, hardware life and performance requirements of the parallel processor are considered to ensure that the adjusted parameters are within a safe range. In addition, the amplitude and speed of parameter adjustment are dynamically adjusted according to the real-time risk assessment results.

[0070] Task Allocation Sub-module: When the parallel processor is about to reach the power consumption limit or performance bottleneck, the system intelligently migrates some high-power-consuming tasks to other idle or low-power-consuming parallel processors according to the importance, urgency, code characteristics, and user operation behavior characteristics of the current tasks. This requires the system to have the task scheduling and communication capabilities across parallel processors and be able to quickly and accurately evaluate the migration costs and benefits of tasks.

[0071] Self-learning Mechanism: The system also has the self-learning ability and can continuously optimize and adjust the parameters of the scheduling strategy according to the actual operation effects. For example, by recording the power consumption changes, performance performance, hardware status, and user feedback after each adjustment, the system can gradually optimize the scheduling strategy and improve its adaptability and robustness to different application scenarios.

[0072] The system adopts an integrated hardware and software design: Hardware Acceleration Module: Integrate a dedicated power consumption optimization acceleration module in the parallel processor hardware. This module can process power consumption prediction and scheduling algorithms in parallel, greatly improving the response speed of the system. At the same time, the acceleration module adopts a low-power design and will not increase too much power consumption while improving performance.

[0073] Intelligent Configuration Engine: Provide an intelligent configuration engine through which users can input their own requirements and preferences, such as preferentially reducing power consumption or preferentially improving performance. The intelligent configuration engine will automatically adjust the parameters of the prediction model and scheduling algorithm according to the user input to achieve personalized power consumption optimization.

[0074] The implementation of the integration and testing of hardware and software includes the following: Firmware and Driver Modification: Integrate the prediction model and scheduling algorithm into the firmware and driver programs of the parallel processor. This requires modifying the underlying control logic of the parallel processor to ensure that the prediction model and scheduling algorithm can obtain the running status of the parallel processor in real time and perform corresponding optimization operations. At the same time, integrate the hardware acceleration module and intelligent configuration engine into the firmware and driver to achieve the coordinated operation of hardware and software.

[0075] System Testing: Conduct system testing in a variety of actual scenarios, including high-load games, large-scale data processing, deep learning training, etc. By comparing the power consumption, performance performance, hardware temperature, and stability indicators before and after optimization, verify the optimization effect of the system. At the same time, collect user feedback and opinions to further optimize and improve the system.

[0076] Performance Evaluation: In addition to the power consumption optimization effect, the system also needs to evaluate its impact on the performance of the parallel processor. By comparing performance metrics such as frame rate, response time, and task completion time before and after optimization, ensure that the system does not significantly degrade the user experience while reducing power consumption. At the same time, evaluate the impact of the system on the hardware lifespan to ensure the stability and reliability of the hardware during long-term operation.

[0077] An embodiment of the present invention also provides a device for dynamically optimizing the power consumption of a parallel processor, including: at least one memory and at least one processor; The at least one memory is used to store machine-readable programs; The at least one processor is used to call the machine-readable program to implement the method for dynamically optimizing the power consumption of the parallel processor described in the above embodiment.

[0078] An embodiment of the present invention also provides a computer-readable medium, on which computer instructions are stored. When the computer instructions are executed by a processor, the processor executes the method for dynamically optimizing the power consumption of the parallel processor described in the above embodiment. Specifically, a system or device equipped with a storage medium can be provided, on which software program code for implementing the functions of any one of the above embodiments is stored, and the computer (or CPU or MPU) of the system or device reads and executes the program code stored in the storage medium.

[0079] In this case, the program code read from the storage medium itself can implement the functions of any one of the above embodiments, so the program code and the storage medium storing the program code constitute a part of the present invention.

[0080] Embodiments of the storage medium for providing program code include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Optionally, the program code can be downloaded from a server computer via a communication network.

[0081] In addition, it should be clear that not only can the actual operations be completed in part or in whole by executing the program code read by the computer, but also by instructions based on the program code to cause an operating system or the like operating on the computer, thereby implementing the functions of any one of the above embodiments.

[0082] In addition, it can be understood that the program code read from the storage medium is written to a memory provided in an expansion board inserted into the computer or written to a memory provided in an expansion unit connected to the computer, and then based on the instructions of the program code, a CPU installed on the expansion board or the expansion unit is enabled to perform part or all of the actual operations, thereby realizing the functions of any of the above-mentioned embodiments.

[0083] The present invention is shown and described in detail above through the accompanying drawings and preferred embodiments. However, the present invention is not limited to these disclosed embodiments. Based on the above multiple embodiments, those skilled in the art can know that the code review methods in the above different embodiments can be combined to obtain more embodiments of the present invention, and these embodiments are also within the protection scope of the present invention.

Claims

1. A method for dynamically optimizing power consumption of a parallel processor, characterized in that: A high-precision power consumption prediction model is implemented based on multimodal feature fusion and attention mechanism enhancement. The multimodal features include application code features and user operation behavior features. The application code features can reflect the complexity and resource requirements of the application, and the user operation behavior features can reflect the user's usage habits and performance requirements. The attention mechanism is enhanced. On the basis of combining CNN and LSTM, an attention mechanism is introduced. The attention mechanism can automatically assign different weights to different input features, so that the model pays more attention to the features that have a great impact on power consumption prediction; The implementation of this method includes the following steps: (1) Data collection and preprocessing; (2) Model training and optimization; (3) Implementation of intelligent adaptive scheduling.

2. The method for dynamic optimization of power consumption of a parallel processor according to claim 1, characterized in that: The data collection and preprocessing specifically include: Data collection: Regularly collect the operation data of the parallel processor. The collection methods include: parallel processor monitoring interface, API provided by the operating system and application log records; the collected information includes: general data, application code features and user operation behavior features; Data cleaning: Data cleaning technology is used to ensure the accuracy and completeness of data. At the same time, the corresponding cleaning algorithm is used to remove invalid information based on the newly collected code features and user operation behavior features. Data normalization: Normalize the cleaned data and scale all feature values ​​to the same scale; for code features and user operation behavior features, use a normalization method that adapts to their data characteristics.

3. The method for dynamic optimization of power consumption of a parallel processor according to claim 2, characterized in that: The conventional data includes: CPU usage, parallel processor usage, video memory occupancy, temperature, voltage, power consumption, and currently running applications and task types; The code feature data of the application program includes code complexity and function call frequency; The user operation behavior characteristic data includes: operation frequency and operation time interval.

4. The method for dynamic optimization of power consumption of a parallel processor according to claim 1, characterized in that: The specific implementation process of the model training and optimization is as follows: (2.1) Data division: The preprocessed data is divided into three parts according to the proportion for model training, verification and testing; (2.2) Model structure design: The model adopts a structure that combines CNN and LSTM, and introduces an attention mechanism. The CNN part contains multiple convolutional layers and pooling layers to extract spatial features. The LSTM part contains multiple LSTM units to capture long-term dependencies in time series. The attention layer comes after the LSTM layer to assign weights to different input features. (2.3) Loss function definition: The mean square error is used as the loss function to evaluate the difference between the model prediction value and the true value; (2.4) Training process: Initialization: Randomly initialize model weights and biases; Forward propagation: input the training set data into the model, perform forward propagation through the CNN and LSTM layers, and calculate the predicted value; Calculate the loss: Use the MSE loss function to calculate the error between the predicted value and the true value; Back propagation: Use the gradient descent method to perform back propagation, calculate the gradient of the loss function with respect to the model parameters, and update the model parameters to minimize the loss; Batch Normalization: Add a batch normalization layer after each convolutional layer or LSTM layer to speed up the training process and improve model stability; Verification and adjustment: After each training cycle, the performance of the model is evaluated using the validation set data. If the loss on the validation set begins to increase, the early stopping method is used to stop the training and return the model with the best validation performance. (2.5) Hyperparameter tuning: Search for the best hyperparameter combination in a predefined parameter space, including learning rate, batch size, number and parameters of CNN and LSTM layers, and parameters of attention layer; Evaluate the performance of the model under different hyperparameter combinations and select the combination with the lowest validation set loss as the final model parameters.

5. The method for dynamic optimization of power consumption of a parallel processor according to claim 1, characterized in that: The intelligent adaptive scheduling implementation includes multi-objective optimization scheduling and real-time risk assessment, The multi-objective optimization scheduling considers the three objectives of power consumption, performance and hardware life at the same time, and finds the optimal balance point among the three objectives through a multi-objective optimization algorithm when dynamically adjusting key parameters of the parallel processor including voltage and frequency and optimizing task allocation; The real-time risk assessment assesses the operational risks of the parallel processor in real time, including overheating and overloading; When potential risks are detected, the scheduling strategy immediately takes corresponding measures, including quickly reducing voltage and frequency and urgently migrating high-risk tasks; at the same time, according to the severity and type of the risk, the parameters of the scheduling strategy are dynamically adjusted to ensure that the parallel processor runs in a safe and stable state.

6. A method for dynamically optimizing power consumption of a parallel processor according to claim 1 or 5, characterized in that: The intelligent adaptive scheduling implementation specifically includes: Power consumption prediction module: During the operation of the parallel processor, the current operation status information is collected in real time and input into the trained power consumption prediction model; the model outputs the predicted power consumption value and its confidence interval within the specified time period in the future; at the same time, the prediction results are further refined based on the newly collected code features and user operation behavior features; Parameter adjustment module: intelligently adjusts the voltage and frequency of the parallel processor according to the predicted power consumption value; if a high power consumption state is predicted, the voltage and frequency are reduced in advance to reduce power consumption; otherwise, they are increased to improve performance; based on the temperature limit, hardware life and performance requirements of the parallel processor, the parameters are adjusted to ensure that the adjusted parameters are within a safe range; at the same time, the amplitude and speed of parameter adjustment are dynamically adjusted according to the real-time risk assessment results; Task allocation module: When a parallel processor reaches the set power consumption limit or performance bottleneck threshold, some high-power consumption tasks are migrated to other idle or low-power consumption parallel processors for execution according to the importance, urgency, code characteristics and user operation behavior characteristics of the current task; Self-learning mechanism: Continuously optimize and adjust the parameters of the scheduling strategy based on the actual operating results.

7. The method for dynamic optimization of power consumption of a parallel processor according to claim 1, characterized in that: The method integrates software and hardware design, integrates a power consumption optimization acceleration module into the parallel processor hardware, and uses the module to implement parallel processing power consumption prediction and scheduling algorithms; provides an intelligent configuration engine, through which users input their own needs and preferences, and the intelligent configuration engine automatically adjusts the parameters of the prediction model and scheduling algorithm according to the user's input to achieve personalized power consumption optimization; the specific implementation is as follows: Firmware and driver modification: Integrate the prediction model and scheduling algorithm into the firmware and driver of the parallel processor. Modify the underlying control logic of the parallel processor to ensure that the prediction model and scheduling algorithm can obtain the operating status of the parallel processor in real time and perform corresponding optimization operations. At the same time, integrate the hardware acceleration module and intelligent configuration engine into the firmware and driver to achieve the collaborative work of software and hardware. System testing: Conduct system testing in a variety of real-world scenarios, including high-load gaming, large-scale data processing, and deep learning training. Verify the optimization effect of the system by comparing power consumption, performance, hardware temperature, and stability indicators before and after optimization. At the same time, collect user feedback and opinions to further optimize and improve the system. Performance evaluation: By comparing the performance indicators before and after optimization, including frame rate, response time, and task completion time, we ensure that the system does not significantly reduce the user experience while reducing power consumption; at the same time, we evaluate the impact of the system on the life of the hardware to ensure the stability and reliability of the hardware during long-term operation.

8. A parallel processor power consumption dynamic optimization system, characterized in that: The system includes a data acquisition and preprocessing module, a model training and optimization module, and an intelligent adaptive scheduling module. The system realizes the training and optimization of a high-precision power consumption prediction model based on multimodal feature fusion and attention mechanism enhancement, and realizes the dynamic optimization of the power consumption of parallel processors based on the high-precision power consumption prediction model. The system specifically realizes dynamic optimization of power consumption of parallel processors through the method described in any one of claims 1 to 7.

9. A device for dynamically optimizing power consumption of a parallel processor, characterized in that: include: at least one memory and at least one processor; The at least one memory is used to store a machine-readable program; The at least one processor is used to call the machine-readable program to implement the method described in any one of claims 1 to 7.

10. A computer-readable medium, characterized in that The computer readable medium stores computer instructions, and when the computer instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • CNN-LSTM fan fault prediction method and CNN-LSTM fan fault prediction system based on attention mechanism

    CN112633317A

  • Scheduling method, device and equipment of photovoltaic energy storage system and storage medium

    CN117578534A

  • Building energy consumption prediction method and system based on TAB-GBDT

    CN118036803A

  • Energy consumption prediction method and device, readable storage medium and processor

    CN118211713A

  • Self-adaptive control method and system for refrigerating machine room

    CN118765083A