Artificial intelligence model energy consumption management method and device

By receiving energy consumption information and threshold relationships, the deployment, simulation, and inference processes of AI/ML models can be monitored and controlled in real time, solving the problem of excessive energy consumption in existing technologies and achieving energy efficiency optimization and energy management.

CN120979958APending Publication Date: 2025-11-18HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410623914.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-15
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing AI/ML management and maintenance workflows have failed to effectively avoid excessive energy consumption during model processing, resulting in a lack of attention and resolution to energy efficiency issues.

Method used

By receiving energy consumption information and threshold relationships, it determines whether to deploy an artificial intelligence model to the inference function, monitors and controls the model's energy consumption in real time, including energy consumption management during deployment, simulation, training and inference processes, and provides instructions and feedback information to optimize energy consumption decisions.

Benefits of technology

This effectively avoids excessive energy consumption during model processing, achieves energy efficiency optimization and management, ensures that model operation complies with equipment energy consumption limits, prevents continued operation when energy consumption is excessive, and reduces overall energy consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120979958A_ABST
    Figure CN120979958A_ABST
Patent Text Reader

Abstract

The invention provides an artificial intelligence model energy consumption management method and device, a first device sends first information to a second device, the first information indicates first energy consumption, and the second device determines whether to deploy a first model to a reasoning function according to a relationship between the first energy consumption and a first threshold value, so that the situation that the first model is deployed to the reasoning function when the energy consumption is too large can be avoided. The second equipment still deploys / loads the first model to the reasoning function, so that the situation that the energy consumption overhead is too large in the deployment stage of the first model is effectively avoided, and meanwhile the situation that the energy efficiency overhead is too large in the subsequent operation process of the first model can be effectively avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technology, and more specifically, to a method and apparatus for energy management of an artificial intelligence model. Background Technology

[0002] To improve network intelligence and automation, artificial intelligence (AI) and machine learning (ML) technologies are being applied in an increasing number of fields. The 3rd Generation Partnership Project (3GPP) working groups have several research topics related to network intelligence. The SA5 working group's R18AIMLMGMT project studied model lifecycle management, mainly including the training, simulation, deployment, and inference phases. The 3GPP SA5 R19AIMLMGMT project will further investigate the sustainability of AI / ML, focusing on energy consumption and efficiency issues.

[0003] Currently, AI / ML management and operation workflows only focus on the data and performance management of AI models. For example, during the inference phase of AI / ML, the inference function provider generates an AIML inference report after performing inference, which includes inference output, such as inference type, inference performance metrics, and output results. Similarly, during the training phase, the training function provider generates an ML training report after performing training, which includes data and performance attributes. Furthermore, because network management service (MnS) providers (producers) incur significant energy consumption during AIML training, inference, deployment, and simulation, and current AI / ML management and operation workflows do not address this energy consumption issue, there is no technology to prevent excessive energy consumption by MnS providers during model processing.

[0004] Therefore, how to avoid excessive energy consumption during model processing is an urgent problem to be solved. Summary of the Invention

[0005] This application provides an energy consumption management method for artificial intelligence models, which can effectively avoid excessive energy consumption during model processing.

[0006] In a first aspect, an artificial intelligence model energy consumption management method is provided, applied to a second device. The method includes: receiving first information, which indicates first energy consumption; and determining whether to deploy a first model to the inference function based on the relationship between the first energy consumption and a first threshold, wherein the first threshold is less than or equal to the maximum energy consumption supported by the second device.

[0007] The maximum energy consumption supported by the second device at different times can be the same or different.

[0008] Optionally, determining whether to deploy the first model to the inference function based on the relationship between the first energy consumption and the first threshold includes: when the first energy consumption is less than or equal to the first threshold, determining to deploy the first model to the inference function; or when the first energy consumption is greater than or equal to the first threshold, determining not to deploy the first model to the inference function.

[0009] It should be understood that when the first energy consumption equals the first threshold, the first model can be deployed to the inference function, or the first model can not be deployed to the inference function.

[0010] Optionally, the first threshold is less than or equal to the maximum power consumption supported by the function (e.g., inference function) in the second device.

[0011] It should be understood that the maximum energy consumption supported by the second device can be the maximum energy consumption that the second device itself can support during operation, or it can be the maximum energy consumption required by the first device. For example, in order to save power, the first device requires the second device / functions in the second device (such as inference functions) to support a maximum energy consumption. This maximum energy consumption is the energy consumption value that the first device requires the second device / functions in the second device to not exceed.

[0012] Based on the above scheme, by issuing a first information indicating the first energy consumption, the second device can determine whether to deploy / load the first model into the inference function based on the first energy consumption and the first threshold. This avoids the second device still deploying / loading the first model into the inference function when the energy consumption is too high, thus effectively avoiding excessive energy consumption during the first model deployment phase, and also effectively avoiding excessive energy consumption during the subsequent operation of the first model.

[0013] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: when it is determined that the first model will not be deployed to the inference function, the method further includes: sending first indication information, the first indication information including a reason for not deploying the first model to the inference function.

[0014] Based on the above scheme, when it is determined that the first model-to-inference function will not be deployed, the second device can also report the reason for not deploying the first model-to-inference function through indication information. For example, when the second device compares the first energy consumption indicated by the first information with the first threshold, if the first energy consumption is greater than the first threshold, the second device determines that the first model-to-inference function will not be deployed. At this time, it needs to send indication information to report that the first energy consumption fails to meet the energy consumption requirement (e.g., the first energy consumption is not less than or equal to the first threshold), so the first model-to-inference function is not deployed. This improves the process of determining whether to deploy the model-to-inference function based on energy consumption.

[0015] In conjunction with the first aspect, in some implementations of the first aspect, when it is determined that the first model is deployed to the inference function, the method further includes: sending a second instruction message, the second instruction message being used to indicate the deployment result of the first model.

[0016] The deployment result includes whether the deployment was successful or failed.

[0017] Based on the above scheme, when it is determined to deploy the first model to the inference function, the second device may be deployed successfully or it may fail during the inference process of deploying the model. Regardless of whether the deployment is successful or not, the second device will send an instruction message to report the result of the model deployment, thereby improving the process of determining the deployment of the model to the inference function based on energy consumption.

[0018] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: when the deployment result is a failure, the second indication information includes the reason for the deployment failure.

[0019] Based on the above scheme, if the deployment fails during the model deployment process, the reason for the deployment failure will also be reported, thereby improving the process of determining the deployment model and inference function based on energy consumption.

[0020] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: obtaining a second energy consumption, which is used to determine the first energy consumption, the second energy consumption being the energy consumption generated by the second device when simulating the first model in a first time period.

[0021] For example, the second energy consumption may be the power consumed by a computing device when it independently runs the inference function of the first model.

[0022] The second energy consumption can be the peak energy consumption obtained by the second device simulating the first model in the first time period, or the total energy consumption of the second device simulating the first model in the first time period, etc., which is not limited in this application.

[0023] Based on the above scheme, the first energy consumption is determined by the second device simulating the first model in the first time period (e.g., the second energy consumption). This allows the second device to determine whether the first model can be deployed to the inference function based on the first energy consumption determined by the second energy consumption and in conjunction with the first threshold. This ensures that the deployment of the first model meets the energy consumption level / energy consumption requirements of the second device or the inference function (e.g., less than or equal to the maximum energy consumption supported by the second device or less than or equal to the maximum energy consumption supported by the inference function). This effectively avoids excessive energy consumption during the subsequent operation of the deployed model, which exceeds the maximum energy consumption supported by the second device or the maximum energy consumption required by the first device.

[0024] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: receiving second information, the second information being used to instruct the second device to calculate and record the energy consumption generated when the first model is simulated.

[0025] Based on the above scheme, the second device can calculate and record the energy consumption generated when the first model is simulated according to the instructions of the instruction information. In other words, the second device can also obtain the second energy consumption by calculating and recording the energy consumption generated when the first model is simulated in the first time period according to the instructions of the received information, and calculating and recording the energy consumption of the first model during simulation, such as the second energy consumption, thereby improving the process of obtaining the second energy consumption.

[0026] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: sending first feedback information for the second information, the first feedback information being used to indicate the second energy consumption.

[0027] Optionally, a first feedback message is sent, which is used to indicate the second energy consumption.

[0028] It should be understood that the first feedback information may be sent in response to the first information or may be sent directly by the second device; this application does not limit this.

[0029] Based on the above scheme, the second device will report the energy consumption (e.g., the second energy consumption) of the first model during simulation through information, so that the reporting device (e.g., the first device) can determine the first energy consumption based on the second energy consumption, so as to indicate the first energy consumption to the second device in the future, so that the second device can determine whether to deploy the first model to the inference function, thus improving the process of the second device deploying the model.

[0030] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: receiving third information, the third information being used to indicate the first threshold.

[0031] The third information and the first information can be sent through the same information or through different information.

[0032] Based on the above scheme, the first threshold can also be indicated by information, thereby improving the process of the second device deployment model.

[0033] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: receiving fourth information, which is used to indicate the type of the reasoning function.

[0034] The fourth and second information can be sent through the same information or through different information.

[0035] Based on the above scheme, the type of inference function to be deployed can also be indicated to the second device through information, so that the second device can deploy the first model to the corresponding type of inference function, thereby improving the process of the second device deploying the model to the inference function of the indicated type.

[0036] In conjunction with the first aspect, in some implementations of the first aspect, when the first model is successfully deployed, the method further includes: receiving fifth information, the fifth information being used to instruct the second device to calculate and record the energy consumption generated when the first model performs inference; and sending second feedback information in response to the fifth information, the second feedback information being used to instruct a third energy consumption, the third energy consumption being the energy consumption generated by the second device when performing inference on the first model during a second time period.

[0037] The third energy consumption can be the peak energy consumption obtained when the first model performs inference in the second time period, or the total energy consumption when the first model performs inference in the second time period, etc., which is not limited in this application.

[0038] Based on the above scheme, when the deployment model is determined and the model is successfully deployed, the second device can also calculate and record the energy consumption generated by the first model during inference in the second time period through the indication information. It can also report the calculated and recorded energy consumption (e.g., third energy consumption) generated by the first model during inference in the second time period through information reporting. This allows the subsequent information reporting device (e.g., the first device) to determine whether to pause, terminate, suspend, or cancel the model inference based on the third energy consumption. This enables the energy consumption of model inference to be monitored in real time, preventing the model from continuing to run / perform inference when the energy consumption is too high, and effectively avoiding excessive energy consumption during model inference / running.

[0039] In conjunction with the first aspect, in some implementations of the first aspect, when the first model is successfully deployed, the method further includes: receiving sixth information, the sixth information being used to request the second device to query the energy consumption generated when the first model performs inference; and sending third feedback information in response to the sixth information, the third feedback information being used to indicate fourth energy consumption, the fourth energy consumption being the energy consumption generated by the second device when performing inference on the first model during a third time period.

[0040] The fourth energy consumption can be the peak energy consumption of the first model during the third time period, or the total energy consumption of the first model during the third time period, etc. This application does not limit it.

[0041] It should be noted that the third time period can be the period before the second time period, or the third time period can be the second time period. When the third time period is the second time period, the fourth energy consumption is the third energy consumption.

[0042] Based on the above scheme, the second device can query the energy consumption generated by the first model during inference in the third time period according to the instructions of the received sixth information. It can be understood that the second device records and calculates the energy consumption generated by the model during inference according to the instructions of the fifth information. The energy consumption generated by the model during inference is saved. Therefore, the second device can query the energy consumption generated by the model during inference in the past according to the instructions of the sixth information, and report the energy consumption in the historical inference process through the feedback information of the sixth information. So that the device that reports the feedback information (such as the first device) can determine whether the inference function of the model can be started at present based on the energy consumption in the historical inference process and the current energy consumption requirement (the energy consumption that the second device can accept) and thus ensure that the energy consumption of the model can meet the current energy consumption requirement. This avoids the model being started and running when the energy consumption requirement is low, which would lead to excessive energy consumption. This effectively avoids excessive energy consumption during the model's inference / running process.

[0043] In conjunction with the first aspect, in some implementations of the first aspect, after the first model is successfully deployed, the method further includes: receiving usage policy information, the usage policy information being used to instruct the second device to stop the inference function when the energy consumption generated by the second device when the first model performs the inference function exceeds a second threshold.

[0044] Optionally, the usage strategy information is used to instruct the second device to shut down, cancel, or suspend the inference function when the energy consumption generated by the first model during inference is greater than a second threshold.

[0045] Based on the above scheme, the second device can determine, according to the instructions of the usage strategy information, that when the energy consumption generated by the first model in performing the inference function exceeds the second threshold, the inference function will be stopped, canceled, or suspended, which can effectively avoid excessive energy consumption of the model during inference / running.

[0046] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: receiving seventh information, the seventh information being used to instruct the second device to calculate and record the energy consumption generated when the first model is trained; and sending fourth feedback information in response to the seventh information, the fourth feedback information being used to instruct fifth energy consumption, the fifth energy consumption being the energy consumption generated by the second device when training the first model during a fourth time period.

[0047] The fifth energy consumption can be the energy consumption generated when the first model performs the nth round of training. The first device can estimate the total energy consumption of N rounds of training based on the energy consumption of the nth round, where 1≤n≤N and N rounds of training is the target total number of training rounds set by the first model.

[0048] When the first model is being initially trained, the fourth time period can be before the first time period. When the first model is being retrained, the fourth time period can be after the first time period. This application does not impose any restrictions on this.

[0049] Based on the above scheme, when the model is trained on the second device, the second device can be instructed by the instruction information to record the energy consumption generated by the model during a certain period of training. This energy consumption can be the estimated total training energy consumption (total energy consumption / peak energy consumption, etc.). The energy consumption generated by the model during training (e.g., the seventh energy consumption) is reported through the fourth feedback information so that the reporting device (e.g., the first device) can determine whether to pause, terminate, suspend, or cancel the training of the model based on the seventh energy consumption. This allows the energy consumption of the model during training to be monitored in real time, preventing the model from continuing to train when the energy consumption is too high, and effectively avoiding excessive energy consumption during the training process.

[0050] In a second aspect, an artificial intelligence model energy consumption management method is provided, characterized in that it is applied to a first device, the method comprising: sending first information, the first information being used to indicate the first energy consumption, the first energy consumption being used in conjunction with a first threshold to determine whether a second device deploys a first model to the inference function, the first threshold being less than or equal to the maximum energy consumption supported by the second device.

[0051] In conjunction with the second aspect, in some implementations of the second aspect, when it is determined that the first model will not be deployed to the inference function, the method further includes: receiving first indication information, the first indication information including a reason for not deploying the first model to the inference function.

[0052] In conjunction with the second aspect, in some implementations of the second aspect, when it is determined that the first model is deployed to the inference function, the method further includes: receiving second indication information, the second indication information being used to indicate the deployment result of the first model.

[0053] In conjunction with the second aspect, in some implementations of the second aspect, when the deployment result is a failure, the second indication information includes the reason for the deployment failure.

[0054] In conjunction with the second aspect, in some implementations of the second aspect, the method further includes: sending second information, which instructs the second device to calculate and record the energy consumption generated when the first model is simulated.

[0055] In conjunction with the second aspect, in some implementations of the second aspect, the method further includes: receiving first feedback information for the second information, the first feedback information being used to indicate a second energy consumption, the second energy consumption being the energy consumption generated by the second device simulating the first model during a first time period, the second energy consumption being used to determine the first energy consumption.

[0056] In conjunction with the second aspect, in some implementations of the second aspect, the method further includes: sending third information, which is used to indicate the first threshold.

[0057] In conjunction with the second aspect, in some implementations of the second aspect, the method further includes: sending a fourth message indicating the type of the reasoning function.

[0058] In conjunction with the second aspect, in some implementations of the second aspect, when the first model is successfully deployed, the method further includes: sending fifth information, the fifth information being used to instruct the second device to calculate and record the energy consumption generated when the first model performs inference; receiving second feedback information in response to the fifth information, the second feedback information being used to indicate third energy consumption, the third energy consumption being the energy consumption generated by the second device when performing inference on the first model during a second time period; and determining whether to suspend the second device's inference on the first model based on the third energy consumption.

[0059] Optionally, based on the third energy consumption, it is determined whether to suspend, cancel, or terminate the second device's inference of the first model.

[0060] Based on the above scheme, when the second device performs inference on the model, it can record the energy consumption generated by the model during a certain period of inference / running through the indication information. This energy consumption can be the estimated total inference energy consumption (total energy consumption / peak energy consumption, etc.). The first device can determine whether to pause the inference of the model based on this energy consumption, so that the energy consumption of the model inference can be monitored in real time, preventing the model from continuing to run when the energy consumption is too high, and effectively avoiding excessive energy consumption during the inference / running process of the model.

[0061] In conjunction with the second aspect, in some implementations of the second aspect, when the first model is successfully deployed, the method further includes: sending a sixth message, the sixth message being used to request the second device to query the energy consumption generated when the first model performs inference; receiving a third feedback message in response to the sixth message, the third feedback message being used to indicate a fourth energy consumption, the fourth energy consumption being the energy consumption generated by the second device when performing inference on the first model during a third time period; and determining whether to enable the second device to perform inference on the first model based on the fourth energy consumption.

[0062] Based on the above scheme, the historical inference energy consumption value of the model is obtained. According to the historical energy consumption value and the current energy consumption requirement (for example, the historical energy consumption value needs to be less than or equal to the third threshold, and the third threshold is less than or equal to the maximum energy consumption supported by the current second device), it is determined whether the inference function of the model can be started at present. This ensures that the energy consumption of the model can meet the current energy consumption requirement, and avoids the model being started and running when the energy consumption requirement is low, which would lead to excessive energy consumption. This effectively avoids excessive energy consumption during the inference / running process of the model.

[0063] In conjunction with the second aspect, in some implementations of the second aspect, after the first model is successfully deployed, the method further includes: sending usage policy information, which is used to instruct the second device to stop the inference function when the energy consumption generated by the first model in performing the inference function exceeds a second threshold.

[0064] Optionally, the usage strategy information is used to instruct the second device to shut down, cancel, or suspend the inference function when the energy consumption generated by the first model during inference is greater than a second threshold.

[0065] The method further includes generating the usage strategy, which is carried in the usage strategy information. Based on the above scheme, by generating a model usage strategy, for example, the usage strategy can specify that when the model's inference / running energy consumption exceeds a certain value, the inference function needs to be stopped, shut down, canceled, or suspended, and the usage strategy is sent to the second device. The second device executes the usage strategy accordingly, which can effectively avoid excessive energy consumption during model operation.

[0066] In conjunction with the second aspect, in some implementations of the second aspect, the method further includes: sending a seventh message, the seventh message being used to instruct the second device to calculate and record a fifth energy consumption, the fifth energy consumption being the energy consumption generated by the second device during training the first model in the fourth time period; receiving a fourth feedback message in response to the seventh message, the fourth feedback message being used to indicate the value of the fifth energy consumption; and determining, based on the value of the fifth energy consumption, whether to pause the training of the first model by the second device.

[0067] Optionally, based on the value of the fifth energy consumption, it can be determined whether to suspend, terminate, or cancel the training of the first model by the second device.

[0068] Based on the above scheme, the system determines whether to pause the training of the model based on the energy consumption reported by the second device when the second device trains the first model. This allows the energy consumption of the model training to be monitored in real time, preventing the model from continuing to train when the energy consumption is too high, and effectively avoiding excessive energy consumption during the training process.

[0069] Thirdly, an apparatus is provided, the apparatus comprising: a processing unit, configured to determine whether to deploy a first model to an inference function based on a relationship between a first energy consumption and a first threshold, wherein the first threshold is less than or equal to the maximum energy consumption supported by a second device;

[0070] A transceiver unit is used to receive first information, which is used to indicate first energy consumption.

[0071] The transceiver unit can perform the receiving and sending processes described in the first aspect above, and the processing unit of the communication device can perform other processes described in the first aspect above besides receiving and sending.

[0072] Fourthly, a communication device is provided, the device comprising: a transceiver unit for transmitting first information, the first information being used to indicate first energy consumption, the first energy consumption being used in conjunction with a first threshold to determine whether a second device deploys a first model to an inference function, the first threshold being less than or equal to the maximum energy consumption supported by the second device;

[0073] The transceiver unit can perform the receiving and sending processes in the second aspect described above, and the processing unit of the communication device can perform other processes in the second aspect described above besides receiving and sending.

[0074] Fifthly, an apparatus is provided, including a processor for executing a computer program, such that the apparatus performs the methods described in the first to second aspects and any possible implementation thereof.

[0075] Optionally, the processor may be one or more.

[0076] Optionally, the communication device further includes a memory for storing the computer program, and the memory may be one or more.

[0077] Optionally, the memory may be integrated with the processor, or the memory may be separate from the processor, or the memory may be located within the processor.

[0078] Optionally, the communication device may also include transceiver circuitry such as a transceiver or input / output circuitry.

[0079] In a sixth aspect, a system is provided, comprising: a first device and a second device, wherein the first device is configured to perform the method in the possible implementation of the second aspect described above, and the second device is configured to perform the method in the possible implementation of the first aspect described above.

[0080] In a seventh aspect, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program or code, which, when executed on a computer, causes the computer to perform the method in any of the possible implementations of the first to second aspects described above.

[0081] Eighthly, a chip is provided, including at least one processor for running a computer program that causes a device on which the chip is mounted to perform the methods of the first to second aspects and any possible implementation thereof.

[0082] The chip may include an output circuit or interface for transmitting information or data, and an input circuit or interface for receiving information or data.

[0083] Ninth aspect, a computer program product is provided, the computer program product comprising: computer program code, which, when run on a communication device, causes the device to perform the methods of the first to second aspects and any possible implementation thereof.

[0084] The chip may include an input circuit or interface for transmitting information or data, and an output circuit or interface for receiving information or data.

[0085] The possible designs and beneficial effects of aspects three through nine can be found in the descriptions of aspects one through two. Attached Figure Description

[0086] Figure 1 This is a schematic diagram of a system architecture applicable to this application.

[0087] Figure 2 This is a flowchart of the general AI / ML operation workflow in the lifecycle of an ML entity.

[0088] Figure 3 This is a schematic flowchart of an artificial intelligence model energy management method 300 provided in an embodiment of this application.

[0089] Figure 4 This is a schematic flowchart of an artificial intelligence model energy management method 400 provided in an embodiment of this application.

[0090] Figure 5 This is a schematic flowchart of an artificial intelligence model energy management method 500 provided in an embodiment of this application.

[0091] Figure 6 This is a schematic flowchart of an artificial intelligence model energy management method 600 provided in an embodiment of this application.

[0092] Figure 7 This is a schematic flowchart of an artificial intelligence model energy management method 700 provided in an embodiment of this application.

[0093] Figure 8 This is a schematic block diagram of the AI ​​model energy management device 1000 provided in the embodiments of this application.

[0094] Figure 9 This is a schematic block diagram of the AI ​​model energy management device 2000 provided in the embodiments of this application.

[0095] Figure 10 This is a schematic block diagram of the chip system 3000 provided in the embodiments of this application. Detailed Implementation

[0096] The technical solutions in this application will now be described with reference to the accompanying drawings.

[0097] The various numerical designations, such as First, Second, #1, #2, etc., are merely for descriptive convenience and are not intended to limit the scope of the embodiments of this application, nor are they intended to indicate order or importance, such as distinguishing different messages or different information. "Predefined" can be implemented by pre-storing corresponding codes, tables, or other methods that can be used to indicate relevant information in the device; this application does not limit the specific implementation method. The "protocol" involved can refer to standard protocols in the field of communication, such as the Long Term Evolution (LTE) protocol, the NR protocol, and related protocols applied to future communication systems; this application does not limit this. Words such as "exemplary," "for example," "exemplarily," and "as (another) example" are used to indicate that something is an example, illustration, or description. Any embodiment or design scheme described as "exemplary" in this application should not be construed as being better or more advantageous than other embodiments or design schemes. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized. "At least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the related objects before and after it are in an "or" relationship. Descriptions involving network element A sending messages, information, or data to network element B, and network element B receiving messages, information, or data from network element A, aim to specify which network element the message, information, or data is to be sent to, without limiting whether they are sent directly or indirectly through other network elements. "Used for indication" can include both direct and indirect indication. When describing an indication used to indicate A, it can include whether the indication directly or indirectly indicates A, but does not necessarily mean that the indication carries A. Descriptions such as "when," "in the case of," "if," and "if" all indicate that the device will take corresponding actions under certain objective circumstances, not a time limit, and do not require the device to perform a judgment action during implementation, nor do they imply any other limitations.

[0098] The technical solutions of this application can be applied to various communication systems, including but not limited to: 5th generation (5G) systems or new radio (NR) systems, long term evolution (LTE) systems, long term evolution-advanced (LTE-A) systems, wireless local area network (WLAN) systems, satellite communication systems, optical communication systems, microwave communication systems, etc. They can also be applied to future communication systems, such as future communication networks, or systems integrating multiple systems. Furthermore, they can be applied to device-to-device (D2D) communication, vehicle-to-everything (V2X) communication, machine-to-machine (M2M) communication, machine-type communication (MTC), and Internet of Things (IoT) communication systems or other communication systems. Furthermore, it can be extended to similar wireless communication systems, such as Wireless-Fidelity (Wi-Fi), Worldwide Interoperability for Microwave Access (WIMAX), and communication systems related to the 3rd Generation Partnership Project (3GPP), without limitation.

[0099] In a communication system, a device can send signals to or receive signals from another device. These signals can include information, signaling, or data. The term "device" can also be replaced by an entity, network entity, communication equipment, communication module, node, communication node, etc. This application uses "device" as an example for description.

[0100] Figure 1 This is a schematic diagram of a system architecture applicable to this application. For example... Figure 1As shown, machine learning (ML) training functions and / or AI (artificial intelligence, AI) / ML inference functions can be located in the radio access network (RAN) domain management services (MnS) consumer, such as a cross-domain management system; or in the management functions of a specific domain management system, such as the RAN or core network (CN); or in network functions.

[0101] For management data analytics (MDA), the ML training function can be located inside or outside the management data analytics function (MDAF). The AI / ML inference function is located within the MDAF.

[0102] For the network data analytics function (NWDAF), the ML training function can be located within the NWDAF or the management system, while the AI / ML inference function is located within the NWDAF.

[0103] For RAN, ML training and AI / ML inference functions can both reside in the next-generation node base station (gNB), or the ML training function can reside in the management system while the AI / ML inference function resides in the gNB.

[0104] Therefore, there may be multiple location schemes for ML training functions and AI / ML inference functions.

[0105] For example, such as Figure 1 As shown in (a), for a RAN domain-specific MDA, the ML training function and AI / ML inference function of the MDA can be located in a RAN domain-specific MDAF.

[0106] For example, such as Figure 1 As shown in (b), the ML training function is located in the 3GPP RAN domain-specific management function, while the AI / ML inference function is located in the gNB.

[0107] For example, such as Figure 1 As shown in (c), both the ML training function and the AI / ML inference function are located in gNB.

[0108] Therefore, in the 3GPP network domain, the network management system (NMS) acts as the consumer of machine learning training (MLT). The MLT producer can be the element management system (EMS) or a network element (NE) managed by the EMS, such as a gNB or CN functional network element (e.g., NWDAF) in the RAN. The MLT consumer can call the model training service provided by the MLT producer. Furthermore, in the ORAN network domain, the service management and orchestration function (SMO) acts as the MLT consumer, and the network elements directly managed by the SMO (which can be heterogeneous, such as EMS, gNB, NWDAF, etc.) act as the MLT producer.

[0109] It should be noted that, apart from training, the providers of ML simulation, deployment, and inference services can be EMS or network elements managed by EMS, or network elements directly managed by SMO. ML training, simulation, deployment, and inference can be completed on the same producer or on different producers. For example, ML training can be performed on EMS, while ML simulation can be performed on network elements managed by EMS; or both ML training and ML simulation can be performed on EMS. This application does not limit this.

[0110] The following section introduces the MLT producer (e.g., EMS, radio access network elements (e.g., gNB), core network elements (e.g., NWDAF), MLT consumer (e.g., NMS, SMO) and their respective functions.

[0111] Machine learning training (MLT) consumer, the caller of model training service.

[0112] Machine learning training provider (MLT producer), a provider of model training services.

[0113] A network management system (NMS) is responsible for the operation, management, and maintenance of a network. It can also be called a cross-domain management system.

[0114] An element management system (EMS) is used to manage one or more network elements of a specific category. It can also be called a domain management system or a single-domain management system.

[0115] A next-generation node base station (gNB) is a device in mobile communications that connects the fixed and wireless components and connects to mobile terminals via an over-the-air wireless channel.

[0116] The network data analytics function (NWDAF) network element has various intelligent computing functions such as AI training and inference.

[0117] The Service Management and Orchestration (SMO) function plays a similar role to the NMS in the network architecture, responsible for the operation, management, and maintenance of various network services and orchestration functions. It can directly manage heterogeneous network elements, such as EMS, gNB, and NWDAF.

[0118] Access network (AN) elements provide network access functionality to authorized users in a specific area and can use transmission tunnels of different quality depending on the user's level and service requirements. Access networks can employ different access technologies. Currently, there are two types of radio access technologies: 3rd Generation Partnership Project (3GPP) access technologies (such as those used in 3G, 4G, or 5G systems) and non-3GPP access technologies. 3GPP access technologies refer to access technologies that conform to 3GPP standards and specifications. Access networks using 3GPP access technologies are called Radio Access Networks (RANs). In 5G systems, access network equipment is called next-generation node base stations (gNBs). Non-3GPP access technologies refer to access technologies that do not conform to 3GPP standards and specifications, such as air interface technologies represented by access points (APs) in Wi-Fi.

[0119] An access network that implements access network functions based on wireless communication technology can be called a radio access network (RAN). RANs can be used for radio resource management, uplink and downlink data classification and quality of service (QoS) applications, as well as signaling processing with control plane functions and data forwarding with user plane functions. RANs can be next-generation (e.g., future communication networks or higher versions) RANs or traditional (e.g., 5G, 4G, 3G, or 2G) RANs. An access network device (or RAN device) is a device that provides wireless communication functions for terminal devices; it can also be called a network device. In the embodiments of this application, the device used to implement the functions of the access network device can be a network device or a device capable of supporting the access network device in implementing these functions. This device can be called a network device or an access network device, such as a chip system or a chip, and can be installed in the access network device. In the embodiments of this application, the chip system can be composed of chips or include chips and other discrete components. For ease of description, in all embodiments of this application, the aforementioned device that provides wireless communication functions for terminal devices is collectively referred to as an access network device or simply RAN or AN.

[0120] For example, the access network device can be a base station, a broadband network gateway (BNG), an aggregation switch, or a non-3GPP access device. Base stations can include various forms, such as macro base stations, micro base stations (also called small stations), relay stations, and access points; this application does not specifically limit these. Devices for terminal access to the core network are uniformly referred to as access network devices in this document. For example, an access network device can be an evolved universal terrestrial radio access network (E-UTRAN) device in a fourth-generation (4G) network, or a next-generation radio access network (NG-RAN) device in a fifth-generation (5G) network.

[0121] In some deployments, the network device mentioned in the embodiments of this application may be a device including a central unit (CU), a distributed unit (DU), or a device including CU and DU, or a device with control plane CU nodes (central unit-control plane (CU-CP)) and user plane CU nodes (central unit-user plane (CU-UP)) and DU nodes. For example, the network device may include gNB-CU-CP, gNB-CU-UP, and gNB-DU.

[0122] In some deployments, multiple RAN nodes collaborate to assist terminals in achieving wireless access, with different RAN nodes implementing some functions of the base station. For example, RAN nodes can be CUs, DUs, CU-CPs, CU-UPs, or radio units (RUs). CUs and DUs can be set up separately or included in the same network element, such as a baseband unit (BBU). RUs can be included in radio equipment or radio units, such as remote radio units (RRUs), active antenna units (AAUs), or remote radio heads (RRHs). In one possible design, the processing unit in the BBU that implements baseband functions is called a baseband high (BBH) unit, and the processing unit in the RRU / AAU / RRH that implements baseband functions is called a baseband low (BBL) unit. In different systems, CUs (or CU-CPs and CU-UPs), DUs, or RUs may have different names, but those skilled in the art will understand their meaning. For example, the radio access network can also be an open radio access network (O-RAN) architecture. In an ORAN system, the CU can also be called an O-CU (open CU), the DU can also be called an O-DU, the CU-CP can also be called an O-CU-CP, the CU-UP can also be called an O-CU-UP, and the RU can also be called an O-RU. Any of the units among the CU (or CU-CP, CU-UP), DU, and RU in this application can be implemented through software modules, hardware modules, or a combination of software modules and hardware modules.

[0123] This application primarily addresses the first device and the second device. The first device is the consumer of the AI ​​model service, and the second device is the producer of the AI ​​model service. The first device can be an NMS or an SMO, and the second device can be an EMS or a network element managed by an EMS or an SMO. Both the first and second devices can be network devices, network apparatuses, or network elements, etc. It should be understood that this application does not limit the specific form of the first and second devices.

[0124] It should be noted that the system architecture described above applicable to the embodiments of this application is merely an example, and the system architecture applicable to this application is not limited thereto. Any system architecture capable of implementing the functions of the aforementioned network elements is applicable to this application. In other words, the system architecture described in this application is for the purpose of more clearly illustrating the technical solutions of this application and does not constitute a limitation on the technical solutions provided in this application. Those skilled in the art will understand that with the evolution of system architecture or the emergence of new business scenarios, the technical solutions provided in this application are also applicable to similar technical problems.

[0125] Furthermore, the network element names and message / information names mentioned in this application are merely examples. In future communication networks, these network elements and messages / information may also use other names, as long as they have the same or similar functions as those described in this application and achieve the same or similar technical objectives, they should all fall within the technical scope covered by this application. For example, in future communication networks, some or all of the names of the aforementioned network elements may be retained from 4G / 5G, or new names may be adopted.

[0126] Currently, AI / ML technologies are widely used in 5G systems (5GS), including the 5G core network (5GC), next-generation radio access network (NG-RAN), and management systems. A general AI / ML operation workflow diagram for the ML entity lifecycle is shown below. Figure 2 As shown. From Figure 2As can be seen, this workflow mainly includes four operational phases: training, simulation, deployment, and inference. The training phase includes ML model training and ML model testing. ML model training primarily involves training one or a set of models and validating ML entities to evaluate their performance on training and validation data. ML testing primarily involves testing the validated ML entities to evaluate the performance of the trained ML model on test data. The simulation phase includes ML simulation, where ML entities are run for inference in a simulation environment to evaluate their inference performance before being applied to the target network or system. The deployment phase includes ML entity loading, which is the process of providing the trained ML entities to the target AI / ML inference function. The inference phase includes AI / ML inference, where the AI / ML inference function uses the trained ML entities for inference. For a detailed description, please refer to 3GPP TS28.105, which will not be repeated here.

[0127] Current AI / ML management and operation workflows focus only on the data and performance management of AI models. For example, during the simulation and inference phases, AIML simulation and inference function providers generate AIML inference reports (AIMLInferenceReport) after performing inference. These reports include inference outputs, representing the output of the inference function. Inference outputs include the following attributes: output identity (ID), inference type (e.g., MDAF, NWDAF, specific function in RAN network elements), output timestamp, inference performance metrics, and output results. During the training phase, AIML training function providers generate MLTrainingReports after performing training. These reports include attributes such as ID, data, and performance. During the deployment phase, AIML deployment functions only include control attributes of the deployment process (e.g., process status, cancellation, pause, resumption) and model identity. Tables 1 and 2 show the attribute information for InferenceOutputs and MLTrainingReport, respectively. In Tables 1 and 2, M represents medium, O represents overwhelming, and CM represents contrasting medium. Support qualifiers are one of the metrics used to evaluate the quality of model predictions. M, O, and CM in the support qualifiers represent a classification of support strength, used to describe the model's confidence and reliability in its predictions. In the table, T and F represent True and False, respectively. For example, the inference output ID is readable but not writable, is a variable, and is notifiable.

[0128] Table 1

[0129]

[0130] Table 2

[0131]

[0132] Furthermore, because MnS producers incur significant energy consumption during AIML training and inference, and current AI / ML management and maintenance workflows do not address this issue, there is no technology to prevent excessive energy consumption by MnS providers during AI model operation. In the future, the 3GPP SA5R19 AIMLMGMT project will further investigate the sustainability of AI / ML, focusing on energy consumption and efficiency. The R19SID Study on AI / ML management-phase 2 will evaluate the energy consumption / efficiency impact of all AIML operation and maintenance phases (training, simulation, deployment, and inference) as one of its research objectives. Therefore, how to avoid excessive energy consumption of AI models during operation and maintenance is a pressing issue.

[0133] Based on the aforementioned technical issues, this application receives first information via a second device, which indicates first energy consumption. The second device determines whether to deploy / load the first model into the inference function based on the relationship between the first energy consumption and a first threshold. This avoids the second device still deploying / loading the first model into the inference function when the energy consumption is too high, thereby effectively avoiding excessive energy consumption during the model deployment stage and also effectively avoiding excessive energy consumption during the subsequent operation of the first model.

[0134] Figure 3 This is a schematic flowchart illustrating an artificial intelligence model energy management method 300 provided in an embodiment of this application. Method 300 includes steps S310 to S320, which will be described in detail below.

[0135] S310, the first device sends information #1 to the second device, which indicates energy consumption #1. Correspondingly, the second device receives information #1 from the first device. Information #1 is an example of first information, and energy consumption #1 is an example of first energy consumption.

[0136] Message #1 also instructs the second device to execute the process of deploying the first AI model. The first AI model is an example of the first model.

[0137] In one approach, when deployment is triggered by a first device, the first device sends an ML entity load request to a second device. Accordingly, the second device receives the ML entity load request from the first device. The ML entity load request is used to indicate energy consumption #1. The ML entity load request can be an MLEntityLoadingRequest. An example of an ML entity load request being information #1 is provided.

[0138] For example, the first device is an NMS and the second device is an EMS. When deployment is triggered by the NMS and the first AI model is deployed on the NE, the NMS sends an ML entity loading request to the EMS, which indicates energy consumption #1. Accordingly, the EMS receives the ML entity loading request from the NMS. This ML entity loading request is also used to request the EMS to perform AI model deployment and deploy the AI ​​model to the NE, which may be managed by the EMS.

[0139] For example, the first device is an SMO (Service Provider Object) and the second device is an NE (Application Provider Object). When deployment is triggered by the SMO and the first AI model is deployed on the NE, the SMO sends an ML (Multi-Level Model) entity loading request to the NE, which indicates energy consumption #1. Correspondingly, the NE receives the ML entity loading request from the SMO. This ML entity loading request also requests the NE to execute the AI ​​model deployment process and deploy the AI ​​model onto the NE, which can be managed by the SMO.

[0140] In another approach, when deployment is triggered by the second device, the first device sends ML entity loading policy information to the second device. Accordingly, the second device receives the ML entity loading policy information from the first device. This ML entity loading policy information is used to indicate energy consumption #1. The ML entity loading policy information can be an MLEntityLoadingPolicy. An example of ML entity loading policy information #1 is provided.

[0141] For example, the first device is an NMS and the second device is an EMS. When deployment is triggered by the EMS, the NMS sends ML entity loading policy information to the EMS, which indicates energy consumption #1. Correspondingly, the EMS receives the ML entity loading policy information from the NMS. This ML entity loading policy information also instructs the EMS or NE to execute the AI ​​model deployment process and deploy the AI ​​model to the NE. The NE can be managed by the EMS.

[0142] For example, the first device is an SMO (System-on-Machine) and the second device is an NE (Network-on-Age). When deployment is triggered by the NE, the SMO sends ML (Multi-Level Model) entity loading strategy information to the NE. This ML entity loading strategy information includes a simulation energy consumption indicator, which indicates the energy consumption value of the NE deploying the model during the simulation phase. Correspondingly, the NE receives the ML entity loading strategy information from the SMO. This ML entity loading strategy information also instructs the NE to execute the AI ​​model deployment process and deploy the AI ​​model to the NE. The NE can be managed by the SMO.

[0143] In one approach, the first device may also send information #2 to the second device, which instructs the second device to calculate and record the energy consumption generated when the first AI model is simulating.

[0144] Information #2 may include an energy consumption record identifier, which can be used to indicate whether the second device should record the energy consumption generated during the simulation of the first AI model. For example, when the energy consumption record identifier is True, the second device needs to record the energy consumption generated during the simulation of the first AI model; when the energy consumption record identifier is False, the second device may not record the energy consumption generated during the simulation of the first AI model. Information #2 can also be used to request the second device to simulate the first AI model.

[0145] Optionally, the information #2 may also include an energy consumption calculation method identifier, which indicates how the second device calculates the energy consumption generated when the first AI model is simulated.

[0146] The methods for calculating energy consumption include, but are not limited to, at least one of the following:

[0147] 1. Calculated by incremental power consumption.

[0148] 2. Estimation: Model training energy consumption ≈ Energy consumption per training round * Number of training rounds; Energy consumption per training round ≈ Execution time of a single training session * Average processor power * Power usage efficiency.

[0149] It should be noted that the unit for energy consumption can be joule, kilowatt-hour, etc., without limitation.

[0150] For example, the first device sends a simulation request to the second device. Correspondingly, the second device receives the simulation request from the first device. The simulation request is an example of message #2.

[0151] The simulation request includes an energy consumption recording identifier. This identifier indicates whether the second device should record the energy consumption generated during the simulation of the first AI model. For example, when the energy consumption recording identifier is True, the second device needs to record the energy consumption generated during the simulation of the first AI model; when the energy consumption recording identifier is False, the second device may not record the energy consumption generated during the simulation of the first AI model. This simulation request can be an EmulationRequest.

[0152] Optionally, the simulation request may include an energy consumption calculation method identifier, which is used to instruct the second device to calculate the energy consumption of the first AI model for simulation.

[0153] For example, the first device is an NMS (Natural Management System) and the second device is an EMS (Energy Management System). The NMS sends a simulation request to the EMS, which includes an energy consumption calculation method identifier and an energy consumption record identifier. Correspondingly, the EMS receives the simulation request from the NMS.

[0154] For example, the first device is an NMS (Natural Management System) and the second device is an NE (Energy Provider Interface). The NMS sends a simulation request to the NE, which includes an energy consumption calculation method identifier and an energy consumption record identifier. Correspondingly, the NE receives the simulation request from the NMS. The NE can be managed by an EMS (Energy Management System).

[0155] For example, the first device is an SMO (Solar Module) and the second device is an NE (Energy Provider Interface). The SMO sends a simulation request to the NE, which includes an energy consumption calculation method identifier and an energy consumption record identifier. Correspondingly, the NE receives the simulation request from the SMO. The NE can be managed by the SMO.

[0156] In one approach, the first device sends information #3 to the second device, whereby information #3 indicates a threshold #1. Accordingly, the second device receives information #3 from the first device. Information #3 is an example of a third type of information.

[0157] It should be noted that information #3 and information #2 can be sent through the same information or through different information; this application does not limit this.

[0158] For example, the simulation request described above is one example of information #3. This simulation request is also used to indicate a threshold #1.

[0159] In one approach, the first device sends information #4 to the second device, whereby information #4 indicates the type of inference function to which the first AI model is deployed. Accordingly, the second device receives information #4 from the first device. Information #4 is an example of a fourth type of information.

[0160] It should be noted that information #4 and information #1 can be sent through the same information or through different information; this application does not limit this.

[0161] For example, the ML entity loading request or ML entity loading strategy involved in this application is an example of information #4. This ML entity loading request or ML entity loading strategy is also used to indicate the type of inference function to which the first AI model is deployed.

[0162] It should be understood that inference functions have various types. For example, the channel state index (CSI) feedback compression enhancement function of the base station is used for channel state measurement between the UE and the base station. Traditionally, a fixed codebook matrix is ​​agreed upon by both parties to indicate the channel state. When the UE establishes a connection with the base station, the base station sends a CSI measurement signal to the UE. The UE performs measurements based on the CSI signal and feeds back a channel state index reference signal (CSI-RS), which contains the agreed-upon codebook. The base station then performs beamforming based on the codebook (channel state) to enhance communication stability with the UE (e.g., reducing noise loss and packet loss). However, due to the large amount of codebook data, it is now possible to reduce the communication overhead of codebook transmission by deploying an AI encoder on the UE side and an AI decoder on the base station side. The UE compresses the codebook, and the base station decodes and recovers the codebook for subsequent tasks. Alternatively, a base station load balancing model can be used to switch some terminals from connecting to high-load base stations to low-load base stations by predicting the number of terminal devices connected to the base station. The various network inference functions specified in 3GPP or ORAN protocols are not limited. Message #4 can indicate the type of inference function, so that after receiving message #4, the second device can deploy the first AI model onto the inference function corresponding to the type of inference function indicated by message #4.

[0163] In one approach, when the second device successfully deploys the first AI model to the inference function, the first device sends message #5 to the second device. Message #5 instructs the second device to calculate and record the energy consumption generated by the first AI model during inference. Message #5 is an example of a fifth message.

[0164] Message #5 may include an energy consumption recording identifier, which can be used to indicate whether the second device records the energy consumption generated during inference of the first AI model. For example, when the energy consumption recording identifier is True, the second device needs to record the energy consumption generated during inference of the first AI model; when the energy consumption recording identifier is False, the second device may not record the energy consumption generated during inference of the first AI model. Message #5 is also used to instruct the second device to execute the inference function of the first AI model.

[0165] Optionally, the information #5 may also include an energy consumption calculation method, which is used to indicate how the second device calculates the energy consumption generated when the first AI model is inferred.

[0166] For example, the first device sends ML inference history control information to the second device, which instructs the second device to record and calculate the total energy consumption generated when the first AI model performs inference. Correspondingly, the second device receives the ML inference history control information from the first device. The ML inference history control information is an example of information #5.

[0167] The ML inference history control information includes an energy consumption record identifier and an energy consumption calculation identifier. The energy consumption record identifier indicates whether the second device records the inference energy consumption. The energy consumption calculation method identifier indicates the method used to calculate energy consumption during the inference phase. This ML inference history control information can be MLInferenceHistoryControl. This ML inference history control information can be used to control the energy consumption calculation and recording method during the model inference process.

[0168] Optionally, the ML inference history control information may include an energy consumption calculation method identifier, which is used to indicate how the second device calculates the energy consumption generated when the first AI model performs inference.

[0169] For example, the first device is an NMS (Non-Managed System) and the second device is an EMS (Electronic Management System). The NMS sends ML (Multi-Level Meter) inference history control information to the EMS, which includes an energy consumption record identifier and an energy consumption calculation identifier. Correspondingly, the EMS receives the ML inference history control information from the NMS.

[0170] For example, the first device is an NMS (Nearest Management System) and the second device is an NE (Energy Provider Interface). The NMS sends ML (Multi-Level Marketing) inference history control information to the NE, which includes energy consumption record identifiers and energy consumption calculation identifiers. Correspondingly, the NE receives the ML inference history control information from the NMS. The NE can be managed by an EMS (Energy Management System).

[0171] For example, the first device is an SMO (Service Management Center) and the second device is an NE (Energy Provider). The SMO sends ML (Multi-Level Marketing) inference history control information to the NE, which includes energy consumption record identifiers and energy consumption calculation identifiers. Correspondingly, the NE receives the ML inference history control information from the SMO. The NE can be managed by the SMO.

[0172] In one approach, when the second device successfully deploys the first AI model to the inference function, the first device sends message #6 to the second device. Message #6 requests the second device to query the energy consumption of the first AI model during inference. Message #6 is an example of a sixth message.

[0173] It should be understood that because the second device recorded and calculated the energy consumption generated by the AI ​​model during inference through the instruction of information #5, the energy consumption generated by the AI ​​model during inference was saved. Information #6 can be used to query the energy consumption generated by the AI ​​model during inference in the historical inference process.

[0174] For example, the first device sends an ML inference history request to the second device. Correspondingly, the second device receives the ML inference history request from the first device. This ML inference history request is used to query the model inference energy consumption during the historical inference process of the first AI model. This ML inference history request can be an MLInferenceHistoryRequest. An ML inference history request is an example of information #6.

[0175] For example, the first device is an NMS and the second device is an EMS. The NMS sends an ML inference history request to the EMS. Correspondingly, the EMS receives the ML inference history request from the NMS.

[0176] For example, the first device is an NMS (Network Management System) and the second device is an NE (Network Provider Interface). The NMS sends an ML (Multi-Level Machine) inference history request to the NE. Correspondingly, the NE receives the ML inference history request from the NMS. This NE can be managed by an EMS (Electronic Management System).

[0177] For example, the first device is an SMO (System-Mounted Module) and the second device is an NE (Network-Element Module). The SMO sends an ML (Multi-Level Machine) inference history request to the NE. Correspondingly, the NE receives the ML inference history request from the SMO. The NE can be managed by the SMO.

[0178] In one approach, the first device generates a strategy for using AIML inference capabilities.

[0179] Specifically, the first device can generate a usage strategy for AIML inference functions based on historical inference energy consumption.

[0180] For example, the usage strategy is to disable the AI ​​model inference function if the energy consumption of the second device per unit time exceeds a certain threshold. The usage strategy can also be to stop, disable, cancel, or suspend the inference function if the energy consumption generated by the second device when the first AI model performs inference exceeds threshold #2.

[0181] For example, the first device is NMS, which generates a usage strategy for AIML inference functionality.

[0182] For example, the first device is an SMO, which generates a usage strategy for AIML inference functionality.

[0183] In one approach, a first device sends usage policy information to a second device, instructing the second device to stop, shut down, cancel, or suspend the inference function when the energy consumption generated by the first AI model performing inference exceeds a threshold #2. Accordingly, the second device receives the usage policy information from the first device. This usage policy information carries the usage policy for the AIML inference function. Threshold #2 is an example of a second threshold.

[0184] For example, the first device is an NMS (Network Management System) and the second device is an EMS (Electronic Management System). The NMS sends usage policy information to the EMS. Correspondingly, the EMS receives the usage policy information from the NMS.

[0185] For example, the first device is an NMS (Network Management System) and the second device is an NE (Network Provider). The NMS sends usage policy information to the NE. Accordingly, the NE receives the usage policy information from the NMS. This NE can be managed by an EMS (Electronic Management System).

[0186] For example, the first device is an SMO (Service Management Center) and the second device is an NE (Network Provider). The SMO sends usage policy information to the NE. Correspondingly, the NE receives the usage policy information from the SMO. The NE can be managed by the SMO.

[0187] In one approach, the first device sends information #7 to the second device, which instructs the second device to calculate and record the energy consumption generated during the training of the first AI model. Information #7 is an example of a seventh message.

[0188] Specifically, information #7 may include an energy consumption recording identifier, which can be used to indicate whether the second device should record the energy consumption generated during the training of the first AI model. For example, when the energy consumption recording identifier is True, the second device needs to record the energy consumption generated during the training of the first AI model; when the energy consumption recording identifier is False, the second device may not record the energy consumption generated during the training of the first AI model. Information #7 can also be used to request the second device to train the first AI model.

[0189] Optionally, the information #7 may also include an energy consumption calculation method, which is used to instruct the second device to calculate the energy consumption generated when the first AI model is trained.

[0190] For example, when model training is triggered by a first device, the first device sends an ML training request to a second device. Accordingly, the second device receives the ML training request from the first device. The ML training request is an example of message #7.

[0191] The ML training request includes an energy consumption recording identifier. This identifier indicates whether the second device should record the energy consumption generated during the training of the first AI model. For example, when the energy consumption recording identifier is True, the second device needs to record the energy consumption generated during the training of the first AI model; when the identifier is False, the second device may not record the energy consumption generated during the training of the first AI model. This training request instructs the second device to perform initial training and / or retraining of the first AI model. The ML training request can be an MLTrainingRequest.

[0192] It should be noted that the models involved in model training, model simulation, model inference, and model deployment in this application can be understood as AI / ML models, and this application does not limit the specific name of the model.

[0193] For example, the first device is NMS and the second device is EMS. Model training is performed in EMS and triggered by NMS. NMS sends ML training requests to EMS, and EMS receives ML training requests from NMS accordingly.

[0194] For example, the first device is an NMS (Network Management System), and the second device is an NE (Network Provider). Model training takes place on the NE and is triggered by the NMS. The NMS sends an ML (Model Learning) training request to the NE, and the NE receives the ML training request from the NMS. The NE can be managed by an EMS (Electronic Management System).

[0195] For example, the first device is an SMO, the second device is an NE, model training takes place on the NE, and model training is triggered by the SMO. The SMO sends an ML training request to the NE, and correspondingly, the NE receives the ML training request from the SMO. The NE can be managed by the SMO.

[0196] S320: The second device determines whether to deploy the first AI model to the inference function based on the relationship between energy consumption #1 and threshold #1, where threshold #1 is less than or equal to the maximum energy consumption supported by the second device.

[0197] Optionally, the first threshold is less than or equal to the maximum power consumption supported by the function (e.g., inference function) in the second device.

[0198] It should be understood that the maximum energy consumption supported by the second device can be the maximum energy consumption that the second device itself can support during operation, or it can be the maximum energy consumption required by the first device. For example, to save power, the first device may require the second device / its functions (such as inference functions) to support a maximum energy consumption, which is the energy consumption value that the first device requires the second device / its functions to not exceed. It should be noted that, in addition to inference functions, it can also be other functions in the second device, such as training functions. Therefore, the maximum energy consumption required by the first device can be the maximum energy consumption supported by one or more functions in the second device, without limitation.

[0199] In one approach, when the relationship between energy consumption #1 and threshold #1 is that energy consumption #1 is less than or equal to threshold #1, the second device determines to deploy the first AI model to the inference function.

[0200] In one approach, when the relationship between energy consumption #1 and threshold #1 is that energy consumption #1 is greater than or equal to threshold #1, the second device determines not to deploy the first AI model to the inference function.

[0201] It should be noted that the threshold #1 can also be understood as the energy consumption requirement or energy consumption level of the second device, reflecting the energy consumption that the second device can accept. This energy consumption can be equal to the maximum energy consumption supported by the second device or the maximum energy consumption that the functions (e.g., inference functions) in the second device can support, or it can be less than the maximum energy consumption supported by the second device or the maximum energy consumption that the functions (e.g., inference functions) in the second device can support. This application does not limit this. The description of energy consumption requirement or energy consumption level in this application can be understood as the energy consumption that the second device can accept, such as threshold #1. This energy consumption requirement or energy consumption level, or threshold #1, can be determined by the second device itself, or it can be indicated to the second device by the first device.

[0202] In one approach, a second device performs a simulation of the first AI model, calculating and recording the energy consumption generated during the simulation.

[0203] For example, the second device is an EMS, which performs a simulation of the first AI model and calculates and records the energy consumption during the simulation process.

[0204] For example, the second device is an NE managed by the EMS. The NE managed by the EMS performs a simulation of the first AI model and calculates and records the energy consumption during the simulation process of the first AI model.

[0205] For example, the second device is an NE managed by the SMO, which performs a simulation of the first AI model and calculates and records the energy consumption during the simulation process.

[0206] In one approach, the second device performs a simulation of the first AI model according to the instruction of information #2, and calculates and records the energy consumption generated when the first AI model is simulated.

[0207] For example, the second device is an EMS (Energy Management System). Based on the simulation request, the EMS performs a simulation of the first AI model, simultaneously calculating and recording the energy consumption during the simulation. The EMS determines the energy consumption calculation method based on the energy consumption calculation method identifier included in the simulation request or its own energy consumption calculation capabilities, and uses this determined method to calculate the energy consumption generated during the simulation of the first AI model. When the energy consumption recording identifier included in the simulation request is True, that is, when the energy consumption recording identifier indicates that the energy consumption generated during the simulation of the first AI model should be recorded, the second device records the energy consumption of the model simulation process while simultaneously measuring and calculating the model simulation energy consumption.

[0208] For example, the second device is a NE managed by the EMS. The NE managed by the EMS executes a simulation of the first AI model according to the simulation request, and simultaneously calculates and records the energy consumption during the simulation process. The NE managed by the EMS determines the energy consumption calculation method based on its own energy consumption calculation capability or based on the energy consumption calculation method identifier included in the simulation request, and uses this method to calculate the energy consumption generated during the simulation of the first AI model. Simultaneously, when the energy consumption recording identifier included in the simulation request is True, the energy consumption generated during the simulation of the first AI model is recorded.

[0209] For example, the second device is an NE managed by the SMO. The NE managed by the SMO executes a simulation of the first AI model according to the simulation request, and simultaneously calculates and records the energy consumption during the simulation process. The NE managed by the SMO determines the energy consumption calculation method based on its own energy consumption calculation capability or based on the energy consumption calculation method identifier included in the simulation request, and uses this method to calculate the energy consumption generated during the simulation of the first AI model. Simultaneously, when the energy consumption recording identifier included in the simulation request is True, the energy consumption generated during the simulation of the first AI model is recorded.

[0210] In one approach, the second device performs inference on the first AI model according to the instructions in information #5, and calculates and records the energy consumption generated when the first AI model performs inference.

[0211] For example, the second device is an EMS (Engineering Management System). The EMS executes the inference of the first AI model based on the ML (Multi-Level Machine) inference history control information, and simultaneously calculates and records the energy consumption of the first AI model during the inference process. The EMS determines the energy consumption calculation method based on the energy consumption calculation method identifier included in the ML inference history control information or its own energy consumption calculation capability, and uses this determined calculation method to calculate the energy consumption generated during the inference process of the first AI model. When the energy consumption record identifier included in the ML inference history control information is True, that is, when the energy consumption record identifier indicates that the energy consumption generated during the inference process of the first AI model should be recorded, the second device measures and calculates the model inference energy consumption while simultaneously recording the energy consumption of the model inference process.

[0212] For example, the second device is a NE managed by the EMS. Based on the ML inference history control information, it executes the first AI model inference, and simultaneously calculates and records the energy consumption of the first AI model during the inference process. The NE managed by the EMS determines the energy consumption calculation method based on its own energy consumption calculation capability, or the NE managed by the EMS determines it based on the energy consumption calculation method identifier included in the ML inference history control information, and uses this method to calculate the energy consumption generated during the first AI model inference process. Simultaneously, when the energy consumption record identifier included in the ML inference history control information is True, the energy consumption generated during the first AI model inference process is recorded.

[0213] For example, the second device is a NE managed by the SMO. Based on the ML inference history control information, it executes the first AI model inference, and simultaneously calculates and records the energy consumption of the first AI model during the inference process. The NE managed by the SMO determines the energy consumption calculation method based on its own energy consumption calculation capability or based on the energy consumption calculation method identifier included in the ML inference history control information, and uses this method to calculate the energy consumption generated during the first AI model inference process. Simultaneously, when the energy consumption record identifier included in the ML inference history control information is True, the energy consumption generated during the first AI model inference process is recorded.

[0214] In one approach, the second device queries the energy consumption of the first AI model during the inference process, based on the instructions in information #6.

[0215] For example, the second device queries the energy consumption of the first AI model during the inference process according to the instruction of the ML inference history request. The energy consumption of the first AI model during the inference process can be the energy consumption of the first AI model in the historical inference process, or it can be the energy consumption of the first AI model when performing inference.

[0216] In one approach, the second device, according to the instructions in information #7, performs training on the first AI model and calculates and records the energy consumption generated during the training of the first AI model.

[0217] For example, the second device is an EMS (Energy Management System). The EMS executes training of the first AI model according to the ML (Multi-Level Machine) training request, and simultaneously calculates and records the energy consumption of the first AI model during training. The EMS determines the energy consumption calculation method based on the energy consumption calculation method identifier included in the ML training request or its own energy consumption calculation capabilities, and uses this determined calculation method to calculate the energy consumption generated during the training of the first AI model. When the energy consumption recording identifier included in the ML training request is True, that is, when the energy consumption recording identifier indicates that the energy consumption generated during the training of the first AI model should be recorded, the second device records the energy consumption of the model training process while simultaneously measuring and calculating the model training energy consumption.

[0218] For example, the second device is a NE managed by EMS. Based on the ML training request, it executes training of the first AI model, and simultaneously calculates and records the energy consumption of the first AI model during training. The NE managed by EMS determines the energy consumption calculation method based on its own energy consumption calculation capability, or the NE managed by EMS determines it based on the energy consumption calculation method identifier included in the ML training request, and uses this method to calculate the energy consumption generated during the training of the first AI model. Simultaneously, when the energy consumption recording identifier included in the ML training request is True, the energy consumption generated during the training of the first AI model is recorded.

[0219] For example, the second device is an NE managed by the SMO. Based on the ML training request, it executes training of the first AI model, simultaneously calculating and recording the energy consumption of the first AI model during training. The NE managed by the SMO determines the energy consumption calculation method based on its own energy consumption calculation capability or based on the energy consumption calculation method identifier included in the ML training request, and uses this method to calculate the energy consumption generated during the training of the first AI model. Simultaneously, when the energy consumption recording identifier included in the ML training request is True, the energy consumption generated during the training of the first AI model is recorded.

[0220] The model inference process can be understood as the process of using / running the model. The energy consumption of inference can be the energy consumption generated by the AI ​​model during use / running. The second device can calculate the energy consumption generated during the use / running of the AI ​​model according to the energy consumption calculation method indicated by the energy consumption calculation method identifier or its own energy consumption calculation capability, and at the same time record the calculated energy consumption value according to the energy consumption record identifier.

[0221] It should be noted that the energy consumption calculation capability of the second device itself can be understood as the capability of its energy consumption calculation methods. For example, the second device may only have the capability of energy consumption calculation method 1, or only have the capability of energy consumption calculation method 2, or may have the capability of both energy consumption calculation methods 1 and 2 simultaneously. This depends on the second device itself, and different second devices have different energy consumption calculation capabilities, which this application does not limit. Different second devices can determine their energy consumption calculation methods based on their own energy consumption calculation capabilities, and use the determined energy consumption calculation methods to calculate the energy consumption consumed in the training, simulation, and inference phases.

[0222] It should also be noted that the energy consumption calculated and recorded above refers to the energy consumption of an AI computing device (e.g., the second device) when it performs training, simulation, and inference of a single AI model (e.g., the first AI model), or the incremental energy consumption (and prediction) when performing some baseline tasks within a certain period of time.

[0223] In one approach, the second device can acquire energy consumption #2, which determines energy consumption #1. Energy consumption #2 is the energy consumption generated by the second device during the simulation of the first AI model within time period #1. Energy consumption #2 is an example of a second energy consumption, and time period #1 is an example of a first time period.

[0224] For example, the second energy consumption may be the power consumed by a computing device (e.g., EMS or NE) when independently running the inference function of the first model. S330, the second device sends instruction information #1 or instruction information #2 to the first device. Correspondingly, the first device receives instruction information #1 or instruction information #2 from the first device.

[0225] In one approach, when it is determined that the first AI model will not be deployed to the inference function, the second device sends instruction information #1 to the first device, which includes the reason for not deploying the first AI model to the inference function. Accordingly, the first device receives instruction information #1 from the second device. Instruction information #1 is an example of a first instruction information.

[0226] For example, when the second device compares the energy consumption #1 indicated by information #1 with the threshold #1, if the energy consumption #1 is greater than the threshold #1, the second device determines that the first AI model will not be deployed to the inference function. At this time, it needs to send indication information #2 to report that the first energy consumption fails to meet the energy consumption requirements. In other words, through indication information #2, the first device knows that the first AI model has not been deployed. The reason for the non-deployment is that the energy consumption #1 is greater than the threshold #1, which is greater than the energy consumption value that the second device can receive.

[0227] In one approach, when it is determined that a first AI model will be deployed to the inference function, the second device sends instruction information #2 to the first device. Instruction information #2 indicates the result of the deployment, including whether the deployment was successful or failed. Accordingly, the first device receives instruction information #2 from the second device. Instruction information #2 is an example of a second instruction information.

[0228] For example, the first device is an NMS (Nearest Management System) and the second device is an EMS (Electronic Management System). The EMS sends deployment result information to the NMS. Correspondingly, the NMS receives the deployment result information from the EMS. This deployment result information is an example of instruction information #1 or instruction information #2. When deployment fails, the deployment result information also includes the reason for the deployment failure.

[0229] For example, when EMS determines that the first AI model will not be deployed to the inference function, the deployment result information is used to indicate deployment failure, and the deployment result information includes the reason for not deploying the first AI model to the inference function, or can be understood as the reason for deployment failure. It should be understood that not deploying the first AI model is also a result of deployment failure.

[0230] For example, when EMS determines to deploy the first AI model to the inference function, the deployment result information is used to indicate whether the deployment failed or succeeded. When deployment fails, the deployment result information also includes the reason for the failure. It should be understood that the deployment failure occurs when the first AI model was determined to be deployed but the deployment failed.

[0231] For example, the first device is an SMO (Service Management Center) and the second device is an NE (Application Provider). The SMO sends deployment result information to the NE. Correspondingly, the NE receives the deployment result information from the SMO. The NE can be managed by the SMO. When deployment fails, the deployment result information also includes the reason for the failure.

[0232] For example, when NE determines not to deploy the first AI model to the inference function, the deployment result information is used to indicate deployment failure, and the deployment result information includes the reason for not deploying the first AI model to the inference function, or can be understood as the reason for deployment failure. It should be understood that not deploying the first AI model is also a result of deployment failure.

[0233] For example, when the NE determines to deploy the first AI model to the inference function, the deployment result information is used to indicate whether the deployment failed or succeeded. When the deployment fails, the deployment result information also includes the reason for the failure. It should be understood that the deployment failure means that the first AI model was determined to be deployed but the deployment failed. It should be noted that indication information #1 and / or indication information #2 can be sent through the same information or through different information; this application does not limit this.

[0234] In one approach, when the deployment fails, the indication message #2 includes the reason for the deployment failure.

[0235] It should be noted that, to distinguish between the two states of non-deployment and deployment, two pieces of information are used to indicate the reason for non-deployment and the result of deployment. It should be understood that when the second device determines not to deploy, it can also be considered a deployment failure. Therefore, a single indication message can be used to indicate the deployment result, which includes deployment success or deployment failure. Deployment failure includes non-deployment or failure during the deployment process when deployment was initially determined. This result can be communicated to the first device through this indication message.

[0236] In one approach, the second device sends feedback information #1 to the first device regarding information #2. This feedback information #1 indicates energy consumption #2, which is the energy consumed by the second device during the simulation of the first AI model within time period #1. Correspondingly, the first device receives the feedback information #1 from the second device. Feedback information #1 is an example of a first feedback information, time period #1 is an example of a first time period, and energy consumption #2 is an example of a second energy consumption. Energy consumption #2 is used to determine energy consumption #1.

[0237] In one approach, the second device sends information #8 to the first device, indicating energy consumption #2, which is the energy consumed by the second device during the simulation of the first AI model in a first time period. Correspondingly, the first device receives information #8 from the second device. Feedback information #8 is an example of a first feedback message. This feedback message #8 is triggered by the second device itself.

[0238] For example, the second device sends a simulation report to the first device. Correspondingly, the first device receives a simulation report from the second device. The simulation report is an example of feedback information #1.

[0239] For example, in response to a simulation request, the second device sends a simulation report to the first device. The simulation report indicates energy consumption #2; that is, the second device feeds back the energy consumption (e.g., energy consumption #2) generated during the simulation of the first AI model within time period #1 to the first device. This simulation report can be an EmulationReport.

[0240] It should be noted that the energy consumption #2 can be the peak energy consumption obtained by the second device when simulating the first AI model during time period #1, or it can be the total energy consumption generated when simulating the first AI model, etc. This application does not limit it in this regard.

[0241] For example, the first device is an NMS and the second device is an EMS. In response to a simulation request, the EMS sends a simulation report to the NMS, which indicates the energy consumption (e.g., energy consumption #2) generated during the simulation of the first AI model in time period #1, calculated and recorded by the second device. Accordingly, the NMS receives the simulation report from the EMS.

[0242] For example, the first device is an NMS (Nearest Management System), and the second device is an NE (Energy Provider). In response to a simulation request, the NE sends a simulation report to the NMS. This report indicates the energy consumption (e.g., energy consumption #2) calculated and recorded by the second device during the simulation of the first AI model in time period #1. Accordingly, the NMS receives the simulation report from the NE. The NE can be managed by an EMS (Energy Management System).

[0243] For example, the first device is an SMO (System-Mounted Machine) and the second device is an NE (Electronic Element). In response to a simulation request, the NE sends a simulation report to the SMO. This report indicates the energy consumption (e.g., energy consumption #2) calculated and recorded by the second device during the simulation of the first AI model in time period #1. Correspondingly, the SMO receives the simulation report from the NE. The NE can be managed by the SMO.

[0244] It should be noted that the determination of energy consumption #2 to energy consumption #1 can be that energy consumption #1 is equal to energy consumption #2, energy consumption #1 is a certain proportion of energy consumption #2, or energy consumption #1 is a value slightly greater than energy consumption #2. This application does not limit this.

[0245] In one embodiment, the second device sends feedback information #2 to the first device in response to information #5. This feedback information #2 indicates energy consumption #3, which is the energy consumed by the second device during inference on the first AI model within time period #2. Correspondingly, the first device receives feedback information #2 from the second device. Feedback information #2 is an example of a second feedback information, time period #2 is an example of a second time period, and energy consumption #3 is an example of a third energy consumption.

[0246] It should be noted that energy consumption #3 can be the peak energy consumption obtained when inferring the first AI model during time period #2, or the total energy consumption generated when inferring the first AI model during time period #2, etc. This application does not limit it in this respect.

[0247] It should be understood that time period #2 can be after time period #1. Only after the second device determines the deployment of the first AI model to the inference function based on the energy consumption #1 determined by the energy consumption #2 obtained in the simulation phase and successfully deploys it can the second device execute the inference function of the first AI model. Therefore, time period #2 can be after time period #1.

[0248] For example, in response to ML inference history control information, the second device sends inference output to the first device. The inference output indicates energy consumption #3; that is, the second device feeds back to the first device the energy consumption (e.g., energy consumption #3) calculated and recorded during inference of the first AI model within time period #2. This inference output can be an InferenceOutput. An example of inference output being feedback information #2 is shown below.

[0249] For example, the second device sends inference output to the first device. Accordingly, the first device receives inference output from the second device.

[0250] For example, the first device is an NMS and the second device is an EMS. In response to ML inference history control information, the EMS feeds back inference output to the NMS. This inference output is used to indicate the energy consumption (e.g., energy consumption #3) generated when the second device calculates and records the inference of the first AI model during time period #2. Accordingly, the NMS receives the inference output from the EMS.

[0251] For example, the first device is an NMS (Network Management System), and the second device is an NE (Network Element). In response to ML inference history control information, the NE feeds back inference output to the NMS. This inference output indicates the energy consumption (e.g., energy consumption #3) generated during inference of the first AI model within time period #2, calculated and recorded by the second device. Accordingly, the NMS receives the inference output from the NE. The NE can be managed by an EMS (Electronic Management System).

[0252] For example, the first device is an SMO (System-Mounted Machine) and the second device is an NE (Element-Neural Network). In response to ML (Multi-Level Machine) inference history control information, the NE feeds back inference output to the SMO. This inference output indicates the energy consumption (e.g., energy consumption #3) generated during inference of the first AI model within time period #2, calculated and recorded by the second device. Correspondingly, the SMO receives the inference output from the NE. The NE can be managed by the SMO.

[0253] In one approach, the first device determines, based on energy consumption #3, whether to pause, terminate, suspend, or cancel the first device's inference on the first AI model.

[0254] For example, when the first device receives feedback information #2, its energy consumption requirement is a threshold #3, which is less than or equal to the maximum energy consumption supported by the second device. When the energy consumption #3 is less than or equal to the threshold #3, the first device determines to pause, terminate, suspend, or cancel the second device's inference on the first AI model. Alternatively, when the energy consumption #3 is greater than or equal to the threshold #3, the first device determines not to pause, terminate, suspend, or cancel the second device's inference on the first AI model.

[0255] It should be noted that the terms "pause," "termination," "suspend," or "cancel" used in this application can be replaced with other terms, such as "stop," "close," etc., and this application does not impose any restrictions.

[0256] In one approach, the feedback information #3 also includes an identifier for the energy consumption calculation method.

[0257] For example, the inference output also includes an identifier for the energy consumption calculation method.

[0258] In one embodiment, the second device sends feedback information #3 to the first device in response to information #6. This feedback information #3 indicates energy consumption #4, which is the energy consumed by the second device during inference on the first AI model within time period #3. Correspondingly, the first device receives feedback information #3 from the second device. Feedback information #3 is an example of a third feedback information, time period #3 is an example of a third time period, and energy consumption #4 is an example of a fourth energy consumption.

[0259] For example, the second device sends an ML inference history report to the first device. Correspondingly, the first device receives the ML inference history report from the second device. The ML inference history report is an example of feedback information #3.

[0260] For example, in response to an ML inference history request, the second device sends an ML inference history report to the first device. This ML inference history report indicates the energy consumption (e.g., energy consumption #4) incurred by the second device during inference on the first AI model in time period #3. This ML inference history report can be an MLInferenceHistoryReport.

[0261] It should be noted that the energy consumption #4 can be the peak energy consumption obtained during inference of the first AI model in time period #3, or the total energy consumption generated during inference of the first AI model in time period #3, etc. This application does not limit it in this way.

[0262] It should be noted that time period #3 can be any past / historical first AI model inference time period. For example, time period #3 can be a time period before time period #2, so that the historical inference energy consumption before time period #2 can be queried. Time period #3 can also be time period #2. When time period #3 is time period #2, energy consumption #4 is energy consumption #3. This application does not limit this.

[0263] For example, the first device is an NMS and the second device is an EMS. In response to an ML inference history request, the EMS sends an ML inference history report to the NMS. This ML inference history report indicates the energy consumption (e.g., energy consumption #4) generated by the second device when performing inference on the first AI model during time period #3. Accordingly, the NMS receives the ML inference history report from the EMS.

[0264] For example, the first device is an NMS and the second device is an NE. In response to an ML inference history request, the NE sends an ML inference history report to the NMS. This ML inference history report indicates the energy consumption (e.g., energy consumption #4) generated by the second device when inferring the first AI model during time period #3. Accordingly, the NMS receives the ML inference history report from the NE.

[0265] For example, the first device is an SMO and the second device is an NE. In response to an ML inference history request, the NE sends an ML inference history report to the SMO. This ML inference history report indicates the energy consumption (e.g., energy consumption #4) generated by the second device when performing inference on the first AI model during time period #3. Accordingly, the SMO receives the ML inference history report from the NE.

[0266] In one approach, the first device determines whether to enable the second device to perform inference on the first AI model based on energy consumption #4.

[0267] For example, when the first device receives feedback information #3, the energy consumption requirement of the second device is a threshold #4, which is less than or equal to the maximum energy consumption supported by the second device. When the energy consumption #3 is less than or equal to the threshold #4, the first device determines to enable the second device to perform inference on the first AI model. Alternatively, when the energy consumption #3 is greater than or equal to the threshold #4, the first device determines not to enable the second device to perform inference on the first AI model.

[0268] In one embodiment, the second device sends feedback information #4 to the first device in response to information #7. This feedback information #4 indicates energy consumption #5, which is the energy consumed by the second device during training the first AI model within time period #4. Correspondingly, the first device receives feedback information #4 from the second device. Feedback information #4 is an example of a fourth feedback information, time period #4 is an example of a fourth time period, and energy consumption #5 is an example of a fifth energy consumption.

[0269] It should be noted that energy consumption #5 can be the peak energy consumption obtained when training the first AI model during time period #4, or the total energy consumption generated when training the first AI model during time period #4, etc. This application does not limit it in this regard.

[0270] For example, in response to an ML training report request, the second device sends an ML training report to the first device. The ML training report indicates energy consumption #5; that is, the second device feeds back the energy consumption (e.g., energy consumption #5) generated during the training of the first AI model within time period #4 to the first device. This ML training report can be an MLTrainingReport. An example of feedback information #4 is provided in the ML training report.

[0271] For example, the second device sends an ML training report to the first device. Accordingly, the first device receives the ML training report from the second device.

[0272] For example, the first device is an NMS and the second device is an EMS. In response to an ML training report request, the EMS sends an ML training report to the NMS. This ML training report is used to indicate the energy consumption (e.g., energy consumption #5) generated during the training of the first AI model in time period #4, calculated and recorded by the second device. Accordingly, the NMS receives the ML training report from the EMS.

[0273] For example, the first device is an NMS (Nearest Management System) and the second device is an NE (Energy Provider). In response to an ML (Multi-Level Machine) training report request, the NE sends an ML training report to the NMS. This ML training report indicates the energy consumption (e.g., energy consumption #5) calculated and recorded by the second device during training the first AI model in time period #4. Accordingly, the NMS receives the ML training report from the NE. The NE can be managed by an EMS (Energy Management System).

[0274] For example, the first device is an SMO, and the second device is an NE. In response to an ML training report request, the NE sends an ML training report to the SMO. This ML training report indicates the energy consumption (e.g., energy consumption #5) calculated and recorded by the second device during training the first AI model in time period #4. Correspondingly, the SMO receives the ML training report from the NE. The NE can be managed by the SMO.

[0275] In one approach, the first device determines whether to pause training of the first AI model based on energy consumption #5.

[0276] For example, when the first device receives feedback information #4, the energy consumption requirement of the second device is a threshold #4, which is less than or equal to the maximum energy consumption supported by the second device. When the energy consumption #5 is less than or equal to the threshold #4, the first device determines to pause the training of the first AI model by the second device. Alternatively, when the energy consumption #5 is greater than or equal to the threshold #4, the first device determines not to pause the training of the first AI model by the second device.

[0277] For example, if the AI ​​model training consists of 100 rounds, and the energy consumption of ML training is reported every 10 rounds, the second device can estimate the total energy consumption of 100 rounds of training based on the energy consumption of the first 10 rounds. The second device reports the estimated total energy consumption of 100 rounds of training. The first device then determines whether to pause the ML training process based on the current energy consumption requirement, or the current energy consumption level. For example, if the current energy consumption level is 1, and the energy consumption of level 1 is 500 joules (an example of threshold #4), and the reported estimated energy consumption of 100 rounds is 600 joules (an example of energy consumption #5), this exceeds the energy consumption value of the current energy consumption level. Therefore, the current ML training process needs to be paused. In other words, because the estimated energy consumption of 100 rounds is higher than the current energy consumption level, the next 90 rounds of training need to be paused.

[0278] It's important to note that this ML / AI training process is only paused. Energy requirements / levels vary at different times; for example, the energy requirement might be 500 joules in the morning and 800 joules at night. This can be understood as the threshold #4 changing over time. Therefore, at night, due to the changed energy requirements, and since the total energy needed for ML / AI training is only 600 joules, lower than the nighttime requirement, the ML / AI training process can resume at the current moment. Similarly, the ML / AI inference process is also only paused and can be resumed when energy requirements are met.

[0279] In one approach, the feedback information #5 also includes an identifier for the energy consumption calculation method.

[0280] For example, the ML training report also includes an identifier of the energy consumption calculation method.

[0281] In one approach, the second device executes the inference function using a strategy during runtime.

[0282] For example, after the first AI model is successfully deployed, if the second device exceeds the threshold #2 when the inference function of the first AI model is executed, as specified in the policy, the inference function of the first AI model is stopped, turned off, canceled, or suspended.

[0283] For example, if the inference energy consumption of the inference process is 600 joules per unit time, and the inference energy consumption per unit time specified in the usage strategy of the inference function is higher than 500 joules (an example of threshold #2), then the inference function needs to be shut down. Therefore, the inference function needs to be shut down within that unit time. If the inference energy consumption per unit time specified in the usage strategy is higher than 700 joules (another example of threshold #2), then the inference function needs to be shut down. Therefore, the inference function does not need to be shut down within that unit time.

[0284] In one possible implementation, the second device is an EMS, which executes the usage strategy of the inference function during runtime.

[0285] In one possible implementation, the second device is an NE, which executes the usage strategy of the inference function during runtime.

[0286] It should be noted that the thresholds #1 to #4 mentioned above may be the same or different, and this application does not limit this.

[0287] It should be noted that the information sent from the first device to the second device in this application can be sent directly or forwarded through a third device. For example, when the first device is an NMS and the second device is an NE, the NMS can send the ML training request to the NE directly or forward it to the NE through an EMS. This application does not limit this.

[0288] It should be noted that the energy consumption values ​​#1 to #5 mentioned above may be the same or different, and this application does not limit this. Energy consumption #2 and #5 can be obtained not only through information instructing the second device to calculate, record, or query the energy consumption generated when processing the first AI model (e.g., training, simulation, inference), but also through other means, such as the second device itself triggering the calculation, recording, or querying of the energy consumption generated when processing the first AI model, and actively reporting the energy consumption, etc., and this application does not limit this either.

[0289] It should also be noted that since the first AI model is pre-trained during simulation, and the second device sends feedback information #4 to the first device after training the first AI model, this feedback information #4 includes an energy consumption calculation method identifier. Therefore, in subsequent stages such as simulation, inference, deployment, and historical inference energy consumption query, the information sent by the first device to the second device may or may not include this energy consumption calculation method identifier. For example, information #2 may include this energy consumption calculation method identifier. Furthermore, the information sent by the second device to the first device may or may not include the energy consumption calculation method identifier. For example, feedback information #1 may include the energy consumption calculation method identifier. This application does not impose any limitations on this.

[0290] Figure 4 This is a schematic flowchart of an artificial intelligence model energy management method 400 provided in an embodiment of this application.

[0291] Figure 4 for Figure 3 One specific embodiment. The following is in conjunction with... Figure 4An artificial intelligence model energy consumption management method 400 is presented, wherein the first device is an NMS and the second device is an EMS. In this embodiment, network service management, model training, and inference capabilities of the NMS are applied in an EMS scenario. Method 400 includes steps S405 to S485, which are described in detail below.

[0292] Optionally, in step S405, when model training is triggered by the NMS, the NMS sends an ML training request to the EMS. Correspondingly, the EMS receives the ML training request from the NMS.

[0293] The ML training request includes an energy consumption record identifier. This identifier indicates whether the EMS records the energy consumption generated during the training of the first AI model.

[0294] Specifically, this training request is used to instruct ML to perform initial training and / or retraining. The ML training request can be an MLTrainingRequest.

[0295] S405 is a specific instance of S310 in method 300, and a detailed description can be found in step S310.

[0296] The S410 and EMS perform training on the first AI model, while simultaneously measuring, calculating, and recording the energy consumption during the model training process.

[0297] Specifically, the EMS executes the training of the first AI model based on the ML training request. It determines the energy consumption calculation method based on its own energy consumption computing capabilities, and uses this method to calculate the energy consumed / generated during training, obtaining the energy consumption value generated during the training of the first AI model. When the energy consumption recording flag included in the ML training request is True, meaning the energy consumption recording flag indicates that training energy consumption should be recorded, the EMS records the energy consumption of the first AI model training process while simultaneously measuring and calculating it.

[0298] S410 is a specific instance of S320 in method 300, and a detailed description can be found in step S320. S415: EMS sends the ML training report to NMS. Correspondingly, NMS receives the ML training report from EMS.

[0299] Specifically, in response to an ML training request, the EMS sends an ML training report to the NMS. This report includes Energy Consumption #5, which represents the energy consumption value (or the quantity of energy consumed during training the first AI model) generated by the EMS during time period #4, and an energy consumption calculation method identifier. In other words, the EMS feeds back the calculated and recorded energy consumption value generated during the training of the first AI model to the NMS. The energy consumption calculation method identifier is used to report the method used by the EMS to calculate the energy consumption value of the first AI model during training. This ML training report can be named MLTrainingReport.

[0300] Based on the energy consumption value and energy consumption calculation method generated during the training of the first reported AI model, NMS determines the overall training energy consumption of the model, and determines whether to pause, terminate, suspend or cancel the training process based on the current energy consumption requirements (or energy consumption level).

[0301] From here on, S405 to S415 constitute the ML training phase. During the ML training phase, the training energy consumption of the AI ​​model is estimated and reported. This helps NMS determine whether to pause, terminate, suspend, or cancel the training process based on the training energy consumption value and energy consumption requirements (or energy consumption level), effectively avoiding the risk of excessive energy consumption during AIML training.

[0302] S415 is a specific instance of S330 in method 300, and a detailed description can be found in step S330. S420: The NMS sends a simulation request to the EMS. Correspondingly, the EMS receives the simulation request from the NMS.

[0303] The simulation request includes an energy consumption calculation method identifier and an energy consumption recording identifier. The energy consumption recording identifier indicates whether the EMS records the energy consumption generated during the simulation of the first AI model. The energy consumption calculation method identifier indicates the method used to calculate the energy consumption generated during the simulation of the first AI model. This simulation request can be an EmulationRequest.

[0304] The simulation request also indicates the type of inference function that EMS should execute for the first AI model. This allows the first AI model to be simulated on the corresponding type of inference function, and also facilitates determining whether to deploy the first AI model on the corresponding type of inference function based on its type.

[0305] S420 is a specific instance of S310 in method 300, and a detailed description can be found in step S310.

[0306] The S425 and EMS perform model simulations, while simultaneously measuring and recording the energy consumption during the model simulation process.

[0307] Specifically, the EMS executes the simulation of the first AI model based on the simulation request. The EMS uses the energy consumption calculation method indicated by the energy consumption calculation method identifier to calculate the energy consumption generated during the simulation of the first AI model, obtaining the energy consumption value generated during the simulation. When the energy consumption record identifier included in the simulation request is True, that is, when the energy consumption record identifier indicates that the energy consumption value generated during the simulation of the first AI model should be recorded, the EMS records the energy consumption value generated during the simulation of the first AI model while simultaneously measuring and calculating it.

[0308] S425 is a specific instance of S320 in method 300, and a detailed description can be found in step S320.

[0309] S430 and EMS send simulation reports to NMS. Correspondingly, NMS receives simulation reports from EMS.

[0310] Specifically, in response to the simulation request, the EMS sends a simulation report to the NMS. This report includes Energy Consumption #2, which represents the energy consumption value generated by the EMS during the simulation of the first AI model within time period #1 (or the quantity of energy consumed during the simulation of the first AI model). In other words, the EMS feeds back the calculated and recorded energy consumption value generated during the simulation of the first AI model to the NMS. This simulation report can be an EmulationReport. Energy Consumption #2 can be used to determine Energy Consumption #1.

[0311] S430 is a specific instance of S330 in method 300, and a detailed description can be found in step S320. Up to this point, S420 to S430 constitute the ML simulation phase. During the ML simulation phase, the simulation energy consumption of the AI ​​model is estimated and reported so that the simulation energy consumption of the first AI model can be indicated to the deployment model during the subsequent deployment phase. The simulation energy consumption of the first AI model can be determined based on the energy consumption value generated during the simulation of the first AI model included in the simulation report (e.g., equal to or less than the energy consumption value generated during the simulation of the first AI model). The deployment device can choose whether to use the AI ​​function based on the indicated simulation energy consumption of the first AI model and its own energy-saving level (or energy consumption level). Simulation energy consumption can be understood as the energy consumption value generated by the simulated first AI model during subsequent operation / inference. Based on this simulated energy consumption value, during the deployment phase, it can be determined whether the inference function of the first AI model needs to be enabled / used in the deployment model, or whether the first AI model needs to be deployed to the inference function.

[0312] Optionally, in step S435, when deployment is triggered by EMS, NMS sends the ML entity loading policy to EMS. Correspondingly, EMS receives the ML entity loading policy from NMS.

[0313] The ML entity loading policy includes a simulation energy consumption indicator, which indicates energy consumption #1. The ML entity loading policy can be MLEntityLoadingPolicy. This simulation energy consumption value can be determined based on energy consumption #2 included in the simulation report in S430.

[0314] The ML entity loading strategy is also used to indicate threshold #1, or in other words, the ML entity loading strategy is also used to indicate energy demand / energy level, which corresponds to threshold #1.

[0315] S435 is a specific instance of S310 in method 300, and a detailed description can be found in step S310. Optionally, in S440, when the deployment is triggered by the NMS, the NMS sends an ML entity loading request to the EMS. Accordingly, the EMS receives the ML entity loading request from the NMS.

[0316] The ML entity loading request includes a simulation energy consumption indicator, which indicates energy consumption #1. The ML entity loading request can be an MLEntityLoadingRequest. This simulation energy consumption value can be determined based on energy consumption #2 included in the simulation report in S430.

[0317] The ML entity load request is also used to indicate threshold #1, or in other words, the ML entity load request is also used to indicate energy demand / energy level corresponding to threshold #1.

[0318] S440 is a specific instance of S310 in method 300, and a detailed description can be found in step S310.

[0319] S442 and EMS determine whether to deploy the first AI model to the inference function based on the energy consumption value and energy consumption requirements of the model simulation stage indicated by the simulation energy consumption indicator.

[0320] Specifically, EMS determines whether to deploy the first AI model to the inference function based on the relationship between the energy consumption value of the model simulation phase indicated by the simulation energy consumption indicator and the threshold #1.

[0321] This energy consumption requirement corresponds to the energy consumption value that the second device can accept, such as threshold #1.

[0322] For example, the simulation request also indicates the type of inference function. When the energy consumption value during the model simulation phase is less than or equal to threshold #1, the EMS determines to deploy the first AI model in the NE to the inference function corresponding to the type indicated by the simulation request. When the energy consumption value during the model simulation phase is greater than or equal to threshold #1, the EMS determines not to deploy the first AI model in the NE to the inference function corresponding to the type indicated by the simulation request. This NE can be managed by the EMS.

[0323] For example, the currently deployed NE has an energy efficiency level of 1, with a maximum energy consumption of 500 joules. When the simulation energy consumption indicator issued by the NMS shows that the simulation energy consumption value of the first AI model is higher than 500 joules, the first AI model will not be used during NE deployment, or its inference function will not be enabled, or it will not be deployed to the inference function. If the simulation energy consumption indicator issued by the first device shows that the simulation energy consumption value of the first AI model is lower than or equal to 500 joules, the EMS can deploy the first AI model to the inference function on the NE. This NE can be managed by the EMS.

[0324] S442 is a specific instance of S320 in method 300, and a detailed description can be found in step S320.

[0325] It's important to note that, similar to the training phase, the energy consumption requirements of the NE (Network Array) may vary at different times, thus its energy efficiency level will change in real time. Therefore, the EMS (Energy Management System) can determine whether to use the first AI model, or whether to deploy the first AI model to the inference function on the NE, based on its real-time energy consumption and the simulation energy consumption value indicated by the simulation energy consumption indicator. Furthermore, because different AI models may generate different energy consumptions during operation, the simulation energy consumption values ​​of different AI models may also differ. For example, a high-precision AI model may generate more energy consumption, while a low-precision AI model may generate less. Similarly, the energy consumption generated during the operation of AI models trained by different optimizers may also differ. Therefore, the calculated and recorded energy consumption represents the energy loss of the AI ​​computing device when performing training, simulation, and inference on a single AI model. Thus, for AI models that may generate different energy consumptions, during the deployment phase, the EMS can determine whether to use the current AI model, or enable the current AI model's function, or deploy the current AI model to the inference function, based on each AI model that may generate different energy consumptions and the EMS's current energy consumption requirements.

[0326] S445. The EMS sends deployment result information to the NMS, which indicates whether the deployment was successful or failed. Correspondingly, the NMS receives the deployment result information from the EMS.

[0327] In one approach, when deployment fails, the deployment result information also includes the reason for the deployment failure.

[0328] For example, when EMS determines that the first AI model will not be deployed to the inference function, the deployment result information is used to indicate deployment failure, and the deployment result information includes the reason for not deploying the first AI model to the inference function, or can be understood as the reason for deployment failure. It should be understood that not deploying the first AI model is also a result of deployment failure.

[0329] For example, when EMS determines to deploy the first AI model to the inference function, the deployment result information is used to indicate whether the deployment failed or succeeded. When deployment fails, the deployment result information also includes the reason for the failure. It should be understood that the deployment failure occurs when the first AI model was determined to be deployed but the deployment failed.

[0330] S445 is a specific instance of S330 in method 300, and a detailed description can be found in step S330.

[0331] S435 to S445 constitute the ML deployment phase. During the ML deployment phase, simulation energy consumption is incorporated into the deployment information. The second device can deploy the AI ​​model based on the simulation energy consumption indication issued by the first device. The EMS can decide whether to deploy the AI ​​model to the inference function based on the simulation energy consumption indication and its own energy-saving level. This effectively avoids deploying AI models with high energy consumption during the AIML deployment phase, and also effectively avoids the risk of excessive energy consumption of the AI ​​model during subsequent operation.

[0332] The following steps are the steps to be performed when the first AI model is determined to be deployed to the inference function and when the first AI model is successfully deployed to the inference function.

[0333] S450 and NMS send ML inference history control information to EMS. Correspondingly, EMS receives ML inference history control information from NMS.

[0334] The ML inference history control information includes an energy consumption record identifier and an energy consumption calculation identifier. The energy consumption record identifier indicates whether the EMS records the energy consumption during the inference process of the first AI model. The energy consumption calculation method identifier indicates the method used to calculate the energy consumption during the inference process of the first AI model. This ML inference history control information can be MLInferenceHistoryControl. This ML inference history control information can be used to control the energy consumption calculation and recording method during the inference process of the first AI model.

[0335] S450 is a specific instance of S310 in method 300, and a detailed description can be found in step S310.

[0336] S455 and EMS perform model inference while simultaneously measuring, calculating, and recording the energy consumption during the model inference process.

[0337] In one approach, the EMS executes the inference function of the first AI model based on ML inference history control information. The EMS determines the energy consumption calculation method based on the energy consumption calculation method identifier included in the ML inference history control information or based on its own energy consumption calculation capabilities, and calculates the energy consumption generated during the inference process to obtain the energy consumption value generated by the first AI model during inference. When the energy consumption record identifier included in the ML inference history control information is True, that is, when the energy consumption record identifier indicates that the energy consumption generated during the inference of the first AI model should be recorded, the EMS records the energy consumption generated by the first AI model during inference while simultaneously measuring and calculating the energy consumption during the inference process.

[0338] S455 is a specific instance of S320 in method 300, and a detailed description can be found in step S320.

[0339] S460 and EMS send inference output to NMS. Correspondingly, NMS receives inference output from EMS.

[0340] The inference output includes energy consumption #3, which represents the energy consumed by the EMS during inference of the first AI model within time period #2. This inference output can be InferenceOutput.

[0341] The NMS determines whether to pause, terminate, suspend, or cancel the inference process based on the energy consumption value generated during inference by the first reported AI model and the current energy consumption requirements (or energy consumption level).

[0342] S460 is a specific instance of S330 in method 300, and a detailed description can be found in step S330.

[0343] S455 to S460 is the ML inference stage. During the ML inference stage, the second device can record and report the inference energy consumption of the AI ​​model. The first device manages the inference process so that the first device can generate the AIML inference function usage strategy based on the inference energy consumption. The first device can also manage the AI ​​model usage strategy of the second device function based on the historical energy consumption information of the second device.

[0344] S465, NMS sends an ML inference history request to EMS. Correspondingly, EMS receives the ML inference history request from NMS.

[0345] This ML inference history request is used to query the energy consumption values ​​generated during the historical inference process of the first AI model. This ML inference history request can be MLInferenceHistoryRequest.

[0346] S465 is a specific instance of S310 in method 300, and a detailed description can be found in step S310.

[0347] S470, EMS sends ML inference history reports to NMS. Correspondingly, NMS receives ML inference history reports from EMS.

[0348] Specifically, in response to the ML inference history request, EMS sends an ML inference history report to NMS. This ML inference history report includes a model inference energy consumption indicator and an energy consumption calculation identifier. The model inference energy consumption indicator indicates energy consumption #4, which is the energy consumption value generated by EMS during inference of the first AI model within time period #3. Time period #3 can be a specific time period in the historical inference of the first AI model. This ML inference history report can be an MLInferenceHistoryReport.

[0349] Based on the energy consumption values ​​generated by the first AI model during historical inference, and according to the current energy consumption requirements (or energy consumption level), NMS determines whether to enable EMS to perform inference on the first AI model.

[0350] It is understandable that S465 and S470 are the historical query and control processes, in which the inference energy consumption value obtained during the historical inference process of the model can be queried to obtain the energy consumption value consumed by the AI ​​model during use / inference in a certain period of time in the past.

[0351] S470 is a specific instance of S330 in method 300, and a detailed description can be found in step S330. S475, the usage strategy of NMS to generate AIML inference function.

[0352] Specifically, NMS can generate a usage strategy for AIML inference functions based on the energy consumption during inference by the first AI model.

[0353] For example, the usage strategy is to stop / close / cancel / suspend the inference function of the first AI model if the energy consumption of the EMS per unit time is higher than the threshold #2.

[0354] S475 is a specific instance of S330 in method 300, and a detailed description can be found in step S330.

[0355] S480 and NMS distribute the AIML inference function usage policy to EMS. Correspondingly, EMS receives the inference function usage policy from NMS.

[0356] S480 is a specific instance of S310 in method 300, and a detailed description can be found in step S310. S485, the usage strategy for executing this inference function during EMS runtime.

[0357] For example, if the inference energy consumption during the first unit of time is 600 joules, and the inference energy consumption during the first unit of time specified in the policy is higher than 500 joules (an example of threshold #2), then the inference function needs to be turned off / stopped / cancelled / suspended. Therefore, the inference function needs to be turned off / stopped / cancelled / suspended in the first unit of time. If the inference energy consumption during the unit of time specified in the policy is higher than 700 joules, then the inference function needs to be turned off / stopped / cancelled / suspended. In this case, the inference function does not need to be turned off / stopped / cancelled / suspended in the first unit of time.

[0358] S485 is a specific instance of S330 in method 300, and a detailed description can be found in step S330.

[0359] Thus, the first device can generate a usage strategy for AIML inference functionality based on inference energy consumption. The second device can execute this usage strategy during operation, allowing the first device to manage the energy consumption generated by the first AI model running on the second device. This prevents the first AI model from generating excessive energy consumption that the deployment device's energy efficiency rating cannot support, allowing it to promptly shut down, stop, cancel, or suspend the inference function of the first AI model. When the deployment device's own energy consumption demand increases, the AI ​​model can continue running, achieving real-time monitoring of energy consumption during AI model usage and effectively avoiding the risk of excessive energy consumption during AI model operation.

[0360] Based on Method 400, the EMS can estimate and report the energy consumption of AI models during the training, simulation, and inference phases. The EMS can deploy AI models based on the simulation energy consumption indicators issued by the NMS. Deploying network elements can decide whether to use the AI ​​model or deploy it to the inference function based on the simulation energy consumption indicators and their own energy-saving levels. Furthermore, the EMS can record and report AI model energy consumption during inference. The NMS manages this process and generates a usage policy for the AIML inference function based on the inference energy consumption. The EMS executes this policy during operation. This execution of the policy by the EMS allows the NMS to manage the energy consumption generated by the AI ​​model during operation, preventing situations where the AI ​​model generates significant energy consumption that the EMS's energy consumption level cannot support. The EMS can promptly shut down / stop / cancel / suspend the AI ​​model's use when the model's energy consumption increases. When the EMS's energy demand increases, the AI ​​model can continue running, achieving real-time monitoring of energy consumption during AI model use and effectively avoiding the risk of excessive energy consumption during AI model operation.

[0361] Figure 5 This is a schematic flowchart of an artificial intelligence model energy management method 500 provided in an embodiment of this application.

[0362] Figure 5 for Figure 3 One specific embodiment. The following is in conjunction with... Figure 5 An artificial intelligence model energy consumption management method 500 is provided, wherein the first device is an NMS (Network Service Management System), the second device is an EMS (Electronic Service Management System) during the AI ​​model training and simulation phase, and the second device is an NE (Network Inference System) during the AI ​​model inference phase. This embodiment addresses a scenario where the NMS manages network services, model training is performed on the EMS, and model inference is conducted on the NE (e.g., gNB, NWDAF, etc.). The NE can be managed by the EMS. Method 500 includes steps S505 to S585, which are described in detail below.

[0363] Optionally, in step S505, when model training is triggered by the NMS, the NMS sends an ML training request to the EMS. Correspondingly, the EMS receives the ML training request from the NMS.

[0364] The ML training request includes an energy consumption recording identifier. This identifier indicates whether the EMS records the energy consumption generated during the training of the first AI model. The ML training request can be MLTrainingRequest.

[0365] S505 is a specific instance of S310 in method 300, and a detailed description can be found in step S310. In S510, the EMS performs training of the first AI model, while measuring, calculating and recording the energy consumption during the model training process.

[0366] Specifically, the EMS executes the training of the first AI model based on the ML training request. It determines the energy consumption calculation method based on its own energy consumption computing capabilities, and uses this method to calculate the energy consumed / generated during training, obtaining the energy consumption value generated during the training of the first AI model. When the energy consumption recording flag included in the ML training request is True, meaning the energy consumption recording flag indicates that training energy consumption should be recorded, the EMS records the energy consumption of the first AI model training process while simultaneously measuring and calculating it.

[0367] S510 is a specific instance of S320 in method 300, and a detailed description can be found in step S320.

[0368] S515 and EMS send ML training reports to NMS. Correspondingly, NMS receives ML training reports from EMS.

[0369] Specifically, in response to the ML training request, EMS sends an ML training report to NMS. This report includes Energy Consumption #5, which represents the energy consumption value (or the quantity of energy consumed during training the first AI model) generated by EMS during time period #4, and an energy consumption calculation method identifier. In other words, EMS feeds back the calculated and recorded energy consumption value generated during the training of the first AI model to NMS. The energy consumption calculation method identifier is used to report to NMS the method used by EMS to calculate the energy consumption value of the first AI model during training. This ML training report can be named MLTrainingReport.

[0370] Based on the energy consumption value generated during the training of the first reported AI model, NMS determines whether to pause, terminate, suspend, or cancel the training process according to the current energy consumption requirements (or energy consumption level).

[0371] S515 is a specific instance of S330 in method 300, and a detailed description can be found in step S330.

[0372] The S520 and NMS send simulation requests to the EMS. Correspondingly, the EMS receives simulation requests from the NMS.

[0373] The simulation request includes an energy consumption calculation method identifier and an energy consumption recording identifier. The energy consumption recording identifier indicates whether the EMS records the energy consumption generated during the simulation of the first AI model. The energy consumption calculation method identifier indicates the method used to calculate the energy consumption generated during the simulation of the first AI model. This simulation request can be an EmulationRequest.

[0374] The simulation request also indicates the type of inference function that EMS should execute for the first AI model. This allows the first AI model to be simulated on the corresponding type of inference function, and also facilitates determining whether to deploy the first AI model on the corresponding type of inference function based on its type.

[0375] S520 is a specific instance of S310 in method 300, and a detailed description can be found in step S310.

[0376] The S525 and EMS perform model simulations, while simultaneously measuring and recording the energy consumption during the model simulation process.

[0377] Specifically, the EMS executes the simulation of the first AI model based on the simulation request. The EMS uses the energy consumption calculation method indicated by the energy consumption calculation method identifier to calculate the energy consumption generated during the simulation of the first AI model, obtaining the energy consumption value generated during the simulation. When the energy consumption record identifier included in the simulation request is True, that is, when the energy consumption record identifier indicates that the energy consumption value generated during the simulation of the first AI model should be recorded, the EMS records the energy consumption value generated during the simulation of the first AI model while simultaneously measuring and calculating it.

[0378] S525 is a specific instance of S320 in method 300; a detailed description can be found in step S320. S530: The EMS sends a simulation report to the NMS. Correspondingly, the NMS receives the simulation report from the EMS.

[0379] Specifically, in response to the simulation request, the EMS sends a simulation report to the NMS. This report includes Energy Consumption #2, which represents the energy consumption value generated by the EMS during the simulation of the first AI model within time period #1 (or the quantity of energy consumed during the simulation of the first AI model). In other words, the EMS feeds back the calculated and recorded energy consumption value generated during the simulation of the first AI model to the NMS. This simulation report can be an EmulationReport. Energy Consumption #2 can be used to determine Energy Consumption #1.

[0380] S530 is a specific instance of S330 in method 300, and a detailed description can be found in step S320.

[0381] Optionally, the S535 and NMS send the ML entity loading strategy to the EMS. Correspondingly, the EMS receives the ML entity loading strategy from the NMS.

[0382] The ML entity loading policy includes a simulation energy consumption indicator, which indicates energy consumption #1. The ML entity loading policy can be MLEntityLoadingPolicy. This simulation energy consumption value can be determined based on energy consumption #2 included in the simulation report in S530.

[0383] The ML entity loading strategy is also used to indicate threshold #1, or in other words, the ML entity loading strategy is also used to indicate energy demand / energy level, which corresponds to threshold #1.

[0384] S535 is a specific instance of S310 in method 300, and a detailed description can be found in step S310.

[0385] Optionally, in step S540, when deployment is triggered by the NMS, the NMS sends an ML entity loading request to the EMS. Correspondingly, the EMS receives the ML entity loading request from the NMS.

[0386] The ML entity loading request includes a simulation energy consumption indicator, which indicates energy consumption #1. The ML entity loading request can be an MLEntityLoadingRequest. This simulation energy consumption value can be determined based on energy consumption #2 included in the simulation report in S530.

[0387] The ML entity load request is also used to indicate threshold #1, or in other words, the ML entity load request is also used to indicate energy demand / energy level corresponding to threshold #1.

[0388] S540 is a specific instance of S310 in method 300, and a detailed description can be found in step S310.

[0389] S542 and EMS determine whether to deploy the first AI model to the inference function based on the energy consumption value and energy consumption requirements of the model simulation stage indicated by the simulation energy consumption indicator.

[0390] Specifically, EMS determines whether to deploy the first AI model to inference function in NE based on the relationship between the energy consumption value of the model simulation phase indicated by the simulation energy consumption indicator and the threshold #1.

[0391] This energy consumption requirement corresponds to the energy consumption value that the second device can accept, such as threshold #1.

[0392] For example, the simulation request also indicates the type of inference function. When the energy consumption value during the model simulation phase is less than or equal to threshold #1, the EMS determines to deploy the first AI model in the NE to the inference function corresponding to the type indicated by the simulation request. When the energy consumption value during the model simulation phase is greater than or equal to threshold #1, the EMS determines not to deploy the first AI model in the NE to the inference function corresponding to the type indicated by the simulation request. This NE can be managed by the EMS.

[0393] S542 is a specific instance of S320 in method 300, and a detailed description can be found in step S320. S545: The EMS sends deployment result information to the NMS, which indicates whether the deployment was successful or failed. Correspondingly, the NMS receives the deployment result information from the EMS.

[0394] In one approach, when deployment fails, the deployment result information also includes the reason for the deployment failure.

[0395] For example, when it is determined that the first AI model will not be deployed to the inference function, the deployment result information is used to indicate deployment failure, and the deployment result information includes the reason for not deploying the first AI model to the inference function, or can be understood as the reason for deployment failure. It should be understood that not deploying the first AI model is also a result of deployment failure.

[0396] For example, when it is determined to deploy the first AI model to the inference function, the deployment result information is used to indicate whether the deployment failed or succeeded. When the deployment fails, the deployment result information also includes the reason for the failure. It should be understood that the deployment failure means that the first AI model was determined to be deployed but the deployment failed.

[0397] S545 is a specific instance of S330 in method 300, and a detailed description can be found in step S330.

[0398] The S550 and NMS send ML inference history control information to the NE. Correspondingly, the NE receives the ML inference history control information from the NMS.

[0399] Optionally, the NMS sends ML inference history control information to the NE via the EMS, and correspondingly, the NE receives the ML inference history control information forwarded by the EMS.

[0400] The ML inference history control information includes an energy consumption record identifier and an energy consumption calculation identifier. The energy consumption record identifier indicates whether the NE records the energy consumption during the inference process of the first AI model. The energy consumption calculation method identifier indicates the method used to calculate the energy consumption during the inference process of the first AI model. This ML inference history control information can be MLInferenceHistoryControl. This ML inference history control information can be used to control the energy consumption calculation and recording method during the inference process of the first AI model.

[0401] S550 is a specific instance of S310 in method 300, and a detailed description can be found in step S310.

[0402] S555 and NE perform model inference while simultaneously measuring, calculating, and recording the energy consumption during the model inference process.

[0403] In one approach, the NE executes the inference function of the first AI model based on ML inference history control information. The NE determines the energy consumption calculation method based on the energy consumption calculation method identifier included in the ML inference history control information or based on its own energy consumption calculation capabilities, and calculates the energy consumption generated during the inference process to obtain the energy consumption value generated by the first AI model during inference. When the energy consumption record identifier included in the ML inference history control information is True, that is, when the energy consumption record identifier indicates that the energy consumption generated during the first AI model's inference should be recorded, the NE records the energy consumption generated by the first AI model during inference while simultaneously measuring and calculating the energy consumption during the first AI model's inference.

[0404] S555 is a specific instance of S320 in method 300, and a detailed description can be found in step S320.

[0405] S560 and NE send inference output to NMS. Correspondingly, NMS receives inference output from NE.

[0406] The NE sends inference output to the NMS via the EMS, and the NMS receives the inference output forwarded by the EMS accordingly.

[0407] The inference output includes energy consumption #3, which represents the energy consumed by the EMS during inference of the first AI model within time period #2. This inference output can be InferenceOutput.

[0408] The NMS determines whether to pause, terminate, suspend, or cancel the inference process based on the energy consumption value generated during inference by the first reported AI model and the current energy consumption requirements (or energy consumption level).

[0409] S560 is a specific instance of S330 in method 300, and a detailed description can be found in step S330.

[0410] S565, NMS sends an ML inference history request to NE. Correspondingly, NE receives the ML inference history request from NMS.

[0411] The NMS sends inference history requests to the NE via the EMS, and the NE receives the inference history requests forwarded by the EMS accordingly.

[0412] This ML inference history request is used to query the energy consumption values ​​generated during the historical inference process of the first AI model. This ML inference history request can be MLInferenceHistoryRequest.

[0413] S565 is a specific instance of S310 in method 300, and a detailed description can be found in step S310.

[0414] S570, NE sends ML inference history reports to NMS. Correspondingly, NMS receives ML inference history reports from NE.

[0415] The NE sends the inference history report to the NMS via the EMS, and the NMS receives the inference history report forwarded by the EMS accordingly.

[0416] Specifically, in response to an ML inference history request, the NE sends an ML inference history report to the NMS. This ML inference history report includes a model inference energy consumption indicator and an energy consumption calculation identifier. The model inference energy consumption indicator is used to indicate energy consumption #4, which is the energy consumption value generated by the EMS during inference of the first AI model within time period #3. Time period #3 can be a specific time period in the historical inference of the first AI model. This ML inference history report can be an MLInferenceHistoryReport.

[0417] Based on the energy consumption value generated by the first AI model during historical inference, and according to the current energy consumption requirements (or energy consumption level), the NMS determines whether to enable NE to infer the first AI model.

[0418] S570 is a specific instance of S330 in method 300, and a detailed description can be found in step S330.

[0419] S575, NMS generation AIML inference function usage strategy.

[0420] Specifically, NMS can generate usage strategies for AIML inference functions based on inference energy consumption.

[0421] For example, the usage strategy is to disable the AI ​​model inference function if the energy consumption of the EMS per unit time exceeds a certain threshold.

[0422] S575 is a specific instance of S330 in method 300, and a detailed description can be found in step S330.

[0423] The S580 and NMS distribute the AIML inference function usage policy to the NE. Correspondingly, the NE receives the inference function usage policy from the NMS.

[0424] Optionally, the NMS distributes the AIML inference function usage policy to the NE via the EMS. Accordingly, the NE receives the inference function usage policy forwarded by the EMS.

[0425] S580 is a specific instance of S310 in method 300, and a detailed description can be found in step S310.

[0426] S585, the usage strategy for executing this inference function during NE runtime.

[0427] For example, if the inference energy consumption during the first unit of time is 600 joules, and the inference energy consumption during the first unit of time specified in the policy is higher than 500 joules (an example of threshold #2), then the inference function needs to be turned off / stopped / cancelled / suspended. Therefore, the inference function needs to be turned off / stopped / cancelled / suspended in the first unit of time. If the inference energy consumption during the unit of time specified in the policy is higher than 700 joules, then the inference function needs to be turned off / stopped, and therefore, the inference function does not need to be turned off / stopped / cancelled / suspended in the first unit of time.

[0428] S585 is a specific instance of S330 in method 300, and a detailed description can be found in step S330.

[0429] Based on Method 500, the system supports EMS in estimating and reporting the energy consumption of AI models during the training and simulation phases, and NE in estimating and reporting energy consumption during the simulation and inference phases. EMS can deploy AI models based on simulation energy consumption indicators issued by NMS, and the deployed NE can decide whether to use the AI ​​model based on the simulation energy consumption indicators and its own energy efficiency level. Furthermore, the system supports NE recording and reporting AI model energy consumption during inference. NMS manages this process through EMS and generates an AIML inference function usage strategy based on inference energy consumption. NE executes this strategy during operation. NE execution of this strategy during operation allows NMS to manage the energy consumption of the AI ​​model during operation within the NE, preventing the AI ​​model from malfunctioning when its energy consumption is too high for the NE's own energy efficiency level to support. The system can then promptly shut down the AI ​​model. When the NE's own energy demand increases, the AI ​​model can continue running, achieving real-time monitoring of energy consumption during AI model use and effectively avoiding the risk of excessive energy consumption during AIML training.

[0430] Figure 6 This is a schematic flowchart of an artificial intelligence model energy management method 600 provided in an embodiment of this application.

[0431] Figure 6 for Figure 3 One specific embodiment. The following is in conjunction with... Figure 6 An energy consumption management method 600 for an artificial intelligence model is provided, wherein the first device is an NMS and the second device is an NE. This embodiment addresses a scenario where the NMS performs network service management, model training, and model inference on the NE (e.g., gNB, NWDAF, etc.). The NE can be managed by an EMS. Method 600 includes steps S605 to S685, which are described in detail below.

[0432] Optionally, in step S605, when model training is triggered by the NMS, the NMS sends an ML training request to the NE. Correspondingly, the NE receives the ML training request from the NMS.

[0433] Optionally, the NMS sends an ML training request to the NE via the EMS. Correspondingly, the NE receives the ML training request forwarded by the EMS.

[0434] The ML training request includes an energy consumption recording identifier. This identifier indicates whether the EMS records the energy consumption generated during the training of the first AI model. The ML training request can be MLTrainingRequest.

[0435] S605 is a specific instance of S310 in method 300, and a detailed description can be found in step S310.

[0436] The S610 and NE perform model training while simultaneously measuring, calculating, and recording the energy consumption during the model training process.

[0437] Specifically, the NE executes the training of the first AI model based on the ML training request. It determines the energy consumption calculation method based on its own energy consumption computing capabilities, and uses this method to calculate the energy consumed / generated during training, obtaining the energy consumption value generated during the training of the first AI model. When the energy consumption recording flag included in the ML training request is True, meaning the energy consumption recording flag indicates that training energy consumption should be recorded, the NE records the energy consumption of the first AI model training process while simultaneously measuring and calculating it.

[0438] S610 is a specific instance of S320 in method 300, and a detailed description can be found in step S320.

[0439] S615, NE sends ML training reports to NMS. Correspondingly, NMS receives ML training reports from NE.

[0440] Optionally, the NE sends the ML training report to the NMS via the EMS. Correspondingly, the NMS receives the ML training report forwarded by the EMS.

[0441] Specifically, in response to an ML training request, the NE sends an ML training report to the NMS. This ML training report includes Energy Consumption #5, which represents the energy consumption value (or the quantity of energy consumed during training the first AI model) generated by the NE during time period #4, and an energy consumption calculation method identifier. In other words, the NE feeds back the calculated and recorded energy consumption value generated during the training of the first AI model to the NMS. The energy consumption calculation method identifier is used to report to the NMS the method used by the NE to calculate the energy consumption value for training the first AI model. This ML training report can be named MLTrainingReport.

[0442] Based on the energy consumption value and energy consumption calculation method generated during the training of the first reported AI model, NMS determines the overall training energy consumption of the model, and determines whether to pause, terminate, suspend or cancel the training process based on the current energy consumption requirements (or energy consumption level).

[0443] S615 is a specific instance of S330 in method 300, and a detailed description can be found in step S330.

[0444] The S620 and NMS send simulation requests to the NE. Correspondingly, the NE receives the simulation requests from the NMS.

[0445] Optionally, the NMS sends a simulation request to the NE via the EMS. Correspondingly, the NE receives the simulation request forwarded by the EMS.

[0446] The simulation request includes an energy consumption calculation method identifier and an energy consumption recording identifier. The energy consumption recording identifier indicates whether the NE records the energy consumption generated during the simulation of the first AI model. The energy consumption calculation method identifier indicates the method for calculating the energy consumption generated during the simulation of the first AI model. This simulation request can be an EmulationRequest.

[0447] The simulation request also indicates the type of inference function the NE executes for the first AI model. This allows the first AI model to be simulated on the corresponding type of inference function, and also facilitates determining whether to deploy the first AI model on the corresponding type of inference function based on its type.

[0448] S620 is a specific instance of S310 in method 300, and a detailed description can be found in step S310.

[0449] The S625 and NE perform model simulations, while simultaneously measuring and recording the energy consumption during the model simulation process.

[0450] Specifically, the NE executes the simulation of the first AI model based on the simulation request. The NE uses the energy consumption calculation method indicated by the energy consumption calculation method identifier to calculate the energy consumption generated during the simulation of the first AI model, obtaining the energy consumption value generated during the simulation. When the energy consumption recording identifier included in the simulation request is True, that is, when the energy consumption recording identifier indicates that the energy consumption value generated during the simulation of the first AI model is to be recorded, the NE records the energy consumption value generated during the simulation of the first AI model while simultaneously measuring and calculating it.

[0451] S625 is a specific instance of S320 in method 300, and a detailed description can be found in step S320.

[0452] S630 and NE send simulation reports to NMS. Correspondingly, NMS receives simulation reports from NE.

[0453] Optionally, the NE sends a simulation report to the NMS via the EMS. Correspondingly, the NMS receives the simulation report forwarded by the EMS.

[0454] Specifically, in response to a simulation request, the NE sends a simulation report to the NMS. This report includes Energy Consumption #2, which represents the energy consumption value generated by the EMS during the simulation of the first AI model within time period #1 (or the quantity of energy consumed during the simulation of the first AI model). In other words, the NE feeds back the calculated and recorded energy consumption value generated during the simulation of the first AI model to the NMS. This simulation report can be an EmulationReport. Energy Consumption #2 can be used to determine Energy Consumption #1.

[0455] S630 is a specific instance of S330 in method 300, and a detailed description can be found in step S320.

[0456] Optionally, in step S635, when deployment is triggered by EMS, NMS sends the ML entity loading policy to EMS. Correspondingly, EMS receives the ML entity loading policy from NMS.

[0457] The ML entity loading policy includes a simulation energy consumption indicator, which indicates energy consumption #1. The ML entity loading policy can be MLEntityLoadingPolicy. This simulation energy consumption value can be determined based on energy consumption #2 included in the simulation report in S630.

[0458] The ML entity loading strategy is also used to indicate threshold #1, or in other words, the ML entity loading strategy is also used to indicate energy demand / energy level, which corresponds to threshold #1.

[0459] S635 is a specific instance of S310 in method 300, and a detailed description can be found in step S310.

[0460] Optionally, in step S640, when deployment is triggered by the NMS, the NMS sends an ML entity loading request to the EMS. Correspondingly, the EMS receives the ML entity loading request from the NMS.

[0461] The ML entity loading request includes a simulation energy consumption indicator, which indicates energy consumption #1. The ML entity loading request can be an MLEntityLoadingRequest. This simulation energy consumption value can be determined based on energy consumption #2 included in the simulation report in S630.

[0462] The ML entity load request is also used to indicate threshold #1, or in other words, the ML entity load request is also used to indicate energy demand / energy level corresponding to threshold #1.

[0463] S640 is a specific instance of S310 in method 300, and a detailed description can be found in step S310.

[0464] S642 and EMS determine whether to deploy the first AI model to the inference function based on the energy consumption value and energy consumption requirements of the model simulation stage indicated by the simulation energy consumption indicator.

[0465] Specifically, EMS determines whether to deploy the first AI model to the inference function based on the relationship between the energy consumption value of the model simulation phase indicated by the simulation energy consumption indicator and the threshold #1.

[0466] This energy consumption requirement corresponds to the energy consumption value that the second device can accept, such as threshold #1.

[0467] For example, the simulation request also indicates the type of inference function. When the energy consumption value during the model simulation phase is less than or equal to threshold #1, the EMS determines to deploy the first AI model in the NE to the inference function corresponding to the type indicated by the simulation request. When the energy consumption value during the model simulation phase is greater than or equal to threshold #1, the EMS determines not to deploy the first AI model in the NE to the inference function corresponding to the type indicated by the simulation request. This NE can be managed by the EMS.

[0468] S642 is a specific instance of S320 in method 300. For a detailed description, please refer to step S320.

[0469] S645, EMS sends deployment result information to NMS, which indicates whether the deployment was successful or failed. Correspondingly, NMS receives the deployment result information from EMS.

[0470] In one approach, when deployment fails, the deployment result information also includes the reason for the deployment failure.

[0471] For example, when it is determined that the first AI model will not be deployed to the inference function, the deployment result information is used to indicate deployment failure, and the deployment result information includes the reason for not deploying the first AI model to the inference function, or can be understood as the reason for deployment failure. It should be understood that not deploying the first AI model is also a result of deployment failure.

[0472] For example, when it is determined to deploy the first AI model to the inference function, the deployment result information is used to indicate whether the deployment failed or succeeded. When the deployment fails, the deployment result information also includes the reason for the failure. It should be understood that the deployment failure means that the first AI model was determined to be deployed but the deployment failed.

[0473] S645 is a specific instance of S330 in method 300, and a detailed description can be found in step S330.

[0474] S650 and NMS send ML inference history control information to NE. Correspondingly, NE receives ML inference history control information from NMS.

[0475] Optionally, the NMS sends ML inference history control information to the NE via the EMS, and correspondingly, the NE receives the ML inference history control information forwarded by the EMS.

[0476] The ML inference history control information includes an energy consumption record identifier and an energy consumption calculation identifier. The energy consumption record identifier indicates whether the NE records the energy consumption during the inference process of the first AI model. The energy consumption calculation method identifier indicates the method used to calculate the energy consumption during the inference process of the first AI model. This ML inference history control information can be MLInferenceHistoryControl. This ML inference history control information can be used to control the energy consumption calculation and recording method during the inference process of the first AI model.

[0477] S650 is a specific instance of S310 in method 300, and a detailed description can be found in step S310.

[0478] S655 and NE perform model inference while measuring and recording the energy consumption of the model inference process.

[0479] For a detailed description, please refer to S555, which will not be repeated here.

[0480] S660 and NE send inference output to NMS. Correspondingly, NMS receives inference output from NE.

[0481] The NE sends inference output to the NMS via the EMS, and the NMS receives the inference output forwarded by the EMS accordingly.

[0482] For a detailed description, please refer to S560, which will not be repeated here.

[0483] S665, NMS sends an ML inference history request to NE. Correspondingly, NE receives the ML inference history request from NMS.

[0484] The NMS sends inference history requests to the NE via the EMS, and the NE receives the inference history requests forwarded by the EMS accordingly.

[0485] For a detailed description, please refer to S565, which will not be repeated here.

[0486] S670, NE sends ML inference history reports to NMS. Correspondingly, NMS receives ML inference history reports from NE.

[0487] The NE sends the inference history report to the NMS via the EMS, and the NMS receives the inference history report forwarded by the EMS accordingly.

[0488] For a detailed description, please refer to S570, which will not be repeated here.

[0489] S675, NMS generation AIML inference function usage strategy.

[0490] Specifically, NMS can generate usage strategies for AIML inference functions based on inference energy consumption.

[0491] For example, the usage strategy is to disable the AI ​​model inference function if the energy consumption of the EMS per unit time exceeds a certain threshold.

[0492] S675 is a specific instance of S330 in method 300, and a detailed description can be found in step S330.

[0493] S680 and NMS distribute the AIML inference function usage policy to NE. Correspondingly, NE receives the inference function usage policy from NMS.

[0494] Optionally, the NMS distributes the AIML inference function usage policy to the NE via the EMS. Accordingly, the NE receives the inference function usage policy forwarded by the EMS.

[0495] S680 is a specific instance of S310 in method 300, and a detailed description can be found in step S310.

[0496] S685, the usage strategy for executing this inference function during NE runtime.

[0497] For a detailed description, please refer to S585, which will not be repeated here.

[0498] Based on Method 600, the NE can estimate and report the energy consumption of AI models during the training, simulation, and inference phases. The NE can deploy AI models based on the simulation energy consumption indicators issued by the NMS, and the deployed NE can decide whether to use the AI ​​model based on the simulation energy consumption indicators and its own energy efficiency level. Furthermore, the NE can record and report the AI ​​model's energy consumption during inference. The NMS manages this process through the EMS and generates an AIML inference function usage strategy based on the inference energy consumption. The NE executes this strategy during operation. This execution of the strategy allows the NMS to manage the energy consumption of the AI ​​model during operation, preventing the AI ​​model from consuming excessive energy when the NE's energy efficiency level cannot support it. The AI ​​model can be shut down promptly. When the NE's own energy demand increases, the AI ​​model can continue running, achieving real-time monitoring of energy consumption during AI model use and effectively avoiding the risk of excessive energy consumption during AIML training.

[0499] Figure 7 This is a schematic flowchart of an artificial intelligence model energy management method 700 provided in an embodiment of this application.

[0500] Figure 7 for Figure 3 One specific embodiment. The following is in conjunction with... Figure 7 An AI model energy management method 700 is provided, where the first device is an SMO and the second device is an NE. This embodiment addresses a scenario where network service management, model training, and model inference are performed on the NE (e.g., gNB, NWDAF, etc.) using the SMO. The NE can be managed by the SMO. Method 700 includes steps S705 to S785, which are described in detail below.

[0501] Optionally, in step S705, when model training is triggered by the SMO, the SMO sends an ML training request to the NE. Correspondingly, the NE receives the ML training request from the SMO.

[0502] The ML training request includes an energy consumption recording identifier. This identifier indicates whether the NE records the energy consumption generated during the training of the first AI model. The ML training request can be an MLTrainingRequest. This request also instructs the ML to perform initial training and / or retraining.

[0503] S705 is a specific instance of S310 in method 300, and a detailed description can be found in step S310.

[0504] The S710 and NE perform model training, measuring, calculating, and recording the energy consumption during the model training process.

[0505] Specifically, the NE executes the training of the first AI model based on the ML training request. It determines the energy consumption calculation method based on its own energy consumption computing capabilities, and uses this method to calculate the energy consumed / generated during training, obtaining the energy consumption value generated during the training of the first AI model. When the energy consumption recording flag included in the ML training request is True, meaning the energy consumption recording flag indicates that training energy consumption should be recorded, the NE records the energy consumption of the first AI model training process while simultaneously measuring and calculating it.

[0506] S710 is a specific instance of S320 in method 300, and a detailed description can be found in step S320.

[0507] S715 and NE send ML training reports to SMO. Correspondingly, SMO receives ML training reports from NE.

[0508] Specifically, in response to an ML training request, the NE sends an ML training report to the SMO. This report includes Energy Consumption #5, which represents the energy consumption value (or the quantity of energy consumed during training the first AI model) generated by the NE during time period #4, and an energy consumption calculation method identifier. In other words, the NE feeds back the calculated and recorded energy consumption value generated during the training of the first AI model to the SMO. The energy consumption calculation method identifier is used to report to the SMO the method used by the NE to calculate the energy consumption value for training the first AI model. This ML training report can be named MLTrainingReport.

[0509] Based on the energy consumption value generated during the training of the first reported AI model and the energy consumption calculation method, SMO determines the overall training energy consumption of the model, and determines whether to pause, terminate, suspend or cancel the training process based on the current energy consumption requirements (or energy consumption level).

[0510] S715 is a specific instance of S330 in method 300, and a detailed description can be found in step S330.

[0511] The S720 and SMO send simulation requests to the NE. Correspondingly, the NE receives the simulation requests from the SMO.

[0512] The simulation request includes an energy consumption calculation method identifier and an energy consumption recording identifier. The energy consumption recording identifier indicates whether the NE records the energy consumption generated during the simulation of the first AI model. The energy consumption calculation method identifier indicates the method for calculating the energy consumption generated during the simulation of the first AI model. This simulation request can be an EmulationRequest.

[0513] The simulation request also indicates the type of inference function the NE executes for the first AI model. This allows the first AI model to be simulated on the corresponding type of inference function, and also facilitates determining whether to deploy the first AI model on the corresponding type of inference function based on its type.

[0514] S720 is a specific instance of S310 in method 300, and a detailed description can be found in step S310.

[0515] The S725 and NE perform model simulations, while simultaneously measuring and recording the energy consumption during the model simulation process.

[0516] Specifically, the NE executes the simulation of the first AI model based on the simulation request. The NE uses the energy consumption calculation method indicated by the energy consumption calculation method identifier to calculate the energy consumption generated during the simulation of the first AI model, obtaining the energy consumption value generated during the simulation. When the energy consumption recording identifier included in the simulation request is True, that is, when the energy consumption recording identifier indicates that the energy consumption value generated during the simulation of the first AI model is to be recorded, the NE records the energy consumption value generated during the simulation of the first AI model while simultaneously measuring and calculating it.

[0517] S725 is a specific instance of S320 in method 300, and a detailed description can be found in step S320.

[0518] The S730 and NE send simulation reports to the SMO. Correspondingly, the SMO receives simulation reports from the NE.

[0519] Specifically, in response to a simulation request, the NE sends a simulation report to the SMO. This report includes Energy Consumption #2, which represents the energy consumption value generated by the NE during the simulation of the first AI model within time period #1 (or the quantity of energy consumed during the simulation of the first AI model). In other words, the NE feeds back the calculated and recorded energy consumption value generated during the simulation of the first AI model to the SMO. This simulation report can be an EmulationReport. Energy Consumption #2 can be used to determine Energy Consumption #1.

[0520] S730 is a specific instance of S330 in method 300, and a detailed description can be found in step S320.

[0521] Optionally, in step S735, when deployment is automatically triggered by the NE, the SMO sends the ML entity loading policy to the NE. Correspondingly, the NE receives the ML entity loading policy from the SMO.

[0522] The ML entity loading policy includes a simulation energy consumption indicator, which indicates energy consumption #1. The ML entity loading policy can be MLEntityLoadingPolicy. This simulation energy consumption value can be determined based on energy consumption #2 included in the simulation report in S730.

[0523] The ML entity loading strategy is also used to indicate threshold #1, or in other words, the ML entity loading strategy is also used to indicate energy demand / energy level, which corresponds to threshold #1.

[0524] S735 is a specific instance of S310 in method 300, and a detailed description can be found in step S310.

[0525] Optionally, in S740, when deployment is triggered by the SMO, the SMO sends an ML entity loading request to the NE. Accordingly, the NE receives the ML entity loading request from the SMO.

[0526] The ML entity loading request includes a simulation energy consumption indicator, which indicates energy consumption #1. The ML entity loading request can be an MLEntityLoadingRequest. This simulation energy consumption value can be determined based on energy consumption #2 included in the simulation report in S730.

[0527] The ML entity load request is also used to indicate threshold #1, or in other words, the ML entity load request is also used to indicate energy demand / energy level corresponding to threshold #1.

[0528] S740 is a specific instance of S310 in method 300, and a detailed description can be found in step S310.

[0529] S742 and NE determine whether to deploy the first AI model to the inference function based on the energy consumption value and energy consumption requirements of the model simulation stage as indicated by the simulation energy consumption indicator.

[0530] Specifically, NE determines whether to deploy the first AI model to the inference function based on the relationship between the energy consumption value of the model simulation phase indicated by the simulation energy consumption indicator and the threshold #1.

[0531] This energy consumption requirement corresponds to the energy consumption value that the second device can accept, such as threshold #1.

[0532] For example, the simulation request also indicates the type of inference function. When the energy consumption value during the model simulation phase is less than or equal to threshold #1, the NE determines to deploy the first AI model to the inference function corresponding to the type indicated by the simulation request. When the energy consumption value during the model simulation phase is greater than or equal to threshold #1, the NE determines not to deploy the first AI model to the inference function corresponding to the type indicated by the simulation request. This NE can be managed by the SMO.

[0533] S742 is a specific instance of S320 in method 300, and a detailed description can be found in step S320.

[0534] S745, the NE sends deployment result information to the SMO, which indicates whether the deployment was successful or failed. Correspondingly, the SMO receives the deployment result information from the NE.

[0535] In one approach, when deployment fails, the deployment result information also includes the reason for the deployment failure.

[0536] For example, when NE determines not to deploy the first AI model to the inference function, the deployment result information is used to indicate deployment failure, and the deployment result information includes the reason for not deploying the first AI model to the inference function, or can be understood as the reason for deployment failure. It should be understood that not deploying the first AI model is also a result of deployment failure.

[0537] For example, when the NE determines to deploy the first AI model to the inference function, the deployment result information is used to indicate whether the deployment failed or succeeded. When the deployment fails, the deployment result information also includes the reason for the failure. It should be understood that the deployment failure means that the first AI model was determined to be deployed but the deployment failed.

[0538] S745 is a specific instance of S330 in method 300, and a detailed description can be found in step S330.

[0539] S750 and SMO send ML inference history control information to NE. Correspondingly, NE receives ML inference history control information from SMO.

[0540] The ML inference history control information includes an energy consumption record identifier and an energy consumption calculation identifier. The energy consumption record identifier indicates whether the NE records the energy consumption during the inference process of the first AI model. The energy consumption calculation method identifier indicates the method used to calculate the energy consumption during the inference process of the first AI model. This ML inference history control information can be MLInferenceHistoryControl. This ML inference history control information can be used to control the energy consumption calculation and recording method during the inference process of the first AI model.

[0541] S750 is a specific instance of S310 in method 300, and a detailed description can be found in step S310.

[0542] The S755 and NE perform model inference while simultaneously measuring, calculating, and recording the energy consumption during the model inference process.

[0543] In one approach, the NE executes the inference function of the first AI model based on ML inference history control information. The NE determines the energy consumption calculation method based on the energy consumption calculation method identifier included in the ML inference history control information or based on its own energy consumption calculation capabilities, and calculates the energy consumption generated during the inference process to obtain the energy consumption value generated by the first AI model during inference. When the energy consumption record identifier included in the ML inference history control information is True, that is, when the energy consumption record identifier indicates that the energy consumption generated during the first AI model's inference should be recorded, the NE records the energy consumption generated by the first AI model during inference while simultaneously measuring and calculating the energy consumption during the first AI model's inference.

[0544] S755 is a specific instance of S320 in method 300, and a detailed description can be found in step S320.

[0545] The S760 and NE send inference output to the SMO. Correspondingly, the SMO receives inference output from the NE.

[0546] The inference output includes energy consumption #3, which is the energy consumed by NE during inference of the first AI model in time period #2. This inference output can be InferenceOutput.

[0547] The SMO determines whether to pause, terminate, suspend, or cancel the inference process based on the energy consumption value generated during inference by the first reported AI model and the current energy consumption requirement (or energy consumption level).

[0548] S760 is a specific instance of S330 in method 300, and a detailed description can be found in step S330.

[0549] S765, SMO sends an ML inference history request to NE. Correspondingly, NE receives the ML inference history request from SMO.

[0550] This ML inference history request is used to query the energy consumption values ​​generated during the historical inference process of the first AI model. This ML inference history request can be MLInferenceHistoryRequest.

[0551] S765 is a specific instance of S310 in method 300, and a detailed description can be found in step S310.

[0552] S770 and NE send ML inference history reports to SMO. Correspondingly, SMO receives ML inference history reports from NE.

[0553] Specifically, in response to the ML inference history request, the NE sends an ML inference history report to the SMO. This report includes a model inference energy consumption indicator and an energy consumption calculation identifier. The energy consumption indicator specifies energy consumption #4, which is the energy consumption value generated by the NE during inference of the first AI model within time period #3. Time period #3 can be a specific time period in the historical inference of the first AI model. This ML inference history report can be an MLInferenceHistoryReport.

[0554] Based on the energy consumption value generated by the first AI model during historical inference, and according to the current energy consumption requirements (or energy consumption level), the SMO determines whether to enable NE to infer the first AI model.

[0555] S770 is a specific instance of S330 in method 300, and a detailed description can be found in step S330.

[0556] S775, SMO generation AIML inference function usage strategy.

[0557] Specifically, SMO can generate usage strategies for AIML inference functions based on the energy consumption during inference by the first AI model.

[0558] For example, the usage strategy is to stop / close / cancel / suspend the inference function of the first AI model if the energy consumption of NE per unit time is higher than the threshold #2.

[0559] S775 is a specific instance of S330 in method 300, and a detailed description can be found in step S330.

[0560] S780 and SMO distribute the usage strategy of AIML inference function to NE. Correspondingly, NE receives the usage strategy of inference function from SMO.

[0561] S780 is a specific instance of S310 in method 300, and a detailed description can be found in step S310.

[0562] S785, the usage strategy for executing this inference function during NE runtime.

[0563] S785 is a specific instance of S330 in method 300, and a detailed description can be found in step S330.

[0564] Based on Method 700, the NE can estimate and report the energy consumption of AI models during the training, simulation, and inference phases. The NE can deploy AI models based on the simulation energy consumption indicators issued by the SMO, and decide whether to use the AI ​​model based on the simulation energy consumption indicators and its own energy efficiency level. Furthermore, the NE can record and report the AI ​​model's energy consumption during inference. The SMO manages this process and generates an AIML inference function usage strategy based on the inference energy consumption. The NE executes this strategy during operation. This allows the SMO to manage the energy consumption of the AI ​​model during operation, preventing the AI ​​model from consuming excessive energy when the NE's energy efficiency level cannot support it. The AI ​​model can be shut down promptly. When the NE's own energy demand increases, the AI ​​model can continue running, achieving real-time monitoring of energy consumption during AI model use and effectively avoiding the risk of excessive energy consumption during AIML training.

[0565] The communication method provided in this application has been described in detail above. The AI ​​model energy management device provided in this application is described below.

[0566] In order to realize the functions of the AI ​​model energy management device (such as the first device, the second device, etc.) in the embodiments of this application, each device can realize the corresponding functions through hardware structure, software module, or hardware structure plus software module.

[0567] Figure 8 This is a schematic block diagram of the AI ​​model energy management device 1000 provided in this application embodiment. Figure 8 As shown, the device 1000 may include a transceiver unit 1010 and a processing unit 1020. The transceiver unit 1010 can communicate with the outside world, and the processing unit 1020 is used for data processing. The transceiver unit 1010 may also be referred to as a communication interface or transceiver unit. The processing unit 1020 can be used for processing.

[0568] Optionally, the device 1000 may further include a storage unit, which can be used to store instructions and / or data, and the processing unit 1020 can read the instructions and / or data in the storage unit to enable the device to implement the aforementioned method embodiments.

[0569] For example, the device 1000 is a first device, which can be an NMS or an SMO, or a device applied to or used in conjunction with an NMS or SMO to implement a method for executing an NMS or SMO, such as a chip, chip system, or circuit. See details below. Figure 10 The chip system shown is described in detail.

[0570] For example, the device 1000 is a second device, which can be an EMS or NE, or a device applied to or used in conjunction with an EMS or NE to implement a method performed by the EMS or NE, such as a chip, chip system, or circuit. See details below. Figure 10 The chip system shown is described in detail.

[0571] In one possible design, the device 1000 can implement the steps or processes corresponding to those performed by the first device in the above method embodiments, wherein the processing unit 1020 is used to perform processing-related operations of the first device in the above method embodiments, and the transceiver unit 1010 is used to perform transceiver-related operations of the first device in the above method embodiments.

[0572] For example, the transceiver unit 1010 is used to send information #1, which is used to indicate energy consumption #1.

[0573] In another possible design, the device 1000 can implement the steps or processes corresponding to those performed by the second device in the above method embodiments, wherein the transceiver unit 1010 is used to perform transceiver-related operations of the second device in the above method embodiments, and the processing unit 1020 is used to perform processing-related operations of the second device in the above method embodiments.

[0574] For example, the transceiver unit 1010 is used to receive information #1, which indicates energy consumption #1; the processing unit 1020 is used to determine whether to deploy the first AI model to the inference function based on the relationship between energy consumption #1 and threshold #1, where threshold #1 is less than or equal to the maximum energy consumption supported by the second device.

[0575] It should be understood that the device 1000 here is embodied in the form of a functional unit. The term "unit" here can refer to an application-specific integrated circuit (ASIC), electronic circuitry, a processor (e.g., a shared processor, a proprietary processor, or a group processor, etc.) and memory for executing one or more software or firmware programs, integrated logic circuitry, and / or other suitable components supporting the described functions. In an alternative example, those skilled in the art will understand that the device 1000 may specifically be the transmitting end in the above embodiments, used to execute the various processes and / or steps corresponding to the transmitting end in the above method embodiments; or, the device 1000 may specifically be the receiving end in the above embodiments, used to execute the various processes and / or steps corresponding to the receiving end in the above method embodiments. To avoid repetition, further details are omitted here.

[0576] The device 1000 in each of the above-described schemes has the function of implementing the corresponding steps performed by the transmitting end in the above-described method, or the device 1000 in each of the above-described schemes has the function of implementing the corresponding steps performed by the receiving end in the above-described method. The function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions; for example, the transceiver unit can be replaced by a transceiver (e.g., the transmitting unit in the transceiver unit can be replaced by a transmitter, and the receiving unit in the transceiver unit can be replaced by a receiver), and other units, such as processing units, can be replaced by processors, respectively executing the transceiver operations and related processing operations in each method embodiment.

[0577] Furthermore, the aforementioned transceiver unit can also be a transceiver circuit (e.g., it may include a receiving circuit and a transmitting circuit), and the processing unit can be a processing circuit. In embodiments of this application, the aforementioned AI model power management device can be the receiving end or transmitting end in the foregoing embodiments, or it can be a chip or a chip system, such as a system on a chip (SoC). The transceiver unit can be an input / output circuit or a communication interface. The processing unit is a processor, microprocessor, or integrated circuit integrated on the chip. No limitations are imposed here.

[0578] Figure 9 This is a schematic block diagram of the AI ​​model energy management device 2000 provided in an embodiment of this application. Figure 9 As shown, the device 2000 includes a processor 2010 and a transceiver 2020. The processor 2010 and the transceiver 2020 communicate with each other through an internal connection path. The processor 2010 is used to execute instructions to control the transceiver 2020 to transmit and / or receive signals.

[0579] Optionally, the device 2000 may further include a memory 2030, which communicates with the processor 2010 and the transceiver 2020 via an internal connection path. The memory 2030 is used to store instructions, and the processor 2010 can execute the instructions stored in the memory 2030.

[0580] For example, the device 2000 is a first device, which can be an NMS or an SMO, or a device applied to or used in conjunction with an NMS or SMO to implement a method for executing an NMS or SMO, such as a chip, chip system, or circuit. See details below. Figure 10 The chip system shown is described in detail.

[0581] For example, the device 2000 is a second device, which can be an EMS or NE, or a device applied to or used in conjunction with an EMS or NE to implement a method performed by the EMS or NE, such as a chip, chip system, or circuit. See details below. Figure 10 The chip system shown is described in detail.

[0582] In one possible implementation, the apparatus 2000 is used to implement the various processes and steps corresponding to the first device in the above method embodiments.

[0583] In another possible implementation, the apparatus 2000 is used to implement the various processes and steps corresponding to the second device in the above method embodiments.

[0584] Optionally, the memory 2030 may include read-only memory and random access memory, and provide instructions and data to the processor. A portion of the memory may also include non-volatile random access memory. For example, the memory may also store device type information. The processor 2010 may be used to execute instructions stored in the memory, and when the processor 2010 executes instructions stored in the memory, the processor 2010 is used to perform the various steps and / or processes of the method embodiments corresponding to the sending end or receiving end described above.

[0585] In implementation, each step of the above method can be completed by integrated logic circuits in the processor's hardware or by instructions in software. The steps of the method disclosed in the embodiments of this application can be directly implemented by a hardware processor, or by a combination of hardware and software modules in the processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, detailed descriptions are omitted here.

[0586] It should be noted that the processor in the embodiments of this application can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method embodiments can be completed by the integrated logic circuitry in the processor's hardware or by instructions in software form. The processor can be a general-purpose processor, digital signal processor, application-specific integrated circuit, field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, or, as mentioned above, a CPU, other general-purpose processor, DSP, ASIC, FPGA or other codeable logic device, or a portion of the circuitry in another chip used for processing functions. The processor in the embodiments of this application can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method.

[0587] It is understood that the memory in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory used in the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0588] In the embodiments of this application, the method described above can be executed by a first device or a second device, or by a chip, chip system, or circuit of the first device or the second device, wherein the chip, chip system, or circuit can be installed in the first device or the second device. Below, in conjunction with... Figure 10 The explanation will take the chip system of the first or second device as an example.

[0589] Figure 10 This is a schematic block diagram of the chip system 3000 provided in an embodiment of this application. Figure 10 As shown, the chip system 3000 (or processing system) includes logic circuitry 3010 and input / output interface 3020.

[0590] The logic circuit 3010 can be a processing circuit in the chip system 3000. The logic circuit 3010 can be coupled to a memory unit, calling instructions from the memory unit, enabling the chip system 3000 to implement the methods and functions of the embodiments of this application. The input / output interface 3020 can be an input / output circuit in the chip system 3000, outputting processed information from the chip system 3000, or inputting data or signaling information to be processed into the chip system 3000 for processing.

[0591] As one approach, the chip system 3000 is used to implement the operations performed by the first device or the second device in the various method embodiments described above.

[0592] For example, logic circuit 3010 is used to implement processing-related operations performed by the first device in the above method embodiments, such as the processing-related operations performed by the first device in the above embodiments; input / output interface 3020 is used to implement sending and / or receiving-related operations performed by the first device in the above method embodiments, such as the sending and / or receiving-related operations performed by the first device in the above embodiments.

[0593] For example, logic circuit 3010 is used to implement processing-related operations performed by the second device in the above method embodiments, such as the processing-related operations performed by the second device in the above embodiments; input / output interface 3020 is used to implement sending and / or receiving-related operations performed by the second device in the above method embodiments, such as the sending and / or receiving-related operations performed by the second device in the above embodiments.

[0594] This application also provides a computer-readable storage medium storing computer instructions for implementing the methods executed by the first device or the second device in the above-described method embodiments.

[0595] This application also provides a computer program product comprising instructions which, when executed by a computer, implement the methods described in the above-described method embodiments by the first device or the second device.

[0596] This application also provides a communication system, which includes the first device or the second device in the above embodiments.

[0597] The explanations and beneficial effects of the relevant contents in any of the devices provided above can be found in the corresponding method embodiments provided above, and will not be repeated here.

[0598] In this application, examples may reference each other without logical contradiction. For example, methods and / or terms between method embodiments may reference each other, functions and / or terms between device embodiments may reference each other, and functions and / or terms between device examples and method examples may reference each other.

[0599] In the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0600] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0601] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0602] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0603] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0604] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0605] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0606] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An energy consumption management method for an artificial intelligence model, characterized in that, Applied to a second device, the method includes: Receive first information, which is used to indicate first energy consumption; Based on the relationship between the first energy consumption and the first threshold, it is determined whether to deploy the first model to the inference function, where the first threshold is less than or equal to the maximum energy consumption supported by the second device.

2. The method according to claim 1, characterized in that, When it is determined that the first model will not be deployed to the inference function, the method further includes: Send a first instruction message, which includes the reason for not deploying the first model to the inference function.

3. The method according to claim 1, characterized in that, When determining to deploy the first model to the inference function, the method further includes: Send a second instruction message, which is used to indicate the deployment result of the first model.

4. The method according to claim 3, characterized in that, When the deployment result is a failure, the second indication information includes the reason for the deployment failure.

5. The method according to any one of claims 1-4, characterized in that, The method further includes: The second energy consumption is obtained, which is used to determine the first energy consumption. The second energy consumption is the energy consumption generated by the second device when simulating the first model in the first time period.

6. The method according to claim 5, characterized in that, The method further includes: Receive second information, which instructs the second device to calculate and record the energy consumption generated when the first model is simulated.

7. The method according to claim 6, characterized in that, The method further includes: Send a first feedback message in response to the second information, the first feedback message being used to indicate the second energy consumption.

8. The method according to any one of claims 1-7, characterized in that, The method further includes: Receive third information, which is used to indicate the first threshold.

9. The method according to any one of claims 1-8, characterized in that, The method further includes: Receive fourth information, which indicates the type of the reasoning function.

10. The method according to any one of claims 1-9, characterized in that, When the first model is successfully deployed, the method further includes: Receive fifth information, the fifth information being used to instruct the second device to calculate and record the energy consumption generated when the first model performs inference; Send a second feedback message in response to the fifth information, the second feedback message being used to indicate a third energy consumption, the third energy consumption being the energy consumption generated by the second device when performing inference on the first model during a second time period.

11. The method according to any one of claims 1-10, characterized in that, When the first model is successfully deployed, the method further includes: Receive a sixth message, the sixth message being used to request the second device to query the energy consumption generated when the first model performs inference; Send a third feedback message in response to the sixth information, the third feedback message being used to indicate a fourth energy consumption, the fourth energy consumption being the energy consumption generated by the second device when performing inference on the first model during the third time period.

12. The method according to any one of claims 1-11, characterized in that, After the first model is successfully deployed, the method further includes: The device receives usage policy information, which instructs the second device to stop the inference function when the energy consumption generated by the first model during inference exceeds a second threshold.

13. An energy consumption management method for an artificial intelligence model, characterized in that, Applied to a first device, the method includes: Send a first message, the first message being used to indicate the first energy consumption, the first energy consumption being used in conjunction with a first threshold to determine whether to deploy the first model to the inference function, the first threshold being less than or equal to the maximum energy consumption supported by the second device.

14. The method according to claim 13, characterized in that, When it is determined that the first model will not be deployed to the inference function, the method further includes: Receive a first instruction message, which includes a reason for not deploying the first model to the inference function.

15. The method according to claim 13, characterized in that, When determining to deploy the first model to the inference function, the method further includes: Receive a second indication message, which indicates the deployment result of the first model.

16. The method according to claim 15, characterized in that, When the deployment result is a failure, the second indication information includes the reason for the deployment failure.

17. The method according to any one of claims 13-16, characterized in that, The method further includes: Send a second message, which instructs the second device to calculate and record the energy consumption generated when the first model is simulated.

18. The method according to claim 17, characterized in that, The method further includes: Receive first feedback information in response to the second information, the first feedback information being used to indicate second energy consumption, the second energy consumption being the energy consumption generated by the second device when simulating the first model within a first time period, the second energy consumption being used to determine the first energy consumption.

19. The method according to any one of claims 13-18, characterized in that, The method further includes: Send a third message, which is used to indicate the first threshold.

20. The method according to any one of claims 13-19, characterized in that, The method further includes: Send a fourth message, which indicates the type of the reasoning function.

21. The method according to any one of claims 13-20, characterized in that, When the first model is successfully deployed, the method further includes: Send a fifth message, which instructs the second device to calculate and record the energy consumption generated when the first model performs inference. Receive second feedback information for the fifth information, the second feedback information being used to indicate third energy consumption, the third energy consumption being the energy consumption generated by the second device when performing inference on the first model during the second time period; Based on the third energy consumption, determine whether to suspend the second device's inference of the first model.

22. The method according to any one of claims 13-21, characterized in that, When the first model is successfully deployed, the method further includes: Send a sixth message, the sixth message being used to request the second device to query the energy consumption generated when the first model performs inference; Receive third feedback information for the sixth information, the third feedback information being used to indicate fourth energy consumption, the fourth energy consumption being the energy consumption generated by the second device when performing inference on the first model during the third time period; Based on the fourth energy consumption, determine whether to enable the second device to perform inference on the first model.

23. The method according to any one of claims 13-22, characterized in that, After the first model is successfully deployed, the method further includes: Send usage policy information, which is used to instruct the second device to stop the inference function when the energy consumption generated by the first model when performing the inference function is greater than a second threshold.

24. An apparatus, characterized in that, The apparatus includes a unit or module for performing the method of any one of claims 1 to 12, or the apparatus includes a unit or module for performing the method of any one of claims 13 to 23.

25. A system, characterized in that, include: A first device and a second device, wherein the first device is used to perform the method as described in any one of claims 13 to 23, and the second device is used to perform the method as described in any one of claims 1 to 12.

26. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when the computer program is run on a computer, cause the computer to perform the method as described in any one of claims 1 to 12, or cause the computer to perform the method as described in any one of claims 13 to 23.

27. A computer program product, characterized in that, The computer program product includes: computer program code, which, when run on a communication device, causes the device to perform the method as described in any one of claims 1 to 12, or causes the device to perform the method as described in any one of claims 13 to 23.