Model inference method and related apparatus
By dynamically adjusting the model segmentation method and adjusting the model parameters according to the uplink bandwidth and performance indicators, the problem of resource waste in end-to-end collaborative inference is solved, and efficient resource utilization is achieved under end-to-end latency requirements.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2025-10-15
- Publication Date
- 2026-05-21
AI Technical Summary
In end-to-end collaborative inference, existing technologies require preloading multiple models when uplink bandwidth changes, resulting in wasted computing resources on both the terminal and network sides, and making it difficult to guarantee end-to-end latency requirements.
By dynamically adjusting the model segmentation method based on the current uplink bandwidth, only necessary model parameters are loaded, avoiding the preloading of the entire model structure and parameter set. The model segmentation method is dynamically adjusted in conjunction with the performance indicators and channel quality information of the terminal and network equipment.
While ensuring end-to-end latency, it reduces the waste of computing resources on the terminal side and network side, and improves resource utilization efficiency.
Smart Images

Figure CN2025127785_21052026_PF_FP_ABST
Abstract
Description
Model reasoning methods and related devices
[0001] This application claims priority to Chinese Patent Application No. 202411617222.8, filed on November 12, 2024, entitled “Model Reasoning Method and Related Apparatus”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of artificial intelligence technology, and in particular to model reasoning methods and related devices. Background Technology
[0003] With the development of artificial intelligence (AI) and machine learning (ML) model inference, applications based on AI / ML models are emerging, such as image classification, speech recognition, and autonomous driving. These applications require increasingly higher computing power to process increasingly complex AI / ML data and models. Furthermore, emerging media services such as virtual reality (VR), augmented reality (AR), and cloud gaming are also developing rapidly. These applications heavily rely on AI / ML technologies to realize their core functions. For example, AR media services based on smart glasses require in-depth analysis of the surrounding environment, and AI / ML technologies can provide the necessary environmental context information by performing object segmentation, recognition, and classification, thereby improving the user experience. However, the massive computational demands of these applications often exceed the computing power of the edge devices. Therefore, splitting AI / ML models between edge devices and the network side to allocate some or even most of the computation to the network side through edge-network collaborative inference has become a solution.
[0004] Some model inference applications, such as autonomous driving, have very high requirements for end-to-end inference latency. While addressing the issue of insufficient computing resources on the device side, end-to-end latency requirements of inference applications must also be met. Currently, to ensure uninterrupted inference service when models are changed due to uplink bandwidth variations, both the device and network sides need to pre-load all multiple models, resulting in significant resource waste on both sides.
[0005] Therefore, how to reduce the waste of computing resources on both the end-to-end and network sides while ensuring end-to-end latency in end-to-end collaborative inference services under the condition of uplink channel bandwidth variation is a hot research topic for those skilled in the art. Summary of the Invention
[0006] This application provides a model inference method and related apparatus that can reduce the waste of computing resources on the terminal side and network equipment side while ensuring end-to-end latency.
[0007] Firstly, this application provides a model inference method, which can be applied to a first communication device. The first communication device can be, for example, a terminal or a module within a terminal (wherein the module includes a communication module and a computing module), or a circuit or chip within the terminal responsible for communication functions (such as a modem chip, also known as a baseband chip, or a system-on-chip (SoC) chip containing a modem core, or a system-in-package (SIP) chip). The method includes: downloading the structure of a first model and a set of model parameters corresponding to the structure of the first model, wherein the first model is used for end-to-end collaborative inference services, and the structure of the first model is the union of the structures of models used by the first communication device under various model segmentation methods. Determining a first model segmentation method corresponding to a first time period based on the current uplink bandwidth, wherein the first model segmentation method is related to the set of model parameters corresponding to the structure of the first model. Loading the first model parameters corresponding to the first model segmentation method, wherein the first model parameters belong to the set of model parameters and are used for the end-to-end collaborative inference service. The first operation is executed repeatedly until the end-to-end collaborative inference service ends within a preset time period, wherein the preset time period consists of at least one time cycle, and the first time cycle belongs to the at least one time cycle. The first operation includes: determining a performance indicator associated with the first communication device based on its own battery level, remaining storage space, CPU load, GPU load, and temperature, wherein the performance indicator is related to the uplink transmission rate. Parameter information is sent to a second communication device, wherein the parameter information includes a Channel Quality Indicator (CQI), the service quality requirements of the end-to-end collaborative inference service, the inference output dimension of the first communication device, the inference frequency, and the performance indicator; the parameter information is used by the second or third communication device to predict the uplink bandwidth of the first communication device. A second model segmentation method corresponding to the next time cycle of the first time cycle is determined, wherein the second model segmentation method is related to the model parameter set corresponding to the first model. A model parameter processing strategy is determined based on the second model segmentation method.
[0008] In this application, taking the first communication device as the terminal and the second communication device as the network device as an example, on the one hand, since some model inference applications have high requirements for end-to-end inference latency, the terminal can determine the first model segmentation method corresponding to the first time period based on the current uplink bandwidth. This first model segmentation method corresponds to a deployment scheme of a model on the terminal side and the network device side. End-to-end latency = terminal-side inference latency + intermediate data transmission latency + network-side inference latency. Since the sum of the terminal-side inference model and the network device-side inference model is equivalent to a complete model, it can be assumed that the sum of the terminal-side inference latency and the network device-side inference latency remains unchanged, and the end-to-end latency is only affected by the intermediate data transmission latency. When the uplink channel bandwidth between the terminal side and the network device side decreases, the intermediate data transmission latency increases, and the end-to-end latency increases. In this case, the terminal can use a model segmentation method with a lower intermediate data output dimension according to the latency requirements when the uplink channel bandwidth decreases, thereby effectively ensuring the end-to-end latency.
[0009] On the other hand, compared to existing solutions where the terminal side and network device side need to preload the structure of the model used under various model partitioning methods and the union of the model parameter sets corresponding to the model structure, and select the appropriate model structure and model parameters to use when the uplink bandwidth changes, this application can load the corresponding model parameters that meet the uplink bandwidth conditions according to the uplink bandwidth as required. There is no need to preload the structure of the model used under various model partitioning methods and the union of the model parameter sets corresponding to the model structure, thereby effectively reducing the waste of computing resources on the terminal side and network device side.
[0010] In summary, this application can reduce the waste of computing resources on the terminal side and network device side while ensuring end-to-end latency.
[0011] In one possible implementation, determining the second model segmentation method corresponding to the next time period of the first time period includes: receiving an uplink bandwidth prediction range from the second communication device; and determining the second model segmentation method corresponding to the next time period of the first time period based on the uplink bandwidth prediction range.
[0012] Optionally, the first time period is the current period or a time period preceding the current period.
[0013] In the above implementation, taking the first communication device as the terminal and the second communication device as the network device as an example, the terminal no longer needs to preload the structure of all models and the set of model parameters corresponding to the model structure. Instead, it loads the structure of the model and the set of model parameters corresponding to the model structure on demand based on the prediction. Specifically, the network device or other devices (such as network management devices) collect data from a global cell perspective and predict the uplink channel bandwidth of the terminal in the next time period of the first time period to obtain the uplink bandwidth prediction range. This uplink bandwidth prediction range can be sent to the terminal by the network device or forwarded to the terminal by the network device. The terminal determines the second model segmentation method corresponding to the next time period of the first time period based on the uplink bandwidth prediction range from the network device or network management device, avoiding the resource waste caused by preloading the structure of all models and the set of model parameters corresponding to the model structure.
[0014] In another possible implementation, determining the second model segmentation method corresponding to the next time period of the first time period includes: receiving the second model segmentation method corresponding to the next time period of the first time period forwarded by the second communication device.
[0015] In the above embodiment, taking the first communication device as the terminal and the second communication device as the network device as an example, the terminal no longer needs to preload the structure of the entire model and the set of model parameters corresponding to the model structure. Instead, it loads the structure of the model and the model parameters corresponding to the model structure on demand based on the prediction. Specifically, the terminal can directly load the corresponding model parameters according to the second model segmentation method corresponding to the next time period of the received first time period, avoiding the resource waste caused by preloading the structure of the entire model and the set of model parameters corresponding to the model structure.
[0016] In another possible implementation, the cyclic execution of the first operation includes: cyclically executing the first operation when a triggering condition is met.
[0017] In the above implementation, the process of executing the first operation in a subsequent loop is initiated only when the triggering condition is met, which can effectively reduce the power consumption of the inference service involving the terminal and achieve a more energy-efficient inference service.
[0018] In another possible implementation, the step of determining the model parameter processing strategy according to the second model segmentation method includes: loading the model structure added by the second model segmentation method relative to the first model segmentation method and the model parameters corresponding to the added model structure in the first time period, and / or unloading the model structure reduced by the second model segmentation method relative to the first model segmentation method and the model parameters corresponding to the reduced model structure in the next time period of the first time period.
[0019] It should be noted that uninstalling in the next time period after the first time period means that uninstallation can be performed immediately after the first time period has ended, without having to wait for the next time period after the first time period to end.
[0020] In the above implementation, this solution can be applied to scenarios with relatively simple segmentation methods. For example, if the first model has already loaded layer X, then layer Y is loaded in the first time period, or layer Z is unloaded in the next time period (where X, Y, and Z are positive integers, and Z < X). In this solution, missing model structures and their corresponding model parameters can be loaded in advance, and redundant model structures and their corresponding model parameters can be unloaded in the next time period. This allows for more adaptive and flexible loading of model structures and their corresponding model parameters, avoiding resource waste caused by preloading the entire model structure and the set of model parameters corresponding to each model structure.
[0021] In another possible implementation, determining the model parameter processing strategy according to the second model segmentation method includes: loading the model parameters corresponding to the second model segmentation method in the first time period, and / or running the terminal-network collaborative inference service on the model parameters corresponding to the second model segmentation method in the next time period of the first time period, and unloading the model parameters corresponding to the first model segmentation method.
[0022] In the above embodiments, this solution can be applied to scenarios with relatively complex segmentation methods. For example, if the first model has already loaded layer X1, then layer X2 is loaded in the first time period, or the system switches to layer X2 for end-to-end collaborative inference service in the next time period after the first time period, and the previous layer X1 is unloaded. In this solution, missing model structures and their corresponding model parameters can be loaded in advance, and redundant model structures and their corresponding model parameters can be unloaded in the next time period. This allows for more adaptive and flexible loading of model structures and their corresponding model parameters, and also avoids the resource waste caused by preloading the entire model structure and the set of model parameters corresponding to the model structure.
[0023] Secondly, embodiments of this application provide a model inference method applied to a second communication device. The second communication device may be, for example, a network device or a module within a network device (wherein the module includes a communication module and a computing module), or a circuit or chip within a network device responsible for communication functions (such as a modem chip, also known as a baseband chip, or a system-on-chip (SoC) chip containing a modem core, or a system-in-package (SIP) chip). The method includes: repeatedly executing a second operation until the end-to-end network collaborative inference service ends within a preset time period, the preset time period consisting of at least one time cycle. The second operation includes: receiving parameter information from a first communication device, wherein the parameter information includes a Channel Quality Indicator (CQI), the Quality of Service (QoS) requirements of the inference service, the inference output dimension of the first communication device, the inference frequency, and the performance indicator. The parameter information is used by the second or third communication device to predict the uplink bandwidth of the first communication device. The channel quality index information of the first communication device is measured, wherein the channel quality index information includes signal-to-interference-plus-noise ratio (SINR), reference signal reception quality (RSRQ), and uplink received signal strength indicator (UL_RSSI). An uplink bandwidth prediction range or a second model segmentation method is sent to the first communication device.
[0024] In this application, taking the first communication device as the terminal and the second communication device as the network device as an example, on the one hand, since some model inference applications have high requirements for end-to-end inference latency, end-to-end latency = terminal-side inference latency + intermediate data transmission latency + network-side inference latency. Since the sum of the terminal-side inference model and the network device-side inference model is equivalent to a complete model, it can be assumed that the sum of the terminal-side inference latency and the network device-side inference latency remains unchanged, and the end-to-end latency is only affected by the intermediate data transmission latency. When the uplink channel bandwidth between the terminal side and the network device side decreases, the intermediate data transmission latency increases, and the end-to-end latency increases. For example, the network device can collect data from the perspective of the entire cell and predict the uplink channel bandwidth of the terminal to obtain the uplink bandwidth prediction range. This uplink bandwidth prediction range can be sent by the network device to the terminal. Alternatively, the uplink bandwidth prediction range can also be forwarded by the network device to the terminal. For example, the network device can also directly forward the second model segmentation method to the terminal. The terminal performs subsequent operations based on the uplink bandwidth prediction range or the second model segmentation method (for example, when the uplink channel bandwidth between the terminal side and the network device side decreases, the intermediate data transmission latency increases, and the end-to-end latency increases. In this case, the terminal can use the model segmentation method with a lower intermediate data output dimension according to the latency requirements when the uplink channel bandwidth decreases, thereby effectively ensuring the end-to-end latency).
[0025] On the other hand, compared to existing solutions where the terminal side and network device side need to preload the structure of the model used under various model partitioning methods and the union of the model parameter sets corresponding to the model structure, and select the appropriate model structure and model parameters to use when the uplink bandwidth changes, this application can load the corresponding model parameters that meet the uplink bandwidth conditions according to the uplink bandwidth as required. There is no need to preload the structure of the model used under various model partitioning methods and the union of the model parameter sets corresponding to the model structure, thereby effectively reducing the waste of computing resources on the terminal side and network device side.
[0026] In summary, this application can reduce the waste of computing resources on the terminal side and network device side while ensuring end-to-end latency.
[0027] In one possible implementation, sending the uplink bandwidth prediction range or the second model segmentation method to the first communication device includes: obtaining the uplink bandwidth prediction range based on the load of the second communication device, the parameter information, and the channel quality index information during a first time period, wherein the first time period belongs to the at least one time period; and sending the uplink bandwidth prediction range to the first communication device.
[0028] In the above embodiment, taking the first communication device as the terminal and the second communication device as the network device as an example, the network device collects data from a global cell perspective and predicts the uplink channel bandwidth of the terminal to obtain the uplink bandwidth prediction range. The terminal performs subsequent operations based on the uplink bandwidth prediction range from the network device, avoiding the resource waste caused by preloading the structure of the entire model and the set of model parameters corresponding to the model structure.
[0029] In another possible implementation, obtaining the uplink bandwidth prediction range based on the load of the second communication device, the parameter information, and the channel quality index information in the first time period includes: inputting the load of the second communication device, the parameter information, and the channel quality index information into the prediction model in the first time period to obtain the uplink bandwidth prediction range.
[0030] In the above implementation, a prediction model is obtained by training multiple batches of data from the entire process. The obtained prediction model provides an accurate mapping from the input to the desired output. The efficiency of obtaining the uplink bandwidth prediction range is improved by utilizing the trained model.
[0031] In another possible implementation, sending the uplink bandwidth prediction range or the second model segmentation method to the first communication device includes: forwarding the parameter information, the channel quality index information, and the load of the second communication device to the third communication device; receiving the uplink bandwidth prediction range from the third communication device; and forwarding the uplink bandwidth prediction range to the first communication device.
[0032] In the above embodiment, taking the first communication device as the terminal, the second communication device as the network device, and the third communication device as the network management device as an example, since the network management device has more abundant computing resources and collects more data, it can collect data from a global cell perspective and run more complex machine learning models to predict the uplink channel bandwidth of the terminal, obtaining the uplink bandwidth prediction range. This uplink bandwidth prediction range is received by the network device and forwarded to the terminal. The terminal performs subsequent operations based on this uplink bandwidth prediction range, avoiding the resource waste caused by preloading the entire model structure and the corresponding model parameter set.
[0033] In another possible implementation, sending the uplink bandwidth prediction range or the second model segmentation method to the first communication device includes: forwarding the parameter information, the channel quality index information, and the load of the second communication device to the third communication device; receiving the second model segmentation method from the third communication device; and forwarding the second model segmentation method to the first communication device.
[0034] In the above implementation, taking the first communication device as the terminal, the second communication device as the network device, and the third communication device as the network management device as an example, since the network management device has more abundant computing resources and collects more data, it can collect data from a global cell perspective and run more complex machine learning models to predict the uplink channel bandwidth of the terminal, thus obtaining the uplink bandwidth prediction range. Then, based on the uplink bandwidth prediction range, a second model segmentation method is given for the terminal in the next time period after the first time period. The terminal can directly perform subsequent operations according to this second model segmentation method, avoiding the resource waste caused by preloading the structure of all models and the set of model parameters corresponding to the model structure.
[0035] Thirdly, embodiments of this application provide a model inference method applied to a third communication device. The third communication device may be, for example, a network management device or a module within the network management device (wherein the module in the network device includes a communication module and a computing module), or a circuit or chip within the network management device responsible for communication functions (such as a modem chip, also known as a baseband chip, or a system-on-chip (SoC) chip containing a modem core, or a system-in-package (SIP) chip). The method includes: cyclically executing a third operation until the end-to-end network collaborative inference service ends within a preset time period, wherein the preset time period consists of at least one time cycle. The third operation includes: receiving parameter information, channel quality index information, and the load of the second communication device from the second communication device. The parameter information includes a channel quality indicator (CQI), the quality of service requirement for inference services, the inference output dimension of the first communication device, the inference frequency, and the performance indicator. This parameter information is used by the third communication device to predict the uplink bandwidth of the first communication device. The channel quality index information includes a signal-to-interference-plus-noise ratio (SINR), a reference signal reception quality (RSRQ), and an uplink received signal strength indicator (UL_RSSI). In a first time period, based on the parameter information of the second communication device, the channel quality index information, and the load of the second communication device, an uplink bandwidth prediction range or a second model segmentation method corresponding to the next time period of the first time period is determined. The first time period belongs to the at least one time period. The uplink bandwidth prediction range or the second model segmentation method is then sent to the second communication device.
[0036] In this application, taking the first communication device as the terminal, the second communication device as the network device, and the third communication device as the network management device as an example, since the network management device has more abundant computing resources and collects more data, it can collect data from a multi-cell global perspective and run more complex machine learning models to predict the uplink channel bandwidth of the first communication device, obtain the uplink bandwidth prediction range, or directly give the terminal the second model segmentation method for the next time period of the first time period. The terminal can perform subsequent operations according to the uplink bandwidth prediction range or the second model segmentation method, avoiding the waste of resources caused by preloading the structure of all models and the model parameter set corresponding to the model structure.
[0037] In one possible implementation, determining the uplink bandwidth prediction range or the second model segmentation method corresponding to the next time period of the first time period based on the parameter information of the second communication device, the channel quality index information, and the load of the second communication device includes: inputting the load of the second communication device, the parameter information, and the channel quality index information into the prediction model to obtain the uplink bandwidth prediction range.
[0038] In the above implementation, a prediction model is obtained by training multiple batches of data from the entire process. The obtained prediction model provides an accurate mapping from the input to the desired output. The efficiency of obtaining the uplink bandwidth prediction range is improved by utilizing the trained model.
[0039] In another possible implementation, determining the uplink bandwidth prediction range or the second model segmentation method corresponding to the next time period of the first time period based on the parameter information of the second communication device, the channel quality index information, and the load of the second communication device includes: inputting the load of the second communication device, the parameter information, and the channel quality index information into the prediction model to obtain the uplink bandwidth prediction range; receiving the maximum computing power resources of the first communication device, wherein the maximum computing power resources are used for end-to-end collaborative inference services; and determining the second model segmentation method corresponding to the next time period of the first time period based on the uplink bandwidth prediction range and the maximum computing power resources.
[0040] In the above embodiment, taking the first communication device as the terminal, the second communication device as the network device, and the third communication device as the network management device as an example, since the network management device has more abundant computing resources and collects more data, it can collect data from a multi-cell global perspective and run more complex machine learning models to predict the uplink channel bandwidth of the first communication device, thus obtaining the uplink bandwidth prediction range. Then, based on the uplink bandwidth prediction range, a second model segmentation method is given for the terminal in the next time period of the first time period. The terminal can directly perform subsequent operations according to this second model segmentation method, avoiding the resource waste caused by preloading the structure of all models and the set of model parameters corresponding to the model structure.
[0041] Fourthly, embodiments of this application provide a model inference device that can be used in a first communication device in the first aspect. The communication device can be a terminal, a device in the terminal (e.g., a chip, a chip system, or a circuit), or a device that can be matched with the terminal, or a logic module or software that can implement all or part of the terminal functions.
[0042] In one possible implementation, the model inference device may include modules or units that perform the methods / operations / steps / actions described in the first aspect. These modules or units may be hardware circuits, software, or a combination of hardware circuits and software.
[0043] Fifthly, embodiments of this application provide a model inference device that can be used in the second communication device of the second aspect. The model inference device can be a network device, a device in the network device (e.g., a chip, a chip system, or a circuit), or a device that can be used in conjunction with the network device, or a logic module or software that can implement all or part of the functions of the network device.
[0044] In one possible implementation, the model inference device may include modules or units that perform the methods / operations / steps / actions described in the second aspect. These modules or units may be hardware circuits, software, or a combination of hardware circuits and software.
[0045] In a sixth aspect, embodiments of this application provide a model inference device that can be used in a third communication device in the third aspect. The model inference device can be a network management device, a device in the network management device (e.g., a chip, a chip system, or a circuit), or a device that can be matched with the network management device. It can also be a logic module or software that can implement all or part of the functions of the network management device.
[0046] In one possible implementation, the model inference device may include modules or units that perform the methods / operations / steps / actions described in the third aspect one by one. These modules or units may be hardware circuits, software, or a combination of hardware circuits and software.
[0047] In a seventh aspect, embodiments of this application provide a model inference apparatus, which includes at least one processor and a communication interface; the communication interface is used for inputting and / or outputting information, and the at least one processor is used for calling a computer program stored in at least one memory to implement the method described in any of the embodiments of the first aspect.
[0048] In one possible implementation, the model inference device further includes at least one of the aforementioned memories. Optionally, the memory and processor are integrated together.
[0049] Eighthly, embodiments of this application provide a model inference apparatus, which includes at least one processor and a communication interface; the communication interface is used for inputting and / or outputting information, and the at least one processor is used for calling a computer program stored in at least one memory to implement the method described in any of the embodiments of the second aspect.
[0050] In one possible implementation, the model inference device further includes at least one of the aforementioned memories. Optionally, the memory and processor are integrated together.
[0051] Ninthly, embodiments of this application provide a model inference apparatus, which includes at least one processor and a communication interface; the communication interface is used for inputting and / or outputting information, and the at least one processor is used for calling a computer program stored in at least one memory to implement the method described in any of the embodiments of the third aspect.
[0052] In one possible implementation, the model inference device further includes at least one of the aforementioned memories. Optionally, the memory and processor are integrated together.
[0053] In a tenth aspect, embodiments of this application provide a model inference apparatus, which includes logic circuitry and an interface, the logic circuitry and the interface being coupled; the interface is used to input and / or output information, and the logic circuitry is used to implement the method described in any of the embodiments of the first to third aspects.
[0054] In one possible implementation of the tenth aspect, the model inference device is a chip or chip system.
[0055] Eleventhly, embodiments of this application provide a model inference system, which includes a first communication device, a second communication device, and a third communication device, which are communicatively connected. The first communication device is used to implement the method of any embodiment of the first aspect, the second communication device is used to implement the method of any embodiment of the second aspect, and the third communication device is used to implement the method of any embodiment of the third aspect.
[0056] In a twelfth aspect, embodiments of this application provide a computer-readable storage medium for storing instructions or computer programs; when the instructions or computer programs are executed, they implement the method of any one of the embodiments of the first to third aspects.
[0057] In a thirteenth aspect, this application provides a computer program product including computer instructions that, when executed on at least one processor, can implement the methods described in any of the first to third aspects or any possible implementations thereof. Exemplarily, the computer program product can be a software installation package, which can be downloaded and executed on a computing device when the aforementioned methods are required.
[0058] The beneficial effects of the technical solutions provided in aspects four to thirteen of this application can be referred to the beneficial effects of the technical solutions in aspects one to three, and will not be repeated here. Attached Figure Description
[0059] Figure 1 is a schematic diagram of the architecture of a model inference system provided in an embodiment of this application;
[0060] Figure 2 is a schematic diagram of the architecture of another model inference system provided in an embodiment of this application;
[0061] Figure 3 is a schematic diagram of the architecture of another model inference system provided in an embodiment of this application;
[0062] Figure 4 is a schematic diagram of an O-RAN system provided in an embodiment of this application;
[0063] Figure 5 is a diagram showing the network element function division and protocol layer structure of an O-RAN system provided in an embodiment of this application;
[0064] Figure 6 is a schematic diagram of an end-to-end collaborative reasoning provided in an embodiment of this application;
[0065] Figure 7 is a flowchart illustrating a model reasoning method provided in an embodiment of this application;
[0066] Figure 8 is a flowchart illustrating another model reasoning method provided in an embodiment of this application;
[0067] Figure 9a is a schematic diagram of a model reasoning process provided in an embodiment of this application;
[0068] Figure 9b is a schematic diagram of another model reasoning process provided in an embodiment of this application;
[0069] Figure 9c is a schematic diagram of another model reasoning process provided in an embodiment of this application;
[0070] Figure 10 is a schematic diagram of the structure of a model inference device 100 provided in an embodiment of this application;
[0071] Figure 11 is a schematic diagram of another model inference device 110 provided in an embodiment of this application;
[0072] Figure 12 is a schematic diagram of the structure of another model inference device 120 provided in an embodiment of this application;
[0073] Figure 13 is a schematic diagram of the structure of another model reasoning device 130 provided in an embodiment of this application. Detailed Implementation
[0074] The embodiments of this application will now be described in detail with reference to the accompanying drawings.
[0075] The system architecture used in the embodiments of this application is described below. It should be noted that the system architecture and business scenarios described in this application are for the purpose of more clearly illustrating the technical solutions of this application, and do not constitute a limitation on the technical solutions provided in this application. As those skilled in the art will know, with the evolution of system architecture and the emergence of new business scenarios, the technical solutions provided in this application are also applicable to similar technical problems.
[0076] Please refer to Figure 1, which is a schematic diagram of the architecture of a model inference system provided in an embodiment of this application. As shown in Figure 1(a), the model inference system includes a first communication device 101 and a second communication device 102. Optionally, the model inference system further includes a third communication device 103.
[0077] Optionally, the first communication device 101, the second communication device 102, and the third communication device 103 can be of the same type or different types. For example, as shown in Figure 1(b), the first communication device 101 is a terminal, the second communication device 102 is a network device, and the third communication device 103 is a network management device. As shown in Figure 1(c), the first communication device 101 is a terminal, the second communication device 102 is a terminal, and the third communication device 103 is a network device. The architecture of the model inference system will be described in detail below, taking the first communication device 101 as a terminal, the second communication device 102 as a network device, and the third communication device 103 as a network management device as an example.
[0078] It is understood that when the model inference system only includes the first communication device 101 and the second communication device 102, the model inference system only shows one terminal and one network device. In actual use, an architecture of at least one terminal and / or at least one network device can be adopted as needed (e.g., the architecture shown in Figure 1(a)). For example, the model inference system shown in Figure 2 includes one network device and multiple terminals, or includes multiple network devices and one terminal. A single terminal can send parameter information to a single network device, and a single network device can send an uplink bandwidth prediction range to a single terminal.
[0079] It is understood that, in the case where the model inference system includes a first communication device 101, a second communication device 102, and a third communication device 103, the model inference system illustrates a terminal, a network device, and a network management device. In actual use, an architecture of at least one terminal and / or at least one network device and / or at least one network device can be adopted as needed. For example, a single terminal (e.g., represented as terminal 1) can send parameter information to a single network device (e.g., represented as network device 1). Network device 1 can send parameter information, measured channel quality index information of terminal 1, and its own load to a single network management device (e.g., represented as network management device 1). Network management device 1 can send the uplink bandwidth prediction range or model segmentation method to network device 1, and then network device 1 forwards the uplink bandwidth prediction range or model segmentation method to terminal 1.
[0080] In this embodiment, the terminal involved may include various handheld devices, vehicle-mounted devices, wearable devices, computing devices, or other processing devices connected to a wireless modem with wireless communication capabilities. The terminal 220 shown in Figure 2 may also be referred to as user equipment (UE), mobile station (MS), mobile terminal (MT), etc., or a device used to provide voice or data connectivity to users, or an Internet of Things (IoT) device. For example, the terminal includes handheld devices and vehicle-mounted devices with wireless connectivity. Currently, terminals can include: mobile phones, tablets, laptops, PDAs, mobile internet devices (MIDs), wearable devices (such as smartwatches, smart bracelets, pedometers, smart glasses, etc.), in-vehicle equipment (such as cars, bicycles, electric vehicles, airplanes, ships, trains, high-speed trains, etc.), satellite terminals, virtual reality (VR) devices, augmented reality (AR) devices, point-of-sale (POS) machines, customer-premises equipment (CPE), light user equipment (UE), reduced capability user equipment (REDCAP UE), wireless terminals in industrial control, smart home devices (such as refrigerators, televisions, air conditioners, electricity meters, etc.), intelligent robots, robotic arms, workshop equipment, wireless terminals in autonomous driving, wireless terminals in telemedicine, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, or wireless terminals in smart homes, and flying equipment (such as intelligent robots, hot air balloons, drones, airplanes), etc. The terminal can also be a vehicle device, such as a vehicle unit, vehicle module, vehicle chip, on-board unit (OBU), or telematics box (T-BOX). The terminal can also be other devices with terminal functions. For example, the terminal can also be a device that plays the role of a terminal in D2D communication.
[0081] Typically, network device 210 can be a node in a radio access network (RAN), such as a wireless relay device and / or a wireless backhaul device (not shown in Figure 2). Network device 210 may also be referred to as an access network device or a RAN node (or device), forming part of a communication system to help terminals achieve wireless access. Network device 210 can also be a 3rd generation partnership project (3GPP) related cellular system, such as a 4th generation (4G) mobile communication system, a 5th generation (5G) mobile communication system, an NTN (non-terrestrial network) system, or a future-oriented evolution system (such as a 6th generation (6G) mobile communication system). Network device 210 can also be an open RAN (O-RAN or ORAN), a cloud radio access network (CRAN), or a wireless fidelity (WiFi) system, or a communication system integrating two or more of the above systems.
[0082] In the communication system 2000, multiple network devices 210 can be nodes of the same type or different types. In some scenarios, the roles of network devices 210 and terminals 220 are relative. For example, in Figure 2, network element 220i can be a helicopter or drone, which can be configured as a mobile base station. For terminals 220j accessing the RAN 200 through network element 220i, network element 220i is a base station; however, for base station 210a, network element 220i is a terminal. Network devices 210 and terminals 220 are sometimes referred to as communication devices. For example, in Figure 2, network elements 210a and 210b can be understood as communication devices with base station functions, and network elements 220a-220j can be understood as communication devices with terminal functions. Terminal 220 connects to network device 210 wirelessly. Network device 210 connects to the core network wirelessly or via a wired connection. The core network equipment and network equipment 210 in the core network can be different physical devices, or they can be the same physical device that integrates core network logical functions and wireless access network logical functions.
[0083] In one possible scenario, network equipment can be a base station, an evolved NodeB (eNodeB), a transmitting and receiving point (TRP), a transmitting point (TP), a next-generation NodeB (gNB), a base station in a future mobile communication system, a satellite, or an access point (AP) in a WiFi system, an integrated access and backhaul (IAB) node, or a network device in a mobile switching center non-terrestrial network (NTN) communication system, i.e., it can be deployed on a high-altitude platform or satellite, etc. Network equipment can be a macro base station (as shown in Figure 2, 210a), a micro base station or indoor station (as shown in Figure 2, 210b), a relay node or donor node, or a wireless controller in a cloud radio access network (CRAN) scenario. Network equipment can also function as a base station in device-to-device (D2D) communication, vehicle-to-everything (V2X) communication, drone communication, and machine-to-machine (M2M) communication. Alternatively, network equipment can also be servers, wearable devices, vehicles, or in-vehicle equipment. For example, the access network equipment in vehicle-to-everything (V2X) technology can be a roadside unit (RSU).
[0084] In another possible scenario, multiple network devices collaborate to assist terminals in achieving wireless access, with each device performing a portion of the base station's functions. For example, these network devices could be a central unit (CU), a distributed unit (DU), a CU-control plane (CP), a CU-user plane (UP), or a radio unit (RU). The CU and DU can be configured separately or included in the same network element, such as the baseband unit (BBU). The CU and DU nodes separate the gNB's protocol layers; some protocol layer functions are centrally controlled by the CU, while the remaining partial or complete protocol layer functions are distributed across the DU, which is centrally controlled by the CU. As one implementation, the CU deploys the Radio Resource Control (RRC) layer, PDCP layer, and Service Data Adaptation Protocol (SDAP) layer in the protocol stack; the DU deploys the Radio Link Control (RLC) layer, Media Access Control (MAC) layer, and Physical Layer (PHY) in the protocol stack. Thus, the CU has the processing capabilities of RRC, PDCP, and SDAP. The DU has the processing capabilities of RLC, MAC, and PHY. It is understood that the above functional division is merely an example and does not constitute a limitation on the CU and DU. The RU can be included in radio equipment or radio units, such as in a remote radio unit (RRU), active antenna unit (AAU), or remote radio head (RRH). It is understood that the network device can be a CU node, a DU node, or a device including both CU and DU nodes. Furthermore, the CU can be classified as a network device in the access network RAN or as a network device in the core network CN; there is no restriction on this.
[0085] In this application, the core network equipment refers to equipment in the core network (CN) that provides service support to the terminal. Examples of core network equipment include: access and mobility management function (AMF) entities, session management function (SMF) entities, user plane function (UPF) entities, etc., which are not listed here. The AMF entity is responsible for terminal access management and mobility management; the SMF entity is responsible for session management, such as user session establishment; and the UPF entity can be a user plane functional entity, primarily responsible for connecting to external networks. It should be noted that in this application, entities can also be referred to as network elements or functional entities. For example, an AMF entity can also be called an AMF network element or an AMF functional entity, and an SMF entity can also be called an SMF network element or an SMF functional entity, etc.
[0086] In this embodiment, the network management device involved can be a network operations administration and maintenance (OAM) network element or a service management and orchestration (SMO) network element. The OAM network element includes a network management system (NMS) and an element management system (EMS). The NMS, also known as a cross-domain management system, is responsible for the operation, management, and maintenance of the network. The EMS, also known as a domain management system or single-domain management system, manages one or more network elements of a specific category. The NMS can directly manage the EMS. The EMS in the RAN domain can directly manage network elements in the RAN domain, such as base stations (gNodeB, gNB). The EMS in the CN domain can directly manage network elements in the CN domain, such as network data analytics function (NWDAF) network elements. The gNB exists in the RAN domain.
[0087] Please refer to Figure 3, which is a schematic diagram of the architecture of another model inference system provided in this application embodiment. As shown in Figure 3, the model inference system includes OAM network element 301 or SMO network element 301, RAN 302 or CN 302, and UE 303. Further, RAN 302 or CN 302 includes gNB, task execution function (TEF) network element, and task control function (TCF) network element.
[0088] In the O-RAN network architecture, the SMO network element can directly manage various heterogeneous network elements such as gNB, NWDAF, TEF, and TCF. Data / model information in the network can be exposed to the UE303's vendor server (over-the-top, OTT) through the NMS / SMO and interact with it. The UE303's data / model information can also interact with the OTT.
[0089] The TEF network element is responsible for executing the computing tasks assigned to it, such as model inference. It is the functional module that actually runs the task or application and exists in the RAN or CN domain. The TCF network element is responsible for managing and controlling the scheduling, allocation, and resource allocation of computing tasks. It is mainly a task management control layer that ensures that tasks run on appropriate resources and coordinates various execution units as needed. It exists in the control plane of the RAN or CN domain.
[0090] Optionally, the method provided in this application embodiment can also be applied to an O-RAN system. Please refer to Figure 4, which is a schematic diagram of an O-RAN system provided in this application embodiment. The O-RAN system may also include other components besides those shown in Figure 4, and this application does not limit this. Optionally, the network device shown in Figure 4 can be an access network device, such as an eNB, gNB, or next-generation access network device. The access network device communicates with the core network (CN) via a backhaul link and with the terminal via an air interface.
[0091] The BBU in the access network equipment communicates with the core network via a backhaul link, and the RU in the access network equipment communicates with at least one terminal via an air interface. The BBU communicates with at least one RU via a fronthaul link. The BBU and RU may or may not be co-located. The BBU includes at least one control unit (CU) and at least one distributed unit (DU), which can communicate via at least one midhaul link.
[0092] Further optionally, please refer to Figure 5. Figure 5 is a diagram illustrating the network element functional division and protocol layer structure of an open radio access network (O-RAN) system provided in an embodiment of this application. As shown in Figure 5, in some examples, the CU is a logical node carrying the RRC layer, Service Data Adaptation Protocol (SDAP) layer, Packet Data Convergence Protocol (PDCP) layer, and other control functions of the access network equipment. The CU is connected to network nodes such as the core network through some interfaces, which may be interfaces such as E2 interfaces. Optionally, the CU may have some functions of the core network, such as the PDCP layer and higher layers. The CU is connected to the DU (e.g., RLC layer and lower layers) through some interfaces, which may be interfaces such as F1 interfaces. In some examples, these interfaces (e.g., the F1 interface) can provide control plane (C-Plane) and user plane (U-Plane) functions (e.g., interface management, system information management, UE context management, RRC message transmission, etc.). F1AP is the application protocol of the F1 interface, and in some examples, the signaling procedures of F1 are defined. The F1 interface supports the control plane F1-C and the user plane F1-U.
[0093] In some examples, the CU can be split into CU-CP (control unit-control plane) and CU-UP (control unit-user plane). CU-CP is a logical node carrying the RRC layer and PDCP-C (control plane part of PDCP) layer, used to implement the CU's control plane functions. CU-CP can interact with network elements in the core network used to implement control plane functions. These network elements in the core network can be access and mobility function (AMF) network elements, such as the access and mobility management function (AMF) in a 5G mobile communication system. AMF network elements are responsible for mobility management in the mobile network, such as terminal location updates, terminal registration with the network, and terminal handover. CU-UP is a logical node carrying the SDAP layer and the PDCP-U (user plane part of PDCP) layer for user plane data, used to implement the CU's user plane functions. CU-UP can interact with network elements in the core network used to implement user plane functions. These network elements in the core network, such as the UPF (user plane function) in a 5G system, are responsible for data forwarding and receiving in terminal devices. It should be understood that the above configurations of CU and DU are merely examples, and the functions of CU and DU can be configured as needed. This application does not impose excessive limitations on this. For example, CU or DU can be configured to have more protocol layer functions, or CU or DU can be configured to have some protocol layer processing functions. Another example is to place some functions of the RLC layer and the protocol layer functions above the RLC layer in the CU, and place the remaining functions of the RLC layer and the protocol layer functions below the RLC layer in the DU. Yet another example is that the functions of CU or DU can be divided according to service type or other system requirements, such as by latency, placing functions that need to meet low latency requirements in the DU, and functions that do not need to meet this latency requirement in the CU.
[0094] In some examples, a DU is a logical node that carries the radio link control (RLC) layer, medium access control (MAC) layer, higher physical layer (PHY) layer, and other functions. In some examples, a DU can control at least one RU. The DU connects to the RU through interfaces, which can be fronthaul interfaces.
[0095] In some examples, the CU may not have a PDCP layer, i.e., it only includes the RRC layer. CU-CP does not have PDCP-C. CU-UP may not have PDCP-U, or may not have CU-UP at all. In some examples, the DU may not have an RLC layer, only a MAC and a higher PHY layer. Furthermore, in some examples, it may not have a CU and may only include the DU.
[0096] In some examples, the higher PHY layer includes parts of the PHY layer that handle processing, such as forward error correction (FEC) encoding and decoding, scrambling, modulation, and demodulation.
[0097] In some examples, the RU is a logical node that carries both lower physical layer (PHY) and radio frequency chain (RF chain) processing. In some examples, the RU can be a 3GPPTRP, a remote radio head (RRH), or other similar functionalities. In some examples, the Low-PHY includes PHY processing functions such as Fast Fourier Transform (FFT), Inverse Fast Fourier Transform (IFFT), digital beamforming, and filtering. The RU communicates with one or more terminals via a wireless link.
[0098] Optionally, the DU and RU may or may not be co-located. The DU and RU exchange control plane information via a fronthaul link through a lower-layer split-control, user plane information (LLS-CUS) and synchronization interface. The LLS-CUS may include LLS-C and LLS-U interfaces that respectively provide the control plane (C-Plane) and user plane (U-Plane). In some examples, the control plane (C-Plane) refers to real-time control between the DU and RU. The DU and RU exchange management information via an LLS-M interface on the fronthaul link; the management plane (M-Plane) refers to non-real-time management operations between the DU and RU.
[0099] Optionally, the DU and RU can cooperate to implement the functions of the PHY layer. A DU can be connected to one or more RUs. The functions of the DU and RU can be configured in various ways depending on the design. For example, the DU can be configured to implement baseband functions, and the RU can be configured to implement mid-RF functions. Alternatively, the DU can be configured to implement higher-level functions in the PHY layer, and the RU can be configured to implement lower-level functions in the PHY layer, or to implement both lower-level and RF functions. Higher-level functions in the physical layer may include a portion of the physical layer's functions that are closer to the MAC layer, while lower-level functions in the physical layer may include another portion of the physical layer's functions that are closer to the mid-RF side.
[0100] In different systems, CU (or CU-CP and CU-UP), DU, or RU may have different names, but those skilled in the art will understand their meaning. For example, in an ORAN system, CU can also be called O-CU (Open CU), DU can also be called O-DU, CU-CP can also be called O-CU-CP, CU-UP can also be called O-CU-UP, and RU can also be called O-RU. For ease of description, this application uses CU, CU-CP, CU-UP, DU, and RU as examples. The network device deployment methods listed here are only examples; as standard technologies evolve, network devices may have other deployment forms.
[0101] To address the issue that the massive computational demands of AI / ML model-based inference applications often exceed the computing capabilities of the edge devices, a current solution involves splitting the AI / ML model between the edge devices and the network side, allocating some or even most of the computation to the network side through edge-network collaborative inference. Some model inference applications, such as autonomous driving, have extremely high requirements for end-to-end inference latency. Edge-network collaborative inference, while addressing the insufficient computing resources on the edge devices, must also meet the end-to-end latency requirements of inference applications.
[0102] The end-to-end latency is the sum of the end-side inference latency, the intermediate data transmission latency, and the network-side inference latency. Since the sum of the end-side inference model and the network-side inference model constitutes a complete model, we can assume that the sum of the end-side inference latency and the network-side inference latency remains constant, and the end-to-end latency is only affected by the intermediate data transmission latency. When the uplink channel bandwidth between the end-side and network sides decreases, the intermediate data transmission latency increases, and the end-to-end latency increases; when the uplink channel bandwidth between the end-side and network sides increases, the intermediate data transmission latency decreases, and the end-to-end latency decreases. To maintain the required end-to-end latency even with reduced uplink channel bandwidth, the amount of intermediate data transmission needs to be reduced to ensure that the intermediate data transmission latency still meets the requirements.
[0103] To ensure that end-to-end latency remains acceptable despite changes in uplink channel bandwidth, multiple model partitioning methods can be devised based on different intermediate data output dimensions. Each model partitioning method corresponds to a deployment scheme for the model on both the endpoint and network sides. Multiple models with different partitioning methods are pre-deployed on both the endpoint and network sides. When uplink channel bandwidth decreases, a model partitioning method with lower intermediate data output dimensions is used based on latency requirements; when uplink channel bandwidth increases, a model with higher intermediate data dimensions is used based on latency requirements.
[0104] Taking Figure 6 as an example, in some solutions, End device model 1 and End device model 2 are pre-loaded on the device side, and Network server model 1 and Network server model 2 are pre-loaded on the network side. When the uplink channel bandwidth is good, both the device and network sides use model 1 for inference, resulting in a higher intermediate output dimension but fewer inference layers on the device side. When the uplink channel bandwidth decreases, both the device and network sides use model 2 for inference. The reason for pre-loading two sets of models instead of temporarily reloading new models when the uplink channel bandwidth changes is that temporarily loading new models takes time, causing inference service interruption. To ensure that the inference service is not interrupted when the model is changed due to uplink bandwidth changes, both the device and network sides need to pre-load all multiple sets of models and select the appropriate model to use when the uplink bandwidth changes. Loading all models when only one model is used at a time results in a significant waste of resources on both the device and network sides.
[0105] In view of this, this application provides a model inference method and related apparatus. Taking a first communication device as the terminal and a second communication device as a network device as an example, on the one hand, since some model inference applications have high requirements for end-to-end inference latency, the terminal can determine the first model segmentation method corresponding to the first time period based on the current uplink bandwidth. This first model segmentation method corresponds to a deployment scheme of a model on the terminal side and the network device side. End-to-end latency = terminal-side inference latency + intermediate data transmission latency + network-side inference latency. Since the sum of the terminal-side inference model and the network device-side inference model is equivalent to a complete model, it can be assumed that the sum of the terminal-side inference latency and the network device-side inference latency remains unchanged, and the end-to-end latency is only affected by the intermediate data transmission latency. When the uplink channel bandwidth between the terminal side and the network device side decreases, the intermediate data transmission latency increases, and the end-to-end latency increases. In this case, the terminal can use a model segmentation method with a lower intermediate data output dimension according to the latency requirements when the uplink channel bandwidth decreases, thereby effectively ensuring end-to-end latency.
[0106] On the other hand, compared to existing solutions where the terminal side and network device side need to preload the structure of the model used under various model partitioning methods and the union of the model parameter sets corresponding to the model structure, and select the appropriate model structure and model parameters to use when the uplink bandwidth changes, this application can load the corresponding model parameters that meet the uplink bandwidth conditions according to the uplink bandwidth as required. There is no need to preload the structure of the model used under various model partitioning methods and the union of the model parameter sets corresponding to the model structure, thereby effectively reducing the waste of computing resources on the terminal side and network device side.
[0107] In summary, this application can reduce the waste of computing resources on the terminal side and network device side while ensuring end-to-end latency.
[0108] In the model reasoning method shown below (as shown in Figure 7 or Figure 8), the specific descriptions of the first communication device, the second communication device, and / or the third communication device can be found in Figures 1 to 5, and will not be detailed here. For ease of description, in the embodiments of this application, specific examples may be used to illustrate the first communication device as a terminal, the second communication device as a network device, and the third communication device as a network management device, but this should not be construed as a limitation on the embodiments of this application.
[0109] Please refer to Figure 7, which is a flowchart illustrating a model inference method provided in an embodiment of this application. Optionally, this method can be applied to a model inference system, such as the model inference systems shown in Figures 1 to 5.
[0110] The method shown in Figure 7 may include steps S701-S705. It should be understood that this application describes the steps in the order of S701-S705 for ease of description, and is not intended to limit the execution to this specific order. This application's embodiments do not limit the order of execution, the execution time, or the number of executions of one or more of the above steps. Steps S701-S705 are as follows:
[0111] Step S701: The first communication device downloads the structure of the first model and the set of model parameters corresponding to the structure of the first model.
[0112] It should be noted that downloading refers to storing the model on the local hard drive, but does not enable inference services. The model must be further loaded after downloading before it can be used for inference services.
[0113] The first model is used for end-to-end collaborative inference services. For example, the sum of the model inferred by the first communication device and the model inferred by the second communication device and / or network nodes that can load the model is equivalent to a complete model. The first model is the model inferred by the first communication device, that is, it is part of a complete model.
[0114] The structure of the first model is the union of the structures of the models used by the first communication device under various model segmentation methods. In other words, the first communication device can first download the union of the structures of the models used for inference under various model segmentation methods and the set of model parameters of the model structure.
[0115] The model segmentation method can also be called the model segmentation point. The model segmentation method is the specific way the model is segmented. For example, the model can be segmented by layers (e.g., the first X layers of the model are deployed on the first communication device, and the remaining X layers are deployed on the second communication device), or it can be segmented diagonally or by modules. It should be noted that the above model segmentation methods are only examples, and this application does not limit the specific method of model segmentation.
[0116] Optionally, the model segmentation method can be specified by the protocol or determined by the first communication device after receiving the uplink channel bandwidth prediction range.
[0117] For example, multiple model segmentation methods include model segmentation method 1, model segmentation method 2, and model segmentation method 3. The model parameter set corresponding to the structure of the first model includes model parameter 1, model parameter 2, and model parameter 3. In this scheme, the first communication device downloads the structure of the model segmented by model segmentation method 1, model segmentation method 2, and model segmentation method 3, as well as the corresponding model parameters 1, 2, and 3. Referring to Figure 6, since the union of end device model1 and end device model2 equals end device model2, the first communication device only needs to download the structure of end device model2 and the corresponding model parameters 1, 2, and 3 (e.g., the first three layers). End device model1 and end device model2 are both part of the first model. During operation, the first communication device loads the first layer (e.g., represented by model parameter 1) if it needs to load one layer, and the first two layers (e.g., represented by model parameter 1 and model parameter 2) if it needs to load two layers.
[0118] Step S702: The first communication device determines the first model segmentation method corresponding to the first time period based on the current uplink bandwidth.
[0119] The first model segmentation method is related to the set of model parameters corresponding to the structure of the first model. For example, referring to Figure 6, when the uplink channel bandwidth between the first and second communication devices decreases, the intermediate data transmission delay increases, and the end-to-end delay increases. In this case, the first model segmentation method corresponding to the first time period can be: the first communication device can use a model segmentation method with a lower intermediate data output dimension according to the delay requirements when the uplink channel bandwidth decreases, that is, use end device model 1 for inference. When the uplink channel bandwidth between the first and second communication devices increases, the intermediate data transmission delay decreases, and the end-to-end delay decreases. In this case, the first model segmentation method corresponding to the first time period can be: the first communication device can use a model segmentation method with a higher intermediate data output dimension according to the delay requirements when the uplink channel bandwidth increases, that is, use end device model 2 for inference, thereby ensuring end-to-end delay.
[0120] Step S703: The first communication device loads the first model parameters corresponding to the first model segmentation method.
[0121] The first model parameter belongs to the model parameter set and is used for end-to-end collaborative inference services. For example, the first model parameter is an exemplary name used to distinguish a particular model parameter. For example, the first model parameter can be model parameter 1 or other model parameters. The first communication device loads model parameter 1 and model parameter 2 corresponding to model segmentation method 1. Taking Figure 6 as an example, the first model segmentation method involves segmentation after the second layer, generating end device model1 and network server model1. Correspondingly, the first model parameter is all the parameters of end device model1. The second model segmentation method involves segmentation after the third layer, generating end device model2 and network server model2. Correspondingly, the second model parameter is all the parameters of end device model2.
[0122] It should be noted that loading refers to loading the model into the running memory or computing units or processors such as GPUs. After downloading, the model must be further loaded before it can perform inference services. Compared to existing solutions where the first and second communication devices need to pre-load the structure of the model used under various model partitioning methods and the union of the model parameter sets corresponding to the model structure, and select the appropriate model structure and corresponding model parameters when the uplink bandwidth changes, this application only needs to load the corresponding model parameters that meet the uplink bandwidth conditions according to the output data dimension. It does not require pre-loading the structure of the model used under various model partitioning methods and the union of the model parameter sets corresponding to the model structure, thereby effectively reducing the waste of computing resources on the terminal side and the network device side.
[0123] Optionally, the first communication device may first download the union of the structures of the models used for inference under multiple model segmentation methods and the set of model parameters for the model structures, and determine the first model segmentation method corresponding to the first time period based on the current uplink bandwidth, and load the first model parameters corresponding to the first model segmentation method. Alternatively, the first communication device may only download the union of the structures of the models used for inference under the current segmentation method and the set of model parameters for the model structures, and determine the first model segmentation method corresponding to the first time period based on the current uplink bandwidth, and load the first model parameters corresponding to the first model segmentation method.
[0124] Step S704: The first communication device repeatedly executes the first operation until the end-to-end collaborative inference service ends within a preset time period.
[0125] The preset time period consists of at least one time cycle, and the first time cycle belongs to at least one time cycle. Optionally, the preset time period can consist of one time cycle. For example, if the preset time period is 10 minutes and one time cycle is 10 minutes, then the first time cycle is also 10 minutes. Optionally, the preset time period can also consist of multiple time cycles. For example, if the preset time period is 10 minutes and consists of 10 time cycles, and one time cycle is 1 minute, then the first time cycle is 1 minute.
[0126] The first operation includes, but is not limited to, the following steps:
[0127] Step 1: The first communication device determines the performance indicator associated with the first communication device based on its own battery power, remaining storage space, CPU load, GPU load, and temperature.
[0128] The performance indicator is related to the uplink transmission rate. For example, factors affecting the uplink transmission rate include: the uplink bandwidth allocated to the first communication device by the second communication device or other devices, and whether the first communication device actively limits the uplink transmission rate (for example, if the battery power of the first communication device is less than a preset value, it may limit the uplink transmission rate to increase standby time and avoid rapid battery depletion). A low value of the performance indicator (e.g., the current value of the performance indicator is less than the preset value of the performance indicator) usually indicates that the first communication device will actively reduce the uplink transmission rate. However, if the value of the performance indicator is high (e.g., the current value of the performance indicator is greater than the preset value of the performance indicator), the first communication device generally will not actively increase the uplink transmission rate, but will need to adjust the uplink transmission rate through the second communication device or other devices.
[0129] For example, during a first time period (e.g., the first time period is 1 minute), the battery power of the first communication device is 10%, the remaining storage space is 55.1G, the CPU load is 30%, the GPU load is 20%, and the temperature of the first communication device is 37.5°C. Based on the above information, the first communication device determines that the value of the performance indicator associated with the first communication device is 29, which is less than the preset value of the performance indicator of 50, indicating that the current value of the performance indicator is low.
[0130] Optionally, the identifier of the performance indicator associated with the first communication device can be configured as 1 to indicate that the value of the performance indicator associated with the first communication device is low; or the identifier of the performance indicator associated with the first communication device can be configured as 0 to indicate that the value of the performance indicator associated with the first communication device is low. It should be understood that the above is only one possible case shown for ease of description and is not intended to limit the specific value of the index in the embodiments of this application.
[0131] Step 2: The first communication device sends parameter information to the second communication device.
[0132] The parameter information includes Channel Quality Indicator (CQI), Service Quality Requirements of the End-to-End Collaborative Inference Service, Inference Output Dimension of the First Communication Device, Inference Frequency, and Performance Indicator. The parameter information is used by the Second Communication Device to predict the uplink bandwidth of the First Communication Device.
[0133] Optionally, the parameter information is information sent when a secure communication connection is established between the first communication device and the second communication device.
[0134] The following are examples of the information in the parameter information:
[0135] (1) CQI is a channel quality indicator, representing the current channel quality and corresponding to the signal-to-noise ratio (SNR) of the channel. Its value typically ranges from 0 to 31. For example, a CQI value of 0 indicates the worst current channel quality, while a CQI value of 31 indicates the best current channel quality. The second communication device can determine the transmitted data block size, the HS-PDSCH channel code value, the encoding method, and the modulation method based on the CQI value.
[0136] (2) The quality of service requirements of the end-to-end network collaborative inference service can be used by the second communication device to allocate bandwidth to various services of the end-to-end network collaborative inference service in a balanced manner under limited bandwidth resources, and provide end-to-end quality of service guarantees for the different needs of various services. In other words, it can be understood as configuring different traffic processing methods according to priority, such as prioritizing important traffic and delaying or discarding unimportant traffic.
[0137] (3) The inference output dimension of the first communication device can be used to predict the uplink channel bandwidth. For example, when the uplink channel bandwidth between the terminal side and the network device side decreases, the intermediate data transmission delay increases and the end-to-end delay increases. In this case, the terminal can use a model segmentation method with a lower intermediate data output dimension according to the delay requirements when the uplink channel bandwidth decreases. When the uplink channel bandwidth between the terminal side and the network device side increases, the intermediate data transmission delay decreases and the end-to-end delay decreases. In this case, the terminal can use a model segmentation method with a higher intermediate data output dimension according to the delay requirements when the uplink channel bandwidth increases. Therefore, the inference output dimension of the first communication device (optionally, the inference output dimension of the first communication device can be the current inference output dimension of the first communication device or the historical inference output dimension of the first communication device) can be used by the second communication device to predict the uplink channel bandwidth in the next period of the first time period.
[0138] (4) Inference frequency refers to the frequency of inference if the inference service is related to video, and inference is performed every frame or every few frames. Therefore, the inference frequency will be higher than that of image recognition-based inference services. Consequently, the required uplink bandwidth will differ when the output dimensions of the first communication devices are similar. Thus, the inference frequency can be uploaded to the second communication device for differentiation. Optionally, the inference frequency can be statistically determined by the first communication device itself based on a certain time range, or it can be an indicator preset by the first communication device according to the service type.
[0139] Step 3: The first communication device determines the second model segmentation method corresponding to the next time period of the first time period.
[0140] The second model segmentation method is related to the set of model parameters corresponding to the first model.
[0141] For example, referring to Figure 6, the first model segmentation method corresponds to end device model 1, and the second model segmentation method corresponds to end device model 2. In the first time period, the first model segmentation method corresponding to end device model 1 is used. When the uplink bandwidth of the predicted next time period decreases within the first time period, in order to meet the end-to-end latency requirement, the second model segmentation method corresponding to end device model 2 can be used in the next time period. In this scheme, the computational load of the first communication device can be reduced while ensuring end-to-end latency. Specifically, when the uplink bandwidth decreases, end-to-end latency can be guaranteed; when the uplink bandwidth increases, the resources of the first communication device can be saved while ensuring end-to-end latency.
[0142] Optionally, the first communication device needs to determine the second model segmentation method corresponding to the next time period of the first time period based on other information. For example, the first communication device receives an uplink bandwidth prediction range from the second communication device and determines the second model segmentation method corresponding to the next time period of the first time period based on the uplink bandwidth prediction range.
[0143] Step 4: The first communication device determines the model parameter processing strategy according to the second model segmentation method.
[0144] The following are two possible implementations of the first communication device determining the model parameter processing strategy based on the second model segmentation method, as exemplified below:
[0145] In the first implementation method, the first communication device loads the model structure and corresponding model parameters of the second model segmentation method that are added to the first model segmentation method in the first time period, and / or unloads the model structure and corresponding model parameters of the second model segmentation method that are reduced in the first model segmentation method in the next time period of the first time period.
[0146] It should be noted that unloading in the next time period after the first time period means that unloading can be done immediately after the first time period ends, without waiting for the next time period to end. For example, this solution can be applied to scenarios with relatively simple segmentation methods. For instance, if the first model has already loaded layer X, then layer Y can be loaded in the first time period, or layer Z can be unloaded in the next time period (where X, Y, and Z are positive integers, and Z < X). In this solution, missing model structures and their corresponding model parameters can be loaded in advance, and redundant model structures and their corresponding model parameters can be unloaded in the next time period. This allows for more adaptive and flexible loading of model structures and their corresponding model parameters, avoiding the resource waste caused by preloading the entire model structure and its corresponding model parameter set.
[0147] Optionally, the added model structure and the corresponding model parameters are minimized to save space resources. In the second implementation, the first communication device loads the model parameters corresponding to the second model segmentation method in the first time period, and / or runs the end-to-end collaborative inference service on the model parameters corresponding to the second model segmentation method in the next time period after the first time period, and unloads the model parameters corresponding to the first model segmentation method.
[0148] For example, this solution can be applied to scenarios with relatively complex segmentation methods. For instance, if the first model has already loaded layer X1, then layer X2 can be loaded in the first time period, or the system can switch to layer X2 for end-to-end collaborative inference service in the next time period, and the previous layer X1 can be unloaded. In this solution, missing model structures and their corresponding model parameters can be loaded in advance, and redundant model structures and their corresponding model parameters can be unloaded in the next time period. This allows for more adaptive and flexible loading of model structures and their corresponding model parameters, and also avoids the resource waste caused by preloading the entire model structure and its corresponding model parameter set.
[0149] It should be noted that the initial loading of the first model parameters corresponding to the first model segmentation method by the first communication device is outside the loop. After the initial loading of the first model parameters corresponding to the first model segmentation method, the loop begins with optimizing model loading based on the uplink channel bandwidth. In other words, the loop starts from the first communication device determining the performance indicator (including this step), followed by information exchange between the first and second communication devices, the distribution of the uplink bandwidth prediction range, and the first communication device determining the loading of model parameters for the next time period based on the uplink bandwidth prediction range. This completes one cycle of the loop. Optionally, the next loop begins when certain conditions are met.
[0150] As one possible implementation, the first operation is executed repeatedly when the triggering condition is met.
[0151] For example, the triggering condition includes at least one of the following:
[0152] (1) The preset triggering period is satisfied. Optionally, the preset triggering period can be a triggering frequency. Further optionally, the preset triggering period can be the same as or different from the first time period. For example, the preset triggering period can be the time period of the terminal-network collaborative inference service or other periods. For example, the preset triggering period can be the first operation executed cyclically every 1 minute, or the first operation executed twice cyclically every 10 minutes.
[0153] (2) Satisfying preset trigger events. For example, a preset trigger event could be the emergence of new service requirements. Or a preset trigger event could be changes in weather conditions (strong winds, heavy rain, etc.) affecting wireless signal propagation, leading to changes in network quality and requiring a reassessment of uplink channel bandwidth. Alternatively, a preset trigger event could be the operator initiating network policy optimization, adjusting bandwidth allocation for different user groups, etc.
[0154] Step S705: The second communication device repeatedly performs the second operation until the end-to-end collaborative inference service ends within a preset time period.
[0155] The preset time period consists of at least one time cycle. Optionally, the preset time period may consist of one time cycle, for example, the preset time period is 10 minutes and one time cycle is 10 minutes. Optionally, the preset time period may also consist of multiple time cycles, for example, the preset time period is 10 minutes and the preset time period consists of 10 time cycles, with one time cycle being 1 minute.
[0156] The second operation includes, but is not limited to, the following steps:
[0157] Step 1: The second communication device receives parameter information from the first communication device.
[0158] The parameter information includes Channel Quality Indicator (CQI), Quality of Service Requirements for Inference Service, Inference Output Dimension of the First Communication Device, Inference Frequency, and Performance Indicator. This parameter information is used by the Second Communication Device to predict the uplink bandwidth of the First Communication Device.
[0159] Optionally, the parameter information is information sent when a secure communication connection is established between the first communication device and the second communication device.
[0160] It should be noted that the relevant content in the parameter information has been explained in detail in step S704, and will not be repeated here.
[0161] Step 2: The second communication device measures and obtains the channel quality index information of the first communication device.
[0162] The channel quality metrics include signal-to-interference-plus-noise ratio (SINR), reference signal reception quality (RSRQ), and uplink received signal strength indicator (UL_RSSI).
[0163] For example, the second communication device can measure the SINR of the first communication device to be 10dB, the RSRQ to be -9.5, and the UL_RSSI to be -65dBm.
[0164] Step 3: The second communication device sends the uplink bandwidth prediction range to the first communication device.
[0165] Optionally, the uplink bandwidth prediction range is determined by the second communication device and sent to the first communication device.
[0166] Specifically, in the first time period, the second communication device obtains the uplink bandwidth prediction range based on the load, parameter information, and channel quality index information of the second communication device, and sends the uplink bandwidth prediction range for the next time period of the first time period to the first communication device.
[0167] The first time period is at least one time period. Optionally, when the preset time period consists of one time period, for example, the preset time period is 10 minutes and one time period is 10 minutes, then the first time period is also 10 minutes. Optionally, when the preset time period consists of multiple time periods, for example, the preset time period is 10 minutes and the preset time period consists of 10 time periods, and one time period is 1 minute, then the first time period is 1 minute.
[0168] Optionally, the second communication device may use a prediction model or other algorithm to predict the uplink bandwidth prediction range, which is not limited in this application.
[0169] Alternatively, the second communication device inputs its load, parameter information, and channel quality index information into the prediction model during the first time period to obtain the uplink bandwidth prediction range.
[0170] The prediction model is trained based on multiple sample data. The sample data includes feature data and label data. The feature data includes the historical load, historical parameter information and historical channel quality index information of the second communication device, and the label data includes the historical uplink bandwidth of the first communication device.
[0171] It should be noted that, since the above sample data cannot be exhaustively listed, the data provided in this application is only an example and should not be construed as a limitation on the sample data.
[0172] In addition, steps S704 and S705 are executed synchronously, that is, the first communication device and the second communication device interact in the same time period. For example, after the first communication device performs the first operation and the second communication device performs the second operation, the interaction is completed in the first time period (this is one cycle), and the interaction continues in the next time period of the first time period (this is the next cycle) until the end-to-end collaborative inference service ends.
[0173] Optionally, since the third communication device has more abundant computing resources and collects more data, it can collect data from a global cell perspective and run more complex machine learning models to predict the uplink channel bandwidth of the first communication device, obtain the uplink bandwidth prediction range, or provide a second model segmentation method for the terminal in the next time period of the first time period. The first communication device can perform subsequent operations according to the uplink bandwidth prediction range or the second model segmentation method, avoiding the waste of resources caused by preloading the structure of all models and the set of model parameters corresponding to the model structure.
[0174] It should be noted that in the process where only the first and second communication devices interact, steps S701-S705 form a loop (i.e., mandatory steps). However, when a third communication device participates in the above information interaction, please refer to Figure 8, which is a flowchart illustrating another model inference method provided in this application embodiment. Optionally, this method can be applied to a model inference system, such as the model inference system shown in Figures 1 to 5.
[0175] The method shown in Figure 8 may include steps S801-S806. It should be understood that this application describes the steps in the order of S801-S806 for ease of description, and is not intended to limit the execution to this specific order. This application's embodiments do not limit the order of execution, the execution time, or the number of executions of one or more of the above steps. Steps S801-S806 are as follows:
[0176] Step S801: The first communication device downloads the structure of the first model and the set of model parameters corresponding to the structure of the first model.
[0177] It should be noted that downloading refers to downloading the model to local hard drive storage, but does not enable inference services. The model must be further loaded after downloading before it can perform end-to-end collaborative inference services.
[0178] The first model is used for end-to-end collaborative inference services. For example, the sum of the model inferred by the first communication device and the model inferred by the second communication device and / or network nodes that can load the model is equivalent to a complete model. The first model is the model inferred by the first communication device, that is, it is part of a complete model.
[0179] The structure of the first model is the union of the structures of the models used by the first communication device under various model segmentation methods. In other words, the first communication device can first download the union of the structures of the models used for inference under various model segmentation methods and the set of model parameters of the model structure.
[0180] The model segmentation method can also be called the model segmentation point. The model segmentation method is the specific way the model is segmented. For example, the model can be segmented by layers (e.g., the first X layers of the model are deployed on the first communication device, and the remaining X layers are deployed on the second communication device), or it can be segmented diagonally or by modules. It should be noted that the above model segmentation methods are only examples, and this application does not limit the specific method of model segmentation.
[0181] Optionally, the model segmentation method can be specified by the protocol or determined by the first communication device after receiving the uplink channel bandwidth prediction range.
[0182] For example, multiple model segmentation methods include model segmentation method 1, model segmentation method 2, and model segmentation method 3. The model parameter set corresponding to the structure of the first model includes model parameter 1, model parameter 2, and model parameter 3. In this scheme, the first communication device downloads the structure of the model segmented by model segmentation method 1, model segmentation method 2, and model segmentation method 3, as well as the corresponding model parameters 1, 2, and 3. Referring to Figure 6, since the union of end device model1 and end device model2 equals end device model2, the first communication device only needs to download the structure of end device model2 and the corresponding model parameters 1, 2, and 3 (e.g., the first three layers). End device model1 and end device model2 are both part of the first model. During operation, the first communication device loads the first layer (e.g., represented by model parameter 1) if it needs to load one layer, and the first two layers (e.g., represented by model parameter 1 and model parameter 2) if it needs to load two layers.
[0183] Step S802: The first communication device determines the first model segmentation method corresponding to the first time period based on the current uplink bandwidth.
[0184] The first model segmentation method is related to the set of model parameters corresponding to the structure of the first model. For example, referring to Figure 6, when the uplink channel bandwidth between the first and second communication devices decreases, the intermediate data transmission delay increases, and the end-to-end delay increases. In this case, the first model segmentation method corresponding to the first time period can be: the first communication device can use a model segmentation method with a lower intermediate data output dimension according to the delay requirements when the uplink channel bandwidth decreases, that is, use end device model 1 for inference. When the uplink channel bandwidth between the first and second communication devices increases, the intermediate data transmission delay decreases, and the end-to-end delay decreases. In this case, the first model segmentation method corresponding to the first time period can be: the first communication device can use a model segmentation method with a higher intermediate data output dimension according to the delay requirements when the uplink channel bandwidth increases, that is, use end device model 2 for inference, thereby ensuring end-to-end delay.
[0185] Step S803: The first communication device loads the first model parameters corresponding to the first model segmentation method.
[0186] The first model parameter belongs to the model parameter set and is used for end-to-end collaborative inference services. For example, the first model parameter is an exemplary name used to distinguish a particular model parameter. For example, the first model parameter can be model parameter 1 or other model parameters. The first communication device loads model parameter 1 and model parameter 2 corresponding to model segmentation method 1. Taking Figure 6 as an example, the first model segmentation method involves segmentation after the second layer, generating end device model1 and network server model1. Correspondingly, the first model parameter is all the parameters of end device model1. The second model segmentation method involves segmentation after the third layer, generating end device model2 and network server model2. Correspondingly, the second model parameter is all the parameters of end device model2.
[0187] It should be noted that loading refers to loading the model into the running memory or computing units or processors such as GPUs. After downloading, the model must be further loaded before it can perform inference services. Compared to existing solutions where the first and second communication devices need to pre-load the structure of the model used under various model partitioning methods and the union of the model parameter sets corresponding to the model structure, and select the appropriate model structure and corresponding model parameters when the uplink bandwidth changes, this application only needs to load the corresponding model parameters that meet the uplink bandwidth conditions according to the output data dimension. It does not require pre-loading the structure of the model used under various model partitioning methods and the union of the model parameter sets corresponding to the model structure, thereby effectively reducing the waste of computing resources on the terminal side and the network device side.
[0188] Optionally, the first communication device may first download the union of the structures of the models used for inference under multiple model segmentation methods and the set of model parameters for the model structures, and determine the first model segmentation method corresponding to the first time period based on the current uplink bandwidth, and load the first model parameters corresponding to the first model segmentation method. Alternatively, the first communication device may only download the union of the structures of the models used for inference under the current segmentation method and the set of model parameters for the model structures, and determine the first model segmentation method corresponding to the first time period based on the current uplink bandwidth, and load the first model parameters corresponding to the first model segmentation method.
[0189] Step S804: The first communication device repeatedly executes the first operation until the end-to-end collaborative inference service ends within a preset time period.
[0190] The preset time period consists of at least one time cycle, and the first time cycle belongs to at least one time cycle. Optionally, the preset time period can consist of one time cycle. For example, if the preset time period is 10 minutes and one time cycle is 10 minutes, then the first time cycle is also 10 minutes. Optionally, the preset time period can also consist of multiple time cycles. For example, if the preset time period is 10 minutes and consists of 10 time cycles, and one time cycle is 1 minute, then the first time cycle is 1 minute.
[0191] The first operation includes, but is not limited to, the following steps:
[0192] Step 1: The first communication device determines the performance indicator associated with the first communication device based on its own battery power, remaining storage space, CPU load, GPU load, and temperature.
[0193] The performance indicator is related to the uplink transmission rate. For example, factors affecting the uplink transmission rate include: the uplink bandwidth allocated to the first communication device by the second communication device or other devices, and whether the first communication device actively limits the uplink transmission rate (for example, if the battery power of the first communication device is less than a preset value, it may limit the uplink transmission rate to increase standby time and avoid rapid battery depletion). A low value of the performance indicator (e.g., the current value of the performance indicator is less than the preset value of the performance indicator) usually indicates that the first communication device will actively reduce the uplink transmission rate. However, if the value of the performance indicator is high (e.g., the current value of the performance indicator is greater than the preset value of the performance indicator), the first communication device generally will not actively increase the uplink transmission rate, but will need to adjust the uplink transmission rate through the second communication device or other devices.
[0194] For example, during a first time period (e.g., the first time period is 1 minute), the battery power of the first communication device is 10%, the remaining storage space is 55.1G, the CPU load is 30%, the GPU load is 20%, and the temperature of the first communication device is 37.5°C. Based on the above information, the first communication device determines that the value of the performance indicator associated with the first communication device is 29, which is less than the preset value of the performance indicator of 50, indicating that the current value of the performance indicator is low.
[0195] Optionally, the identifier of the performance indicator associated with the first communication device can be configured as 1 to indicate that the value of the performance indicator associated with the first communication device is low; or the identifier of the performance indicator associated with the first communication device can be configured as 0 to indicate that the value of the performance indicator associated with the first communication device is low. It should be understood that the above is only one possible case shown for ease of description and is not intended to limit the specific value of the index in the embodiments of this application.
[0196] Step 2: The first communication device sends parameter information to the second communication device.
[0197] The parameter information includes CQI, the service quality requirements of the end-to-end collaborative inference service, the inference output dimension of the first communication device, the inference frequency, and the performance indicator. The parameter information is used by the third communication device to predict the uplink bandwidth of the first communication device.
[0198] Optionally, the parameter information is information sent when a secure communication connection is established between the first communication device and the second communication device.
[0199] The following are examples of the information in the parameter information:
[0200] (1) CQI is a channel quality indicator, representing the current channel quality and corresponding to the signal-to-noise ratio (SNR) of the channel. Its value typically ranges from 0 to 31. For example, a CQI value of 0 indicates the worst current channel quality, while a CQI value of 31 indicates the best current channel quality. The second communication device can determine the transmitted data block size, the HS-PDSCH channel code value, the encoding method, and the modulation method based on the CQI value.
[0201] (2) The quality of service requirements of the end-to-end network collaborative inference service can be used by the second communication device to allocate bandwidth to various services of the end-to-end network collaborative inference service in a balanced manner under limited bandwidth resources, and provide end-to-end quality of service guarantees for the different needs of various services. In other words, it can be understood as configuring different traffic processing methods according to priority, such as prioritizing important traffic and delaying or discarding unimportant traffic.
[0202] (3) The inference output dimension of the first communication device can be used to predict the uplink channel bandwidth. For example, when the uplink channel bandwidth between the terminal side and the network device side decreases, the intermediate data transmission delay increases and the end-to-end delay increases. In this case, the terminal can use a model segmentation method with a lower intermediate data output dimension according to the delay requirements when the uplink channel bandwidth decreases. When the uplink channel bandwidth between the terminal side and the network device side increases, the intermediate data transmission delay decreases and the end-to-end delay decreases. In this case, the terminal can use a model segmentation method with a higher intermediate data output dimension according to the delay requirements when the uplink channel bandwidth increases. Therefore, the inference output dimension of the first communication device (optionally, the inference output dimension of the first communication device can be the current inference output dimension of the first communication device or the historical inference output dimension of the first communication device) can be used by the second communication device to predict the uplink channel bandwidth in the next period of the first time period.
[0203] (4) Inference frequency refers to the frequency of inference if the inference service is related to video, meaning inference is performed every frame or every few frames. This inference frequency is higher than that of image recognition-based inference services. Therefore, when the output dimensions of the two first communication devices are similar, the required uplink bandwidth will differ. Thus, the inference frequency can be uploaded to the second communication device for differentiation. Optionally, the inference frequency can be statistically determined by the first communication device itself based on a certain time range, or it can be an indicator pre-set by the first communication device according to the service type. Step 3: The first communication device determines the second model segmentation method corresponding to the next time period of the first time period.
[0204] The second model segmentation method is related to the set of model parameters corresponding to the first model.
[0205] For example, referring to Figure 6, the first model segmentation method corresponds to end device model 1, and the second model segmentation method corresponds to end device model 2. In the first time period, the first model segmentation method corresponding to end device model 1 is used. When the uplink bandwidth of the predicted next time period decreases within the first time period, in order to meet the end-to-end latency requirement, the second model segmentation method corresponding to end device model 2 can be used in the next time period. In this scheme, the computational load of the first communication device can be reduced while ensuring end-to-end latency. Specifically, when the uplink bandwidth decreases, end-to-end latency can be guaranteed; when the uplink bandwidth increases, the resources of the first communication device can be saved while ensuring end-to-end latency.
[0206] Optionally, the method by which the first communication device determines the second model segmentation corresponding to the next time period of the first time period can be described in detail through the following two implementation methods.
[0207] In one implementation method, the first communication device needs to determine the second model segmentation method corresponding to the next time period of the first time period based on other information. For example, the first communication device receives the uplink bandwidth prediction range forwarded by the second communication device and determines the second model segmentation method corresponding to the next time period of the first time period based on the uplink bandwidth prediction range.
[0208] In the second implementation method, the first communication device directly receives the second model segmentation method corresponding to the next time period of the first time period forwarded by the second communication device.
[0209] Step 4: The first communication device determines the model parameter processing strategy according to the second model segmentation method.
[0210] The following are two possible implementations of the first communication device determining the model parameter processing strategy based on the second model segmentation method, as exemplified below:
[0211] In the first implementation method, the first communication device loads the model structure and corresponding model parameters of the second model segmentation method that are added to the first model segmentation method in the first time period, and / or unloads the model structure and corresponding model parameters of the second model segmentation method that are reduced in the first model segmentation method in the next time period of the first time period.
[0212] It should be noted that uninstalling in the next time period after the first time period means that uninstallation can be performed immediately after the first time period has ended, without having to wait for the next time period after the first time period to end.
[0213] For example, this solution can be applied to scenarios with relatively simple segmentation methods. For example, if the first model has already loaded layer X, then layer Y is loaded in the first time period, or layer Z is unloaded in the next time period (where X, Y, and Z are positive integers, and Z < X). In this solution, missing model structures and their corresponding model parameters can be loaded in advance, and redundant model structures and their corresponding model parameters can be unloaded in the next time period. This allows for more adaptive and flexible loading of model structures and their corresponding model parameters, avoiding the resource waste caused by preloading the entire model structure and its corresponding model parameter set.
[0214] In the second implementation method, the first communication device loads the model parameters corresponding to the second model segmentation method in the first time period, and / or runs the terminal network collaborative inference service on the model parameters corresponding to the second model segmentation method in the next time period of the first time period, and unloads the model parameters corresponding to the first model segmentation method.
[0215] For example, this solution can be applied to scenarios with relatively complex segmentation methods. For instance, if the first model has already loaded layer X1, then layer X2 can be loaded in the first time period, or the system can switch to layer X2 for end-to-end collaborative inference service in the next time period, and the previous layer X1 can be unloaded. In this solution, missing model structures and their corresponding model parameters can be loaded in advance, and redundant model structures and their corresponding model parameters can be unloaded in the next time period. This allows for more adaptive and flexible loading of model structures and their corresponding model parameters, and also avoids the resource waste caused by preloading the entire model structure and its corresponding model parameter set.
[0216] It should be noted that the initial loading of the first model parameters corresponding to the first model segmentation method by the first communication device is outside the loop. After the initial loading of the first model parameters corresponding to the first model segmentation method, the loop begins with optimizing model loading based on the uplink channel bandwidth. In other words, the loop starts from the first communication device determining the performance indicator (including this step), followed by information exchange between the first and second communication devices, the distribution of the uplink bandwidth prediction range, and the first communication device determining the loading of model parameters for the next time period based on the uplink bandwidth prediction range. This completes one cycle of the loop. Optionally, the next loop begins when certain conditions are met.
[0217] As one possible implementation, the first operation is executed repeatedly when the triggering condition is met.
[0218] For example, the triggering condition includes at least one of the following:
[0219] (1) The preset triggering period is satisfied. Optionally, the preset triggering period can be a triggering frequency. Further optionally, the preset triggering period can be the same as or different from the first time period. For example, the preset triggering period can be the time period of the terminal-network collaborative inference service or other periods. For example, the preset triggering period can be the first operation executed cyclically every 1 minute, or the first operation executed twice cyclically every 10 minutes.
[0220] (2) Satisfying preset trigger events. For example, a preset trigger event could be the emergence of new service requirements. Or a preset trigger event could be changes in weather conditions (strong winds, heavy rain, etc.) affecting wireless signal propagation, leading to changes in network quality and requiring a reassessment of uplink channel bandwidth. Alternatively, a preset trigger event could be the operator initiating network policy optimization, adjusting bandwidth allocation for different user groups, etc.
[0221] Step S805: The second communication device repeatedly performs the second operation until the end-to-end collaborative inference service ends within a preset time period.
[0222] The preset time period consists of at least one time cycle. Optionally, the preset time period may consist of one time cycle, for example, the preset time period is 10 minutes and one time cycle is 10 minutes. Optionally, the preset time period may also consist of multiple time cycles, for example, the preset time period is 10 minutes and the preset time period consists of 10 time cycles, with one time cycle being 1 minute.
[0223] The second operation includes, but is not limited to, the following steps:
[0224] Step 1: The second communication device receives parameter information from the first communication device.
[0225] The parameter information includes Channel Quality Indicator (CQI), Quality of Service Requirements for Inference Service, Inference Output Dimension of the First Communication Device, Inference Frequency, and Performance Indicator. This parameter information is used by the Second Communication Device to predict the uplink bandwidth of the First Communication Device.
[0226] Optionally, the parameter information is information sent when a secure communication connection is established between the first communication device and the second communication device.
[0227] It should be noted that the relevant content in the parameter information has been explained in detail in step S704, and will not be repeated here.
[0228] Step 2: The second communication device measures and obtains the channel quality index information of the first communication device.
[0229] The channel quality metrics include signal-to-interference-plus-noise ratio (SINR), reference signal reception quality (RSRQ), and uplink received signal strength indicator (UL_RSSI).
[0230] For example, the second communication device can measure the SINR of the first communication device to be 10dB, the RSRQ to be -9.5, and the UL_RSSI to be -65dBm.
[0231] Step 3: The second communication device sends the uplink bandwidth prediction range or the second model segmentation method to the first communication device.
[0232] The following are two possible implementations of a second communication device sending an uplink bandwidth prediction range or a second model segmentation method to a first communication device, as exemplified below:
[0233] In Implementation Method 1, the uplink bandwidth prediction range is determined by the third communication device, and then received by the second communication device and forwarded to the first communication device.
[0234] Specifically, the second communication device forwards parameter information, channel quality index information, and the load of the second communication device to the third communication device, receives the uplink bandwidth prediction range from the third communication device, and forwards the uplink bandwidth prediction range to the first communication device.
[0235] In the second implementation method, the second model segmentation method is further determined by the third communication device after determining the uplink bandwidth prediction range, and then received by the second communication device and forwarded to the first communication device.
[0236] Specifically, the second communication device forwards parameter information, channel quality index information, and the load of the second communication device to the third communication device, receives the second model segmentation method from the third communication device, and forwards the second model segmentation method to the first communication device.
[0237] Step S806: The third communication device repeatedly performs the third operation until the end-to-end collaborative inference service ends within a preset time period.
[0238] The preset time period consists of at least one time cycle. Optionally, the preset time period may consist of one time cycle, for example, the preset time period is 10 minutes and one time cycle is 10 minutes. Optionally, the preset time period may also consist of multiple time cycles, for example, the preset time period is 10 minutes and the preset time period consists of 10 time cycles, with one time cycle being 1 minute.
[0239] The third operation includes, but is not limited to, the following steps:
[0240] Step 1: The third communication device receives parameter information, channel quality index information, and the load of the second communication device from the second communication device.
[0241] The parameter information includes CQI, the quality of service requirement for inference service, the inference output dimension of the first communication device, the inference frequency, and the performance indicator. The parameter information is used by the third communication device to predict the uplink bandwidth of the first communication device. The channel quality index information includes SINR, RSRQ, and UL_RSSI.
[0242] Step 2: In the first time period, the third communication device determines the uplink bandwidth prediction range or the second model segmentation method corresponding to the next time period of the first time period based on the parameter information, channel quality index information, and load of the second communication device.
[0243] The first time period is at least one time period. Optionally, when the preset time period consists of one time period, for example, the preset time period is 10 minutes and one time period is 10 minutes, then the first time period is also 10 minutes. Optionally, when the preset time period consists of multiple time periods, for example, the preset time period is 10 minutes and the preset time period consists of 10 time periods, and one time period is 1 minute, then the first time period is 1 minute.
[0244] The following are two possible implementation methods for a third communication device to determine the uplink bandwidth prediction range or the second model segmentation method corresponding to the next time period of the first time period based on the parameter information of the second communication device, the channel quality index information, and the load of the second communication device in the first time period, as exemplarily described below:
[0245] In one implementation method, the third communication device may use a prediction model or other algorithm to predict the uplink bandwidth prediction range, but this application does not limit this.
[0246] Optionally, the third communication device inputs the load, parameter information, and channel quality index information of the second communication device into the prediction model during the first time period to obtain the uplink bandwidth prediction range.
[0247] The prediction model is trained based on multiple sample data. The sample data includes feature data and label data. The feature data includes the historical load, historical parameter information and historical channel quality index information of the second communication device, and the label data includes the historical uplink bandwidth of the first communication device.
[0248] It should be noted that, since the above sample data cannot be exhaustively listed, the data provided in this application is only an example and should not be construed as a limitation on the sample data.
[0249] In the second implementation method, the third communication device inputs the load, parameter information and channel quality index information of the second communication device into the prediction model to obtain the uplink bandwidth prediction range, receives the maximum computing power resources of the first communication device, and determines the second model segmentation method corresponding to the next time period of the first time period based on the uplink bandwidth prediction range and the maximum computing power resources.
[0250] The largest computing resources are used for end-to-end collaborative inference services.
[0251] Step 3: The third communication device sends the uplink bandwidth prediction range or the second model segmentation method to the second communication device.
[0252] Accordingly, the second communication device receives the uplink bandwidth prediction range or the second model segmentation method.
[0253] In addition, steps S804, S805 and S806 are executed synchronously, that is, the first communication device, the second communication device and the third communication device interact in the same time period. For example, after the first communication device performs the first operation, the second communication device performs the second operation and the third communication device performs the third operation in the first time period (this is one cycle), they continue to interact in the next time period of the first time period (this is the next cycle) until the end-to-end collaborative inference service ends.
[0254] In this application, taking the first communication device as the terminal and the second communication device as the network device as an example, on the one hand, since some model inference applications have high requirements for end-to-end inference latency, the terminal can determine the first model segmentation method corresponding to the first time period based on the current uplink bandwidth. This first model segmentation method corresponds to a deployment scheme of a model on the terminal side and the network device side. End-to-end latency = terminal-side inference latency + intermediate data transmission latency + network-side inference latency. Since the sum of the terminal-side inference model and the network device-side inference model is equivalent to a complete model, it can be assumed that the sum of the terminal-side inference latency and the network device-side inference latency remains unchanged, and the end-to-end latency is only affected by the intermediate data transmission latency. When the uplink channel bandwidth between the terminal side and the network device side decreases, the intermediate data transmission latency increases, and the end-to-end latency increases. In this case, the terminal can use a model segmentation method with a lower intermediate data output dimension according to the latency requirements when the uplink channel bandwidth decreases, thereby effectively ensuring the end-to-end latency.
[0255] On the other hand, compared to existing solutions where the terminal side and network device side need to preload the structure of the model used under various model partitioning methods and the union of the model parameter sets corresponding to the model structure, and select the appropriate model structure and model parameters to use when the uplink bandwidth changes, this application can load the corresponding model parameters that meet the uplink bandwidth conditions according to the uplink bandwidth as required. There is no need to preload the structure of the model used under various model partitioning methods and the union of the model parameter sets corresponding to the model structure, thereby effectively reducing the waste of computing resources on the terminal side and network device side.
[0256] In summary, this application can reduce the waste of computing resources on the terminal side and network device side while ensuring end-to-end latency.
[0257] The embodiment shown in Figure 7 provides a detailed explanation of the main interaction principle between the first communication device and the second communication device. The embodiment shown in Figure 8 provides a detailed explanation of the main interaction principle between the first communication device, the second communication device, and the third communication device. For ease of understanding, specific examples of the three model reasoning methods are illustrated below with reference to Figures 9a-9c.
[0258] Please refer to Figure 9a, which is a schematic flowchart of a model reasoning process provided in an embodiment of this application. It should be understood that, for ease of description, this application describes the process in the order of steps 11-20, and does not intend to limit the execution to this specific order. This application embodiment does not limit the order of execution, execution time, or number of executions of one or more of the above steps. As shown in Figure 9a, the specific steps of Case 1 are as follows:
[0259] Step 11: The first communication device downloads the structure of the first model and the set of model parameters corresponding to the structure of the first model.
[0260] Step 12: The first communication device determines the first model segmentation method corresponding to the first time period based on the current uplink bandwidth.
[0261] Step 13: The first communication device loads the first model parameters corresponding to the first model segmentation method.
[0262] Step 14: The first communication device determines the performance indicator associated with the first communication device based on its own battery power, remaining storage space, CPU load, GPU load, and temperature.
[0263] Step 15: The first communication device sends parameter information to the second communication device.
[0264] Accordingly, the second communication device receives the parameter information.
[0265] The parameter information includes CQI, the service quality requirements of the end-to-end collaborative inference service, the inference output dimension of the first communication device, the inference frequency, and the performance indicator. The parameter information is used by the second communication device to predict the uplink bandwidth of the first communication device.
[0266] Step 16: The second communication device measures and obtains the channel quality index information of the first communication device.
[0267] The channel quality metrics include SINR, RSRQ, and UL_RSSI.
[0268] Step 17: The second communication device obtains the uplink bandwidth prediction range based on the load, parameter information and channel quality index information of the second communication device in the first time period.
[0269] Step 18: The second communication device sends the uplink bandwidth prediction range for the next time period of the first time period to the first communication device.
[0270] Accordingly, the first communication device receives the uplink bandwidth prediction range.
[0271] Step 19: The first communication device determines the second model segmentation method based on the uplink bandwidth prediction range to ensure that the lower limit requirement of uplink bandwidth is met.
[0272] Step 20: The first communication device determines the model parameter processing strategy according to the second model segmentation method.
[0273] In this scheme, the second communication device collects data from the perspective of the entire cell and predicts the uplink channel bandwidth of the first communication device in the next time period of the first time period to obtain the uplink bandwidth prediction range. The first communication device determines the model segmentation method corresponding to the next time period of the first time period based on the uplink bandwidth prediction range, thus avoiding the waste of resources caused by preloading the structure of the entire model and the set of model parameters corresponding to the model structure.
[0274] It should be noted that detailed explanations of steps 11-20 above can be found in the embodiment shown in Figure 7, and will not be repeated here.
[0275] Please refer to Figure 9b, which is a schematic flowchart of another model reasoning provided in an embodiment of this application. It should be understood that, for ease of description, this application describes the process in the order of steps 31-42, and does not intend to limit the execution to the above order. This application embodiment does not limit the order of execution, execution time, or number of executions of one or more of the above steps. As shown in Figure 9b, the specific steps of Case Two are as follows:
[0276] Step 31: The first communication device downloads the structure of the first model and the set of model parameters corresponding to the structure of the first model.
[0277] Step 32: The first communication device determines the first model segmentation method corresponding to the first time period based on the current uplink bandwidth.
[0278] Step 33: The first communication device loads the first model parameters corresponding to the first model segmentation method.
[0279] Step 34: The first communication device determines the performance indicator associated with the first communication device based on its own battery power, remaining storage space, CPU load, GPU load, and temperature.
[0280] Step 35: The first communication device sends parameter information to the second communication device.
[0281] Accordingly, the second communication device receives the parameter information.
[0282] The parameter information includes CQI, the service quality requirements of the end-to-end collaborative inference service, the inference output dimension of the first communication device, the inference frequency, and the performance indicator. The parameter information is used by the third communication device to predict the uplink bandwidth of the first communication device.
[0283] Step 36: The second communication device measures and obtains the channel quality index information of the first communication device.
[0284] The channel quality metrics include SINR, RSRQ, and UL_RSSI.
[0285] Step 37: The second communication device sends the load, parameter information and channel quality index information of the second communication device to the third communication device.
[0286] Accordingly, the third communication device receives the load, parameter information and channel quality index information of the second communication device.
[0287] Step 38: The third communication device obtains the uplink bandwidth prediction range based on the load, parameter information and channel quality index information of the second communication device during the first time period.
[0288] Step 39: The third communication device sends the uplink bandwidth prediction range for the next time period of the first time period to the second communication device.
[0289] Accordingly, the second communication device receives the uplink bandwidth prediction range.
[0290] Step 40: The second communication device forwards the uplink bandwidth prediction range to the first communication device.
[0291] Accordingly, the first communication device receives the uplink bandwidth prediction range.
[0292] Step 41: The first communication device determines the second model segmentation method based on the uplink bandwidth prediction range to ensure that the lower limit requirement of uplink bandwidth is met.
[0293] Step 42: The first communication device determines the model parameter processing strategy according to the second model segmentation method.
[0294] In this scheme, since the third communication device has more computing resources and collects more data, it can collect data from a multi-cell global perspective and run more complex machine learning models to predict the uplink channel bandwidth of the first communication device, thereby obtaining the uplink bandwidth prediction range. The first communication device then determines the model segmentation method corresponding to the next time period of the first time period based on the uplink bandwidth prediction range, thus avoiding the waste of resources caused by preloading the structure of all models and the set of model parameters corresponding to the model structure.
[0295] It should be noted that detailed explanations of steps 31-42 above can be found in the embodiment shown in Figure 8, and will not be repeated here.
[0296] Please refer to Figure 9c, which is a flowchart illustrating another model reasoning process provided in this application embodiment. It should be understood that, for ease of description, this application describes the process in the order of steps 51-62, and does not intend to limit the execution to this specific order. This application embodiment does not limit the order of execution, execution time, or number of executions of one or more of the above steps. As shown in Figure 9c, the specific steps of Case 3 are as follows:
[0297] Step 51: The first communication device downloads the structure of the first model and the set of model parameters corresponding to the structure of the first model.
[0298] Step 52: The first communication device determines the first model segmentation method corresponding to the first time period based on the current uplink bandwidth.
[0299] Step 53: The first communication device loads the first model parameters corresponding to the first model segmentation method.
[0300] Step 54: The first communication device determines the performance indicator associated with the first communication device based on its own battery power, remaining storage space, CPU load, GPU load, and temperature.
[0301] Step 55: The first communication device sends parameter information to the second communication device.
[0302] Accordingly, the second communication device receives the parameter information.
[0303] The parameter information includes CQI, the service quality requirements of the end-to-end collaborative inference service, the inference output dimension of the first communication device, the inference frequency, and the performance indicator. The parameter information is used by the third communication device to predict the uplink bandwidth of the first communication device.
[0304] Step 56: The second communication device measures and obtains the channel quality index information of the first communication device.
[0305] The channel quality metrics include SINR, RSRQ, and UL_RSSI.
[0306] Step 57: The second communication device sends the load, parameter information and channel quality index information of the second communication device to the third communication device.
[0307] Accordingly, the third communication device receives the load, parameter information and channel quality index information of the second communication device.
[0308] Step 58: The third communication device obtains the uplink bandwidth prediction range based on the load, parameter information and channel quality index information of the second communication device during the first time period.
[0309] Step 59: The third communication device determines the second model segmentation method corresponding to the next time period of the first time period based on the uplink bandwidth prediction range and the maximum computing power resources of the first communication device.
[0310] Step 60: The third communication device sends the second model segmentation method to the second communication device.
[0311] Accordingly, the second communication device receives the second model segmentation method.
[0312] Step 61: The second communication device forwards the second model segmentation method to the first communication device.
[0313] Accordingly, the first communication device receives the second model segmentation method.
[0314] Step 62: The first communication device determines the model parameter processing strategy according to the second model segmentation method.
[0315] In this scheme, since the third communication device has more abundant computing resources and collects more data, it can collect data from a multi-cell global perspective and run more complex machine learning models to predict the uplink channel bandwidth of the first communication device, obtain the uplink bandwidth prediction range, and then determine the model segmentation method corresponding to the next time period of the first time period based on the uplink bandwidth prediction range. This avoids the waste of resources caused by preloading the structure of the entire model and the model parameter set corresponding to the model structure.
[0316] It should be noted that detailed explanations of steps 51-62 above can be found in the embodiment shown in Figure 8, and will not be repeated here.
[0317] The methods of the embodiments of this application have been described in detail above. The apparatus of the embodiments of this application is provided below.
[0318] It should be understood that the division of units in the apparatus provided in this application embodiment is only a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, the units in the apparatus can be implemented by a processor calling software. For example, the apparatus includes a processor connected to a memory, which stores instructions. The processor calls the instructions stored in the memory to implement any of the above methods or to implement the functions of each unit of the apparatus. The processor is, for example, a general-purpose processor, such as a central processing unit (CPU) or a microprocessor, and the memory is either internal or external to the apparatus.
[0319] Alternatively, the units in the device can be implemented as hardware circuits. The functionality of some or all of the units can be achieved through the design of these hardware circuits, which can be understood as one or more processors. For example, in one implementation, the hardware circuit is an application-specific integrated circuit (ASIC). The functionality of some or all of the above units is achieved through the design of the logical relationships between the components within the circuit. In another implementation, the hardware circuit can be implemented using a programmable logic device (PLD). Taking a field-programmable gate array (FPGA) as an example, it can include a large number of logic gates. The connection relationships between the logic gates are configured through a configuration file, thereby achieving the functionality of some or all of the above units.
[0320] In the embodiments of this application, each unit in the device may be one or more processors (or processing circuits) configured to implement the above methods, such as: CPU, graphics processing unit (GPU), neural network processing unit (NPU), tensor processing unit (TPU), deep learning processing unit (DPU), microprocessor unit (MPU), digital signal processor (DSP), ASIC, FPGA, or a combination of at least two of these processor forms.
[0321] Furthermore, the units in the above devices can be integrated in whole or in part, or they can be implemented independently. In one implementation, these units are integrated together as a system-on-a-chip (SOC). The SOC may include at least one processor for implementing any of the above methods or for implementing the functions of the units in the device. The at least one processor can be of different types, such as including a CPU and an FPGA, or including a CPU and an AI processor, or including a CPU and a GPU, etc. Several possible devices are listed below.
[0322] Please refer to Figure 10, which is a schematic diagram of the structure of a model inference device 100 provided in an embodiment of this application. Optionally, the model inference device 100 can be a first communication device or a component within the first communication device, such as a chip or integrated circuit. The model inference device 100 is used to implement the aforementioned model inference method, such as the model inference method shown in Figure 7 or Figure 8.
[0323] In one possible design, the model inference device 100 includes a processing unit 1001 and a communication unit 1002. The model inference device 100 is used to implement the aforementioned model inference method, such as the model inference method shown in FIG. 7 or FIG. 8. Exemplarily, the model inference device is used, for example, to execute the method executed by the first communication device.
[0324] In one possible implementation, the processing unit 1001 is configured to download the structure of a first model and the set of model parameters corresponding to the structure of the first model, wherein the first model is used for end-to-end collaborative inference service, and the structure of the first model is the union of the structures of the models used by the first communication device under various model segmentation methods. The processing unit 1001 is further configured to determine a first model segmentation method corresponding to a first time period based on the current uplink bandwidth, wherein the first model segmentation method is related to the set of model parameters corresponding to the structure of the first model. The processing unit 1001 is further configured to load the first model parameters corresponding to the first model segmentation method, wherein the first model parameters belong to the set of model parameters, and the first model parameters are used for the end-to-end collaborative inference service. The processing unit 1001 is further configured to repeatedly execute the first operation until the end-to-end collaborative inference service ends within a preset time period, wherein the preset time period consists of at least one time period, and the first time period belongs to the at least one time period. The first operation includes:
[0325] Based on the battery level, remaining storage space, CPU load, GPU load, and temperature of the first communication device, a performance indicator associated with the first communication device is determined, wherein the performance indicator is related to the uplink transmission rate. Parameter information is sent to the second communication device, including Channel Quality Indicator (CQI), the Quality of Service requirement of the end-to-end collaborative inference service, the inference output dimension and inference frequency of the first communication device, and the performance indicator. This parameter information is used by the second or third communication device to predict the uplink bandwidth of the first communication device. A second model segmentation method corresponding to the next time period of the first time period is determined, wherein the second model segmentation method is related to the model parameter set corresponding to the first model. A model parameter processing strategy is determined based on the second model segmentation method. The communication unit 1002 is used for sending and receiving data.
[0326] In another possible implementation, in determining the second model segmentation method corresponding to the next time period of the first time period, the processing unit 1001 is specifically configured to: receive an uplink bandwidth prediction range from the second communication device, and determine the second model segmentation method corresponding to the next time period of the first time period based on the uplink bandwidth prediction range.
[0327] In another possible implementation, in determining the second model segmentation method corresponding to the next time period of the first time period, the processing unit 1001 is specifically configured to: receive the second model segmentation method corresponding to the next time period of the first time period forwarded by the second communication device.
[0328] In yet another possible implementation, in terms of cyclically executing the first operation, the processing unit 1001 is specifically configured to: cyclically execute the first operation when a triggering condition is met.
[0329] In another possible implementation, in determining the model parameter processing strategy according to the second model segmentation method, the processing unit 1001 is specifically used to: load the model structure added by the second model segmentation method relative to the first model segmentation method and the model parameters corresponding to the added model structure in the first time period, and / or unload the model structure reduced by the second model segmentation method relative to the first model segmentation method and the model parameters corresponding to the reduced model structure in the next time period of the first time period.
[0330] In another possible implementation, in determining the model parameter processing strategy according to the second model segmentation method, the processing unit 1001 is specifically used to: load the model parameters corresponding to the second model segmentation method in the first time period, and / or run the terminal network collaborative inference service on the model parameters corresponding to the second model segmentation method in the next time period of the first time period, and unload the model parameters corresponding to the first model segmentation method.
[0331] The embodiments of this application and the method embodiments shown above are based on the same concept and have the same technical effects. For the specific principles, please refer to the description of the embodiments shown above, which will not be repeated here.
[0332] Please refer to Figure 11, which is a schematic diagram of another model inference device 110 provided in an embodiment of this application. Optionally, the model inference device 110 can be a second communication device, or a component within the second communication device, such as a chip or integrated circuit. The model inference device 110 is used to implement the aforementioned model inference method, such as the model inference method shown in Figure 7 or Figure 8.
[0333] In one possible design, the model inference device 110 includes a processing unit 1101 and a communication unit 1102. The model inference device 110 is used to implement the aforementioned model inference method, such as the model inference method shown in FIG. 7 or FIG. 8. Exemplarily, the model inference device may be used to execute, for example, the method executed by the second communication device.
[0334] In one possible implementation, the processing unit 1101 is used to repeatedly execute the second operation until the end-to-end collaborative inference service ends within a preset time period, the preset time period consisting of at least one time cycle. The second operation includes: receiving parameter information from the first communication device, wherein the parameter information includes a Channel Quality Indicator (CQI), the Quality of Service (QoS) requirement for the inference service, the inference output dimension of the first communication device, the inference frequency, and the performance indicator; the parameter information is used by the second or third communication device to predict the uplink bandwidth of the first communication device. Measuring the channel quality index information of the first communication device, wherein the channel quality index information includes a Signal-to-Interference-plus-Noise Ratio (SINR), a Reference Signal Received Quality (RSRQ), and an Uplink Received Signal Strength Indicator (UL_RSSI). Sending the uplink bandwidth prediction range or a second model segmentation method to the first communication device. The communication unit 1102 is used to send and receive data.
[0335] In another possible implementation, regarding the transmission of the uplink bandwidth prediction range or the second model segmentation method to the first communication device, the communication unit 1102 is specifically configured to: obtain the uplink bandwidth prediction range based on the load of the second communication device, the parameter information, and the channel quality index information during a first time period, wherein the first time period belongs to the at least one time period; and transmit the uplink bandwidth prediction range to the first communication device.
[0336] In another possible implementation, in the process of obtaining the uplink bandwidth prediction range based on the load of the second communication device, the parameter information, and the channel quality index information during the first time period, the processing unit 1101 is specifically configured to: input the load of the second communication device, the parameter information, and the channel quality index information into the prediction model during the first time period to obtain the uplink bandwidth prediction range.
[0337] In another possible implementation, regarding the transmission of the uplink bandwidth prediction range or the second model segmentation method to the first communication device, the communication unit 1102 is specifically configured to: forward the parameter information, the channel quality index information, and the load of the second communication device to the third communication device; receive the uplink bandwidth prediction range from the third communication device; and forward the uplink bandwidth prediction range to the first communication device.
[0338] In another possible implementation, regarding the transmission of the uplink bandwidth prediction range or the second model segmentation method to the first communication device, the communication unit 1102 is specifically configured to: forward the parameter information, the channel quality index information, and the load of the second communication device to the third communication device; receive the second model segmentation method from the third communication device; and forward the second model segmentation method to the first communication device.
[0339] The embodiments of this application and the method embodiments shown above are based on the same concept and have the same technical effects. For the specific principles, please refer to the description of the embodiments shown above, which will not be repeated here.
[0340] Please refer to Figure 12, which is a schematic diagram of another model inference device 120 provided in an embodiment of this application. Optionally, the model inference device 120 can be a third communication device, or a component within the third communication device, such as a chip or integrated circuit. The model inference device 120 is used to implement the aforementioned model inference method, such as the model inference method shown in Figure 7 or Figure 8.
[0341] In one possible design, the model inference device 120 includes a processing unit 1201 and a communication unit 1202. The model inference device 120 is used to implement the aforementioned model inference method, such as the model inference method shown in FIG. 7 or FIG. 8. Exemplarily, the model inference device may be used to execute a method executed by a third communication device.
[0342] In one possible implementation, the processing unit 1201 is used to repeatedly execute a third operation until the end-to-end collaborative inference service ends within a preset time period, the preset time period consisting of at least one time cycle. The third operation includes: receiving parameter information, channel quality index information, and the load of the second communication device from the second communication device. The parameter information includes a channel quality indicator (CQI), the quality of service requirement for the inference service, the inference output dimension of the first communication device, the inference frequency, and the performance indicator. The parameter information is used by the third communication device to predict the uplink bandwidth of the first communication device. The channel quality index information includes a signal-to-interference-plus-noise ratio (SINR), a reference signal reception quality (RSRQ), and an uplink received signal strength indicator (UL_RSSI). In a first time cycle, based on the parameter information of the second communication device, the channel quality index information, and the load of the second communication device, an uplink bandwidth prediction range or a second model segmentation method corresponding to the next time cycle of the first time cycle is determined, wherein the first time cycle belongs to the at least one time cycle. The uplink bandwidth prediction range or the second model segmentation method is sent to the second communication device. The communication unit 1202 is used to send and receive data.
[0343] In another possible implementation, in determining the uplink bandwidth prediction range or the second model segmentation method corresponding to the next time period of the first time period based on the parameter information of the second communication device, the channel quality index information and the load of the second communication device, the processing unit 1201 is specifically used to: input the load of the second communication device, the parameter information and the channel quality index information into the prediction model to obtain the uplink bandwidth prediction range.
[0344] In another possible implementation, regarding the determination of the uplink bandwidth prediction range or the second model segmentation method corresponding to the next time period of the first time period based on the parameter information of the second communication device, the channel quality index information, and the load of the second communication device, the processing unit 1201 is specifically configured to: input the load of the second communication device, the parameter information, and the channel quality index information into the prediction model to obtain the uplink bandwidth prediction range; receive the maximum computing power resources of the first communication device, wherein the maximum computing power resources are used for end-to-end collaborative inference services; and determine the second model segmentation method corresponding to the next time period of the first time period based on the uplink bandwidth prediction range and the maximum computing power resources.
[0345] The embodiments of this application and the method embodiments shown above are based on the same concept and have the same technical effects. For the specific principles, please refer to the description of the embodiments shown above, which will not be repeated here.
[0346] Please refer to Figure 13, which is a schematic diagram of the structure of another model inference device 130 provided in an embodiment of this application. The model inference device 130 can be a standalone device, such as a first, second, or third communication device, or it can be a component included in a standalone device, such as a chip, software module, or integrated circuit. The model inference device 130 can include at least one processor 1301 and a communication interface 1302. Optionally, it can also include at least one memory 1303. Further optionally, it can also include a connection line 1304, wherein the processor 1301, the communication interface 1302, and / or the memory 1303 are connected through the connection line 1304, and / or communicate with each other through the connection line 1304 to transmit control signals and / or data signals.
[0347] Wherein: processor 1301 is a module that performs arithmetic and / or logical operations, and may specifically include one or more of the following modules: filter, modem, power amplifier, low noise amplifier (LNA), baseband processor, radio frequency processor, radio frequency circuit, CPU, AP, microcontroller unit (MCU), electronic control unit (ECU), GPU, MPU, ASIC, image signal processor (ISP), DSP, FPGA, complex programmable logic device (CPLD), or coprocessor, etc.
[0348] The communication interface 1302 can be used to provide information input or output to at least one processor, or to receive signals sent externally and / or send signals to externally.
[0349] For example, the communication interface 1302 may include interface circuitry, such as input / output interfaces, chip pins, etc.
[0350] For example, the communication interface 1302 may include a wired link interface such as an Ethernet cable, or a wireless link interface (Wi-Fi, Bluetooth, general wireless transmission, vehicle short-range communication technology and other short-range wireless communication technologies, etc.).
[0351] Optionally, the communication interface 1302 may also include a radio frequency transmitter, an antenna, etc. When the communication interface 1302 includes an antenna, the number of antennas can be one or more.
[0352] As one possible design, if the model inference device 130 is a standalone device, the communication interface 1302 may include a receiver and a transmitter. The receiver and transmitter may be the same component or different components. When the receiver and transmitter are the same component, this component may be referred to as a transceiver.
[0353] As another possible design, if the model inference device 130 is a chip or circuit, the communication interface 1302 may include an input interface and an output interface, which may be the same interface or different interfaces.
[0354] Alternatively, the functions of the communication interface 1302 can be implemented by a transceiver circuit or a dedicated transceiver chip.
[0355] The memory 1303 provides storage space, in which data such as the operating system and computer programs can be stored. The memory 1303 can be one or a combination of several of the following: cache, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), compact disc read-only memory (CD-ROM), synchronous dynamic random access memory (SDRAM), hard disk drive (HDD), solid-state drive (SSD), etc. Memory is any other medium capable of carrying or storing desired program code in the form of instructions or data structures, and accessible by a computer, but is not limited thereto. The memory in this embodiment can also be a circuit or any other device capable of implementing storage functions, used to store computer programs or instructions, and / or data.
[0356] The functions and actions of each module or unit in the model reasoning device 130 listed above are merely illustrative examples.
[0357] Each functional unit in the model inference device 130 can be used to implement the aforementioned model inference method, such as the model inference method shown in FIG7 or FIG8, for example, to execute the method executed by the first communication device, or to execute the method executed by the second communication device, or to execute the method executed by the third communication device.
[0358] Optionally, the processor 1301 may be a processor specifically designed to perform the aforementioned methods (for ease of distinction, referred to as a dedicated processor), or a processor that performs the aforementioned methods by calling a computer program (for ease of distinction, referred to as a dedicated processor). Optionally, at least one processor may include both dedicated processors and general-purpose processors.
[0359] Optionally, if the model inference apparatus 130 includes at least one memory 1303, and the processor 1301 implements the aforementioned model inference method by calling a computer program, the computer program can be stored in the memory 1303.
[0360] This application also provides a chip, which includes logic circuitry and a communication interface. The communication interface is used to receive or transmit signals; the logic circuitry is used to receive or transmit signals through the communication interface. The chip is used to implement the aforementioned model inference method, such as the model inference method shown in FIG7 or FIG8, for example, to execute a method executed by a first communication device, or a method executed by a second communication device, or a method executed by a third communication device.
[0361] This application also provides a computer-readable storage medium storing instructions that, when executed on at least one processor (or model inference device), implement the aforementioned model inference method, such as the model inference method shown in FIG7 or FIG8, for example, for executing a method executed by a first communication device, or for executing a method executed by a second communication device, or for executing a method executed by a third communication device.
[0362] This application also provides a computer program product, which includes computer instructions for implementing the aforementioned model reasoning method, such as the model reasoning method shown in FIG7 or FIG8, for example, for executing a method executed by a first communication device, or for executing a method executed by a second communication device, or for executing a method executed by a third communication device.
[0363] It should be noted that, in the embodiments of this application, the words "exemplarily" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design scheme described as "exemplarily" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of the words "exemplarily" or "for example" is intended to present the relevant concepts in a specific manner.
[0364] In the embodiments of this application, "at least one" refers to one or more items, and "more than one" refers to two or more items. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items.
[0365] For example, at least one of a, b, or c can be represented as: a, b, c, (a and b), (a and c), (b and c), or (a and b and c), where a, b, and c can be single or multiple. "AND / OR" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects have an "OR" relationship.
[0366] Furthermore, unless otherwise stated, the use of ordinal numbers such as "first" and "second" in the embodiments of this application is for distinguishing multiple objects and is not for limiting the order, sequence, priority, or importance of multiple objects. Similarly, terms like "first node" and "second node" are merely for convenience in describing new parameters in different implementations and do not indicate differences in their execution operations, importance, structure, etc.
[0367] In the above embodiments, the term "when..." can be interpreted, depending on the context, as meaning "if...", "before...", "determined...", or "detected...". The above descriptions are merely optional embodiments of this application and are not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the concept and principles of this application should be included within the protection scope of this application.
[0368] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
Claims
1. A model reasoning method, characterized in that, Applied to a first communication device, the method includes: Download the structure of the first model and the set of model parameters corresponding to the structure of the first model, wherein the first model is used for terminal network collaborative reasoning service, and the structure of the first model is the union of the structures of the models used by the first communication device under various model splitting methods; The first model segmentation method corresponding to the first time period is determined based on the current uplink bandwidth, wherein the first model segmentation method is related to the model parameter set corresponding to the structure of the first model; Load the first model parameters corresponding to the first model segmentation method, wherein the first model parameters belong to the model parameter set and are used to perform the terminal network collaborative inference service; The first operation is executed repeatedly until the terminal-network collaborative inference service ends within a preset time period, wherein the preset time period consists of at least one time cycle, and the first time cycle belongs to the at least one time cycle. The first operation includes: A performance indicator associated with the first communication device is determined based on the first communication device's own battery power, remaining storage space, CPU load, GPU load, and temperature, wherein the performance indicator is related to the uplink transmission rate. Send parameter information to the second communication device, wherein the parameter information includes a channel quality indicator (CQI), the quality of service requirement of the end-to-end collaborative inference service, the inference output dimension of the first communication device, the inference frequency, and the performance indicator. The parameter information is used by the second or third communication device to predict the uplink bandwidth of the first communication device. Determine the second model segmentation method corresponding to the next time period of the first time period, wherein the second model segmentation method is related to the model parameter set corresponding to the first model; The model parameter processing strategy is determined based on the second model segmentation method.
2. The method according to claim 1, characterized in that, The method for determining the second model segmentation corresponding to the next time period of the first time period includes: Receive the uplink bandwidth prediction range from the second communication device; The second model segmentation method corresponding to the next time period of the first time period is determined based on the uplink bandwidth prediction range.
3. The method according to claim 1, characterized in that, The method for determining the second model segmentation corresponding to the next time period of the first time period includes: Receive the second model segmentation method corresponding to the next time period of the first time period forwarded from the second communication device.
4. The method according to any one of claims 1-3, characterized in that, The loop executes the first operation, including: If the triggering condition is met, the first operation is executed repeatedly.
5. The method according to any one of claims 1-4, characterized in that, The step of determining the model parameter processing strategy based on the second model segmentation method includes: In the first time period, the model structure added by the second model segmentation method compared to the first model segmentation method and the model parameters corresponding to the added model structure are loaded, and / or in the next time period of the first time period, the model structure reduced by the second model segmentation method compared to the first model segmentation method and the model parameters corresponding to the reduced model structure are unloaded.
6. The method according to any one of claims 1-4, characterized in that, The step of determining the model parameter processing strategy based on the second model segmentation method includes: In the first time period, load the model parameters corresponding to the second model segmentation method, and / or in the next time period of the first time period, run the terminal network collaborative inference service on the model parameters corresponding to the second model segmentation method, and unload the model parameters corresponding to the first model segmentation method.
7. A model reasoning method, characterized in that, Applied to a second communication device, the method includes: The second operation is executed repeatedly until the end-to-end collaborative inference service ends within a preset time period, which consists of at least one time cycle. The second operation includes: Receive parameter information from the first communication device, wherein the parameter information includes a channel quality indicator (CQI), a quality of service requirement for inference services, an inference output dimension of the first communication device, an inference frequency, and a performance indicator. The parameter information is used by the second or third communication device to predict the uplink bandwidth of the first communication device. The channel quality index information of the first communication device is measured, wherein the channel quality index information includes signal interference plus noise ratio (SINR), reference signal reception quality (RSRQ), and uplink received signal strength indicator (UL_RSSI). Send the uplink bandwidth prediction range or the second model segmentation method to the first communication device.
8. The method according to claim 7, characterized in that, Sending the uplink bandwidth prediction range or the second model segmentation method to the first communication device includes: In a first time period, the uplink bandwidth prediction range is obtained based on the load of the second communication device, the parameter information, and the channel quality index information, wherein the first time period belongs to at least one time period; The uplink bandwidth prediction range is sent to the first communication device.
9. The method according to claim 8, characterized in that, The process of obtaining the uplink bandwidth prediction range based on the load of the second communication device, the parameter information, and the channel quality index information during the first time period includes: During the first time period, the load of the second communication device, the parameter information, and the channel quality index information are input into the prediction model to obtain the uplink bandwidth prediction range.
10. The method according to claim 7, characterized in that, Sending the uplink bandwidth prediction range or the second model segmentation method to the first communication device includes: The parameter information, the channel quality index information, and the load of the second communication device are forwarded to the third communication device. Receive the uplink bandwidth prediction range from the third communication device; The uplink bandwidth prediction range is forwarded to the first communication device.
11. The method according to claim 7, characterized in that, Sending the uplink bandwidth prediction range or the second model segmentation method to the first communication device includes: The parameter information, the channel quality index information, and the load of the second communication device are forwarded to the third communication device. Receive the second model segmentation method from the third communication device; The second model segmentation method is forwarded to the first communication device.
12. A model reasoning method, characterized in that, Applied to a third communication device, the method includes: The third operation is executed repeatedly until the end-to-end collaborative inference service ends within a preset time period, which consists of at least one time cycle. The third operation includes: The third communication device receives parameter information, channel quality index information, and the load of the second communication device. The parameter information includes a channel quality indicator (CQI), a quality of service requirement for inference services, an inference output dimension of the first communication device, an inference frequency, and a performance indicator. The parameter information is used by the third communication device to predict the uplink bandwidth of the first communication device. The channel quality index information includes a signal-to-interference-plus-noise ratio (SINR), a reference signal reception quality (RSRQ), and an uplink received signal strength indicator (UL_RSSI). In the first time period, the uplink bandwidth prediction range or the second model segmentation method corresponding to the next time period of the first time period is determined based on the parameter information of the second communication device, the channel quality index information and the load of the second communication device, wherein the first time period belongs to the at least one time period; Send the uplink bandwidth prediction range or the second model segmentation method to the second communication device.
13. The method according to claim 12, characterized in that, The step of determining the uplink bandwidth prediction range or the second model segmentation method corresponding to the next time period of the first time period based on the parameter information of the second communication device, the channel quality index information, and the load of the second communication device includes: The load of the second communication device, the parameter information, and the channel quality index information are input into the prediction model to obtain the uplink bandwidth prediction range.
14. The method according to claim 12, characterized in that, The step of determining the uplink bandwidth prediction range or the second model segmentation method corresponding to the next time period of the first time period based on the parameter information of the second communication device, the channel quality index information, and the load of the second communication device includes: The load of the second communication device, the parameter information, and the channel quality index information are input into the prediction model to obtain the uplink bandwidth prediction range. Receive the maximum computing power resources of the first communication device, wherein the maximum computing power resources are used for end-to-end collaborative inference services; The second model segmentation method corresponding to the next time period of the first time period is determined based on the uplink bandwidth prediction range and the maximum computing power resources.
15. A model reasoning device, characterized in that, The model inference apparatus includes a communication unit and a processing unit, which are used to perform the method as described in any one of claims 1-6.
16. A model reasoning device, characterized in that, The model inference apparatus includes a communication unit and a processing unit, which are used to perform the method as described in any one of claims 7-11.
17. A model reasoning device, characterized in that, The model inference apparatus includes a communication unit and a processing unit, which are used to perform the method as described in any one of claims 12-14.
18. A model reasoning device, characterized in that, The model inference device includes a processor; When the processor invokes a computer program or instruction in memory, it causes the model inference device to implement the method as described in any one of claims 1-6.
19. A model reasoning device, characterized in that, The model inference device includes a processor; When the processor invokes a computer program or instruction in memory, it causes the model inference device to implement the method as described in any one of claims 7-11.
20. A model reasoning device, characterized in that, The model inference device includes a processor; When the processor invokes a computer program or instruction in memory, it causes the model inference device to implement the method as described in any one of claims 12-14.
21. A model reasoning device, characterized in that, Includes logic circuits and interfaces. The logic circuit and the interface are coupled; The interface is used for inputting and / or outputting information, and the logic circuit is used to enable the model inference device to implement the method as described in any one of claims 1-14.
22. The apparatus according to claim 21, characterized in that, The model inference device is a chip or chip system.
23. A model reasoning system, characterized in that, The model inference system includes the model inference apparatus as described in claim 15, the model inference apparatus as described in claim 16, and the model inference apparatus as described in claim 17; or The model inference system includes the model inference apparatus as described in claim 18, the model inference apparatus as described in claim 19, and the model inference apparatus as described in claim 20.
24. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store instructions or computer programs; When the instructions or the computer program are executed, the method described in any one of claims 1-14 is implemented.
25. A computer program product, characterized in that, include: Instructions or computer programs; When the instructions or the computer program are executed, the method described in any one of claims 1-14 is implemented.