Communication method and communication apparatus
By obtaining computation time requirements on the terminal side and exchanging information with the network side, adjusting model parameters and datasets, the problem of model performance degradation caused by differences between the sender and receiver was solved, achieving effective model deployment and performance improvement.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-19
- Publication Date
- 2026-04-02
AI Technical Summary
During the deployment of dual-end AI models, the performance of the models degrades due to differences between the datasets or models of the sender and receiver.
By obtaining computation time requirements on the terminal side, it is ensured that the model's inference task meets business needs. Information exchange between the terminal side and the network side is carried out to adjust model parameters and datasets to adapt to the computation time, thereby achieving effective model deployment.
This improved the performance of the dual-end AI model, avoiding the problem of model unavailability caused by insufficient computation time to meet business needs.
Smart Images

Figure CN2025122488_02042026_PF_FP_ABST
Abstract
Description
Communication method and communication apparatus
[0001] This application claims priority to the Chinese patent application No. 202411371031.8, filed on September 29, 2024, with the State Intellectual Property Office of China, and the Chinese patent application No. 202411371031.8 has the title of “Communication method and communication apparatus”, the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0002] The present application relates to the field of wireless communication, and in particular, to a communication method and a communication apparatus. BACKGROUND
[0003] At present, there are some solutions that use an artificial intelligence (AI) model to compress and reconstruct (or recover) channel state information (CSI). The sender (such as a terminal device) of a CSI report can compress the CSI by using the AI model, and send the obtained CSI report. The receiver (such as a network device) of the CSI report can reconstruct the channel measurement result based on the received CSI report by using the AI model.
[0004] In order to support the docking and model development or deployment of the double-end AI model (such as the sender and the receiver of the CSI report), it is necessary to provide the sender and / or the receiver with a dataset or a model. During the model development or deployment, there may be differences between the dataset or model of the sender and the status of the receiver, which may cause the performance of the double-end AI model to decrease. SUMMARY
[0005] The present application provides a communication method and a communication apparatus to improve the performance of the double-end AI model.
[0006] In a first aspect, a communication method is provided. The method can be applied to a terminal side, for example, to a terminal device such as the terminal device itself, or a component deployed in the terminal device, such as a circuit or a chip (such as a modem chip, also known as a baseband chip, or a system on chip (SoC) chip or a system in package (SIP) chip containing a modem core, etc.) inside the terminal device, etc.; or, the method can be applied to a device other than the terminal device, such as a host or a cloud server of an over the top (OTT) system, or a component deployed in the device other than the terminal device, such as a circuit or a chip, etc.; or, the method can be applied to a logic module or software, etc. capable of implementing all or part of the functions of the terminal side. The present application does not limit this.
[0007] For example, the method comprises obtaining a computing time requirement, the computing time requirement being a range that a time length for performing an inference task of a first model should satisfy.
[0008] In this regard, the computing time can be replaced by computing time, computing delay, time length, delay, time, time length, etc.
[0009] Based on the above scheme, the terminal side can determine the range that the time length for performing the inference task of the first model should satisfy by obtaining the computing time requirement, and then perform the inference task of the first model with reference to the range. This is conducive to meeting the computing time requirement of the business of the model deployed on the terminal side, improving the performance of the model, and avoiding the model deployed on the terminal side from being unusable due to the computing time being too large or not meeting the computing time requirement of the business.
[0010] In combination with the first aspect, in some implementations of the first aspect, the method further comprises: performing the inference task of the first model if the time length of the inference task of the first model satisfies the computing time requirement.
[0011] In the case where the computing time requirement is satisfied, the inference task of the first model is performed, so that the performance of the model on the terminal side is guaranteed.
[0012] Optionally, the method further comprises: sending first information, the first information being used to indicate a first computing time, the first computing time being an actual time length required for performing the inference task of the first model.
[0013] The first computing time can be an actual time length detected by the terminal side after performing the inference task of the first model. By reporting the first computing time, the network side can configure the reporting time of the CSI report for the terminal device based on the first computing time.
[0014] Optionally, the inference task is an inference of compressing CSI of R transmission layers, R is a maximum number of transmission layers supported by the terminal device, R is a positive integer; and the first information is further used to indicate R second calculation time lengths, a rth second calculation time length in the R second calculation time lengths is actually required for compressing CSI of an rth transmission layer in the R transmission layers.
[0015] In other words, the R second calculation time lengths correspond to the R transmission layers one by one. The terminal side can serially perform CSI compression inferences of the R transmission layers, and report actual time lengths of the CSI compression inferences of each transmission layer to the network side. In this way, the network side can also obtain the second calculation time lengths corresponding to the R transmission layers respectively, and further determine the first calculation time length. Therefore, the R second calculation time lengths can be regarded as an indication manner of the first calculation time length.
[0016] By indicating the R second calculation time lengths, the network side can have more understanding of the execution of the inference task of the first model by the terminal side, so as to better configure CSI reporting of the terminal device.
[0017] With reference to the first aspect, in some implementations of the first aspect, the method further includes: in a case where the time length of the inference task of the first model does not satisfy the calculation time length requirement, sending second information, the second information being used to indicate one or more of the following: that the time length of performing the inference task of the first model does not satisfy the calculation time length requirement; that the inference task of the first model is requested to be closed; that the first model is requested to be replaced; that model parameters used for model development are requested to be replaced, the model development including one or more of the following: model training, model adaptation or model enhancement; that a data set used for model training is requested to be replaced; or that the calculation time length requirement is requested to be replaced.
[0018] In this way, the network side can respond in a case where the inference task of the first model does not satisfy the calculation time length requirement, such as sending a simpler model or model parameters, sending a simpler data set, or sending a calculation time length requirement that is easier to satisfy. The response of the network side can be a response made after learning the status of the terminal side, and therefore the model, model parameters, data set, calculation time length requirement, etc. sent can better adapt to the capability of the terminal side, thereby facilitating improvement of performance of the double-end model.
[0019] With reference to the first aspect, in some implementations of the first aspect, the method further includes: receiving third information, the third information being used to indicate the calculation time length requirement.
[0020] That is, the calculation duration requirement can be indicated by the network side. In other words, the calculation duration requirement can be flexibly adjusted. For example, the network side can adjust the calculation duration requirement according to the calculation duration requirement of the service, according to the capability of the terminal side, and the like, thereby facilitating the performance of the double-end model connection.
[0021] In combination with the first aspect, in some implementations of the first aspect, before receiving the third information, the method further includes: sending capability information, the capability information being used to indicate one or more of the following: the maximum number of transmission layers supported by the terminal device, the storage capability of the terminal side, or the calculation capability of the terminal side.
[0022] The storage capability can refer to the maximum value of the storage space, or the upper limit of the storage space. The storage capability of the terminal side can refer to the storage capability of the device for deploying the model on the terminal side, which can be a device that is about to deploy the first model or has deployed the first model. For example, the device for deploying the model on the terminal side is a terminal device, and the storage capability of the terminal side can refer to the storage capability of the terminal device; for another example, the device for deploying the model on the terminal side is a host or a cloud server of an OTT system, and the storage capability of the terminal side can refer to the storage capability of the host or the cloud server of the OTT system.
[0023] The calculation capability can also be referred to as computing power, which can be represented by parameters such as calculation speed and processing time per unit data volume. The calculation capability of the terminal side can refer to the calculation capability of the device for deploying the model on the terminal side, which can be a device that is about to deploy the first model or has deployed the first model. For example, the device for deploying the model on the terminal side is a terminal device, and the calculation capability of the terminal side can refer to the calculation capability of the terminal device; for another example, the device for deploying the model on the terminal side is a host or a cloud server of an OTT system, and the calculation capability of the terminal side can also refer to the calculation capability of the host or the cloud server of the OTT system.
[0024] If the method of the first aspect is performed by a device for deploying the model, such as a host or a cloud server of an OTT system, the device can obtain the maximum number of transmission layers supported by the terminal device from the terminal device, and then send the maximum number of transmission layers to the network side through the third information. Alternatively, the maximum number of transmission layers supported by the terminal device can also be sent by the terminal device to the network side through another information over the air interface. The present application does not limit this.
[0025] The terminal side reports the capability information to the network side, which facilitates the network side to understand the status of the terminal side, and then can issue a calculation duration requirement that can be adapted to the capability of the terminal side, can issue a data set that can be adapted to the capability of the terminal side, and can issue a model or model parameter that can be adapted to the capability of the terminal side.
[0026] In conjunction with the first aspect, in some implementations of the first aspect, the inference task is to perform compressed inference on the CSI of multiple transport layers, and the aforementioned computation time requirement includes [related to R]. max R corresponding to each transport layer max A range, with respect to R max The range corresponding to the i-th transport layer in the 1-transport layer is the range that the inference time for compressing the CSI of the i-th transport layer must satisfy, R. max R is the maximum number of transport layers supported by the terminal device. max It is a positive integer.
[0027] Among them, R max It could be reported by the terminal side through capability information. In this case, the R... max The value of R is the same as the maximum number of transmission layers R supported by the aforementioned terminal devices. max It can also be predefined, such as a protocol predefined one. In this case, the maximum number of transport layers supported by the terminal device can be the maximum value defined considering terminal devices with multiple capabilities. Therefore, this R... max The value may be the same as or different from the value of the maximum number of transmission layers R supported by the aforementioned terminal device.
[0028] By indicating the range corresponding to each transport layer, the terminal side can use this as a reference to perform CSI compression inference for each transport layer. This helps the model deployed on the terminal side meet the computation time requirements of the service, improves model performance, and avoids the model becoming unusable due to excessive computation time or failure to meet the computation time requirements of the service.
[0029] In conjunction with the first aspect, in some implementations of the first aspect, the first model is obtained through model development, which includes one or more of the following: model training, model adaptation, or model enhancement.
[0030] In one possible design, the first model is obtained through model development, including: the first model is obtained by training the model based on the dataset.
[0031] Accordingly, the method also includes receiving a dataset for model training.
[0032] Specifically, the dataset can refer to the training dataset. Taking encoder training as an example, the training dataset can include the encoder's input, output, and quantization method. Training can include initial training or retraining.
[0033] The terminal can train the model based on the received dataset to obtain the first model.
[0034] In another possible design, the first model is obtained through model development, including: the first model is obtained through model development on the received second model.
[0035] Accordingly, the method further includes: receiving fourth information, the fourth information indicating one or more models, the one or more models including the second model.
[0036] With reference to the first aspect, in some implementations of the first aspect, the first model is received.
[0037] Accordingly, the method further includes: receiving fourth information, the fourth information indicating one or more models, the one or more models including the first model.
[0038] As can be seen from the above, the fourth information can be used to indicate one or more models, for example, including one or more model files and / or one or more sets of model parameters. Therefore, alternatively, the method further includes: receiving fourth information, the fourth information including one or more model files and / or one or more sets of model parameters.
[0039] With reference to the first aspect, in some implementations of the first aspect, the inference task is inference for compressing CSI, a start point of a time length of the inference task is one of: a configuration time of a reference signal resource, a transmission time of a reference signal, or a configuration time of a CSI report, and an end point of the time length of the inference task is a time of reporting the CSI report; wherein the reference signal is used for inference for compressing CSI, and the reference signal resource is used for transmitting the reference signal.
[0040] The time length of the inference task can also be regarded as the aforementioned first calculation time length. The definition of the start point and the end point of the first calculation time length should be aligned with the definition of the start point and the end point of the calculation time length requirement, so that the terminal side can determine the relationship between the calculation time length requirement and the time length of performing the inference task of the first model based on the definition, and the network side can also configure the reporting time of the CSI report based on the first calculation time length reported by the terminal side.
[0041] In a second aspect, a communication method is provided, which can be used on the network side, for example, can be applied to a network device, such as the network device itself, or a component deployed in the network device, such as a circuit or a chip (such as a modem chip, or a SoC chip or a SIP chip containing a modem core, etc.) with a near real-time radio access network (RAN) intelligent control function inside the network device, etc.; or, can also be applied to a device other than the network device, such as an intelligent network element with a near real-time RAN intelligent control function, etc.; or, can also be applied to a logic module or software, etc. capable of realizing all or part of the functions of the network side. The present application does not limit this.
[0042] Exemplarily, the method comprises: determining third information, the third information being used to indicate a calculation duration requirement, the calculation duration requirement being a range that a duration for the terminal side to perform an inference task of the first model needs to satisfy; and sending the third information to the terminal side.
[0043] Based on the above scheme, the network side determines the calculation duration requirement and sends the third information indicating the calculation duration requirement to the terminal side, so that the terminal side can determine a range that a duration for performing an inference task of the first model should satisfy, and then can perform the inference task of the first model with reference to the range. Thus, it is beneficial to make the model deployed on the terminal side satisfy the calculation duration requirement of the service, and it is beneficial to improve the model performance, and thus avoid the model deployed on the terminal side from being unusable due to the calculation duration being too large or not satisfying the calculation duration requirement of the service.
[0044] In combination with the second aspect, in some implementations of the second aspect, the inference task is an inference of compressing CSI of a plurality of transmission layers, and the calculation duration requirement comprises R ranges corresponding to R transmission layers, and an i-th range corresponding to an i-th transmission layer in the R transmission layers is a range that a duration for the inference of compressing CSI of the i-th transmission layer needs to satisfy, R is a maximum value of a number of transmission layers supported by a terminal device of the terminal side, and R is a positive integer. max max max max max
[0045] In combination with the second aspect, in some implementations of the second aspect, before determining the third information, the method further comprises: receiving capability information, the capability information being used to indicate one or more of a maximum number of transmission layers supported by a terminal device of the terminal side, a storage capability of the terminal side, or a calculation capability of the terminal side.
[0046] Optionally, the determining the third information comprises: determining the third information according to the capability information.
[0047] In combination with the second aspect, in some implementations of the second aspect, the method further comprises: receiving first information, the first information being used to indicate the first calculation duration, the first calculation duration being an actual duration needed by the terminal side to perform the inference task of the first model.
[0048] In combination with the second aspect, in some implementations of the second aspect, the inference task is an inference of compressing CSI, and the first information is used to indicate R second calculation durations, an r-th second calculation duration in the R second calculation durations being an actual duration needed by an r-th transmission layer to compress CSI, R being a maximum number of transmission layers supported by the terminal device, and R being a positive integer.
[0049] With reference to the second aspect, in some implementations of the second aspect, the method further includes receiving second information, the second information being used for one or more of the following: indicating that the duration of the terminal-side performing the inference task of the first model does not meet the computation duration requirement; requesting to close the inference task of the first model; requesting to replace the first model; requesting to replace model parameters, the model parameters being used for model development, the model development including one or more of the following: model training, model adaptation, or model enhancement; requesting to replace a dataset, the dataset being used for model training; or requesting to replace the computation duration requirement.
[0050] With reference to the second aspect, in some implementations of the second aspect, the method further includes sending a dataset, the dataset being used for the model training, the model training being used to obtain the first model.
[0051] With reference to the second aspect, in some implementations of the second aspect, the method further includes sending fourth information, the fourth information indicating one or more models. Alternatively, the fourth information includes one or more model files and / or one or more sets of model parameters.
[0052] Optionally, the one or more models include the first model.
[0053] Optionally, the one or more models include a second model, and the first model is obtained based on model development on the second model.
[0054] With reference to the second aspect, in some implementations of the second aspect, the inference task is inference for compressing CSI, the start point of the duration of the inference task is one of the following: a configuration time of a reference signal resource, a sending time of a reference signal, or a configuration time of a CSI report, and the end point of the duration of the inference task is a time of reporting the CSI report; wherein the reference signal is used for inference for compressing CSI, and the reference signal resource is used for transmitting the reference signal.
[0055] It should be understood that the method provided by the second aspect corresponds to the first aspect, and the corresponding descriptions and technical effects of the implementations of the second aspect can be referred to the descriptions of the first aspect, and will not be repeated here.
[0056] In a third aspect, an apparatus is provided. The apparatus can include a function module corresponding to each of the methods / operations / steps / actions described in any possible implementation of the first aspect, or a function module corresponding to each of the methods / operations / steps / actions described in any aspect of the second aspect. The module can be a hardware circuit, or software, or a combination of hardware circuit and software.
[0057] In one design, the apparatus can include a processing module and a communication module. The communication module can be configured to perform the transmitting and receiving actions performed by the network side in the method described in the second aspect above, and the processing module can be configured to perform the processing actions performed by the network side in the method described in the second aspect above.
[0058] In one design, the apparatus can be a terminal device, or a device, module, circuit, or chip configured to be deployed in a terminal device, or a device that can be used in conjunction with a terminal device, such as an OTT host or a cloud server.
[0059] In one design, the apparatus can include a processing module and a communication module. The communication module can be configured to perform the transmitting and receiving actions performed by the network side in the method described in the second aspect above, and the processing module can be configured to perform the processing actions performed by the network side in the method described in the second aspect above.
[0060] In one design, the apparatus can be a network device, or a device, module, circuit, or chip configured to be deployed in a network device, or a device that can be used in conjunction with a network device, such as a smart network element deployed with a radio access network (RAN) intelligent controller (RIC).
[0061] In a fourth aspect, an apparatus is provided, which includes a processor and a storage medium storing instructions that, when executed by the processor, cause the method in the first aspect or any possible implementation of the first aspect to be implemented, or cause the method in the second aspect or any possible implementation of the second aspect to be implemented.
[0062] In a fifth aspect, an apparatus is provided, which includes a processing circuitry configured to process data and / or information, so that the method in the first aspect or any possible implementation of the first aspect is implemented, or the method in the second aspect or any possible implementation of the second aspect is implemented.
[0063] The processing circuitry can include one or more processors, or all or a part of circuitry for controlling or processing functions in the one or more processors.
[0064] Optionally, the apparatus can further include a memory configured to store programs or instructions, and the processor can be configured to execute the programs or instructions, so that the method in the first aspect or any possible implementation of the first aspect is implemented, or the method in the second aspect or any possible implementation of the second aspect is implemented.
[0065] Optionally, the apparatus can further include the transceiver circuit, or the input / output interface.
[0066] In a sixth aspect, a chip is provided, including processing circuitry configured to execute programs or instructions to cause the method in the first aspect or any possible implementation of the first aspect to be implemented, or to cause the method in the second aspect or any possible implementation of the second aspect to be implemented.
[0067] Optionally, the chip can further include a memory configured to store the programs or instructions.
[0068] Optionally, the chip can further include the transceiver circuit, or the input / output interface.
[0069] In a seventh aspect, a computer readable storage medium is provided, including instructions, when the instructions are executed by a processor, causing the method in the first aspect or any possible implementation of the first aspect to be implemented, or causing the method in the second aspect or any possible implementation of the second aspect to be implemented.
[0070] In an eighth aspect, a computer program product is provided, including computer program codes or instructions, when the computer program codes or instructions are executed, causing the method in the first aspect and any possible implementation of the first aspect to be implemented, or causing the method in the second aspect or any possible implementation of the second aspect to be implemented.
[0071] In a ninth aspect, a communication system is provided, including the apparatus in the first aspect and any possible implementation of the first aspect, or including the apparatus in the second aspect and any possible implementation of the second aspect.
[0072] It should be understood that the third aspect to the ninth aspect of the present application correspond to the technical solutions of the first aspect to the second aspect of the present application, and the beneficial effects obtained by each aspect and the corresponding possible implementation manner are similar, which will not be described again. BRIEF DESCRIPTION OF DRAWINGS
[0073] Fig. 1 is a schematic diagram of a communication system suitable for the communication method of the embodiments of the present application;
[0074] Fig. 2 is a schematic diagram of another communication system suitable for the communication method of the embodiments of the present application;
[0075] Fig. 3 is a schematic diagram of a possible application framework in the communication system;
[0076] Fig. 4 is a schematic diagram of another possible application framework in the communication system;
[0077] FIG. 5 is a schematic diagram of CSI feedback using an auto-encoder (AE) model according to an embodiment of the present application;
[0078] FIG. 6 is a schematic diagram of a deep neural network (DNN);
[0079] FIG. 7 shows an example of a neuron structure;
[0080] FIG. 8 is a schematic diagram of data set interfacing between a network side and a terminal side according to an embodiment of the present application;
[0081] FIG. 9 is a schematic diagram of model interfacing between a network side and a terminal side according to an embodiment of the present application;
[0082] FIG. 10 is a schematic flow chart of a communication method according to an embodiment of the present application;
[0083] FIG. 11 is a schematic diagram of a calculation time length according to an embodiment of the present application;
[0084] FIG. 12 is another schematic flow chart of a communication method according to an embodiment of the present application;
[0085] FIG. 13 is yet another schematic flow chart of a communication method according to an embodiment of the present application;
[0086] FIG. 14 is still another schematic flow chart of a communication method according to an embodiment of the present application;
[0087] FIG. 15 is a schematic block diagram of a communication apparatus according to an embodiment of the present application;
[0088] FIG. 16 is another schematic block diagram of a communication apparatus according to an embodiment of the present application. DETAILED DESCRIPTION
[0089] The technical solutions in the present application will be described below with reference to the accompanying drawings.
[0090] For the convenience of understanding the embodiments of the present application, the following points are first explained:
[0091] First, in this application, the terminal side can also be referred to as the UE (user equipment, UE) side, including: a terminal device, a component (such as a circuit or chip inside the terminal device, etc.) deployed in the terminal device, a device (such as an OTT host or a cloud server) deployed outside the terminal device or a component (such as a circuit or chip inside the device, etc.) deployed in the device outside the terminal device. The network (network, NW) side (NW side) includes: a network device in communication with the terminal device, a component (such as a circuit or chip inside the network device with near real-time RAN intelligent control function, etc.) deployed in the network device, a device (such as an intelligent network element, such as an intelligent network element with near real-time RAN intelligent control function) deployed outside the network device or a component (such as a circuit or chip inside the intelligent network element, etc.) deployed in the intelligent network element. Among them, the network device can include: an access network device, a core network device, or an operation administration and maintenance (OAM).
[0092] Second, in this application, indication includes direct indication (also known as explicit indication) and indirect indication (also known as implicit indication). Among them, directly indicating information A means including the information A; indirectly indicating information A can mean indicating information A through the correspondence between information A and information B and directly indicating information B; or indicating information A through a preset rule that can be used to determine A according to B and directly indicating information B. Among them, the correspondence between information A and information B and the preset rule can be predefined, pre-stored, pre-burned, or pre-configured.
[0093] Third, in this application, "at least one" means one or more, and "multiple" means two or more. The "and / or" describes the association relationship between the associated objects, which means that there can be three kinds of relationships, for example, A and / or B can represent: A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after it, but does not rule out the case that the associated objects before and after it represent an "and" relationship, and the meaning expressed can be understood in conjunction with the context. "At least one of the following" or similar expressions means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b, or c can mean: a, b, c; a and b; a and c; b and c; or a and b and c. Where a, b, and c can be single or multiple.
[0094] Fourthly, in the present application, the use of prefixes such as "first", "second" and the like is merely intended to differentiate between different objects belonging to the same category, and does not imply an order, a magnitude or a quantity. For example, "first information" and "second information" are merely different information, and do not limit the quantity, the order of transmission or the priority relationship.
[0095] Fifthly, in the present application, "sending" and "receiving" represent the direction of signal transmission. For example, "sending information to XX" can be understood as that the destination of the information is XX, which can include direct sending through the air interface, or indirect sending through the air interface by other units or modules. "Receiving information from YY" can be understood as that the source of the information is YY, which can include direct receiving from YY through the air interface, or indirect receiving from YY through the air interface by other units or modules. "Sending" can also be understood as "output" of a chip interface, and "receiving" can also be understood as "input" of a chip interface. In other words, sending and receiving can be between devices, such as between a terminal device and a computing node, or within a device, such as between components, modules, chips, software modules or hardware modules within a device through a bus, wire or interface.
[0096] Sixthly, in the embodiments of the present application, "when", "if" and "whether" all refer to the objective situation that the device will make corresponding processing, and are not limited to time, and do not require the device to have a judgment action when implemented, nor mean that there are other limitations.
[0097] Seventhly, in the present application, the words "example", "exemplary", "for example", "for instance" or "such as" are used to represent an example, an illustration or a description. Any embodiment or design solution described as "example", "exemplary", "for example", "for instance" or "such as" in the present application should not be interpreted as more preferred or more advantageous than other embodiments or design solutions. Rather, the words "example", "exemplary", "for example", "for instance" or "such as" are intended to present the relevant concept in a specific manner.
[0098] Eighthly, for the convenience of understanding the method provided in the present application, the following text refers to numbers in many places, such as the i-th transmission layer, the r-th transmission layer and the like. These numbers can be numbered starting from 0, or can be numbered starting from 1, or can be numbered starting from other preset values, and the present application does not limit this.
[0099] The technical solutions provided in the present application can be applied to various communication systems, for example, a 5th generation (5G) or new radio (NR) system, a long term evolution (LTE) system, an LTE frequency division duplex (FDD) system, an LTE time division duplex (TDD) system, a wireless local area network (WLAN) system, a satellite communication system, a future communication system, or a fusion system of multiple systems, and the like. The technical solutions provided in the present application can also be applied to device to device (D2D) communication, vehicle-to-everything (V2X) communication, machine to machine (M2M) communication, machine type communication (MTC), and an internet of things (IoT) communication system or other communication systems.
[0100] A network element in a communication system can send a signal to another network element or receive a signal from another network element. The signal can include information, signaling, data, and the like. The network element can also be replaced by an entity, a network entity, a device, a communication device, a communication module, a node, a communication node, and the like. For example, the communication system can include at least one terminal device and at least one network device. The network device can send a downlink signal to the terminal device, and / or the terminal device can send an uplink signal to the network device. It can be understood that the terminal device in the present disclosure can be replaced by a first network element, and the network device can be replaced by a second network element, both of which perform the corresponding communication method in the present disclosure.
[0101] FIG. 1 is a schematic diagram of a communication system suitable for a communication method according to an embodiment of the present application. As shown in FIG. 1, the communication system 100A can include at least one access network device, for example, the access network device 110 shown in FIG. 1; the communication system 100A can also include at least one terminal device, for example, the terminal device 120 and the terminal device 130 shown in FIG. 1. The access network device 110 and the terminal device (such as the terminal device 120 and the terminal device 130) can communicate through a wireless link. The communication devices in the communication system, for example, the access network device 110 and the terminal device 120, can communicate through multi-antenna technology.
[0102] In a wireless communication network, such as a mobile communication network, the services supported by the network are increasingly diverse, and thus the requirements to be met are increasingly diverse. For example, the network needs to be able to support ultra-high rates, ultra-low latencies, and / or ultra-large connections. This feature makes network planning, network configuration, and / or resource scheduling increasingly complex. In addition, as the functions of the network become increasingly powerful, such as supporting increasingly high frequency spectrums, supporting high-order multiple input multiple output (MIMO) technology, supporting beam forming (BF), supporting beam management, and other new technologies, network energy saving has become a hot research topic. These new requirements, new scenarios, and new features bring unprecedented challenges to network planning, operation and maintenance, and efficient operation. To meet this challenge, artificial intelligence technology can be introduced into the wireless communication network, thereby realizing network intelligence. In order to support AI technology in the wireless network, an AI node can also be introduced into the network.
[0103] FIG. 2 is a schematic diagram of another communication system suitable for the communication method of the embodiments of the present application. Compared with the communication system 100A shown in FIG. 1, the communication system 100B shown in FIG. 2 further includes an AI network element 140. The AI network element 140 is configured to perform AI-related operations, such as constructing a training data set or training an AI model. The AI network element can also be referred to simply as an intelligent network element.
[0104] In a possible implementation, the access network device 110 can send data related to the training of the AI model to the AI network element 140, and the AI network element 140 constructs a training data set and trains an AI model. For example, the data related to the training of the AI model can include data reported by the terminal device. The AI network element 140 can send the result of the AI model-related operation to the access network device 110 and forward it to the terminal device through the access network device 110. For example, the result of the AI model-related operation can include at least one of the following: a trained AI model, an evaluation result or a test result of the model, and the like. Illustratively, part of the trained AI model can be deployed on the access network device 110, and the other part can be deployed on the terminal device 120 and / or the terminal device 130. Alternatively, the trained AI model can be deployed on the access network device 110. Or, the trained AI model can be deployed on the terminal device 120 and / or the terminal device 130.
[0105] It should be understood that FIG. 2 is only used as an example for illustrating that the AI network element 140 is directly connected with the access network device 110, and in other scenarios, the AI network element 140 can also be connected with a terminal device. Alternatively, the AI network element 140 can be simultaneously connected with the access network device 110 and the terminal device. Alternatively, the AI network element 140 can also be connected with the access network device 110 through a third-party network element. The connection relationship between the AI network element and other network elements is not limited in the embodiments of the present application. For example, the AI network element 140 can also be arranged as a module in the access network device and / or the terminal device, for example, arranged in the access network device 110 or the terminal device shown in FIG. 1.
[0106] It should be noted that FIG. 1 and FIG. 2 are only simplified schematic diagrams for example and for understanding, for example, the communication system can further include other devices, such as wireless relay devices and / or wireless backhaul devices, which are not shown in FIG. 1 and FIG. 2. In actual application, the communication system can include multiple access network devices, and can also include multiple terminal devices. The number of access network devices and terminal devices included in the communication system is not limited in the embodiments of the present application.
[0107] In the embodiments of the present application, the terminal device can also be referred to as a UE, an access terminal, a user unit, a user station, a mobile station, a mobile station, a remote station, a remote terminal, a mobile device, a user terminal, a terminal, a wireless communication device, a user agent or a user equipment.
[0108] The terminal device can be a device providing voice / data, for example, a handheld device with wireless connection function, a vehicle-mounted device, etc. At present, some examples of terminals are: mobile phone, tablet computer, notebook computer, palm computer, mobile internet device (MID), wearable device, virtual reality (VR) device, augmented reality (AR) device, wireless terminal in industrial control, wireless terminal in self-driving, wireless terminal in remote medical surgery, wireless terminal in smart grid, wireless terminal in transportation safety, wireless terminal in smart city, wireless terminal in smart home, cellular phone, cordless phone, session initiation protocol (SIP) phone, wireless local loop (WLL) station, personal digital assistant (PDA), handheld device with wireless communication function, computing device or other processing device connected to a wireless modem, wearable device, terminal device in a 5G network, or terminal device in a future evolved public land mobile network (PLMN), etc. The embodiments of the present application are not limited thereto.
[0109] By way of example and not limitation, in the embodiments of the present application, the terminal device can also be a wearable device. The wearable device can also be referred to as a wearable smart device, which is a general term for devices that are designed and developed by applying wearable technology to daily wear, such as glasses, gloves, watches, clothing, and shoes. The wearable device is a portable device that is directly worn on the body or integrated into the user's clothes or accessories. The wearable device is not only a hardware device, but also a device that realizes powerful functions through software support and data interaction and cloud interaction. The general wearable smart device includes devices with full functions, large size, and the ability to realize complete or partial functions without relying on a smart phone, such as smart watches or smart glasses, and devices that focus on a certain application function and need to be used in cooperation with other devices, such as smart phones, such as various smart wristbands and smart jewelry for monitoring vital signs.
[0110] In the embodiments of the present application, the apparatus for implementing the function of the terminal device can be a terminal device, or an apparatus capable of supporting the terminal device to implement the function, for example, a chip system, which can be installed in the terminal device or used in matching with the terminal device. In the embodiments of the present application, the chip system can be composed of a chip, or can include the chip and other discrete devices. In the embodiments of the present application, only the apparatus for implementing the function of the terminal device is taken as an example for illustration, and the scheme of the embodiments of the present application is not limited.
[0111] The access network device in the embodiments of the present application can be a device for communicating with the terminal device, and the access network device can also be referred to as a network device, for example, the access network device can be a base station. The access network device in the embodiments of the present application can refer to a RAN node (or device) for accessing the terminal device to a wireless network. The base station can broadly cover the following various names, or be replaced by the following names, such as: Node B (NodeB), evolved Node B (eNB), next generation Node B (gNB), relay station, access point, transmitting and receiving point (TRP), transmitting point (TP), primary station, secondary station, motor slide retainer (MSR) node, home base station, network controller, access node, wireless node, access point (AP), transmission node, transceiver node, baseband unit (BBU), remote radio unit (RRU), active antenna unit (AAU), remote radio head (RRH), central unit (CU), distributed unit (DU), radio unit (RU), positioning node, etc. The base station can be a macro base station, a micro base station, a relay node, a donor node or the like, or a combination thereof. The base station can also refer to a communication module, modem or chip for being arranged in the foregoing devices or apparatuses. The base station can also be a mobile switching center, a device assuming a base station function in D2D, V2X, M2M communication, a device assuming a base station function in a future communication system, etc. The base station can support networks of the same or different access technologies. Optionally, the RAN node can also be a server, a wearable device, a vehicle or a vehicle-mounted device, etc. For example, the access network device in the V2X technology can be a road side unit (RSU). The embodiments of the present application do not limit the specific technology and specific device form adopted by the access network device.
[0112] A base station can be fixed, or mobile. For example, a helicopter or unmanned aerial vehicle can be configured to act as a mobile base station, one or more cells can move according to the location of the mobile base station. In other examples, a helicopter or unmanned aerial vehicle can be configured to act as a device that communicates with another base station.
[0113] In some deployments, the access network device mentioned in the embodiments of the present application can be a device including a CU, or a DU, or a device including a CU and a DU, or a control plane CU node (central unit-control plane (CU-CP)) and a user plane CU node (central unit-user plane (CU-UP)) and a DU node. For example, the access network device can include a gNB-CU-CP, a gNB-CU-UP and a gNB-DU.
[0114] In some deployments, a plurality of RAN nodes cooperate to assist a terminal to implement wireless access, and different RAN nodes respectively implement part of the functions of a base station. For example, the RAN node can be a CU, a DU, a CU-CP, a CU-UP, or an RU, etc. The CU and the DU can be separately arranged, or can also be included in the same network element, for example, in a BBU. The RU can be included in a radio frequency device or a radio frequency unit, for example, included in an RRU, an AAU or an RRH.
[0115] The RAN node can support one or more types of fronthaul interface, different fronthaul interfaces respectively corresponding to DUs and RUs having different functions. If the fronthaul interface between the DU and the RU is a common public radio interface (CPRI), the DU is configured to implement one or more of baseband functions, and the RU is configured to implement one or more of radio frequency functions. If the fronthaul interface between the DU and the RU is another interface, relative to the CPRI, one or more of the partial baseband functions of the downlink and / or uplink, such as, for the downlink, one or more of precoding, digital beamforming (BF), or inverse fast Fourier transform (IFFT) / add cyclic prefix (CP), are moved from the DU to the RU for implementation, and for the uplink, one or more of digital beamforming (BF), or fast Fourier transform (FFT) / remove cyclic prefix (CP) are moved from the DU to the RU for implementation. In a possible implementation, the interface can be an enhanced common public radio interface (eCPRI). Under the eCPRI architecture, the splitting manner between the DU and the RU is different, corresponding to different categories (Cat) of eCPRI, such as eCPRI Cat A, B, C, D, E, F.
[0116] Taking eCPRI Cat A as an example, for downlink transmission, with layer mapping as the cut, the DU is configured to implement layer mapping and one or more functions (i.e., one or more of encoding, rate matching, scrambling, modulation, layer mapping) before layer mapping, and other functions (e.g., one or more of resource element (RE) mapping, digital beamforming (BF), or inverse fast Fourier transform (IFFT) / adding CP) after layer mapping are implemented in the RU. For uplink transmission, with de-RE mapping as the cut, the DU is configured to implement de-mapping and one or more functions (i.e., one or more of decoding, de-rate matching, de-scrambling, de-modulation, inverse discrete Fourier transform (IDFT), channel equalization, de-RE mapping) before de-mapping, and other functions (e.g., one or more of digital BF or FFT / CP removal) after de-mapping are implemented in the RU. It can be understood that, for the function description of the DU and the RU corresponding to various types of eCPRI, reference can be made to the eCPRI protocol, which is not described herein.
[0117] In a possible design, the processing unit in the BBU for implementing baseband functions is referred to as a base band high (BBH) unit, and the processing unit in the RRU / AAU / RRH for implementing baseband functions is referred to as a base band low (BBL) unit.
[0118] In different systems, the CU (or CU-CP and CU-UP), DU or RU can also have different names, but those skilled in the art can understand their meanings. For example, in an open RAN (ORAN) architecture, the CU can also be referred to as an open-CU (O-CU), the DU can also be referred to as an open-DU (O-DU), the CU-CP can also be referred to as an open-CU-CP (O-CU-CP), the CU-UP can also be referred to as an open-CU-UP (O-CU-UP), and the RU can also be referred to as an open-RU (O-RU). Any of the CU (or CU-CP, CU-UP), DU and RU in this application can be implemented by a software module, a hardware module, or a combination of a software module and a hardware module.
[0119] In the embodiments of the present application, the apparatus for implementing the function of the network device can be a network device, or can be an apparatus capable of supporting the network device to implement the function, such as a chip system, a hardware circuit, a software module, or a hardware circuit plus a software module. The apparatus can be installed in the network device or used in combination with the network device. In the embodiments of the present application, only the apparatus for implementing the function of the network device is taken as an example for description, and the scheme of the embodiments of the present application is not limited in this way.
[0120] The network device and / or the terminal device can be deployed on land, including indoors or outdoors, handheld or vehicle-mounted; can also be deployed on water surface; and can also be deployed on airplanes, balloons and satellites in the air. The scenarios in which the network device and the terminal device are located are not limited in the embodiments of the present application. In addition, the terminal device and the network device can be hardware devices, or can be software functions running on special hardware, software functions running on general hardware, such as virtualized functions instantiated on a platform (for example, a cloud platform), or entities including special or general hardware devices and software functions. The specific forms of the terminal device and the network device are not limited in the present application.
[0121] Optionally, the AI node can be deployed in one or more of the following positions in the communication system: an access network device, a terminal device, or a network element of a core network, and the like. Optionally, the AI node can also be deployed separately, for example, in a position other than any of the above devices, such as a host or a cloud server of an over the top (OTT) system. The AI node can communicate with other devices in the communication system, which can be one or more of the following: an access network device, a terminal device, or a network element of a core network, and the like.
[0122] It can be understood that the number of AI nodes is not limited in the present application. For example, when there are multiple AI nodes, the multiple AI nodes can be divided based on functions, such as different AI nodes being responsible for different functions.
[0123] It can also be understood that the AI node can be a device independent of each other, or can be integrated into the same device to implement different functions, or can be a network element in a hardware device, or can be a software function running on special hardware, or can be a virtualized function instantiated on a platform (for example, a cloud platform), and the specific forms of the AI node are not limited in the present application.
[0124] The AI node can be an AI network element or an AI module.
[0125] FIG. 3 is a schematic diagram of a possible application framework in a communication system. As shown in FIG. 3, network elements in the communication system are connected through interfaces (e.g., next generation (NG) interface, Xn interface), or air interfaces. The NG interface is an interface between a radio access network and a 5G core network. The Xn interface is an interface between access network devices, and the air interface is an interface between an access network device and a terminal device. One or more AI modules (only one is shown in FIG. 3 for clarity) are deployed in one or more of the network element nodes, such as a core network device, an access network node (RAN node), a terminal, or an OAM device. The access network node can be a single RAN node or can include multiple RAN nodes, such as a CU and a DU. The CU and / or the DU can also be provided with one or more AI modules. Optionally, the CU can be further split into a CU-CP and a CU-UP. One or more AI models are deployed in the CU-CP and / or the CU-UP.
[0126] The AI module is used to implement a corresponding AI function. The AI modules deployed in different network elements can be the same or different. The AI module can implement different functions according to different parameter configurations of the model of the AI module. The model of the AI module can be configured based on one or more of the following parameters: a structural parameter (such as at least one of a number of neural network layers, a neural network width, a connection relationship between layers, a weight of a neuron, an activation function of a neuron, or a bias in the activation function), an input parameter (such as a type of input parameter and / or a dimension of the input parameter), or an output parameter (such as a type of output parameter and / or a dimension of the output parameter). The bias in the activation function can also be referred to as a bias of the neural network.
[0127] One AI module can have one or more models. One model can infer an output including one parameter or multiple parameters. The learning process, the training process, or the inference process of different models can be deployed in different nodes or devices, or can be deployed in the same node or device.
[0128] The network device can be a network device provided with one or more AI modules. The network device can include one or more of the core network device, the access network node (RAN node), or the operation administration and maintenance (OAM) shown in FIG. 3. For example, the AI module can be a RAN intelligent controller (RIC) shown in FIG. 4, such as a near-real time RIC (near-RT RIC) or a non-real time RIC (Non-RT RIC). For example, the near-real time RIC is provided in the RAN node (for example, in the CU, the DU), and the non-real time RIC is provided in the OAM, the cloud server, the core network device, or other access network device. The RIC can obtain a subset of data from multiple terminal devices from the RAN node (for example, the CU, the CU-CP, the CU-UP, the DU, and / or the RU), reorganize the subset of data into a training data set, and train based on the training data set. For example, the near-real time RIC and the non-real time RIC can also be provided as a network element, respectively. The access network device can be the near-real time RIC or the non-real time RIC.
[0129] FIG. 4 is a schematic diagram of another possible application framework in a communication system. In addition to the access network node (CUs, DUs, and RUs are shown in the figure) and the terminal, the communication system shown in FIG. 4 also includes a RIC. For example, the RIC can be an AI module shown in FIG. 3, which can be used to implement AI-related functions. The RIC includes a near-real time RIC and a non-real time RIC. The non-real time RIC mainly processes non-real time information, such as data that is not sensitive to latency, which can be seconds. The real-time RIC mainly processes near-real time information, such as data that is relatively sensitive to latency, which can be tens of milliseconds.
[0130] The near-real time RIC is used for model training and inference. For example, the AI model is trained, and inference is performed using the AI model. The near-real time RIC can obtain network side and / or terminal side information from the RAN node (for example, the CU, the CU-CP, the CU-UP, the DU, and / or the RU) and / or the terminal. The information can be used as training data or inference data. Optionally, the near-real time RIC can submit the inference result to the RAN node and / or the terminal. Optionally, the inference result can be exchanged between the CU and the DU, and / or between the DU and the RU. For example, the near-real time RIC submits the inference result to the DU, and the DU sends it to the RU.
[0131] The non-real-time RIC is also used for model training and inference. For example, for training an AI model, inference is performed using the model. The non-real-time RIC can obtain network-side and / or terminal-side information from the RAN node (for example, one or more of a CU, a CU-CP, a CU-UP, a DU, or an RU) and / or a terminal. This information can be used as training data or inference data, and the inference result can be delivered to the RAN node and / or the terminal. Alternatively, the inference result can be exchanged between the CU and the DU, and / or between the DU and the RU, for example, the non-real-time RIC delivers the inference result to the DU, which then sends it to the RU.
[0132] The near-real-time RIC and the non-real-time RIC can also be separately provided as a network element. Alternatively, the near-real-time RIC and the non-real-time RIC can also be part of other devices, for example, the near-real-time RIC is provided in the RAN node (for example, in the CU or the DU), and the non-real-time RIC is provided in the OAM, the cloud server, the core network device, or other network devices.
[0133] With the development of wireless communication technology, more and more services are supported, and higher requirements are put forward for the communication system in terms of system capacity, communication delay, and the like. Among them, a massive MIMO system can achieve spatial diversity gain and significantly increase system capacity by configuring a large-scale antenna array at the transceiver end. For example, an access network device can simultaneously send data to multiple terminal devices using the same time-frequency resource, that is, multi-user MIMO (MU-MIMO), or the access network device can also simultaneously send multiple data streams to the same terminal device, that is, single-user MIMO (SU-MIMO). The data between the multiple terminal devices or the multiple data streams of the same terminal device are spatially multiplexed, so it becomes a key direction in the evolution of communication systems.
[0134] The access network device needs to obtain the CSI of the downlink channel to determine the configuration of the downlink data channel of the terminal device, such as the resource, the modulation and coding scheme (MCS), and the precoding.
[0135] Taking precoding as an example, in massive MIMO, an access network device needs to precode downlink data by using a precoding matrix. The access network device can use precoding technology to realize spatial division multiplexing between terminal devices or between data streams, that is, data between different terminal devices or between different data streams of the same terminal device is isolated in space, so as to reduce interference between different terminal devices or between different data streams and improve the signal to interference plus noise ratio (SINR) of the terminal device. In order to calculate the precoding matrix, the access network device needs to obtain the CSI of the downlink channel, and determine the precoding matrix according to the CSI.
[0136] In a time division duplexing (TDD) system, since the uplink and downlink channels are reciprocal, the access network device can obtain the uplink CSI by measuring the uplink reference signal, and then infer the more accurate downlink CSI, for example, use the uplink CSI as the downlink CSI. However, in a frequency division duplexing (FDD) system, the uplink and downlink reciprocity cannot be guaranteed, and the downlink CSI is obtained by the terminal device measuring the downlink reference signal, such as measuring the channel state information reference signal (CSI-RS) or the synchronization signal block (SSB) to obtain the downlink CSI. Therefore, the terminal device needs to generate a CSI report in a manner of predefinition by a protocol or configuration by the access network device, and feed back the generated CSI report to the access network device, so that the access network device obtains the downlink CSI.
[0137] In the FDD system, an important part of the CSI feedback is the precoding matrix indicator (PMI), that is, 0-1 bits are used to quantize the channel matrix or the precoding matrix in the CSI. The design of the PMI (also referred to as the codebook design) is a basic problem in a mobile communication system. The traditional codebook design method is to predefine (agree) a series of precoding matrices and corresponding numbers in the protocol, and these precoding matrices are called code words. The linear combination of the predefined code words or a plurality of predefined code words can approximate the channel matrix or the precoding matrix. Therefore, the terminal device can feed back one or more of the numbers corresponding to the code words and the weighting coefficients to the access network device through the PMI, so as to reconstruct the channel matrix or the precoding matrix by the access network device.
[0138] With the increasing size of the antenna array of the MIMO system, the number of supportable antenna ports increases, and the dimension of the corresponding channel matrix and precoding matrix grows. In order to enable the terminal device to estimate (or measure) the downlink channel, the access network device increases the overhead of the reference signal. At the same time, the error of approximating the large-scale channel matrix and precoding matrix with a limited number of predefined codebooks increases. One method to improve the accuracy of channel reconstruction is to increase the number of codebooks in the codebook, but this will also increase the overhead of the CSI feedback (including one or more of the codebook corresponding number and the weighting coefficient), thereby reducing the available resources for data transmission and causing a loss of system capacity. In summary, it is necessary to study how to more effectively compress the channel information and how to more effectively reconstruct the channel according to the feedback information without increasing the overhead of the reference signal and the overhead of the CSI feedback. There is a correlation between different elements in the downlink channel matrix between the access network device and the terminal device, and there is a correlation between the downlink channel matrices of different time slots. For example, the correlation between different elements in the channel matrix means that there is a set of bases (which can be represented by matrices U1 and U2), and when the channel matrix H is projected onto the set of bases, a sparse equivalent channel H' can be obtained, that is, H' = U1H H U2 is a sparse matrix, where the superscript H represents the conjugate transpose operation. In theory, only the non-zero elements in H' need to be estimated and fed back through the reference signal to reconstruct the channel matrix H. Therefore, there is a compression space for the overhead of the reference signal and the CSI feedback. However, the traditional CSI feedback scheme, such as the above-mentioned feedback method based on the codebook, does not fully utilize the channel compression space, and the channel compression process can cause a large amount of information loss. The method of machine learning (such as deep learning (DL)) has stronger nonlinear feature extraction capability, so it can more effectively extract the correlation between channel matrices, and thus can more effectively compress the channel information and more effectively reconstruct the channel information according to the feedback information compared with the traditional scheme.
[0139] To facilitate understanding of the embodiments of the present application, the terms involved in the present application are first explained.
[0140] 1、CSI: The meaning of CSI is broader than that in the traditional scheme, including but not limited to channel quality indication (CQI), PMI, rank indicator (RI), CSI-RS resource indicator (CRI), and can also include one or more of the following: channel response information (such as channel response matrix, frequency domain channel response information, time domain channel response information), weight information corresponding to the channel response, reference signal receiving power (RSRP) or signal to interference plus noise ratio (SINR) and the like.
[0141] 2、AE model: It can generally refer to a network structure composed of two models, for example, composed of an encoder and a decoder. Each model can be an AI model. The AE model can also be called a bilateral model, a double-end model, a cooperative model, etc. The encoder and the decoder of the AE are usually trained together and can be used together.
[0142] In this application, the feedback of the CSI report can be implemented based on the AI model of the AE. FIG. 5 is a schematic diagram of CSI feedback using an AE model according to an embodiment of the present application.
[0143] As shown in the figure, the CSI obtained by the terminal side measurement can be denoted as V, and the measurement obtained CSI (i.e., V) can be used as the input of the encoder. The encoder on the terminal side can compress the measurement obtained CSI, and output the compressed CSI, denoted as C. The terminal side can quantize the compressed CSI to obtain the feedback CSI. The feedback CSI can be carried in the CSI report. The compression of the measurement obtained CSI by the encoder on the terminal side can be understood as an example of an inference task on the terminal side. The terminal side can send the feedback CSI to the network side.
[0144] The network side can first dequantize the received feedback CSI to obtain the compressed CSI, denoted as The compressed CSI (i.e., C) is the input of the decoder on the network side. The decoder on the network side can reconstruct the CSI based on the compressed CSI to obtain the reconstructed CSI (or recovered CSI), denoted as Wherein, the CSI can include a channel matrix or a precoding matrix, etc., without limitation.
[0145] It should be noted that the encoder on the terminal side can be deployed inside the terminal device, or in other devices outside the terminal device, such as the aforementioned OTT host or cloud server, etc.; the decoder on the network side can be deployed inside the network device, or in other devices outside the network device, such as the aforementioned intelligent network element.
[0146] It should be understood that the AE model is only one possible model for implementing CSI compression and quantization on the terminal side and dequantization and CSI reconstruction on the network side, and should not constitute any limitation on the present application. The AE model can also be replaced by other AI models capable of achieving the same or similar functions.
[0147] It should also be understood that although the encoder and the decoder are shown in the figure, this is only a model division from the functional point of view, and should not constitute a limitation on the number of models included in the AE model.
[0148] 3. Model development: Model development refers to constructing and training a model through AI technology to solve a specific inference task. Illustratively, model development can include one or more of the following: model training, model adaptation, or model enhancement. Model training includes one or more of the following: initial training of the model, retraining of the model, fine-tuning of the model, updating of the model. Model adaptation refers to the adaptation of the model format, i.e., converting the received model parameters and / or model files into a model format executable by the device to perform model inference on the device. Model enhancement refers to enhancing the functionality of the model through optimization of the model structure and model training, for example, adding other functions or neural networks to the existing functions of the model, so that the enhanced model can be applied to more scenarios or more diverse needs. Model deployment refers to applying the developed model to actual scenarios.
[0149] Model development can also be embodied as offline engineering in the standard.
[0150] It should be understood that the terms model development, model training, model adaptation, model enhancement, etc. described herein are only introduced for the purpose of facilitating understanding of offline engineering, and in specific implementations, the operations involved in these terms can not be explicitly divided, and the operations involved in each term can include but are not limited to the content listed herein, or offline engineering can also include part of the operations listed above, or it can also include other operations, which are not limited by the present application.
[0151] 4. Model deployment: Model deployment refers to deploying the trained model to a specific device for model inference.
[0152] 5、Reference model: can be a predefined AI model, such as an AI model defined in a standard. The reference model can be determined by the structure of the reference model and / or the parameters of the reference model. The structure of the reference model can specifically refer to the type of neural network, for example, convolutional neural network (CNN), recurrent neural network (RNN), feedforward neural network (FNN), etc. The structure of the reference model can be standard predefined, can be pre-negotiated, can be pre-configured, or can be indicated by a model file, which can have a fixed format, such as a standard predefined format, or a format pre-negotiated by both ends of the interface. The parameters of the reference model can include parameters in the reference model, for example, including but not limited to, the number of layers of the neural network, the type and weight of neurons in each neural network layer, etc. The present application does not limit the way the parameters of the reference model are issued.
[0153] Take DNN as an example. DNN has multiple neural network layers, including an input layer, one or more hidden layers (or called, implicit layers), and an output layer. Generally, the first layer is the input layer, the last layer is the output layer, and the layers in between are hidden layers. Each layer includes multiple neurons. The layers are fully connected. That is, any neuron in the i-th layer is connected to any neuron in the (i+1)-th layer. The input layer can pass the received values (i.e., the input of the DNN) to the intermediate hidden layer after processing by the neurons. Similarly, the hidden layer can pass the calculation result to the last output layer to produce the output of the DNN. FIG. 6 shows an example of DNN. The DNN model shown in FIG. 6 has three neural network layers, namely an input layer, a hidden layer, and an output layer.
[0154] Each neuron can perform a weighted sum operation on its input and generate an output through a nonlinear function. FIG. 7 shows an example of a neuron structure. The input of the neuron shown in FIG. 7 is x = [x0 x1 … xN-1], the weight corresponding to the input is w = [w0 w1 … wN-1], the bias of the weighted sum is b, and the form of the nonlinear function f() can be diversified, for example, the nonlinear function f() is the maximum function max{0, x}. The effect of the execution of a neuron is N-1 N-1 where N is a positive integer; n is a positive integer greater than or equal to 0 and less than or equal to (N-1).
[0155] Suppose the reference model is the DNN shown in FIG. 6, the structure of the reference model is a DNN; the parameters of the reference model include the number of layers of the neural network in the DNN, the number of neurons in each layer, and the values of the weights of each neuron, etc.
[0156] It should be understood that the above examples in connection with FIG. 6 and FIG. 7 are only shown for the convenience of understanding, and should not constitute any limitation on the present application. The present application does not limit the structure and parameters used by the reference model.
[0157] In some possible implementations, the indication of the reference model includes an indication of the structure of the reference model and an indication of the parameters of the reference model. Therefore, sending the reference model can include sending a model file and model parameters of the reference model.
[0158] In another possible implementation, the structure of the reference model can be predefined, such as protocol predefined, and the indication of the reference model includes an indication of the parameters of the reference model. Therefore, sending the reference model can include sending the model parameters of the reference model.
[0159] It should be understood that the present application does not limit the way of indicating the reference model.
[0160] 6、Model parameters: Model parameters can refer to parameters in a neural network model, for example, including but not limited to, the number of layers of the neural network, the type and weight of neurons in each neural network layer, etc. The present application does not limit the way of issuing the parameters of the reference model. In the embodiments of the present application, the model parameters can include the parameters of the reference model, and / or the parameters used for model development of the reference model.
[0161] 7、Double-end interfacing: Refers to the interfacing between the sender (such as the network side) and the receiver (such as the terminal side), which can be the interfacing of the data set or the interfacing of the model. Taking the inference task of CSI compression as an example, the interfacing of the data set and the interfacing of the model are described respectively.
[0162] The interfacing of the data set mainly refers to that the sender provides the data set for the receiver for model training (or model development) and model deployment. FIG. 8 is a schematic diagram of the data set interfacing between the network side and the terminal side.
[0163] As shown in FIG. 8, the network side can first obtain a data set in a joint training manner. Illustratively, the network side can employ a virtual CSI compression model to compress the target CSI, and then employ a CSI reconstruction model to reconstruct the target CSI. That is, the target CSI is the input of the CSI compression model, and the output of the CSI compression model can be compressed CSI. The compressed CSI is the input of the CSI reconstruction model, and the reconstructed CSI is the output of the CSI reconstruction model. In this way, the data set can be obtained. Illustratively, the data set includes the input and output of the CSI compression model, for example, including the target CSI and the compressed CSI.
[0164] Optionally, the compressed CSI can be further quantized to obtain quantized CSI. In a specific implementation, the quantized CSI can be used for CSI reporting. Accordingly, the data set can include, for example, the target CSI, the quantized CSI, and the quantization manner.
[0165] The network side can distribute the data set to the terminal side. The terminal side can perform model training according to the received data set to obtain a CSI compression model of the terminal side.
[0166] The docking of the model mainly refers to that the sender provides the model file and / or model parameter for the receiver, for model deployment of the receiver, or for model development and model deployment of the receiver. FIG. 9 is a schematic diagram of model docking between the network side and the terminal side.
[0167] As shown in FIG. 9, the network side can first obtain a CSI compression model in a joint training manner. The specific process of joint training of the network side can refer to the description of FIG. 8 above, and will not be described herein. The network side can distribute the CSI compression model obtained by joint training to the terminal side. For example, the network side can send the model file and / or model parameter of the CSI compression model to the terminal side. The terminal side can perform model deployment according to the received model file and / or model parameter, or perform model development and model deployment.
[0168] In order to support the docking and model development or deployment of the double-end AI model, it is necessary to provide a data set or a model for one or both of the double-end AI model. For example, the network side can provide a data set for the terminal side, for model development and deployment; for another example, the network side can provide a reference model and / or model parameter for the terminal side, for model development and / or deployment. However, since the network side can not be aware of the capability of the terminal side, such as the computing capability, the storage capability, etc., there can be a difference between the data set, the reference model, the model parameter, etc. sent by the network side and the condition of the terminal side. The performance of the model developed and / or deployed by the terminal side based on the received data set, the reference model, the model parameter, etc. can not achieve the expected effect.
[0169] The model performance can be controlled by monitoring the model performance. The model monitoring refers to monitoring the performance of the model to determine whether the model works normally. If the model performance is poor, the non-AI mode can be switched to, i.e., the model is not used or the model is switched or the model is updated, etc. The model monitoring can be determined by monitoring the accuracy of the AI model output or by monitoring the system performance. Among them, the accuracy of the AI model output is determined by comparing the difference between the output of the AI model and the corresponding label or ground truth to determine whether the performance of the AI model meets the requirements. The system performance is monitored to determine whether the AI model meets the requirements. Exemplarily, the model performance monitoring can be evaluated by key performance indicators (KPIs). For monitoring the AI model output, the KPIs can be referred to as intermediate KPIs, such as but not limited to, generalized cosine similarity (GCS), square generalized cosine similarity (SGCS), normalized mean square error (NMSE), etc. For monitoring the system performance, the KPIs can be referred to as final KPIs, such as but not limited to, throughput, spectral efficiency, transmission rate, block error rate (BLER), hypothetical BLER, hybrid automatic repeat request (HARQ) feedback, etc.
[0170] As can be seen, the above monitoring of the model performance mainly focuses on the accuracy of the model. However, the model performance is not limited to the accuracy dimension. Simply focusing on the accuracy of the model and ignoring the performance of other dimensions may cause the performance of the model to be lost or even cause the model to be unusable. For example, if the developed model has a large processing delay, even if the accuracy meets the performance indicators, it may not meet the business computing time requirements and cannot be used.
[0171] Therefore, the present application provides a method for providing a delay requirement, so that a device (such as a terminal device) to which a first model is about to be deployed or has been deployed can use the delay requirement as a reference to perform an inference task of the first model. Thus, it is beneficial to meet the business computing time requirements and improve the model performance, thereby avoiding the situation that the model deployed on the device is unusable due to the large computing time or the failure to meet the business computing time requirements.
[0172] The method provided in the present application will be described in detail below in combination with the accompanying drawings.
[0173] FIG. 10 is a schematic flowchart of a communication method 800 provided by an embodiment of the present application. FIG. 10 shows the flow of the method 800 from the perspective of the interaction between the terminal side and the network side. As shown in the figure, the communication method 800 shown in FIG. 10 can include step 810. Optionally, the method 800 further includes one or more of steps 820 to 850. Each step in the method 800 will be described in detail below.
[0174] In step 810, the terminal side acquires a computing time requirement, which is a range that the time length for the terminal side to perform an inference task of a first model should satisfy.
[0175] In the present application, the computing time requirement refers to a range, or a recommended range, that the time length for the terminal side to perform an inference task of a first model should satisfy. That is, the time length for the terminal side to perform an inference task of a first model should be controlled as much as possible within the computing time requirement.
[0176] In the present application, the computing time requirement refers to a range, or a recommended range, that the time length for the terminal side to perform an inference task of a first model should satisfy. That is, the time length for the terminal side to perform an inference task of a first model should be controlled as much as possible within the computing time requirement.
[0177] Exemplarily, the computing time requirement can only include an upper limit, without including a lower limit; or can include an upper limit and a lower limit, which is not limited in the present application. For example, the computing time requirement is [t1, t2], and t2 is greater than t1. Then the actual time length required by the terminal side to perform an inference task of the first model should be greater than or equal to t1 and less than or equal to t2. For another example, the computing time requirement is less than or equal to t2, then the actual time length required by the terminal side to perform an inference task of the first model should be less than or equal to t2.
[0178] In another implementation, the computation duration requirement can also be a range of duration of processing unit data volume or a range of computation speed that can derive a range of duration of performing the inference task. It can be understood that in the case of the inference task, the data volume that needs to be processed for performing the inference task (hereinafter referred to as processing data volume) can also be determined, and therefore, based on the range of duration of processing unit data volume or the range of computation speed, the range of duration of performing the inference task on the terminal side can also be derived, and therefore, the range of duration of processing unit data volume or the range of computation speed can be regarded as another form of the computation duration requirement.
[0179] It should be understood that the above examples of the computation duration requirement are only for the convenience of understanding, and the specific form and range of the computation duration requirement are not limited in the present application.
[0180] As an example, the inference task can be the inference of CSI compression. The terminal side performing the inference task of the first model can mean that the terminal side performs the inference task using the first model, such as the terminal side performing the inference of CSI compression using the first model. Since the terminal side (more specifically, the terminal device) can support data transmission of one or more (assuming R, R is a positive integer) transmission layers, the inference task of the first model is the inference of CSI compression of R transmission layers. For more detailed description of the inference of CSI compression, please refer to the description of FIG. 5 above, which will not be repeated here.
[0181] Optionally, the computation duration requirement can also be indicated by the network side, such as the network side can send third information to the terminal side to indicate the computation duration requirement. Accordingly, the step 810 specifically includes 810a: the terminal side receives the third information, and the third information is used to indicate the computation duration requirement. Accordingly, the network side sends the third information.
[0182] As an example, the terminal side receiving the third information can be that the terminal device receives the third information from the network device, or the terminal device receives the third information from the host or cloud server of the OTT system, and the computation duration requirement can be received by the host or cloud server of the OTT system from the intelligent network element of the network side. Alternatively, the terminal device receiving the third information can be that the host or cloud server of the OTT system receives the third information from the intelligent network element of the network side, or the host or cloud server of the OTT system receives the third information from the terminal device, and the computation duration requirement can be received by the terminal device from the network device.
[0183] It can be seen that the calculation duration requirement is sent by the network side to the terminal side through the third information, regardless of whether it is received from the network device or the intelligent network element. Alternatively, before step 810a, the method further includes that the network side determines the calculation duration requirement.
[0184] The first calculation duration may, for example, be determined by the network side according to the service requirement, such as according to the calculation duration requirement of the service or according to the end-to-end (E2E) duration requirement of the service, and the present application does not limit this. Alternatively, the network side may, in combination with the capability information of the terminal side, determine the calculation duration requirement, in which case the network side may receive the capability information of the terminal side before determining the calculation duration requirement. Based on this, the network side may determine a calculation duration requirement that can be adapted to the capability of the terminal side.
[0185] Alternatively, the calculation duration requirement may be predefined, such as predefined by a protocol. Therefore, the calculation duration requirement may be pre-stored in the memory when the device of the terminal side, such as the terminal device, is manufactured, and the calculation duration requirement is read from the memory when there is a use requirement. Accordingly, step 810 may specifically include 810b: the terminal side obtains the calculation duration requirement from the local terminal.
[0186] Exemplarily, step 810b includes that the terminal device obtains the calculation duration requirement from the local terminal (corresponding to the case that the terminal device stores the calculation duration requirement), or the host or cloud server of the OTT system obtains the calculation duration requirement from the terminal device (corresponding to the case that the terminal device stores the calculation duration requirement), or the host or cloud server of the OTT system obtains the calculation duration requirement from the local terminal (corresponding to the case that the host or cloud server of the OTT system stores the calculation duration requirement), or the terminal device obtains the calculation duration requirement from the host or cloud server of the OTT system (corresponding to the case that the host or cloud server of the OTT system stores the calculation duration requirement).
[0187] It should be understood that the calculation duration requirement is a parameter or a reference value for the terminal side to perform the inference task of the first model for model performance monitoring, and how the terminal side uses the calculation duration requirement after obtaining the calculation duration requirement is an internal implementation of the device of the terminal side, which is not limited by the present application.
[0188] The above description of the calculation duration requirement is only a range, and in specific implementation, the start point and the end point of the calculation duration may be further defined, so that the network side and the terminal side describe the calculation duration based on the same definition. The network side may determine the calculation duration requirement based on the start point and the end point of the calculation duration. The terminal side may determine the calculation duration based on the start point and the end point of the calculation duration.
[0189] It can be understood that, since the calculation duration in the present application refers to the duration of performing the inference task of the first model, the start point and the end point of the calculation duration are also the start point and the end point of the duration of the inference task of the first model, or in other words, the start point and the end point of the duration of the inference task.
[0190] Optionally, the start point of the calculation duration is one of the following: the configuration time of the reference signal resource, the transmission time of the reference signal, or the configuration time of the CSI report, and the end point of the calculation duration is the transmission time of the CSI report; wherein the reference signal is used for CSI inference, and the reference signal resource is used for transmitting the reference signal.
[0191] The configuration time of the reference signal resource can refer to the transmission time of the configuration information of the reference signal (corresponding to the network side) or the reception time of the configuration information of the reference signal (corresponding to the terminal side). The configuration time of the CSI report can refer to the transmission time of the configuration information of the CSI report (corresponding to the network side) or the reception time of the configuration information of the CSI report (corresponding to the terminal side). The transmission time of the reference signal (corresponding to the network side) can also be the reception time of the reference signal (corresponding to the terminal side).
[0192] For the convenience of understanding and description, the calculation duration determined based on different start points and end points is described below by taking FIG. 11 as an example.
[0193] FIG. 11 shows the configuration time T1 of the reference signal resource, the transmission time T2 of the reference signal, the configuration time T3 of the CSI report, and the transmission time T4 of the CSI report, respectively, in the order of execution. For the network side, the configuration time T1 of the reference signal resource can be the transmission time of the configuration information of the reference signal resource; for the terminal side, the configuration time T1 of the reference signal resource can be the reception time of the configuration information of the reference signal resource. For the network side, T2 can be the transmission time of the reference signal; for the terminal side, T2 can be the reception time of the reference signal. For the network side, the configuration time T3 of the CSI report can be the transmission time of the configuration information of the CSI report; for the terminal side, the configuration time T3 of the CSI report can be the reception time of the configuration information of the CSI report. For the network side, T4 can be the reception time of the CSI report; for the terminal side, T4 can be the transmission time of the CSI report.
[0194] As shown in the figure, if the start point of the calculation duration is the configuration moment of the reference signal resource, and the end point is the sending moment of the CSI report, the calculation duration #1 can be obtained as the difference between T4 and T1; if the start point of the calculation duration is the sending moment of the reference signal, and the end point is the sending moment of the CSI report, the calculation duration #2 can be obtained as the difference between T4 and T2; if the start point of the calculation duration is the configuration moment of the CSI report, and the end point is the sending moment of the CSI report, the calculation duration #3 can be obtained as the difference between T4 and T3.
[0195] Exemplarily, the reference signal can be a channel state information reference signal (CSI-RS) for downlink channel measurement, and the reference signal resource can be a CSI-RS resource. The configuration information of the reference signal resource can be information for configuring the CSI-RS resource, for example, can be a CSI resource configuration (CSI-ResourceConfig) defined in the third generation partnership project (3GPP) technical specification (TS) 38.214. The configuration information of the CSI report can be information for configuring the reporting resource of the CSI report, for example, can be a CSI reporting configuration (CSI-ReportConfig) defined in the 3GPP TS 38.214. rd Exemplarily, the reference signal can be a channel state information reference signal (CSI-RS) for downlink channel measurement, and the reference signal resource can be a CSI-RS resource. The configuration information of the reference signal resource can be information for configuring the CSI-RS resource, for example, can be a CSI resource configuration (CSI-ResourceConfig) defined in the third generation partnership project (3GPP) technical specification (TS) 38.214. The configuration information of the CSI report can be information for configuring the reporting resource of the CSI report, for example, can be a CSI reporting configuration (CSI-ReportConfig) defined in the 3GPP TS 38.214.
[0196] It should be understood that the CSI-RS, the CSI-RS resource and the associated CSI resource configuration, the CSI report and the associated CSI reporting configuration are only examples for facilitating understanding, and should not constitute any limitation on the present application.
[0197] Optionally, the calculation duration requirement includes R max ranges corresponding to R max transmission layers, and the range corresponding to the i-th transmission layer in the R max transmission layers is a range that needs to be met by the inference duration of compressing the CSI of the i-th transmission layer, R max is the maximum value of the number of transmission layers supported by the terminal device, and R max is a positive integer.
[0198] wherein R max is the maximum value of the number of transmission layers supported by the terminal device, and R max is a positive integer. R maxThe value of R can be predefined, such as predefined by a protocol, or reported by the terminal side through capability information, such as reported in step 850, which is not limited in the present application. It can be understood that if R is predefined, the value of R can be different from the maximum number of transmission layers R actually supported by the terminal side, for example, R can be greater than R. max If R is predefined, the value of R can be different from the maximum number of transmission layers R actually supported by the terminal side, for example, R can be greater than R. max If R is predefined, the value of R can be different from the maximum number of transmission layers R actually supported by the terminal side, for example, R can be greater than R. max
[0199] The terminal side can perform inference of CSI compression of each transmission layer with reference to the range corresponding to each transmission layer, so as to facilitate the model deployed on the terminal side to meet the computing time length requirement of the service, improve the performance of the model, and avoid the model deployed on the terminal side from being unavailable due to too long computing time length or failure to meet the computing time length requirement of the service.
[0200] Optionally, the method further includes: determining, by the terminal side, whether the time length of performing the inference task of the first model meets the computing time length requirement according to the computing time length requirement.
[0201] Before the terminal side determines whether the time length of the inference task of the first model meets the computing time length requirement, the first model can have been obtained through model development, or the first model can have been deployed on the terminal side, or the first model can not have been obtained, such as not having been developed or deployed. Therefore, the time length of performing the inference task of the first model by the terminal side can specifically refer to the time length of performing the inference task of the first model predicted by the terminal side (which can correspond to the case that the first model has not been developed or deployed), or the time length of performing the inference task of the first model measured by the terminal side (which can correspond to the case that the first model has been deployed on the terminal side).
[0202] Correspondingly, the terminal side can determine whether the time length of performing the inference task meets the computing time length requirement according to the predicted time length of performing the inference task of the first model, or according to the measured time length of performing the inference task of the first model.
[0203] Exemplarily, one possible implementation manner that the terminal side determines whether the time length for performing the inference task of the first model meets the computation time length requirement according to the predicted time length for performing the inference task of the first model is that the terminal side can determine (or predict) whether the time length for performing the inference task of the first model meets the computation time length requirement according to the storage capability, the computation capability and the like of the device (such as a terminal device or a host or a cloud server of an OTT system) to which the first model is deployed. For example, in the case of the inference task, the processing data amount of the inference task can also be determined, and the terminal side can predict the time length for completing the inference task according to the processing data amount and the computation speed (which is an example of the computation capability), and further predict whether the time length for performing the inference task of the first model meets the computation time length requirement. For another example, the computation capability of the device (such as a terminal device or a host or a cloud server of an OTT system) to which the first model is deployed is strong, and the terminal side can also predict the time length for completing the inference task according to the maximum value of the storage space of the device and the processing data amount of the inference task, and further predict whether the time length for performing the inference task of the first model meets the computation time length requirement.
[0204] Another possible implementation manner that the terminal side determines whether the time length for performing the inference task of the first model meets the computation time length requirement according to the predicted time length for performing the inference task of the first model is that, since the computation time length requirement can also be represented by the range of the time length for processing a unit data amount or the range of the computation speed that needs to be met, and the range of the time length for performing the inference task can be derived from the parameters, the terminal side can also directly predict the time length for performing the inference task according to the computation capability of the device to which the model is deployed, and further predict whether the time length for performing the inference task of the first model meets the computation time length requirement.
[0205] It should be understood that the specific manner that the terminal side determines whether the time length for performing the inference task of the first model meets the computation time length requirement exemplified above is only given for the purpose of understanding, and the application does not limit the specific manner that the terminal side determines whether the time length for performing the inference task of the first model meets the computation time length requirement.
[0206] In addition, the terminal side determining whether the time length for performing the inference task of the first model meets the computation time length requirement is implemented internally in the device of the terminal side, and in the specific implementation process, it can also be combined with step 820 or 840 below, and is not executed as a separate step. Herein, it is only described for the purpose of understanding.
[0207] It should also be understood that the step of determining whether the duration of the terminal side performing the inference task of the first model meets the computation duration requirement can be performed before the model development of the first model, can be performed after the model development of the first model, can be performed before the model deployment of the first model, and can be performed after the model deployment of the first model. The present application does not limit the time sequence of the step and the model development or the model deployment.
[0208] The duration of the terminal side performing the inference task of the first model can meet the computation duration requirement or can not meet the computation duration requirement. The following will be described in combination with the two cases of meeting the computation duration requirement and not meeting the computation duration requirement.
[0209] Optionally, the method 800 further includes step 820: in the case where the duration of the terminal side performing the inference task of the first model meets the computation duration requirement, performing the inference task of the first model.
[0210] In the case where the duration of the terminal side performing the inference task of the first model meets the computation duration requirement, the terminal side can perform the inference task of the first model.
[0211] Different devices have different computing capabilities. The actual duration required for performing the inference task of the first model on different terminal side devices can be different. Therefore, the terminal side can report the actual duration required for performing the inference task of the first model to the network side, so that the network side configures the reporting time of the CSI report according to the duration reported by the terminal side.
[0212] Optionally, the method 800 can further include step 830: the terminal side sends first information, the first information being used to indicate the first computation duration, the first computation duration being the actual duration required for performing the inference task of the first model. Correspondingly, the network side receives the first information.
[0213] Taking CSI compression as an example of the inference task, the first computation duration can refer to the actual duration required for the first model to perform the inference of CSI compression, or the actual duration required for the first model to perform CSI compression.
[0214] The terminal side can indicate the first computation duration by the first information in various ways. For example, the terminal side can directly indicate the value of the duration by the first information, or can indicate the difference between the first computation duration and the upper limit or lower limit of the computation duration requirement by the first information, etc. The present application does not limit this.
[0215] Since the terminal side can support data transmission of R transmission layers, the inference task of the first model is the inference of CSI compression of the R transmission layers. When the terminal side performs the inference task of the first model, the inference of CSI compression of the R transmission layers can be performed in series, for example, the inference of CSI compression of the R transmission layers can be performed by one model; or the inference of CSI compression of the R transmission layers can be performed in parallel, for example, the inference of CSI compression of the R transmission layers can be performed by R models. In the two modes of series and parallel, the actual time length required by the terminal side to perform the inference task of the first model is also different.
[0216] An example of serial processing is as follows: the terminal side can perform the inference of CSI compression of the R transmission layers by one model, that is, compress the CSI of the R transmission layers by one model. In this case, the first model is one model for compressing the CSI of the R transmission layers. The above execution first calculation time length can be the total time length actually required by the first model to perform the inference of CSI compression of the R transmission layers, or in other words, the total time length actually required by the first model to compress the CSI of the R transmission layers.
[0217] The network side can refer to the first calculation time length indicated by the first information reported by the terminal side to configure the CSI reporting time of the terminal side, for example, the terminal side can report the CSI of the R transmission layers by one CSI report. It can be understood that the CSI of the R transmission layers is the CSI after compression and quantization. The network side can refer to the first calculation time length to configure the reporting time of the terminal side to perform CSI reporting, for example, according to the starting point of the predefined calculation time length, the time interval between the reporting time of the terminal side to perform CSI reporting and the starting point is not less than the first calculation time length.
[0218] Optionally, the first information is also used to indicate R second calculation time lengths, the R second calculation time lengths correspond to the R transmission layers, and the rth second calculation time length in the R second calculation time lengths can indicate the time length actually required for the inference of CSI compression of the rth transmission layer in the R transmission layers, or in other words, the rth second calculation time length can indicate the time length actually required for compressing the CSI of the rth transmission layer.
[0219] It can be understood that if the first information indicates the R second calculation time lengths, the first calculation time length can be obtained by adding the R second calculation time lengths, in other words, the indication of the R second calculation time lengths can also be regarded as an implicit indication of the first calculation time length. Therefore, the first information can indicate the R second calculation time lengths without indicating the first calculation time length. Or it can also be said that the first information indicates the R second calculation time lengths, that is, another way of indicating the first calculation time length.
[0220] In this case, the network side can parse R second calculation durations from the first information according to the number of transmission layers R reported by the terminal side (for example, reported through RI), and further obtain the first calculation duration. The network side can configure the reporting time of the CSI report of the terminal side with the first calculation duration as a reference.
[0221] An example of parallel processing is as follows: the terminal side can perform the inference of the CSI compression of the R transmission layers through R models, that is, compress the CSI of the R transmission layers through the R models respectively. In this case, the first model includes the R models, which correspond to the R transmission layers one by one. Each model can be used to perform the inference of the CSI compression of the corresponding transmission layer, or in other words, each model is used to compress the CSI of the corresponding transmission layer. For example, the r th model in the R models is used to perform the inference of the CSI compression of the r th transmission layer in the R transmission layers, or in other words, is used to compress the CSI of the r th transmission layer. The above-mentioned first calculation duration can be the actual time required for the R models to perform the inference of the CSI compression of the R transmission layers, or in other words, the actual time required for the R models to compress the CSI of the R transmission layers respectively.
[0222] It should be noted that since the R models can perform the inference task in parallel, the first calculation duration is not necessarily the sum of the time required for the R models to perform the inference task respectively.
[0223] The network side can configure the reporting time of the CSI report of the terminal side with the first calculation duration indicated by the first information reported by the terminal side as a reference.
[0224] It should be further noted that the terminal side sending the first information can be that the terminal device sends the first information to the network device, and the first information can be determined by the terminal device by performing the inference task of the first model; or the inference task of the first model is performed on the host or cloud server of the OTT system, so that the first calculation duration (optionally, the R second calculation durations) can be obtained by the terminal device from the host or cloud server of the OTT system and sent to the network device through the first information. The terminal side sending the first information can also be that the host or cloud server of the OTT system sends the first information to the intelligent network element of the network side, and the first information can be determined by the host or cloud server of the OTT system by performing the inference task of the first model; or the inference task of the first model is performed on the terminal device, so that the first calculation duration (optionally, the R second calculation durations) can be obtained by the host or cloud server of the OTT system from the terminal device and sent to the intelligent network element through the first information.
[0225] Optionally, the method 800 further includes step 840: the terminal side sends second information in a case that the duration of performing the inference task of the first model does not satisfy the computation duration requirement. Correspondingly, the network side receives the second information.
[0226] Exemplarily, the second information can be used for one or more of the following:
[0227] a. indicating that the duration of performing the inference task of the first model does not satisfy the computation duration requirement;
[0228] b. requesting to close the inference task of the first model;
[0229] c. requesting to change the computation duration requirement;
[0230] d. requesting to change the first model;
[0231] e. requesting to change the model parameters; or
[0232] f. requesting to change the dataset used for the model training of the first model.
[0233] For example, the terminal side can request to close the inference task of the first model through the second information. The terminal side can also indicate through the second information that the duration of performing the inference task of the first model does not satisfy the computation duration requirement.
[0234] Closing the inference task of the first model can be understood as that the terminal side no longer performs the inference task of the first model. Based on this, the network side can no longer perform relevant configurations for the inference task of the first model. For example, the inference task is the inference of CSI compression of R transmission layers, and since the terminal side does not satisfy the computation duration requirement in performing the inference task of the first model, the terminal side can no longer be configured to report the CSI report of the R transmission layers, and the like. It should be understood that the operation of the network side after closing the inference task of the first model can be determined by the network side itself, and the present application does not limit this.
[0235] For another example, since the terminal side has limited computing capability, the computation duration requirement is difficult for the terminal side to satisfy, and therefore the computation duration requirement can be requested to be changed through the second information, for example, the computation duration requirement is requested to be changed to a larger duration range. For example, the upper limit of the duration range is increased. The terminal side can also indicate through the second information that the duration of performing the inference task of the first model does not satisfy the computation duration requirement.
[0236] For another example, the first model is received by the terminal side from the network side, and therefore the terminal side can request to change the first model through the second information, for example, the first model is requested to be changed to a simpler model, such as a model with lower computational complexity. The terminal side can also indicate through the second information that the duration of performing the inference task of the first model does not satisfy the computation duration requirement.
[0237] For example, the first model is developed by the terminal side based on the model parameters received from the network side, and thus the terminal side can request to replace the model parameters through the second information, for example, replace the model parameters with simpler model parameters, such as fewer layers of neural networks. The terminal side can also indicate through the second information that the time length of the inference task of the first model does not meet the calculation time length requirement.
[0238] For another example, the first model is trained by the terminal side according to the data set received from the network side, and thus the terminal side can request to replace the data set through the second information, for example, request to replace the data set with a simpler data set, such as a data set with lower computational complexity, or a data set with lower calculation accuracy, etc. The terminal side can also indicate through the second information that the time length of the inference task of the first model does not meet the calculation time length requirement.
[0239] For another example, the terminal side can also indicate through the second information that the time length of the inference task of the first model does not meet the calculation time length requirement without indicating other information (such as any one of b to f listed above). How the network side responds after receiving the second information can be determined by the network side itself, or can also be predefined by the protocol. For example, the protocol can define that in the case where the time length of the inference task of the first model does not meet the calculation time length requirement, the inference task of the first model is closed, or the calculation time length requirement is replaced, etc., which are not listed here.
[0240] The terminal side sends the second information, which can be that the terminal device sends the second information to the network device, or that the host or cloud server of the OTT system sends the second information to the intelligent network element. The present application does not limit this.
[0241] It should be understood that steps 820, 830 and step 840 are steps respectively performed by the terminal side in the case where the time length of the inference task of the first model meets or does not meet the calculation time length requirement, and the terminal side can perform one of them according to the actual situation, and does not necessarily have to perform all of them.
[0242] Optionally, before step 810, the method 800 further includes step 850, the terminal side sends capability information, the capability information indicating one or more of the following: the maximum number of transmission layers supported by the terminal device, the storage capability supported by the terminal side, or the calculation capability supported by the terminal side. Correspondingly, the network side receives the capability information.
[0243] The storage capability can refer to the maximum value of the storage space, or the upper limit of the storage space. The computing capability can also be referred to as computing power, which can be represented by parameters such as computing speed and processing time per unit data volume. The storage capability supported by the terminal side and the computing capability supported by the terminal side can refer to the storage capability and the computing capability supported by the device (such as a terminal device or a host or a cloud server of an OTT system) for deploying the first model.
[0244] In one example, the terminal device is used to perform the inference task of the first model, or in other words, the terminal device is the device for deploying the first model. The terminal side can send the capability information, which can be that the terminal device sends the capability information to the network device, and the capability information indicates one or more of the following: the maximum number of transmission layers supported by the terminal device, the storage capability supported by the terminal device, or the computing capability supported by the terminal device.
[0245] In another example, the host or the cloud server of the OTT system is used to perform the inference task of the first model, or in other words, the host or the cloud server of the OTT system is the device for deploying the first model. The capability information indicates one or more of the following: the maximum number of transmission layers supported by the terminal device, the storage capability supported by the host or the cloud server of the OTT system, or the computing capability supported by the host or the cloud server of the OTT system. The terminal side can send the capability information, which can be that the terminal device sends the capability information to the network device, in which case the terminal device can obtain its storage capability and / or computing capability from the host or the cloud server of the OTT system. Alternatively, the terminal side can also send the capability information, which can be that the host or the cloud server of the OTT system sends the capability information to the intelligent network element of the network device, in which case the host or the cloud server of the OTT system can obtain the maximum number of transmission layers supported by the terminal device from the terminal device.
[0246] The network side can also obtain the capability information of the terminal side in advance, and determine the computing duration requirement according to the capability information of the terminal side, so that the computing duration requirement issued to the terminal side can match the capability of the terminal side. For example, the network side can also predict the duration required for performing the inference task of the first model in the manner listed in the foregoing step 810, and determine the computing duration requirement with reference to this.
[0247] In addition, the network side can also determine the data for model development and / or model deployment to be sent to the terminal side according to the capability information of the terminal side, so that the data sent can adapt to the capability of the terminal side, or in other words, the time length of the terminal side performing the inference task meets the calculation time length requirement. Thus, it can also be avoided that the data for model development and / or model deployment sent does not adapt to the capability of the terminal side, resulting in that the terminal side does not perform model development and / or model deployment, or resulting in that the terminal side re-requests the data for model development and / or model deployment from the network side, causing a lengthy process of model development and / or model deployment.
[0248] In one possible design (e.g., referred to as design one), the first model is obtained through model development.
[0249] In one example (e.g., referred to as design one A), the first model is obtained through model training. Optionally, the method further includes: sending, by the network side, a data set, the data set being used for model training of the first model. Correspondingly, the terminal side receives the data set.
[0250] The network side sending the data set can be that the network device sends the data set to the terminal device, or the intelligent network element of the network side sends the data set to the host or cloud server of the OTT system of the terminal side, which is not limited in the present application.
[0251] If the device for deploying the first model is the terminal device, the terminal side receiving the data set can be that the terminal device receives the model from the network device, or the terminal device receives the model from the intelligent network element from the host or cloud server of the OTT system. If the device for deploying the first model is the host or cloud server of the OTT system, the terminal side receiving the data set can be that the host or cloud server of the OTT system receives the data set from the intelligent network element, or the host or cloud server of the OTT system receives the model from the network device from the terminal device.
[0252] Exemplarily, the terminal side can initialize a model locally, and then perform model training based on the data set from the network side to obtain the first model.
[0253] The specific process of the terminal side performing model training of the first model can refer to the prior art, which is not described in detail herein.
[0254] In another example (e.g., referred to as design one B), the first model is obtained through model development on a received second model. Optionally, the method further includes: sending, by the network side, the second model. Correspondingly, the terminal side receives the second model.
[0255] It should be understood that the second model can be a reference model for performing model development to obtain the first model. Therefore, the second model can be one model, or Rmax A model, with R max Each transport layer corresponds to one of them. Corresponding to the second model, the terminal side can develop a model based on one model, or it can be based on R. max Model development is carried out on R models out of the given models. Among them, regarding R... max The determination of the value of and its relationship with the maximum number of transmission layers R supported by the terminal device can be found in the detailed explanation in step 810 above, and will not be repeated here.
[0256] One possible way to develop a second model is to train a first model using the second model as the initial model. In this case, the method further includes: the network side sending a dataset for training the first model; and the terminal side receiving the dataset accordingly.
[0257] Another possible approach to developing the second model is to adjust its parameters to obtain the first model. In this case, the method further includes: the network side sending model parameters for model development; and the terminal side receiving the model parameters accordingly.
[0258] In another possible design (e.g., Design 2), the first model is received.
[0259] For example, the first model is sent from the network side to the terminal side. For a more detailed explanation of how the network side sends the model to the terminal side, please refer to the reference model and model parameters in the terminology introduction above, which will not be repeated here.
[0260] Optionally, the method further includes: the network side sending a first model. Correspondingly, the terminal side receiving the first model.
[0261] The first model can be a single model or R. max A model, with R max Each transport layer corresponds to one of them. Among them, regarding R... max The determination of the value of and its relationship with the maximum number of transmission layers R supported by the terminal device can be found in the detailed explanation in step 810 above, and will not be repeated here.
[0262] In summary, the network side can send models and model parameters to the terminal side. Therefore, optionally, the method further includes: the network side sending fourth information indicating one or more models, which may include a first model, or the one or more models may include a second model. For example, the fourth information may include one or more model files and / or one or more sets of model parameters.
[0263] The network side sending model can be that the network device sends a model to the terminal device, or the intelligent network element at the network side sends a model to the host of the OTT system or the cloud server at the terminal side. The present application does not limit this.
[0264] If the device for deploying the first model is the terminal device, the terminal side receiving model can be that the terminal device receives a model from the network device, or the terminal device receives a model from the intelligent network element from the host of the OTT system or the cloud server. If the device for deploying the first model is the host of the OTT system or the cloud server, the terminal side receiving model can be that the host of the OTT system or the cloud server receives a model from the intelligent network element, or the host of the OTT system or the cloud server receives a model from the network device.
[0265] As described above, the terminal side can send the capability information to the network side, so the network side can issue a data set matching the capability of the terminal side, or a model file and / or model parameters matching the capability of the terminal side to the terminal side according to the capability of the terminal side.
[0266] Based on the above technical solution, the terminal side can obtain the calculation time length requirement, that is, determine the range that the time length for executing the inference task of the first model should meet, and then use this as a reference to execute the inference task of the first model. For example, the inference task of the first model can be executed in the case where the calculation time length requirement can be met, or reported to the network side in the case where the calculation time length requirement cannot be met, the inference task of the first model is closed, or a simpler model or data set is requested, and the like, thereby facilitating the model deployed at the terminal side to meet the calculation time length requirement of the business, facilitating the improvement of the model performance, and thereby avoiding the model deployed at the terminal side from being unusable due to the calculation time length being too large, not meeting the calculation time length requirement of the business, and the like.
[0267] In order to better understand the method provided by the present application, the above method will be described below in combination with a more specific flow. FIG. 12 and FIG. 13 take the terminal device as an example of the terminal side and the network device as an example of the network side, and respectively describe the specific implementation flow of the method 800 in the two scenarios of data set docking and model docking.
[0268] FIG. 12 is another schematic flowchart of a communication method provided by an embodiment of the present application. The method 1200A shown in FIG. 12 includes the following steps:
[0269] In step 1201, the terminal device sends capability information to the network device, the capability information indicating one or more of the following: the maximum number of transmission layers supported by the terminal device, the storage capability of the terminal device, or the calculation capability of the terminal device.
[0270] The details of step 1201 can be referred to the related description of step 850 in method 800 above, and will not be repeated here.
[0271] In step 1202a, the network device sends a data set for model development and model deployment. Accordingly, the terminal device receives the data set.
[0272] The details of step 1202a can be referred to the related description of design A in method 800 above, and the related description of data set docking in the term introduction, and will not be repeated here.
[0273] In step 1203, the network device sends third information for indicating a calculation duration requirement. Accordingly, the terminal device receives the third information.
[0274] The details of step 1203 can be referred to the related description of step 810a in method 800 above, and will not be repeated here.
[0275] In step 1204a, the terminal device develops and deploys a model according to the data set.
[0276] The terminal device can train a model according to the data set, and after completing the training, deploy the first model obtained by training to an actual scene.
[0277] It should be understood that step 1204a can be performed after step 1202, or after step 1203, which is not limited in the present application.
[0278] In step 1205, the terminal device determines whether the duration of performing an inference task of the first model meets the calculation duration requirement.
[0279] The details of step 1205 can be referred to the related description of determining whether the duration of performing an inference task of the first model meets the calculation duration requirement on the terminal side in method 800 above, and will not be repeated here.
[0280] It should be understood that step 1205 can be performed after step 1203, or after step 1204a, which is not limited in the present application.
[0281] In step 1206, the terminal device performs an inference task of the first model when the duration of performing the inference task of the first model meets the calculation duration requirement.
[0282] In step 1207, the terminal device sends first information to the network device, the first information being used for indicating a first calculation duration.
[0283] Step 1208. In a case where the time length for performing the inference task of the first model does not satisfy the computing time length requirement, the terminal device sends second information, the second information being used for one or more of the following: indicating that the time length for performing the inference task of the first model does not satisfy the computing time length requirement, requesting to close the inference task of the first model, requesting to change the computing time length requirement, or requesting to change the data set.
[0284] The detailed description of steps 1206 to 1208 can refer to the related description of steps 820 to 840 in method 800 above, and will not be repeated here.
[0285] It should be understood that steps 1206 and 1207 and step 1208 can be executed alternatively according to the result determined in step 1205, and do not necessarily have to be executed all.
[0286] FIG. 12 describes a communication method provided by the present application, taking the docking of a data set between a terminal device and a network device as an example. The technical solution shown in FIG. 12 corresponds to the technical solution shown in FIG. 10 above, and thus the beneficial effects achieved are similar, and will not be repeated here.
[0287] FIG. 13 is another schematic flowchart of a communication method provided by an embodiment of the present application. The method 1200B shown in FIG. 13 includes the following steps:
[0288] Step 1201. The terminal device sends capability information to the network device, the capability information indicating one or more of the following: the maximum number of transmission layers supported by the terminal device, the storage capability of the terminal device, or the computing capability of the terminal device.
[0289] The detailed description of step 1201 can refer to the related description of step 850 in method 800 above, and will not be repeated here.
[0290] Step 1202b. The network device sends fourth information, the fourth information including one or more model files and / or one or more sets of model parameters. Correspondingly, the terminal device receives the fourth information.
[0291] The detailed description of step 1202b can refer to the related description of design one B and design two in method 800 above, and the related description of model docking in the term introduction, and will not be repeated here.
[0292] Step 1203. The network device sends third information, the third information being used for indicating a computing time length requirement. Correspondingly, the terminal device receives the third information.
[0293] The detailed description of step 1203 can refer to the related description of step 810a in method 800 above, and will not be repeated here.
[0294] At step 1204b, the terminal device performs model development and / or model deployment according to the fourth information.
[0295] As described above, the fourth information can include one or more model files and / or one or more sets of model parameters, and the terminal device can perform model development and / or model deployment according to the fourth information.
[0296] In one example, the fourth information indicates one or more reference models, and the terminal device can take part or all of the one or more reference models as the first model according to the fourth information.
[0297] In another example, the fourth information indicates one or more reference models and one or more sets of model parameters, and the terminal device can take part or all of the one or more reference models as the initial model, and then perform parameter optimization based on the one or more sets of model parameters indicated by the fourth information to obtain the first model.
[0298] The above examples of the terminal device performing model development and / or model deployment according to the fourth information are only for understanding and should not constitute any limitation on the present application. The present application includes but is not limited to the above examples.
[0299] It should be understood that step 1204a can be performed after step 1202 or after step 1203, and the present application does not limit this.
[0300] At step 1205, the terminal device determines whether the time length for performing the inference task of the first model meets the calculation time length requirement.
[0301] The detailed description of step 1205 can refer to the description of the terminal side determining whether the time length for performing the inference task of the first model meets the calculation time length in the above method 800, and will not be repeated here.
[0302] It should be understood that step 1205 can be performed after step 1203 or after step 1204a, and the present application does not limit this.
[0303] At step 1206, the terminal device performs the inference task of the first model when the time length for performing the inference task of the first model meets the calculation time length requirement.
[0304] At step 1207, the terminal device sends the first information to the network device, and the first information is used to indicate the first calculation time length.
[0305] In step 1208, the terminal device sends second information in a case where the time length for performing the inference task of the first model does not meet the calculation time length requirement, the second information being used for one or more of the following: indicating that the time length for performing the inference task of the first model does not meet the calculation time length requirement, requesting to close the inference task of the first model, requesting to change the calculation time length requirement, or requesting to change the data set.
[0306] The details of steps 1206 to 1208 can be referred to the related description of steps 820 to 840 in method 800 above, and will not be repeated here.
[0307] It should be understood that steps 1206 and 1207 and step 1208 can be executed alternatively according to the result determined in step 1205, and do not necessarily have to be executed all.
[0308] FIG. 13 describes a communication method provided by the present application, taking the docking of a model between a terminal device and a network device as an example. The technical solution shown in FIG. 13 corresponds to the technical solution shown in FIG. 10 above, and thus the beneficial effects achieved are similar, and will not be repeated here.
[0309] It should also be understood that in the embodiments shown in FIGS. 12 and 13 above, the terminal side can be a terminal device, and the network side can also be a network device. In another design, the terminal side can also include a terminal device and a host or a cloud server of an OTT system, and the network side can include a network device and an intelligent network element. In this case, there can also be interactions between the devices on the terminal side, and there can also be interactions between the devices on the network side. The specific implementation process of the method provided by the present application will be shown below from the perspective of the interactions between a host or a cloud server of an OTT system (hereinafter referred to as an OTT system), a terminal device, a network device, and an intelligent network element, through method 1400 shown in FIG. 14.
[0310] FIG. 14 is still another schematic flowchart of a communication method provided by an embodiment of the present application. The method 1400 shown in FIG. 14 can include the following steps:
[0311] In step 1401, a terminal device sends capability information to a network device, the capability information indicating one or more of the following: a maximum number of transmission layers supported by the terminal device, a storage capability of an OTT system, or a calculation capability of the OTT system.
[0312] The details of step 1401 can be referred to the related description of step 850 in method 800 above, and will not be repeated here.
[0313] In step 1402a, an intelligent network element sends a data set to a network device, the data set being used for model development and model deployment. Accordingly, the network device receives the data set from the intelligent network element.
[0314] At step 1403a, the network device sends the data set to the terminal device. Accordingly, the terminal device receives the data set from the network device.
[0315] At step 1404a, the terminal device sends the data set to the OTT system. Accordingly, the OTT system receives the data set from the terminal device.
[0316] At step 1405a, the OTT system performs model development and model deployment based on the data set.
[0317] Steps 1402a-1405a show a process of data set interfacing. The data set can be sent out by the intelligent network element through the network device, and the terminal device can forward the received data set to the OTT system. For details of the data set and the use of the data set for model development and model deployment, refer to the description of the above method 800 related to the design of A, and the description of the data set interfacing in the terminology introduction, which will not be repeated here.
[0318] At step 1402b, the intelligent network element sends one or more model files and / or one or more sets of model parameters to the network device. Accordingly, the network device receives the one or more model files and / or the one or more sets of model parameters.
[0319] At step 1403b, the network device sends fourth information to the terminal device, where the fourth information includes the one or more model files and / or the one or more sets of model parameters. Accordingly, the terminal device receives the fourth information.
[0320] At step 1404b, the terminal device sends the one or more model files and / or the one or more sets of model parameters to the OTT system. Accordingly, the OTT system receives the one or more model files and / or the one or more sets of model parameters.
[0321] At step 1405b, the OTT system performs model development and / or model deployment based on the one or more model files and / or the one or more sets of model parameters.
[0322] Steps 1402b-1405b show a process of model interfacing. The intelligent network element can send one or more model files and / or one or more sets of model parameters to the network device, and the terminal device can forward the received one or more model files and / or one or more sets of model parameters to the OTT system. The information of the intelligent network element sending the one or more model files and / or one or more sets of model parameters to the network device can be the fourth information, or other information, and the information of the terminal device sending the one or more model files and / or one or more sets of model parameters to the OTT system can be the fourth information, or other information, which is not limited in the present application.
[0323] The detailed description of the model file, the model parameter, and the model file and / or the model parameter used for model development and / or model deployment can refer to the description of design one and design two in the method 800 and the description of the data set docking in the term introduction, and will not be repeated.
[0324] At step 1406, the network device sends third information to the terminal device, where the third information is used to indicate the calculation time requirement. Accordingly, the terminal device receives the third information from the network device.
[0325] At step 1407, the terminal device determines whether the time length of performing the inference task of the first model meets the calculation time requirement.
[0326] The detailed description of steps 1406 and 1407 can refer to the description of steps 810a and the determination of whether the time length of performing the inference task of the first model meets the calculation time in the method 800, and will not be repeated.
[0327] At step 1408, the terminal device instructs the OTT system to perform the inference task of the first model in the case that the time length of performing the inference task of the first model meets the calculation time requirement.
[0328] At step 1409, the OTT system performs the inference task of the first model.
[0329] At step 1410, the OTT system sends information indicating the first calculation time to the terminal device. Accordingly, the terminal device receives the information indicating the first calculation time from the OTT system.
[0330] At step 1411, the terminal device sends first information to the network device, where the first information is used to indicate the first calculation time. Accordingly, the network device receives the first information from the terminal device.
[0331] Steps 1408 to 1411 show the operation of the terminal side in the case that the time length of performing the inference task of the first model meets the calculation time requirement. In step 1410, the information indicating the first calculation time sent by the OTT system to the terminal can be the first information or other information, which is not limited in the present application.
[0332] At step 1412a, the terminal device sends second information to the network device in the case that the time length of performing the inference task of the first model does not meet the calculation time requirement, where the second information is used to request to replace the data set.
[0333] At step 1413a, the network device requests the intelligent network element to replace the data set.
[0334] Steps 1412a and 1413a correspond to steps 1402a to 1405a above. The terminal device can request a replacement of the data set through the second information. Of course, the terminal device can also indicate through the second information that the duration of performing the inference task of the first model does not meet the calculation duration requirement, or request the network device to replace the calculation duration requirement, or request the network device to close the inference task of the first model, etc., without limitation.
[0335] Step 1412b, the terminal device sends second information to the network device in the case that the duration of performing the inference task of the first model does not meet the calculation duration requirement, the second information being used to request a replacement of the model or the model parameter.
[0336] Step 1413b, the network device replaces the model or the model parameter to the intelligent network element.
[0337] Steps 1412b and 1413b correspond to steps 1402b to 1405b above. The terminal device can request a replacement of the model or the model parameter through the second information. Of course, the terminal device can also indicate through the second information that the duration of performing the inference task of the first model does not meet the calculation duration requirement, or request the network device to replace the calculation duration requirement, or request the network device to close the inference task of the first model, etc., without limitation.
[0338] For detailed description of the second information, please refer to the relevant description of step 840 of method 800 above, which will not be repeated here.
[0339] The above-mentioned flow is exemplified by the interaction between the terminal device and the network device, the interaction between the network device and the intelligent network element, and the interaction between the terminal device and the OTT system. In specific implementation, the interaction mode between devices is not limited thereto. For example, the OTT system can also communicate directly with the intelligent network element, such as wired communication, without transferring data through the air interface between the network device and the terminal device. In this case, the process of sending the data set from the network side to the terminal side shown in steps 1402a to 1404a above can be replaced by step 1414a shown in dashed line in the figure, in which the intelligent network element sends the data set to the OTT system; the process of sending one or more model files and / or one or more sets of model parameters from the network side to the terminal side shown in steps 1402b to 1404b can be replaced by step 1414b shown in dashed line in the figure, in which the intelligent network element sends fourth information to the OTT system, the fourth information including one or more model files and / or one or more sets of model parameters. FIG. 14 shows the communication method provided by the present application in the case of data set docking and model docking, taking the interaction between the OTT system, the terminal device, the network device and the intelligent network element as an example. The technical solution shown in FIG. 14 corresponds to the technical solution shown in FIG. 10 above, and therefore the beneficial effects obtained are similar, which will not be repeated here.
[0340] FIG. 14 shows interactions between an OTT system (e.g., a host or cloud server of the OTT system), a terminal device, a network device, and a smart network element, but should not be construed as limiting the present application. For example, the terminal device and the OTT system in FIG. 14 can be replaced by a terminal device, in which case the interaction between the terminal device and the OTT system is internal interaction of the device. Or, the network device and the smart network element in FIG. 14 can be replaced by a network device, in which case the interaction between the network device and the smart network element is internal interaction of the device. For brevity, no further description will be given.
[0341] It should be understood that, in the various embodiments shown in the above in conjunction with the multiple drawings, the magnitude of the serial numbers of the steps does not mean the order of execution, and the order of execution of the steps should be determined according to their functions and inherent logic, and should not be construed as limiting the implementation process of the embodiments of the present application.
[0342] In the various embodiments of the present application, the terms and / or descriptions of different embodiments are consistent and can be mutually referred to if there is no special description and no logical conflict, and the technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationship.
[0343] In some of the above embodiments, the devices in the current network architecture are mainly taken as examples for illustrative description (such as terminal devices, network devices, hosts or cloud servers of OTT systems, smart network elements, etc.), and the specific form of the devices is not limited in the embodiments of the present application. For example, devices that can realize the same functions in the future can also be applicable to the methods provided by the embodiments of the present application.
[0344] It can be understood that, in the various method embodiments described above, the methods and operations implemented by the devices (such as terminal devices, network devices, hosts or cloud servers of OTT systems, smart network elements) can also be implemented by components (such as chips or circuits) of the devices.
[0345] The above provides a detailed description of the method provided by the embodiments of the present application in conjunction with the multiple drawings. The following describes the apparatus provided by the embodiments of the present application in conjunction with the drawings.
[0346] FIGS. 15 and 16 are schematic block diagrams of possible apparatuses provided by the embodiments of the present application. These apparatuses can be used to implement the functions of the terminal side or the network side in the above method embodiments, and thus can also achieve the beneficial effects possessed by the above method embodiments.
[0347] FIG. 15 is a schematic block diagram of an apparatus provided by the embodiments of the present application. The apparatus 1500 shown in FIG. 15 can include a processing module 1510 and a communication module 1520.
[0348] In one possible design, the apparatus 1500 can be used to implement the communication method performed by the terminal side in any of the embodiments of FIG. 10, FIG. 12, and FIG. 14. For example, the processing module 1510 can be used to implement the steps related to processing performed by the terminal side in each of the method embodiments, such as obtaining a computation duration requirement, determining whether a computation duration of performing an inference task of a first model meets the computation duration requirement, performing the inference task of the first model, and the like. The communication module 1520 can be used to implement the steps performed by the terminal side in each of the method embodiments, such as one or more of transmitting the first information, transmitting the second information, receiving the third information, or receiving the fourth information.
[0349] For example, the processing module 1510 can be used to obtain a computation duration requirement, the computation duration requirement being a range that a duration of performing an inference task of a first model should meet.
[0350] Optionally, the communication module 1520 can also be used to receive third information, the third information being used to indicate the computation duration requirement.
[0351] Optionally, the processing module 1510 can also be used to perform the inference task of the first model.
[0352] Optionally, the communication module 1520 can also be used to transmit first information, the first information being used to indicate a first computation duration.
[0353] Optionally, the communication module 1520 can also be used to transmit second information, the second information being used to one or more of the following: indicate that a duration of performing the inference task of the first model does not meet the computation duration requirement; request to close the inference task of the first model; request to replace the first model; request to replace a model parameter, the model parameter being used for model development, the model development including one or more of the following: model training, model adaptation, or model enhancement; request to replace a data set, the data set being used for model training; or request to replace the computation duration requirement.
[0354] Optionally, the communication module 1520 can also be used to receive a data set, the data set being used for model training.
[0355] Optionally, the communication module 1520 can also be used to receive fourth information, the fourth information including one or more model files and / or one or more sets of model parameters. Alternatively, the fourth information can be used to indicate one or more models.
[0356] More detailed descriptions of the processing module 1510 and the communication module 1520 can be directly obtained by referring to the related descriptions in the method embodiments of FIG. 10, FIG. 12, and FIG. 14, which are not repeated here.
[0357] In another possible design, the apparatus 1500 can be configured to implement the communication method performed by the network side in any of the embodiments of FIG. 10, FIG. 12, and FIG. 14. For example, the processing module 1510 can be configured to implement the processing-related steps performed by the network side in each of the method embodiments, such as generating the data set, generating one or more model files and / or one or more sets of model parameters, etc. The communication module 1520 can be configured to implement the steps performed by the network side in each of the method embodiments, such as one or more of receiving the first information, receiving the second information, sending the third information, or sending the fourth information.
[0358] For example, the communication module 1520 can also be configured to send the third information, where the third information is configured to indicate the computation duration requirement, which is a range that a duration of performing an inference task of a first model should satisfy.
[0359] Optionally, the communication module 1520 can also be configured to receive the first information, where the first information is configured to indicate a first computation duration.
[0360] Optionally, the communication module 1520 can also be configured to receive the second information, where the second information is configured to indicate one or more of the following: that a duration of performing the inference task of the first model does not satisfy the computation duration requirement; that the inference task of the first model is requested to be closed; that the first model is requested to be replaced; that model parameters used for model development are requested to be replaced, where the model development includes one or more of the following: model training, model adaptation, or model enhancement; that a data set used for model training is requested to be replaced; or that the computation duration requirement is requested to be replaced.
[0361] Optionally, the communication module 1520 can also be configured to send a data set, where the data set is configured to be used for model training.
[0362] Optionally, the communication module 1520 can also be configured to send the fourth information, where the fourth information includes one or more model files and / or one or more sets of model parameters. Alternatively, the fourth information is configured to indicate one or more models.
[0363] For more detailed description of the processing module 1510 and the communication module 1520, please refer to the related description in the method embodiments of FIG. 10, FIG. 12, and FIG. 14, which will not be repeated here.
[0364] It should be noted that the communication module can also be referred to as a transceiver module, a transceiver unit, a transceiver, a transceiver, or a transceiver device, etc. The processing module can also be referred to as a processor, a processing board, a processing unit, or a processing device, etc. Optionally, the communication module is used to perform the sending operation and the receiving operation of the first communication device or the second communication device in the above method, the device in the communication module for realizing the receiving function can be regarded as a receiving module, and the device in the communication module for realizing the sending function can be regarded as a sending module, that is, the communication module can include a receiving module and a sending module.
[0365] It should be further noted that, in a possible design, the foregoing processing module and / or communication module can be implemented through a virtual module, for example, the processing module can be implemented through a software function unit or a virtual device, and the communication module can be implemented through a software function or a virtual device. In another possible design, the processing module or the communication module can also be implemented through an entity device, for example, if the device is implemented by using a chip / chip circuit, the communication module can be an input / output circuit and / or a communication interface, and performs an input operation (corresponding to the foregoing receiving operation) and an output operation (corresponding to the foregoing sending operation); and the processing module can be an integrated processor or a microprocessor or an integrated circuit.
[0366] The division of the modules in the embodiments of the present application is illustrative, and is merely a logical function division. In actual implementation, another division manner can be used. In addition, each function module in each example in the embodiments of the present application can be integrated in one processor, or can be a separate physical entity, or two or more modules can be integrated in one module. The integrated module can be implemented in the form of hardware or in the form of a software function module.
[0367] FIG. 16 is a structural schematic diagram of a communication device provided by another embodiment of the present application. As shown in FIG. 16, the device 1600 includes processing circuitry 1610 and communication circuitry 1620. The processing circuitry 1610 and the communication circuitry 1620 are coupled to each other.
[0368] It can be understood that the processing circuitry 1610 can be one or more processors, or can be all or part of a circuit having a processing function in the one or more processors.
[0369] It can be understood that the communication circuitry 1620 can be a transceiver or an input / output interface.
[0370] Optionally, the device 1600 can further include a memory 1630, used to store instructions executed by the processing circuitry 1610 or to store input data required by the processing circuitry 1610 for running instructions or to store data generated after the processing circuitry 1610 runs instructions.
[0371] It is to be understood that the memory 1630 can be located either external to or internal to the processing circuitry 1610.
[0372] As an example, the processing circuitry 1610 is configured to perform the functions of the above-mentioned processing module 1510, and the communication circuitry 1620 is configured to perform the functions of the above-mentioned communication module 1520.
[0373] As an example, the apparatus 1600 can be a communication device, or a chip applied to a communication device.
[0374] When the apparatus 1600 is a communication device, the communication circuitry can be a transceiver; when the apparatus 1600 is a chip, the communication circuitry can be an input / output circuit, a bus, a pin, or other types of communication interfaces, wherein the input circuit in the input / output circuit can be used for receiving, and the output interface can be used for transmitting.
[0375] The application also provides a computer program product, which, when running on a processor, can implement the communication method performed by the terminal side or the communication method performed by the network side in the above method embodiments.
[0376] The application also provides a computer readable storage medium, which contains computer instructions, which, when running on a processor, can implement the communication method performed by the terminal side or the communication method performed by the network side in the above method embodiments.
[0377] The application also provides a communication system, which includes the terminal side and the network side described above, the terminal side can be used to implement the communication method performed by the terminal side in the above method embodiments, and the network side can be used to implement the communication method performed by the network side in the above method embodiments.
[0378] It is to be understood that the processor in the embodiments of the application can be a device or all or part of the circuit for processing functions in the device: a central processing unit (CPU), which can also be other general-purpose processors, digital signal processors (DSPs), field programmable gate arrays (FPGAs) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. The general-purpose processor can be a microprocessor or any conventional processor.
[0379] The terms "unit", "module" and the like used in the specification can be used to represent computer-related entities, hardware, combinations of hardware and software, software, or software in execution.
[0380] Those skilled in the art can clearly understand that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0381] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working processes of the above-described system, device and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be repeated here.
[0382] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other ways. For example, the above-described device embodiments are only schematic, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.
[0383] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. According to actual needs, some or all of the units can be selected to achieve the purpose of the embodiment.
[0384] In addition, each functional unit in each embodiment of the present application can be integrated into a processing unit, or each unit can exist physically independently, or two or more units can be integrated into one unit.
[0385] In the above embodiments, the functions of the various functional units can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented by software, the software can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions (programs). When the computer program instructions (programs) are loaded and executed on a computer, the whole or part of the processes or functions described in the embodiments of the present application are generated. The computer can be a general purpose computer, a special purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer readable storage medium or transferred from one computer readable storage medium to another computer readable storage medium, for example, the computer instructions can be transferred from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media. The available media can be magnetic media (such as floppy disk, hard disk, magnetic tape), optical media (such as digital video disc (DVD)), or semiconductor media (such as solid state disk (SSD)) and the like.
[0386] The functions, if implemented in the form of software functional units and sold or used as independent products, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part of the prior art or the part of the technical solutions of the present application can be embodied in the form of software product, which is stored in a storage medium and includes a plurality of instructions for making a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the method described in the embodiments of the present application. The foregoing storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk and various media that can store program codes.
[0387] The above is only a specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A communication method characterized by comprising: The method comprises: obtaining a computation duration requirement, the computation duration requirement being a range that a duration for performing an inference task of a first model should satisfy.
2. The method of claim 1, wherein, The method further comprises: in a case where the duration for performing the inference task of the first model satisfies the computation duration requirement, performing the inference task of the first model.
3. The method of claim 2, wherein, The method further comprises: sending first information, the first information being used to indicate a first computation duration, the first computation duration being a duration that is actually required for performing the inference task of the first model.
4. The method of claim 3, wherein, The inference task is inference for compressing channel state information (CSI) of R transmission layers, R being a maximum number of transmission layers supported by a terminal device, R being a positive integer. The first information is further used to indicate R second computation durations, an rth second computation duration in the R second computation durations being a duration that is actually required for compressing CSI of the rth transmission layer.
5. The method of claim 1, wherein, The method further comprises: in a case where the duration for performing the inference task of the first model does not satisfy the computation duration requirement, sending second information, the second information being used to perform one or more of the following: indicating that the duration for performing the inference task of the first model does not satisfy the computation duration requirement; requesting to close the inference task of the first model; requesting to replace the first model; requesting to replace a model parameter, the model parameter being used for model development, the model development comprising one or more of the following: model training, model adaptation, or model enhancement; requesting to replace a data set, the data set being used for model training; or requesting to replace the computation duration requirement.
6. The method of any one of claims 1 to 5, wherein, The obtaining a computation duration requirement comprises: receiving third information, the third information being used to indicate the computation duration requirement.
7. The method of any one of claims 1 to 6, wherein, The inference task is inference of compressing CSI of multiple transmission layers, and the calculation time length requirement includes R max ranges corresponding to R max transmission layers, an i-th range corresponding to an i-th transmission layer in the R max transmission layers is a range that needs to be met by an inference time length of compressing CSI of the i-th transmission layer, R max is a maximum value of a number of transmission layers supported by a terminal device, and R max is a positive integer.
8. The method of claim 6 or 7, wherein, Before the receiving third information, the method further comprises: sending capability information, the capability information being used to indicate one or more of the following: a maximum number of transmission layers supported by a terminal device, a storage capability on a terminal side, or a computation capability on the terminal side.
9. The method of any one of claims 1 to 8, wherein, The first model is obtained through model development, the model development comprising one or more of the following: model training, model adaptation, or model enhancement.
10. The method of claim 9, wherein, The first model is obtained through model development, comprising: the first model is obtained through model training according to a data set.
11. The method of claim 10, wherein, The method further comprises: receiving the data set, the data set being used for the model training.
12. The method of claim 9, wherein, The first model is obtained through model development, comprising:
13. The method of claim 12, wherein, the first model is obtained through the model development on a received second model. The method further comprises:
14. The method of any one of claims 1 to 8, wherein, receiving fourth information, the fourth information indicating one or more models, the one or more models comprising the second model.
15. The method of claim 14, wherein, The first model is received. The method further comprises: receiving fourth information, the fourth information indicating one or more models, the one or more models comprising the first model.
16. The method of any one of claims 1 to 15, wherein, The inference task is inference for compressing CSI, a start of a time length of the inference task is one of a configuration time of a reference signal resource, a transmission time of a reference signal, or a configuration time of a CSI report, and an end of the time length of the inference task is a time of reporting the CSI report; wherein the reference signal is used for inference for compressing the CSI, and the reference signal resource is used for transmitting the reference signal.
17. A method of communication, comprising: Comprise: determining third information, the third information being used for indicating a calculation time length requirement, the calculation time length requirement being a range that needs to be met by a time length of performing an inference task of a first model on a terminal side; sending the third information to the terminal side.
18. The method of claim 17, wherein, The inference task is inference of compression of channel state information (CSI) of multiple transmission layers, the calculation time length requirement includes R max ranges corresponding to R max transmission layers, an i-th range corresponding to an i-th transmission layer in the R max transmission layers is a range that needs to be met by an inference time length of compression of the CSI of the i-th transmission layer, R max is a maximum value of a number of transmission layers supported by the terminal device, and R max is a positive integer.
19. The method of claim 17 or 18, wherein, Before the determining third information, the method further comprises: receiving capability information, the capability information being used for indicating one or more of a maximum transmission layer number supported by a terminal device, a storage capability of the terminal side, or a calculation capability of the terminal side.
20. The method of any one of claims 17 to 19, wherein, The method further comprises: receiving first information, the first information being used for indicating a first calculation time length, the first calculation time length being an actual time length required by the terminal side for performing the inference task of the first model.
21. The method of claim 20, wherein, The inference task is inference for compressing CSI; the first information is used for indicating R second calculation time lengths, an rth second calculation time length in the R second calculation time lengths being an actual time length required for compressing CSI of the rth transmission layer, R being a maximum transmission layer number supported by a terminal device of the terminal side, and R being a positive integer.
22. The method of any one of claims 17 to 21, wherein, The method further comprises: receiving second information, the second information being used for one or more of: indicating that a time length of performing the inference task of the first model on the terminal side does not meet the calculation time length requirement; requesting to close the inference task of the first model; requesting to replace the first model; requesting to replace a model parameter, the model parameter being used for model development, the model development comprising one or more of model training, model adaptation, or model enhancement; requesting to replace a data set, the data set being used for model training; or requesting to replace the calculation time length requirement.
23. The method of any one of claims 17 to 22, wherein, The method further comprises: sending a data set, the data set being used for the model training, the model training being used to obtain the first model.
24. The method of any one of claims 17 to 22, wherein, The method further comprises: sending fourth information, the fourth information indicating one or more models; the one or more models comprising the first model, or the one or more models comprising a second model, the first model being obtained based on model development on the second model, the model development comprising one or more of model training, model adaptation, or model enhancement.
25. The method of any one of claims 17 to 24, wherein, The inference task is inference for compressing CSI, a start of a time length of the inference task is one of a configuration time of a reference signal resource, a transmission time of a reference signal, or a configuration time of a CSI report, and an end of the time length of the inference task is a time of reporting the CSI report; wherein the reference signal is used for inference for compressing the CSI, and the reference signal resource is used for transmitting the reference signal.
26. A communications device, characterized by comprising functional means for implementing the method according to any one of claims 1 to 25.
27. A communications device, characterized by comprising one or more processors and communication circuitry for at least one of input or output of signals by the communication device; the one or more processors configured to implement the method according to any one of claims 1 to 25.
28. A computer-readable storage medium, characterized in that, The computer program, which when executed by a processor, causes the method according to any one of claims 1 to 25 to be performed.
29. A computer program product, characterised in that, The computer program, which when executed by a processor, causes the method according to any one of claims 1 to 25 to be performed.
Citation Information
Patent Citations
Information processing device, information processing method, and program
CN115989481A
Information transmission method and apparatus, and communication device
CN116208493A
Wireless communication method, terminal device and network device
CN118556412A
Methods and systems for artificial intelligence based architecture in wireless network
US20230319585A1
Techniques for balancing dynamic inferencing by machine learning models
US20240231928A1