Communication method and related device
By exchanging intermediate model information between communication devices, the problems of AI model processing interruption and large storage space consumption are solved, achieving more efficient model processing and storage optimization.
Patent Information
- Application Number
- CN202410522823.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-28
- Publication Date
- 2025-10-28
AI Technical Summary
During the long AI model processing process, communication equipment may experience processing interruptions, resulting in reduced model processing efficiency and large storage space consumption.
The processed intermediate model information is backed up to the second communication device through the first communication device. The backup of the intermediate model is achieved by using information interaction, which avoids repeated processing and saves storage space.
It improves the processing efficiency of AI models, reduces storage space consumption, and enhances the flexibility and efficiency of model processing.
Smart Images

Figure CN120856576A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communications, and more particularly to a communication method and related apparatus. Background Technology
[0002] With the development of communication technology, communication equipment in communication systems can now perform not only traditional communication services but also other new types of services, such as artificial intelligence (AI) services. Generally, a communication system capable of handling AI services can also be called an AI system.
[0003] Currently, communication devices can serve as participating nodes in AI systems, providing their own computing power and data. For example, a communication device can perform multiple model processing steps (such as model training, model updating, or model fine-tuning) on an AI model based on local data to obtain another AI model. Furthermore, increasing the number of parameters in an AI model can effectively improve its performance. Consequently, the number of parameters in the AI model processed by the communication device may gradually increase, which could make the aforementioned model processing process potentially take a considerable amount of time (e.g., days, weeks, or even months).
[0004] However, during the aforementioned lengthy model processing, the communication equipment may experience processing interruptions after processing a portion of the data. Following these interruptions, the same portion of data needs to be reprocessed, leading to reduced model processing efficiency. Therefore, improving model processing efficiency is a pressing technical problem that needs to be solved. Summary of the Invention
[0005] This application provides a communication method and related apparatus for improving the processing efficiency of AI models, saving storage space consumption of AI processing nodes, and reducing storage requirements for AI processing nodes.
[0006] The first aspect of this application provides a communication method executed by a first communication device. The first communication device can be a communication equipment (such as a terminal device or network device), or it can be a component of a communication equipment (such as a processor, chip, or chip system), or it can be a logic module or software capable of implementing all or part of the functions of the communication equipment. In this method, the first communication device determines first information, which includes N model information pieces. These N model information pieces are used to indicate information about N intermediate models obtained after a first AI model undergoes first local data processing, where N is a positive integer. The N intermediate model information pieces are used to determine a second AI model. The first communication device sends the first information to a second communication device, which backs up the first information.
[0007] Based on the above scheme, the first communication device can process the first AI model based on the first local data to obtain information on N intermediate models. Subsequently, the first communication device can send first information containing the information of the N models to a second communication device used for backup. These N model information items respectively indicate the information of the N models. In other words, after processing the first AI model based on the first local data to obtain intermediate models, the first communication device can back up the intermediate models through information exchange. Therefore, in the event of an interruption in AI model processing, the second communication device can provide the backed-up intermediate model information. Since the intermediate models have already learned the knowledge or features corresponding to the first local data, the first communication device does not need to repeat the model processing process based on the already processed first wireless data, thus improving the processing efficiency of the AI model.
[0008] Furthermore, in the above scheme, the first communication device serves as the AI processing node for processing the AI model. By backing up the information of the intermediate model through the second communication device, the AI processing node does not need to store the information of the intermediate model locally, which can save the storage space consumption of the AI processing node and reduce the storage requirements of the AI processing node.
[0009] In this application, the terms AI model, neural network model, AI neural network model, machine learning model, and AI processing model can be used interchangeably.
[0010] It should be understood that the first communication device can process the first AI model to obtain N intermediate models, and the processing may include one or more of the following: model training, updating, iteration, and optimization.
[0011] Furthermore, the N intermediate models obtained by the first communication device can be AI models obtained through N different time information. For example, the first communication device processes the first AI model through the first time information in the N time information to obtain the first AI model. Then, the first communication device processes the first intermediate model to obtain the second intermediate model, and so on. After obtaining the i-th intermediate model (i takes values from 2 to N-1), the first communication device processes the i-th intermediate model to obtain the (i+1)-th intermediate model.
[0012] Optionally, the first communication device can obtain multiple intermediate models based on the first AI model, and the first communication device can send the multiple intermediate models in various ways.
[0013] For example, the first communication device can send information about the multiple intermediate models based on a predefined time period. That is, the first communication device sends information about the intermediate models obtained within that time period at regular intervals. Accordingly, in the above scheme, the information about the N models included in the first information can be information about the intermediate models obtained within a certain time period.
[0014] For example, the first communication device can send the information of the multiple intermediate models based on a predefined number of times; that is, after obtaining a number of model information each time, the first communication device sends the information of a number of intermediate models. Accordingly, in the above scheme, the information of the N models included in the first information can be the information of the N intermediate models obtained in a certain instance.
[0015] Furthermore, the first AI model can be implemented in various ways, such as being a small model or a large model. In the case of a large model, due to the extremely large number of model parameters, the above solution allows the first communication device to significantly reduce its storage space consumption by backing up intermediate model information through a second communication device. This is beneficial for scenarios where the AI processing node of the large model is deployed on the edge (e.g., the first communication device is a terminal). Moreover, it makes extensive use of the rich data resources on the edge, while eliminating the need to upload edge data, effectively protecting privacy.
[0016] For example, a large model can refer to a machine learning model with a large number of parameters and a complex structure, capable of processing massive amounts of data and completing various complex tasks, such as natural language processing, computer vision, and speech recognition.
[0017] Alternatively, large models are typically built from deep neural networks and have billions or even hundreds of billions or more parameters.
[0018] Alternatively, the design purpose of large models may be to improve the expressive power and predictive performance of the models, enabling them to handle more complex tasks and data.
[0019] Alternatively, large models can learn complex patterns and features by training on massive amounts of data, resulting in stronger generalization capabilities and the ability to make accurate predictions on unprocessed data.
[0020] In contrast, a small model can refer to a model with fewer parameters and a shallower number of layers. Generally, compared to a small model, a large model usually has more parameters and a deeper number of layers, possessing stronger expressive power and higher accuracy, but also requiring more computing resources and time for training and inference. It is suitable for scenarios with large amounts of data and abundant computing resources, such as cloud computing, high-performance computing, and artificial intelligence. For example, small models have advantages such as being lightweight, efficient, and easy to deploy, making them suitable for scenarios with smaller amounts of data and limited computing resources, such as mobile applications, embedded devices, and the Internet of Things (IoT).
[0021] It should be understood that the above solution can be applied to various AI scenarios. The first AI model will be used as an example below.
[0022] Scenario 1: Federated Learning Scenario.
[0023] In this scenario, the third communication device can be a central node, and multiple first communication devices can be distributed nodes. Correspondingly, the first AI model can be a general large model. Multiple first communication devices can process the general large model based on local data (for example, the local data may include vertical domain indications) to obtain an intermediate model. Subsequently, the third communication device can perform one or more processing based on the intermediate model to obtain a global fine-tuning model. This fine-tuning model can be understood as fine-tuning the large model.
[0024] As an example, in Scenario 1, the third communication device can perform model processing once based on intermediate models sent by multiple first communication devices to obtain a globally fine-tuned model. That is, the third communication device can only perform model processing after receiving intermediate models from all the first communication devices. In this case, the intermediate models sent by different first communication devices are processed through the same procedure; the model processing of the third communication device can be understood as a synchronous federated learning (SFL) process.
[0025] As another example, in Scenario 1, the third communication device can perform model processing two or more times based on intermediate models sent by multiple first communication devices to obtain a globally fine-tuned model. That is, the third communication device can perform model processing after receiving an intermediate model processed by any of the first communication devices. In this case, the intermediate models obtained by different first communication devices are processed through different processes by the third communication device, and the model processing process of the third communication device can be understood as an asynchronous federated learning (AFL) process. Furthermore, in the implementation of AFL, since the third communication device does not need to wait for all the first communication devices to send intermediate models, it can effectively reduce processing latency and improve model processing efficiency.
[0026] Optionally, any first communication device can be connected to one or more subordinate nodes, and the first communication device can be the central node of the one or more subordinate nodes. The first communication device and the one or more subordinate nodes can perform model processing using either SFL or AFL; no limitation is made here.
[0027] Furthermore, when model processing between the first communication device and the one or more subordinate nodes can be performed using SFL (Small Federation), the relationship between the first communication device and the one or more subordinate nodes can be referred to as a small federation (or a cluster), and the relationship between the third communication device and multiple first communication devices can be referred to as a large federation. In this implementation, the aggregation model (cluster model) of each small federation can learn data from both strong and weak devices equally, enhancing the availability of weak device data within the small federation. Simultaneously, using AFL (Advanced Framework) between different clusters avoids excessively long training times and low training efficiency. Thus, it balances improving the utilization rate of weak device data with system training efficiency.
[0028] Scenario 2: The scenario where the first AI model fine-tunes the model itself.
[0029] In scenario two, the first communication device can perform one or more model processing steps on its own deployed first AI model to obtain a fine-tuned second AI model.
[0030] Optionally, the second communication device can be an entity / network element / device used to back up model information. For example, the second communication device can be a data storage function (DSF) entity / network element, or other names, which are not limited here.
[0031] Optionally, in addition to including the information of the N models, the first information may also include other information, such as one or more of the index / identifier of the N intermediate models, the indication information indicating the number of times the model is processed corresponding to the N intermediate models, and the indication information indicating the period corresponding to the N intermediate models, so that the second communication device can store the information of the N models based on the one or more of the information.
[0032] In one possible implementation of the first aspect, the method further includes: the first communication device sending second information to the second communication device, the second information being used to request first model information, the first model information being one of the N model information; and the first communication device receiving the first model information from the second communication device.
[0033] Based on the above scheme, the first communication device can send second information to the second communication device to request the first model information, thereby enabling the second communication device to send the first model information to the first communication device. In the event of a model processing interruption, the first communication device can obtain intermediate model information through the second communication device, allowing it to locally recover the intermediate model and improve the processing efficiency of the AI model.
[0034] Optionally, the N model information pieces correspond to N time information pieces, and the first model information piece is the latest model information piece corresponding to the time information of the N model information pieces; or, the second information piece includes the index of the first model information piece.
[0035] In one possible implementation of the first aspect, the method further includes: the first communication device sending third information to the second communication device, the third information being used to request the second communication device to send second model information to the third communication device, the second model information being one of the N model information.
[0036] Based on the above scheme, the first communication device can send a third message to the second communication device requesting the second communication device to send the second model information to the third communication device, thereby enabling the second communication device to send the second model information to the third communication device. In the event of an interruption in model processing, the first communication device can provide the intermediate model it processed to the third communication device through the second communication device. This allows the third communication device to obtain the information of the backup intermediate model and further process it, thereby improving the processing efficiency of the AI model.
[0037] Optionally, the N model information pieces correspond to N time information pieces respectively, and the second model information is the latest model information corresponding to the time information of the N model information pieces (i.e., the first model information and the second model information can be the same); or, the third information includes the index of the second model information.
[0038] In one possible implementation of the first aspect, the first communication device sends third information, including: when a first condition is met, the first communication device sends the third information, the first condition including any one of the following:
[0039] The model processing of the intermediate model corresponding to the first AI model (such as any intermediate model obtained by the first communication device, the current intermediate model, the most recently processed intermediate model, etc.) is interrupted;
[0040] Determine the current intermediate model of the first AI model to be nearing its expiration date (or about to expire, about to become invalid, etc.).
[0041] Based on the above scheme, when the first communication device determines that the model processing of the intermediate model corresponding to the first AI model is interrupted, the first communication device can request the second communication device to provide the intermediate model processed by the first communication device to the third communication device via a third information request. This allows the third communication device to perform model processing based on the information of the intermediate model sent by the second communication device. Therefore, the latency and model expiration issues caused by the first communication device's fine-tuning interruption and retraining can be resolved, improving the utilization rate of local data.
[0042] And / or, if the first communication device determines that the current intermediate model of the first AI model is nearing expiration, the first communication device can request the second communication device to provide the intermediate model processed by the first communication device to the third communication device via a third information request. This allows the third communication device to perform model processing based on the intermediate model information sent by the second communication device. Therefore, the potential problem of model expiration can be solved, effectively preventing model expiration and improving the utilization rate of the local data of the first communication device.
[0043] In one possible implementation of the first aspect, the first communication device determines the current intermediate model period of the first AI model by: the first communication device receiving indication information for indicating the global number of model processing times; and when the difference information between the local number of model processing times and the global number of model processing times satisfies a second condition, the first communication device determines the current intermediate model period of the first AI model.
[0044] Based on the above scheme, when the difference between the local model processing count and the global model processing count meets the second condition, the first communication device can determine that it is a device with weaker capabilities. That is, the processing speed of the first communication device in processing AI models is lower than that of other communication devices at the same level. Therefore, while other communication devices have uploaded multiple models and performed multiple global model updates, the local model update of the first communication device has not yet ended. To this end, the first communication device can determine the current intermediate model expiration of the first AI model and trigger the third communication device to obtain the intermediate model information of the first communication device through the third information, so as to make full use of the local data of each device to participate in the global model processing.
[0045] Optionally, the indication information used to indicate the number of global model processing times can be periodically sent information or information triggered based on a request from the first communication device; no limitation is made here.
[0046] In one possible implementation of the first aspect, the third information includes any of the following:
[0047] The first indication information is used to indicate that the third information is triggered by a model processing interruption based on the intermediate model corresponding to the first AI model;
[0048] The second instruction information is used to indicate that the third information is triggered based on the current intermediate model of the first AI model.
[0049] Based on the above scheme, the third information sent by the first communication device may include any of the above-mentioned indication information, so that the recipient of the third information can determine the reason why the first communication device sent the third information based on the third information.
[0050] In one possible implementation of the first aspect, the method further includes: the first communication device receiving model information of the second AI model; and the first communication device performing model processing on the second AI model based on second local data.
[0051] Based on the above scheme, after obtaining the global model (i.e., the second AI model), the third communication device can also send the model information of the second AI model to the first communication device, so that the first communication device can obtain the global model.
[0052] In one possible implementation of the first aspect, the method further includes: the first communication device sending M model information to the second communication device, the M model information being used to indicate information of M intermediate models obtained after the second AI model undergoes second local data processing, where M is a positive integer.
[0053] Based on the above scheme, after receiving the model information of the second AI model, the first communication device can also process the second AI model based on local data (this processing can be understood as model fine-tuning, secondary fine-tuning, etc.). Similarly, the first communication device can also back up the intermediate model obtained from the processing to the second communication device to improve the processing efficiency of the AI model and reduce the storage requirements of the first communication device.
[0054] Similarly, the information on the M models can be contained within a certain information, which may include other information besides the information on the M models. For example, one or more of the following: the index / identifier of the M intermediate models, the indication information indicating the number of model processing times corresponding to the M intermediate models, and the indication information indicating the period corresponding to the M intermediate models, so that the second communication device can store the information on the M models based on the one or more of this information.
[0055] In one possible implementation of the first aspect, the method further includes: the first communication device receiving indication information for indicating the address of the second communication device.
[0056] Based on the above scheme, the first communication device can obtain the address of the second communication device based on the instructions of other devices, so that the first communication device can back up the intermediate model obtained locally by the first communication device to the designated backup node (i.e., the second communication device).
[0057] Optionally, the indication information used to indicate the address of the second communication device may come from the DSF entity or from other entities / network elements / devices, such as data operation (DA) entities / network elements, data control (DC) entities / network elements, etc.
[0058] Optionally, the address of the second communication device includes at least one of the following fields:
[0059] The first field is used to indicate the geographical location of the second communication device;
[0060] The second field is used to indicate the identifier of the bearer transmitting the first information;
[0061] The third field is used to indicate the identifier of the first communication device.
[0062] A second aspect of this application provides a communication method executed by a second communication device. The second communication device can be a communication equipment, or it can be a component of a communication equipment (e.g., a processor, chip, or chip system), or it can be a logic module or software capable of implementing all or part of the functions of the communication equipment. In this method, the second communication device receives first information from a first communication device. This first information includes N model information pieces, each of which indicates information about N intermediate models obtained after the first AI model undergoes first local data processing, where N is a positive integer. The N intermediate model information pieces are used to determine a second AI model. The second communication device stores the first information.
[0063] Based on the above scheme, the second communication device can receive and store first information, which includes N model information pieces. These N model information pieces are used to indicate the information of N intermediate models obtained after the first AI model has undergone first local data processing. In other words, after processing the first AI model based on the first local data to obtain intermediate models, the first communication device can back up the intermediate models through information exchange. Therefore, in the event of an interruption in AI model processing, the second communication device can provide the backed-up intermediate model information. Since the intermediate models have already learned the knowledge or features corresponding to the first local data, the first communication device does not need to repeat the model processing process based on the already processed first wireless data, thus improving the processing efficiency of the AI model.
[0064] Furthermore, in the above scheme, the first communication device serves as the AI processing node for processing the AI model. By backing up the information of the intermediate model through the second communication device, the AI processing node does not need to store the information of the intermediate model locally, which can save the storage space consumption of the AI processing node and reduce the storage requirements of the AI processing node.
[0065] In one possible implementation of the second aspect, the method further includes: the second communication device receiving second information from the first communication device, the second information being used to request first model information, the first model information being one of the N model information; and the second communication device sending the first model information.
[0066] Based on the above scheme, the second communication device can receive second information from the first communication device requesting the first model information, enabling the second communication device to send the first model information to the first communication device. In the event of a model processing interruption, the first communication device can obtain intermediate model information through the second communication device, allowing it to locally recover the intermediate model and improve the processing efficiency of the AI model.
[0067] Optionally, the N model information pieces correspond to N time information pieces, and the first model information piece is the latest model information piece corresponding to the time information of the N model information pieces; or, the second information piece includes the index of the first model information piece.
[0068] In one possible implementation of the second aspect, the method further includes: the second communication device sending second model information to the third communication device, the second model information being one of the N model information.
[0069] Based on the above scheme, the second communication device can provide the intermediate model processed by the first communication device to the third communication device, so that the third communication device can obtain the information of the backup intermediate model and further process the intermediate model to improve the processing efficiency of the AI model.
[0070] Optionally, the second model information may be included in a certain information, which may include other information besides the second model information. For example, it may include one or more of the following: the index / identifier of the second intermediate model, the indication information of the number of model processing times corresponding to the second intermediate model, the indication information of the period corresponding to the second intermediate model, the identifier of the second communication device, the address of the second communication device, the address of the first communication device, and the identifier of the first communication device, so that the third communication device can process the second model based on the one or more of the information.
[0071] In one possible implementation of the second aspect, before the second communication device sends the second model information to the third communication device, the method further includes: the second communication device receiving third information from the first communication device, the third information being used to request the second communication device to send the first model information to the third communication device.
[0072] Based on the above scheme, the first communication device can send a third message to the second communication device requesting the second communication device to send the second model information to the third communication device, so that the second communication device sends the second model information to the third communication device based on the request.
[0073] In one possible implementation of the second aspect, the second communication device sends the second model information, including: the second communication device receiving indication information for indicating the global number of model processing times; and the second communication device sending the second model information when the difference information between the number of model processing times corresponding to the N model information and the global number of model processing times satisfies a second condition.
[0074] Based on the above scheme, when the second communication device determines that the intermediate model of the first AI model is nearing expiration, the second communication device can provide the intermediate model processed by the first communication device to the third communication device. This allows the third communication device to perform model processing based on the information of the intermediate model sent by the second communication device. Therefore, the potential problem of model expiration can be solved, effectively preventing model expiration and improving the utilization rate of the local data of the first communication device.
[0075] In one possible implementation of the second aspect, the method further includes: the second communication device sending instruction information to the first communication device to instruct the cessation of model processing of the current intermediate model corresponding to the first AI model.
[0076] Based on the above scheme, when the second communication device determines that the current intermediate model of the first AI model is nearing its end, the third communication device will perform model processing based on the intermediate model information sent by the second communication device. To this end, the second communication device can send an instruction message to the first communication device, so that the first communication device can stop the model processing of the current intermediate model corresponding to the first AI model based on the instruction message, so as to avoid unnecessary processing overhead.
[0077] A third aspect of this application provides a communication method executed by a third communication device. The third communication device can be a communication equipment, or it can be a component of a communication equipment (e.g., a processor, chip, or chip system), or it can be a logic module or software capable of implementing all or part of the functions of the communication equipment. In this method, the third communication device receives second model information, which indicates information about an intermediate model obtained after a first AI model undergoes local data processing by the first communication device. Based on the second model information, the third communication device processes the first AI model to obtain a first global model. The third communication device also receives third model information, which indicates information about an intermediate model obtained after the first AI model undergoes local data processing by a fourth communication device. Based on the third model information, the third communication device processes the first global model to obtain a second global model. The second global model is used to determine a second AI model.
[0078] Based on the above scheme, the third communication device can perform model processing two or more times based on the intermediate model obtained from multiple communication devices (such as the first communication device, the fourth communication device, etc.) to obtain a globally fine-tuned model. In this case, the intermediate model obtained from multiple communication devices can be processed through different processes of the third communication device, and the model processing process of the third communication device can be understood as an asynchronous federated learning (AFL) process. In this way, during the implementation of AFL, since the third communication device does not need to wait for all communication devices to send intermediate models, it can effectively reduce processing latency and improve model processing efficiency.
[0079] Optionally, during the process of the third communication device receiving the second model information, the second model information may come from the first communication device or from a DSF entity / network element used to back up the model information of the first communication device (such as the second communication device described above). Similarly, during the process of the third communication device receiving the third model information, the second model information may come from the fourth communication device or from a DSF entity / network element used to back up the model information of the fourth communication device.
[0080] A fourth aspect of this application provides a communication device, which is a first communication device, comprising a transceiver unit and a processing unit; the processing unit is used to determine first information, the first information including N model information, the N model information being used to indicate information of N intermediate models obtained after the first AI model undergoes first local data processing, where N is a positive integer; the N intermediate model information is used to determine a second AI model; the transceiver unit is used to send the first information to a second communication device, and the second communication device is used to back up the first information.
[0081] In the fourth aspect of this application, the constituent modules of the communication device can also be used to perform the steps executed in various possible implementations of the first aspect and achieve the corresponding technical effects. For details, please refer to the first aspect, which will not be repeated here.
[0082] The fifth aspect of this application provides a communication device, which is a first communication device, comprising a transceiver unit and a processing unit; the transceiver unit is used to receive first information from the first communication device, the first information including N model information, the N model information being used to indicate information of N intermediate models obtained after the first AI model undergoes first local data processing, where N is a positive integer; the N intermediate model information is used to determine a second AI model; the processing unit is used to store the first information.
[0083] In the fifth aspect of this application, the constituent modules of the communication device can also be used to perform the steps executed in various possible implementations of the second aspect and achieve the corresponding technical effects. For details, please refer to the second aspect, which will not be repeated here.
[0084] A sixth aspect of this application provides a communication device, which is a first communication device. The communication device includes a transceiver unit and a processing unit. The transceiver unit is used to receive second model information, which indicates information of an intermediate model obtained after a first AI model undergoes local data processing by the first communication device. The processing unit is used to process the first AI model based on the second model information to obtain a first global model. The transceiver unit is also used to receive third model information, which indicates information of an intermediate model obtained after the first AI model undergoes local data processing by a fourth communication device. The processing unit is also used to process the first global model based on the third model information to obtain a second global model. The second global model is used to determine a second AI model.
[0085] In the sixth aspect of this application, the constituent modules of the communication device can also be used to perform the steps executed in various possible implementations of the third aspect and achieve the corresponding technical effects. For details, please refer to the third aspect, which will not be repeated here.
[0086] A seventh aspect of this application provides a communication device including at least one processor coupled to a memory; the memory is used to store a program or instructions; the at least one processor is used to execute the program or instructions to enable the communication device to implement the method described in any possible implementation of any of the first to third aspects. Optionally, the communication device may include the memory.
[0087] The eighth aspect of this application provides a communication device including at least one logic circuit and an input / output interface; the logic circuit is used to perform the method as described in any one of the possible implementations of the first to third aspects described above.
[0088] The ninth aspect of this application provides a communication system that includes at least two of the first communication device, the second communication device, and the third communication device described above.
[0089] The tenth aspect of this application provides a computer-readable storage medium for storing one or more computer-executable instructions, which, when executed by a processor, perform the method as described in any possible implementation of any of the first to third aspects described above.
[0090] The eleventh aspect of this application provides a computer program product (or computer program) that, when executed by a processor, performs the method described in any possible implementation of any of the first to third aspects described above.
[0091] The twelfth aspect of this application provides a chip system including at least one processor for supporting a communication device in implementing the method described in any possible implementation of any of the first to third aspects.
[0092] In one possible design, the chip system may further include a memory for storing program instructions and data necessary for the communication device. The chip system may be composed of chips or may include chips and other discrete devices. Optionally, the chip system may also include interface circuitry that provides program instructions and / or data to the at least one processor.
[0093] The technical effects of any of the design methods in aspects four through twelfth can be found in the technical effects of the different design methods in aspects one through three above, and will not be repeated here. Attached Figure Description
[0094] Figures 1a to 1c A schematic diagram of the communication system provided in this application;
[0095] Figures 2a to 2g This is a schematic diagram of the AI processing involved in this application;
[0096] Figure 3 An interactive schematic diagram of the communication method provided in this application;
[0097] Figure 4 Another interactive schematic diagram of the communication method provided in this application;
[0098] Figures 5 to 9 A schematic diagram of the communication device provided in this application. Detailed Implementation
[0099] First, some terms used in the embodiments of this application will be explained to facilitate understanding by those skilled in the art.
[0100] (1) Terminal device: can be a wireless terminal device that can receive network device scheduling and instruction information. The wireless terminal device can be a device that provides voice and / or data connectivity to the user, or a handheld device with wireless connection function, or other processing device connected to a wireless modem.
[0101] Terminal devices can communicate with one or more core networks or the Internet via a radio access network (RAN). Terminal devices can be mobile terminal devices, such as mobile phones (or "cellular" phones), computers, and data cards. For example, they can be portable, pocket-sized, handheld, computer-embedded, or vehicle-mounted mobile devices that exchange voice and / or data with the RAN. Examples include personal communication service (PCS) phones, cordless phones, session initiation protocol (SIP) phones, wireless local loop (WLL) stations, personal digital assistants (PDAs), tablets, and computers with wireless transceiver capabilities. Wireless terminal equipment can also be referred to as a system, subscriber unit, subscriber station, mobile station (MS), remote station, access point (AP), remote terminal, access terminal, user terminal, user agent, subscriber station (SS), customer premises equipment (CPE), terminal, user equipment (UE), mobile terminal (MT), etc.
[0102] By way of example and not limitation, in this embodiment, the terminal device can also be a wearable device. Wearable devices, also known as wearable smart devices or smart wearable devices, are a general term for devices that utilize wearable technology to intelligently design and develop everyday wearables, such as glasses, gloves, watches, clothing, and shoes. Wearable devices are portable devices that are worn directly on the body or integrated into the user's clothing or accessories. Wearable devices are not merely hardware devices, but also achieve powerful functions through software support, data interaction, and cloud interaction. Broadly speaking, wearable smart devices include those that are feature-rich, large in size, and can achieve complete or partial functions without relying on a smartphone, such as smartwatches or smart glasses, as well as those that focus on a specific type of application function and require the use of other devices such as smartphones, such as various smart bracelets, smart helmets, and smart jewelry for vital sign monitoring.
[0103] Terminals can also be drones, robots, devices in device-to-device (D2D) communication, vehicles to everything (V2X) communication, virtual reality (VR) terminal devices, augmented reality (AR) terminal devices, wireless terminals in industrial control, wireless terminals in self-driving, wireless terminals in telemedicine or telehealth services, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, etc.
[0104] Furthermore, terminal devices can also be terminal devices in communication systems evolved from fifth-generation (5G) communication systems (such as 5G Advanced or sixth-generation (6G) communication systems), or terminal devices in future public land mobile networks (PLMNs). For example, 5G Advanced or 6G networks can further expand the form and function of 5G communication terminals; 6G terminals include, but are not limited to, vehicles, cellular network terminals (integrating satellite terminal functions), drones, and Internet of Things (IoT) devices.
[0105] In this embodiment, the terminal device can also obtain artificial intelligence (AI) services provided by the network device. Optionally, the terminal device can also have AI processing capabilities.
[0106] (2) Network equipment: This can be equipment within a wireless network. For example, network equipment can be a RAN node (or device) that connects terminal devices to the wireless network, and can also be called a base station. Currently, some examples of RAN equipment include: base station, evolved NodeB (eNodeB), gNB (gNodeB) in 5G communication systems, transmission reception point (TRP), evolved Node B (eNB), radio network controller (RNC), Node B (NB), home base station (e.g., home-evolved Node B, or home Node B, HNB), base band unit (BBU), or wireless fidelity (Wi-Fi) access point (AP), etc. In addition, in a network architecture, network equipment can include central unit (CU) nodes, distributed unit (DU) nodes, or RAN equipment including both CU and DU nodes.
[0107] Optionally, RAN nodes can also be macro base stations, micro base stations, indoor stations, relay nodes, donor nodes, or radio controllers in cloud radio access network (CRAN) scenarios. RAN nodes can also be servers, wearable devices, vehicles, or in-vehicle equipment. For example, the access network equipment in vehicle-to-everything (V2X) technology can be a roadside unit (RSU).
[0108] In another possible scenario, multiple RAN nodes collaborate to assist the terminal in achieving wireless access, with different RAN nodes each implementing some of the base station's functions. For example, RAN nodes can be CUs, DUs, CUs (control plane, CP), CUs (user plane, UP), or radio units (RUs). CUs and DUs can be set up separately or included in the same network element, such as a baseband unit (BBU). RUs can be included in radio equipment or radio units, such as remote radio units (RRUs), active antenna units (AAUs), radio heads (RHs), or remote radio heads (RRHs).
[0109] In different systems, CU (or CU-CP and CU-UP), DU, or RU may have different names, but those skilled in the art will understand their meaning. For example, in an open access network (open RAN, O-RAN, or ORAN) system, CU can also be called O-CU (open CU), DU can also be called O-DU, CU-CP can also be called O-CU-CP, CU-UP can also be called O-CU-UP, and RU can also be called O-RU. For ease of description, this application uses CU, CU-CP, CU-UP, DU, and RU as examples. Any of the units among CU (or CU-CP, CU-UP), DU, and RU in this application can be implemented through software modules, hardware modules, or a combination of software modules and hardware modules.
[0110] Communication between access network devices and terminal devices follows a specific protocol layer structure. This protocol layer may include a control plane protocol layer and a user plane protocol layer. The control plane protocol layer may include at least one of the following: radio resource control (RRC) layer, packet data convergence protocol (PDCP) layer, radio link control (RLC) layer, media access control (MAC) layer, or physical (PHY) layer, etc. The user plane protocol layer may include at least one of the following: service data adaptation protocol (SDAP) layer, PDCP layer, RLC layer, MAC layer, or physical layer, etc.
[0111] The correspondence between network elements and their achievable protocol layer functions in the ORAN system can be found in Table 1 below.
[0112] Table 1
[0113] ORAN network elements 3GPP protocol layer functions O-CU-CP RRC+PCDP - Control Plane (PDCP-C) O-CU-UP SDAP+PCDP - User Plane (PDCP-U) O-DU RLC+MAC+PHY-high O-RU PHY-low
[0114] Network devices can be other devices that provide wireless communication functions for terminal devices. The embodiments of this application do not limit the specific technology or form of the network device. For ease of description, the embodiments of this application are not limited.
[0115] Network equipment may also include core network equipment, such as the Mobility Management Entity (MME), Home Subscriber Server (HSS), Serving Gateway (S-GW), Policy and Charging Rules Function (PCRF), and Public Data Network Gateway (PDN gateway or P-GW) in 4th generation (4G) networks; and access and mobility management function (AMF), user plane function (UPF), or session management function (SMF) in 5G networks. Furthermore, this core network equipment may also include other core network equipment in 5G networks and next-generation networks of 5G networks.
[0116] In this embodiment of the application, the network device may also have network nodes with AI capabilities, which can provide AI services to terminals or other network devices. For example, it may be an AI node, computing node, RAN node with AI capabilities, or core network element with AI capabilities on the network side (access network or core network).
[0117] In this application embodiment, the device for implementing the function of the network device can be the network device itself, or it can be a device capable of supporting the network device in implementing the function, such as a chip system. This device can be disposed within the network device. In the technical solutions provided in this application embodiment, the example of a network device being used to implement the function of the network device is used to describe the technical solutions provided in this application embodiment.
[0118] (3) Configuration and Pre-configuration: In this application, both configuration and pre-configuration are used. Configuration refers to the network device / server sending configuration information or parameter values to the terminal via messages or signaling, so that the terminal can determine communication parameters or resources for transmission based on these values or information. Pre-configuration is similar to configuration; it can be parameter information or parameter values pre-negotiated between the network device / server and the terminal device, parameter information or parameter values specified by standard protocols for use by the base station / network device or terminal device, or parameter information or parameter values pre-stored in the base station / server or terminal device. This application does not limit this.
[0119] Furthermore, these values and parameters can be changed or updated.
[0120] (4) The terms "system" and "network" in the embodiments of this application can be used interchangeably. "Multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, or B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the related objects before and after are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, "at least one of A, B and C" includes A, B, C, AB, AC, BC or ABC. And, unless otherwise specified, the ordinal numbers such as "first" and "second" mentioned in the embodiments of this application are used to distinguish multiple objects and are not used to limit the order, sequence, priority or importance of multiple objects.
[0121] (5) In the embodiments of this application, "send" and "receive" indicate the direction of signal transmission. For example, "send information to XX" can be understood as the destination of the information being XX, which may include sending directly through the air interface or sending indirectly through the air interface by other units or modules. "Receive information from YY" can be understood as the source of the information being YY, which may include receiving directly from YY through the air interface or receiving indirectly from YY through the air interface by other units or modules. "Send" can also be understood as the "output" of the chip interface, and "receive" can also be understood as the "input" of the chip interface.
[0122] In other words, sending and receiving can occur between devices, such as between network devices and terminal devices, or within a device, such as between components, modules, chips, software modules, or hardware modules within the device via buses, wiring, or interfaces.
[0123] It is understandable that information may undergo necessary processing, such as encoding and modulation, between the source and destination, but the destination can understand the valid information from the source. Similar statements in this application can be interpreted in a similar way and will not be elaborated further.
[0124] (6) In the embodiments of this application, "instruction" may include direct instruction and indirect instruction, as well as explicit instruction and implicit instruction. The information indicated by a certain piece of information (as described below, the instruction information) is called the information to be instructed. In the specific implementation process, there are many ways to indicate the information to be instructed, such as, but not limited to, directly indicating the information to be instructed, such as the information to be instructed itself or its index. It can also indirectly indicate the information to be instructed by indicating other information, where there is an association between the other information and the information to be instructed; or it can only indicate a part of the information to be instructed, while the other parts of the information to be instructed are known or pre-agreed upon. For example, the instruction can be implemented by using a pre-agreed (e.g., protocol predefined) arrangement order of various information, thereby reducing the instruction overhead to a certain extent. This application does not limit the specific method of instruction. It is understood that for the sender of the instruction information, the instruction information can be used to indicate the information to be instructed, and for the receiver of the instruction information, the instruction information can be used to determine the information to be instructed.
[0125] In this application, unless otherwise specified, the same or similar parts between the various embodiments can be referred to each other. In the various embodiments of this application, and the various methods / designs / implementations within each embodiment, unless otherwise specified or logically conflicting, the terminology and / or descriptions between different embodiments and between the various methods / designs / implementations within each embodiment are consistent and can be mutually referenced. The technical features in different embodiments and the various methods / designs / implementations within each embodiment can be combined to form new embodiments, methods, or implementations based on their inherent logical relationships. The following descriptions of the embodiments of this application do not constitute a limitation on the scope of protection of this application.
[0126] This application can be applied to long-term evolution (LTE) systems, new radio (NR) systems, or communication systems evolving after 5G (such as 6G). The communication system includes at least one network device and / or at least one terminal device.
[0127] Please see Figure 1a This is a schematic diagram of a communication system in this application. Figure 1aThe example illustrates one network device and six terminal devices, namely terminal device 1, terminal device 2, terminal device 3, terminal device 4, terminal device 5, and terminal device 6. Figure 1a The example shown uses terminal device 1 as a smart teacup, terminal device 2 as a smart air conditioner, terminal device 3 as a smart fuel dispenser, terminal device 4 as a vehicle, terminal device 5 as a mobile phone, and terminal device 6 as a printer for illustration.
[0128] like Figure 1a As shown, the entity sending the AI configuration information can be a network device. The entity receiving the AI configuration information can be terminal devices 1-6. In this case, the network device and terminal devices 1-6 form a communication system. In this communication system, terminal devices 1-6 can send data to the network device, and the network device can receive the data sent by terminal devices 1-6. The network device can also send configuration information to terminal devices 1-6.
[0129] For example, in Figure 1a In this system, terminal devices 4 to 6 can also form a communication system. Terminal device 5 acts as a network device, i.e., the entity sending AI configuration information; terminal devices 4 and 6 act as terminal devices, i.e., the entities receiving AI configuration information. For example, in a vehicle-to-everything (V2X) system, terminal device 5 sends AI configuration information to terminal devices 4 and 6 respectively, and receives data sent by terminal devices 4 and 6; correspondingly, terminal devices 4 and 6 receive the AI configuration information sent by terminal device 5 and send data back to terminal device 5.
[0130] by Figure 1a Taking the communication system shown as an example, in addition to performing communication-related services, different devices (including network devices and terminal devices, and / or terminal devices and terminal devices) may also perform AI-related services.
[0131] like Figure 1b As shown, taking a network device as a base station as an example, a base station can perform communication-related services and AI-related services with one or more terminal devices, and different terminal devices can also perform communication-related services and AI-related services.
[0132] like Figure 1c As shown, taking terminal devices including TVs and mobile phones as an example, TVs and mobile phones can also perform communication-related services and AI-related services.
[0133] The technical solution provided in this application can be applied to wireless communication systems (e.g.) Figure 1a , Figure 1b or Figure 1c The system shown, for example, the communication system provided in this application, can incorporate AI network elements to implement some or all AI-related operations. AI network elements can also be called AI nodes, AI devices, AI entities, AI modules, AI models, or AI units, etc. The AI network element can be built into a network element within the communication system. For example, an AI network element can be an AI module built into: access network equipment, core network equipment, cloud server, or operation, administration, and maintenance (OAM) to implement AI-related functions. The OAM can act as the network management system for the core network equipment and / or the access network equipment. Alternatively, the AI network element can also be a network element independently set up in the communication system. Optionally, the terminal or its built-in chip can also include an AI entity to implement AI-related functions.
[0134] The following is a brief introduction to the concepts that may be involved in this application.
[0135] AI can endow machines with human-like intelligence, for example, allowing them to use computer hardware and software to simulate certain intelligent human behaviors. To achieve artificial intelligence, machine learning methods can be employed. In machine learning, machines learn (or train) a model using training data. This model represents the mapping between inputs and outputs. The learned model can be used for reasoning (or prediction), that is, it can be used to predict the output corresponding to a given input. This output can also be called the reasoning result (or prediction result).
[0136] Machine learning can include supervised learning, unsupervised learning, and reinforcement learning. Unsupervised learning can also be called learning without supervision.
[0137] Supervised learning, based on collected sample values and labels, uses machine learning algorithms to learn the mapping relationship between sample values and labels, and then expresses this learned mapping relationship using an AI model. The process of training the machine learning model is the process of learning this mapping relationship. During training, sample values are input into the model to obtain the model's predicted values, and the model parameters are optimized by calculating the error between the model's predicted values and the sample labels (ideal values). After the mapping relationship is learned, it can be used to predict new sample labels. The mapping relationship learned in supervised learning can include linear or non-linear mappings. Based on the type of label, the learning task can be divided into classification tasks and regression tasks.
[0138] Unsupervised learning relies on collected sample values to discover inherent patterns within the samples themselves. One type of unsupervised learning algorithm uses the samples themselves as supervisory signals, meaning the model learns the mapping relationship from sample to sample; this is called self-supervised learning. During training, model parameters are optimized by calculating the error between the model's predictions and the samples themselves. Self-supervised learning can be used for signal compression and decompression recovery applications; common algorithms include autoencoders and generative adversarial networks.
[0139] Reinforcement learning, unlike supervised learning, is a type of algorithm that learns problem-solving strategies through interaction with the environment. Unlike supervised and unsupervised learning, reinforcement learning problems do not have explicit "correct" action labels. The algorithm needs to interact with the environment to obtain reward signals from the environment, and then adjust its decision actions to obtain a larger reward signal value. For example, in downlink power control, the reinforcement learning model adjusts the downlink transmission power of each user based on the total system throughput feedback from the wireless network, aiming to achieve a higher system throughput. The goal of reinforcement learning is also to learn the mapping relationship between the environment state and a better (e.g., optimal) decision action. However, because the label of the "correct action" cannot be obtained in advance, the network cannot be optimized by calculating the error between the action and the "correct action." Reinforcement learning training is achieved through iterative interaction with the environment.
[0140] Neural networks (NNs) are a specific model in machine learning techniques. According to the general approximation theorem, neural networks can theoretically approximate any continuous function, thus enabling them to learn arbitrary mappings. Traditional communication systems rely on extensive expert knowledge to design communication modules, while deep learning communication systems based on neural networks can automatically discover hidden pattern structures from large datasets, establish mapping relationships between data, and achieve performance superior to traditional modeling methods.
[0141] The idea behind neural networks comes from the neuronal structure of the brain. For example, each neuron performs a weighted summation of its input values and outputs the result through an activation function.
[0142] like Figure 2a The diagram shown is a schematic representation of a neuron structure. Assume the neuron's input is x = [x0, x1, ..., x...]. n The weights corresponding to each input are w = [w0, w1, ..., w1]. n ], where n is a positive integer, w i and x i It can be any possible type, such as a decimal, an integer (e.g., 0, a positive integer, or a negative integer), or a complex number. i As x i The weights are used to assign weights to x.i Weighting is applied. The bias for the weighted summation of the input values is, for example, b. Activation functions can take many forms. Assuming a neuron's activation function is y = f(z) = max(0, z), then the neuron's output is: For example, if the activation function of a neuron is y = f(z) = z, then the output of that neuron is: Here, b can be any possible type, such as a decimal, an integer (e.g., 0, a positive integer, or a negative integer), or a complex number. The activation functions of different neurons in a neural network can be the same or different.
[0143] Furthermore, neural networks generally consist of multiple layers, each of which may include one or more neurons. Increasing the depth and / or width of a neural network can improve its expressive power, providing more powerful information extraction and abstract modeling capabilities for complex systems. The depth of a neural network can refer to the number of layers it includes, and the number of neurons in each layer can be called the width of that layer. In one implementation, a neural network includes an input layer and an output layer. The input layer processes the received input information through neurons and passes the processing result to the output layer, which then obtains the output of the neural network. In another implementation, a neural network includes an input layer, hidden layers, and an output layer. The input layer processes the received input information through neurons and passes the processing result to the hidden layer. The hidden layer calculates the received processing result and passes the calculation result to the output layer or the next adjacent hidden layer, ultimately obtaining the output of the neural network. A neural network may include one hidden layer or multiple sequentially connected hidden layers, without limitation.
[0144] Neural networks, for example, are deep neural networks (DNNs). Depending on how the network is constructed, DNNs can include feedforward neural networks (FNNs), convolutional neural networks (CNNs), and recurrent neural networks (RNNs).
[0145] Figure 2b This is a schematic diagram of a Free-Nearest Neural Network (FNN). A key characteristic of FNNs is that neurons in adjacent layers are completely connected pairwise. This characteristic makes FNNs typically require a large amount of storage space, leading to high computational complexity.
[0146] CNNs are neural networks specifically designed to process data with a grid-like structure. For example, time-series data (e.g., discrete sampling along a time axis) and image data (e.g., two-dimensional discrete sampling) can both be considered grid-like data. CNNs do not use all the input information at once for computation; instead, they use a fixed-size window to extract a portion of the information for convolution operations, which significantly reduces the computational cost of model parameters. Furthermore, depending on the type of information extracted by the window (e.g., people and objects in an image represent different types of information), each window can use different convolution kernels, allowing CNNs to better extract features from the input data.
[0147] Recurrent Neural Networks (RNNs) are a type of neural network that utilizes feedback time-series information. The input to an RNN includes the current input value and its own output value from the previous time step. RNNs are suitable for acquiring temporally correlated sequence features, and are applicable to applications such as speech recognition and channel coding / decoding.
[0148] In the model training process described above, a loss function can be defined. The loss function describes the difference between the model's output value and the ideal target value. The loss function can be expressed in various forms, and there are no restrictions on its specific form. The model training process can be viewed as follows: by adjusting some or all of the model's parameters, the value of the loss function is made to be less than a threshold or to meet the target requirement.
[0149] A model can also be called an AI model, a rule, or other names. An AI model can be considered a specific method for implementing AI functions. An AI model represents the mapping relationship or function between the model's input and output. AI functions can include one or more of the following: data collection, model training (or model learning), model information dissemination, model inference (or model reasoning, inference, or prediction, etc.), model monitoring or model validation, or inference result publication, etc. AI functions can also be called AI (related) operations or AI-related functions.
[0150] The implementation process of the neural network will be described below with reference to the accompanying drawings.
[0151] 1. Fully connected neural network, also known as multilayer perceptron (MLP).
[0152] like Figure 2c As shown, an MLP consists of an input layer (left side), an output layer (right side), and multiple hidden layers (middle). Each layer of an MLP contains several nodes, called neurons. Neurons in adjacent layers are connected pairwise.
[0153] Optionally, considering neurons in two adjacent layers, the output h of a neuron in the next layer is the weighted sum of all neurons x in the previous layer connected to it, processed by an activation function, and can be expressed as:
[0154] h = f(wx + b).
[0155] Where w is the weight matrix, b is the bias vector, and f is the activation function.
[0156] Alternatively, the output of the neural network can be recursively expressed as:
[0157] y = f z (w z f z-1 (...)+b z ).
[0158] Where z is the index of the neural network layer, z is greater than or equal to 1 and z is less than or equal to Z, where Z is the total number of layers in the neural network.
[0159] In other words, a neural network can be understood as a mapping from an input data set to an output data set. Neural networks are typically initialized randomly; the process of obtaining this mapping from random values w and b using existing data is called training the neural network.
[0160] Optionally, the training method involves using a loss function to evaluate the output of the neural network.
[0161] like Figure 2d As shown, the error can be backpropagated, and the neural network parameters (including w and b) can be iteratively optimized using gradient descent until the output of the loss function reaches its minimum value. Figure 2d The term "relative advantage (e.g., optimal advantage)" is used. This is understandable. Figure 2d The neural network parameters corresponding to the "better points (e.g., the best points)" in the data can be used as neural network parameters in the trained AI model information.
[0162] Alternatively, the gradient descent process can be represented as:
[0163]
[0164] Where θ represents the parameters to be optimized (including w and b), L is the loss function, and η is the learning rate, controlling the step size of gradient descent. This represents the differentiation operation. This indicates taking the derivative of θ with respect to L.
[0165] Alternatively, the backpropagation process can utilize the chain rule for partial derivatives.
[0166] like Figure 2e As shown, the gradient of the parameters in the previous layer can be recursively calculated from the gradient of the parameters in the next layer, and can be expressed as:
[0167]
[0168] Among them, w ij Let s be the weight of the connection between node j and node i. i The weighted sum of the inputs at node i.
[0169] 2. Federated Learning (FL).
[0170] The concept of federated learning effectively addresses the current challenges in the development of artificial intelligence. While fully protecting user data privacy and security, it enables various edge devices and central servers to collaborate efficiently to complete the model's learning task.
[0171] like Figure 2f As shown, the FL architecture is currently the most widely used training architecture in the FL field, and the FedAvg algorithm is the foundational algorithm of FL. The FedAvg algorithm flow is roughly as follows:
[0172] (1) Initialize the model to be trained at the center end. And broadcast it to all clients.
[0173] (2) In the t∈[1,T] round, the client k∈[1,K] is based on the local dataset. For the received global model Perform E epochs of training to obtain the local training results. Report it to the central node. Figure 2f In the example shown, the local training results sent by distributed nodes n, k, and m are denoted as G, respectively. n G k G m .
[0174] (3) The central node collects local training results from all (or some) clients. Assume the set of clients uploading local models in round t is... The central server will use the number of samples from the corresponding client as weights to calculate the new global model. The specific update rule is as follows: Then the central end will send the latest version of the global model. The broadcast is sent to all clients for a new round of training.
[0175] (4) Repeat steps (2) and (3) until the model finally converges or the number of training rounds reaches the upper limit.
[0176] Optionally, in addition to reporting the local model, the client can also... It can also train local gradients The central node averages the local gradients reported by all clients and updates the global model based on this average gradient.
[0177] As can be seen in the FL framework, the dataset resides on distributed nodes (such as clients). These distributed nodes collect their local datasets, perform local training, and report the local results (model or gradients) to the central node. The central node itself does not have a dataset; it is only responsible for fusing the training results from the distributed nodes to obtain a global model, which is then distributed back to the distributed nodes.
[0178] In one implementation example, federated learning can be implemented using SFL (Symptom-Failure Learning). Taking a central node as a cloud server and distributed nodes as terminals as an example, in SFL, after the cloud server distributes the global model to each terminal, the terminal processes the global model using local data to obtain a local model. Subsequently, each terminal uploads its local model to the cloud server, which then performs model aggregation on all the local models uploaded by the terminals, forming a round. This process is repeated periodically until the training requirements are met.
[0179] In another implementation example, federated learning can be implemented using AFL (Automatic Learning Framework). The difference between AFL and SFL is that each terminal can upload its local model after completing local learning. The cloud server can then update itself based on the uploaded local model from any terminal, without waiting for all terminals to upload their local models. For example, the cloud server's update method satisfies:
[0180] X t =(1-α) t )X t-1 +α t X new α t =s(t-τ);
[0181] Among them, X t Let X represent the global model obtained after the t-th update (t > 1). t-1 Let S represent the global model obtained after the (t-1)th update. new This represents a local model uploaded by the terminal; α t It is determined by a monotonically decreasing function s(·), where t is the epoch maintained by the cloud server for counting updates to the global model (e.g., epoch + 1 for each update to the global model), and τ is the epoch maintained by the terminal for counting updates to the local model (when the terminal receives the global model, it can update it, for example, to τ = t).
[0182] 3. Decentralized learning.
[0183] like Figure 2g The diagram shows a fully distributed system without a central node. The design goal f(x) of a decentralized learning system is generally the goal f of each node. i The mean of (x), i.e. Where n is the number of distributed nodes, and x is the parameter to be optimized; in machine learning, x is the parameter of the machine learning model (such as a neural network). Each node utilizes local data and its local target f. i (x) Calculate the local gradient Then it is sent to its communicatively reachable neighboring nodes. Upon receiving the gradient information from its neighbor, any node can update the parameters x of its local model according to the following formula:
[0184]
[0185] in, This represents the parameters of the local model after the (k+1)th update (k is a natural number) in the i-th node. This represents the parameters of the local model after the k-th update in the i-th node (if k is 0, then it represents...). (where α is the parameter of the local model of the i-th node that is not involved in the update) k N represents the tuning coefficient. i It is the set of neighboring nodes of node i, |N i | represents the number of elements in the set of neighboring nodes of node i, that is, the number of neighboring nodes of node i. Through information interaction between nodes, the decentralized learning system will eventually learn a unified model.
[0186] The technical solution provided in this application can be applied to communication systems (e.g.) Figure 1a or Figure 1b or Figure 1c In a communication system (as shown in the diagram), communication nodes typically possess both signal transmission and reception capabilities and computational capabilities. Taking a network device with computational capabilities as an example, the network device's computational capabilities primarily provide computing power support for signal transmission and reception capabilities (e.g., processing signals for transmission and reception) to enable the network device to perform communication tasks with other communication nodes.
[0187] With the development of communication technology, communication equipment in communication systems can now perform not only traditional communication services but also other new types of services, such as artificial intelligence (AI) services. Generally, a communication system capable of handling AI services can also be called an AI system.
[0188] Currently, communication devices can serve as participating nodes in AI systems, providing their own computing power and data. For example, a communication device can perform multiple model processing steps (such as model training, model updating, or model fine-tuning) on an AI model based on local data to obtain another AI model (such as the implementation process of SFL and AFL mentioned above).
[0189] In the field of AI, increasing the number of parameters in an AI model can effectively improve its performance; the emergence of large-scale models is a typical application. The rapid development of large-scale models has spurred applications across various industries, with edge-side large-scale models becoming a definite trend. In the future, they will have an unprecedented impact on the information and communication technology (ICT) application ecosystem and meet users' needs for protecting edge-side privacy data. Edge-cloud data collaboration technology is expected to facilitate the training and promotion of edge-side large-scale models. Generally, before implementing an edge-side large-scale model, it is necessary to fine-tune the general large-scale model using vertical domain knowledge. Federated fine-tuning, as a method for fine-tuning large-scale models, can extensively utilize rich edge-side data resources, while eliminating the need to upload edge-side data, effectively protecting privacy.
[0190] Furthermore, when large models are applied to communication devices, the number of parameters in the AI models processed by the communication devices may gradually increase, which may cause the above model processing process to take a long time (e.g., days, weeks, or even months). During the long model processing process, the communication devices may experience processing interruptions after processing a portion of the data, and after the processing interruption, the model needs to be reprocessed on that portion of data, resulting in a decrease in model processing efficiency.
[0191] As a way to improve model processing efficiency, checkpoints can be used to periodically store large model parameters for fine-tuning. Taking low-rank adaptation (LoRA) as an example, the AI processing node (such as the terminal in SFL mentioned earlier) stores the dimensionality reduction and dimensionality increase matrices for fine-tuning. The fine-tuned parameters account for a certain proportion of the total parameters, and this proportion is adjustable. For example, the proportion of adjustable parameters to the total parameters is [0.1-1]. A proportion of 0.1 means that the number of fine-tuned parameters accounts for one-tenth of the total parameters, and a proportion of 1 means that all parameters are fine-tuned. In addition, checkpoints need to be triggered periodically throughout the fine-tuning process to back up intermediate model parameters multiple times, in order to adapt to situations where the fine-tuning results are not good and it may be necessary to restart the fine-tuning from a specified checkpoint. However, storing the parameters backed up by checkpoints locally will lead to a significant consumption of storage space on the AI processing node, especially in the application scenario of large models, where the storage space of the AI processing node may be insufficient, making it impossible to implement checkpoint backups on the AI processing node.
[0192] For example, taking generative pre-trained transformer 3 (GPT3) as an example, the number of parameters is 175B (B represents billion, i.e., 175B represents 175 billion), stored as single-precision floating-point numbers. Assuming one single-precision floating-point number occupies 32 bits = 4 bytes of storage space, the 175B parameters occupy 175 x 4 = 700GB of storage space. The number of fine-tuning parameters for the LoRA-based large model is 1B, occupying 4GB of storage space. The original data size for the LoRA-based large model is 10B, occupying 40GB of storage space. Assuming the embedding vector occupies ten times the space of the original data, i.e., 40GB x 10 = 400GB. Assuming each terminal (as a client) has the same amount of raw data, 0.4GB, then the space required for the relevant data in each terminal is 0.4GB (raw) + 4GB (embedding vector) + 4GB (fine-tuning parameters) + 700GB (raw model) = 708.4GB. Such a huge amount of storage space is already beyond the capacity of ordinary terminals (such as mobile phones, tablets, etc.). In addition, if there are 100 checkpoints, and each checkpoint backs up 4GB of fine-tuning parameters, an additional 100x4 = 400GB of storage space is required. Terminals are even less able to store such a large number of backup parameters.
[0193] Therefore, while improving model processing efficiency, how to reduce the storage overhead of AI processing nodes (such as terminals) is a technical problem that urgently needs to be solved.
[0194] To address the aforementioned problems, this application provides a communication method and related apparatus, which will be described in detail below with reference to the accompanying drawings.
[0195] Please see Figure 3 This is a schematic diagram of an implementation of the communication method provided in this application, which includes the following steps.
[0196] It should be noted that in the following text, Figure 3 This application illustrates the method by using a first communication device and other communication devices (such as a second communication device and a third communication device) as examples of the execution subjects of the interaction, but it does not limit the execution subjects of the interaction. For example, the communication device can be a communication equipment (such as a terminal device or a network device), or a chip, baseband chip, modem chip, system-on-chip (SoC) chip containing a modem core, system-in-package (SIP) chip, communication module, chip system, processor, logic module, or software in the communication equipment.
[0197] As an example, the first communication device can be a terminal device, an access network device, an Internet of Things device (including sensors and having certain computing capabilities), or an edge cloud server (or a cloud server instance deployed at the edge), etc.
[0198] As an example, the second communication device can be a DSF entity / network element, or other entities / network elements / devices with storage / backup capabilities, such as cloud-side databases or other devices with large storage space.
[0199] As an example, the third communication device can be an access network device, a core network device, a cloud computing instance (such as a virtual machine instance, container instance, or server instance provided by a public cloud service provider), a private cloud server (such as a private cloud server built within an enterprise or organization), a dedicated federated learning server (which can be pre-configured with the software framework, tools, and optimization algorithms required for federated learning), or a high-performance computing cluster (composed of multiple computing, storage, and network nodes).
[0200] S301. The first communication device sends first information, and correspondingly, the second communication device receives the first information. The second communication device is used to back up the first information, which includes N model information pieces. These N model information pieces are used to indicate information about N intermediate models obtained after the first AI model undergoes first local data processing, where N is a positive integer. The N intermediate model information pieces are used to determine the second AI model.
[0201] In one possible implementation, prior to S301, the method further includes: the first communication device receiving indication information indicating the address of the second communication device. In other words, the first communication device can obtain the address of the second communication device based on the indication from other devices, enabling the first communication device to back up the intermediate model obtained locally by the first communication device to a designated backup node (i.e., the second communication device).
[0202] For example, among the N model information items, each model information item may include the model's network structure parameters, model training parameters, or other model-related parameters.
[0203] For example, the network structure parameters of a model may include one or more of the following parameters used to determine the network structure of the model: neural network type identifier / index, number of layers in the neural network structure, weight index, weight value, parameter type identifier, and parameter value.
[0204] For example, the training parameters of a model include one or more of the following parameters used for model training: learning rate and scheduling method, batch size, number of training rounds, loss function, optimizer, and data preprocessing method.
[0205] Optionally, the indication information used to indicate the address of the second communication device may come from a data storage function (DSF) entity or from other entities / network elements / devices, such as a data operation (DA) entity / network element, a data control (DC) entity / network element, etc.
[0206] Optionally, the address indication information of the second communication device includes at least one of the following fields:
[0207] The first field is used to indicate the geographical location of the second communication device;
[0208] The second field is used to indicate the identifier of the bearer transmitting the first information;
[0209] The third field is used to indicate the identifier of the first communication device.
[0210] Optionally, the first communication device can determine the address of the second communication device through pre-configuration to save costs.
[0211] Optionally, in S301, the first information sent by the first communication device may include other information besides the N model information, such as one or more of the following: the index of the N intermediate models, the identifier of the N intermediate models, the indication information indicating the number of times the model is processed corresponding to the N intermediate models, and the indication information indicating the period corresponding to the N intermediate models.
[0212] S302. The second communication device stores the first information.
[0213] In this application, the terms AI model, neural network model, AI neural network model, machine learning model, and AI processing model can be used interchangeably.
[0214] It should be understood that the first communication device can process the first AI model to obtain N intermediate models, and the processing may include one or more of the following: model training, updating, iteration, and optimization.
[0215] Furthermore, the N intermediate models obtained by the first communication device can be AI models obtained through N different time information. For example, the first communication device processes the first AI model through the first time information in the N time information to obtain the first AI model. Then, the first communication device processes the first intermediate model to obtain the second intermediate model, and so on. After obtaining the i-th intermediate model (i takes values from 2 to N-1), the first communication device processes the i-th intermediate model to obtain the (i+1)-th intermediate model.
[0216] Optionally, the first communication device can obtain multiple intermediate models based on the first AI model, and the first communication device can send the multiple intermediate models in various ways.
[0217] For example, the first communication device can send information about the multiple intermediate models based on a predefined time period. That is, the first communication device sends information about the intermediate models obtained within that time period at regular intervals. Accordingly, in the above scheme, the information about the N models included in the first information can be information about the intermediate models obtained within a certain time period.
[0218] For example, the first communication device can send the information of the multiple intermediate models based on a predefined number of times; that is, after obtaining a number of model information each time, the first communication device sends the information of a number of intermediate models. Accordingly, in the above scheme, the information of the N models included in the first information can be the information of the N intermediate models obtained in a certain instance.
[0219] Furthermore, the first AI model can be implemented in various ways, such as being a small model or a large model. In the case of a large model, due to the extremely large number of model parameters, the above solution allows the first communication device to significantly reduce its storage space consumption by backing up intermediate model information through a second communication device. This is beneficial for scenarios where the AI processing node of the large model is deployed on the edge (e.g., the first communication device is a terminal). Moreover, it makes extensive use of the rich data resources on the edge, while eliminating the need to upload edge data, effectively protecting privacy.
[0220] For example, a large model can refer to a machine learning model with a large number of parameters and a complex structure, capable of processing massive amounts of data and completing various complex tasks, such as natural language processing, computer vision, and speech recognition.
[0221] Alternatively, large models are typically built from deep neural networks and have billions or even hundreds of billions or more parameters.
[0222] Alternatively, the design purpose of large models may be to improve the expressive power and predictive performance of the models, enabling them to handle more complex tasks and data.
[0223] Alternatively, large models can learn complex patterns and features by training on massive amounts of data, resulting in stronger generalization capabilities and the ability to make accurate predictions on unprocessed data.
[0224] In contrast, a small model can refer to a model with fewer parameters and a shallower number of layers. Generally, compared to a small model, a large model usually has more parameters and a deeper number of layers, possessing stronger expressive power and higher accuracy, but also requiring more computing resources and time for training and inference. It is suitable for scenarios with large amounts of data and abundant computing resources, such as cloud computing, high-performance computing, and artificial intelligence. For example, small models have advantages such as being lightweight, efficient, and easy to deploy, making them suitable for scenarios with smaller amounts of data and limited computing resources, such as mobile applications, embedded devices, and the Internet of Things (IoT).
[0225] based on Figure 3 In the illustrated scheme, the first communication device can process the first AI model based on the first local data to obtain information on N intermediate models. Subsequently, in step S301, the first communication device can send first information containing the information of the N models to a second communication device used for backup. These N model information items respectively indicate the information of the N models. In other words, after processing the first AI model based on the first local data to obtain intermediate models, the first communication device can back up the intermediate models through information exchange. Therefore, in the event of an interruption in AI model processing, the second communication device can provide the backed-up intermediate model information. Since the intermediate models have already learned the knowledge or features corresponding to the first local data, the first communication device does not need to perform repeated model processing based on the already processed first wireless data, thus improving the processing efficiency of the AI model.
[0226] Furthermore, in the above scheme, the first communication device serves as the AI processing node for processing the AI model. By backing up the information of the intermediate model through the second communication device, the AI processing node does not need to store the information of the intermediate model locally, which can save the storage space consumption of the AI processing node and reduce the storage requirements of the AI processing node.
[0227] exist Figure 3 In one possible implementation, the method further includes: the first communication device receiving information about the second AI model; and the first communication device performing model processing on the second AI model based on second local data. Specifically, after obtaining the global model (i.e., the second AI model), the third communication device can also send information about the second AI model to the first communication device, enabling the first communication device to obtain the global model.
[0228] Optionally, the information of the second AI model may include model information of the second AI model. Optionally, the information of the second AI model may also include one or more of the following: the identifier of the second AI model, and indication information indicating the global processing cycle corresponding to the second model.
[0229] Optionally, after receiving the model information of the second AI model, the first communication device can also process the second AI model based on local data (this processing can be understood as model fine-tuning, secondary fine-tuning, etc.). Similarly, the first communication device can also back up the intermediate model obtained from the processing to the second communication device to improve the processing efficiency of the AI model and reduce the storage requirements of the first communication device.
[0230] Similarly, the information about the M models can be contained within a certain information that includes not only the information about the M models but also other information. For example, it could include one or more of the following: the index of the M intermediate models, the identifier of the M intermediate models, an indication of the number of model processing steps corresponding to the M intermediate models, and an indication of the corresponding period for the M intermediate models.
[0231] It should be understood that Figure 3 The proposed solution can be applied to various AI scenarios. The following example will use the first AI model as the main model.
[0232] Scenario 1: Federated Learning Scenario.
[0233] In this scenario, the third communication device can be a central node, and multiple first communication devices can be distributed nodes. Correspondingly, the first AI model can be a general large model. Multiple first communication devices can process the general large model based on local data (for example, the local data may include vertical domain indications) to obtain an intermediate model. Subsequently, the third communication device can perform one or more processing based on the intermediate model to obtain a global fine-tuning model. This fine-tuning model can be understood as fine-tuning the large model.
[0234] As an example of scenario one, the third communication device can perform model processing once based on intermediate models sent by multiple first communication devices to obtain a globally fine-tuned model. That is, the third communication device can only perform model processing after receiving intermediate models sent by all the first communication devices. In this case, the intermediate models sent by different first communication devices are processed through the same process; the model processing process of the third communication device can be understood as a synchronous federated learning (SFL) process.
[0235] Taking the first communication device as the terminal and the third communication device as the server as an example, what problems might arise during the implementation based on SFL?
[0236] Problem 1. Terminals with different capabilities need to upload their local models at different times. The server can only perform model aggregation after all terminals have uploaded their local models. This approach can be called a "greedy strategy." If one terminal trains slowly, other terminals need to wait for it to finish training, resulting in excessively long training times and low training efficiency in the SFL implementation.
[0237] Question 2. When a weak computing terminal slows down the overall training progress, a common method is to discard slow terminals, i.e., user scheduling. Each round of federated learning involves terminal selection, choosing only high-capacity terminals as participants. However, this approach essentially discards high-quality data from weaker terminals; selecting devices solely based on computing power leads to data loss. High-quality data is crucial for fine-tuning large models, which makes the SFL implementation unfavorable for the effectiveness of federated fine-tuning.
[0238] As another example of Scenario 1, the third communication device can perform model processing two or more times based on intermediate models sent by multiple first communication devices to obtain a globally fine-tuned model. That is, the third communication device can receive intermediate models processed by any of the first communication devices and then perform model processing. In this case, the intermediate models obtained by different first communication devices are processed through different processes by the third communication device. The model processing process of the third communication device can be understood as an asynchronous federated learning (AFL) process. Furthermore, in the implementation of AFL, since the third communication device does not need to wait for all the first communication devices to send intermediate models, it can effectively reduce processing latency and improve model processing efficiency.
[0239] For example, in Figure 3 The method may further include: a third communication device receiving second model information, which indicates information about an intermediate model obtained after the first AI model undergoes local data processing by the first communication device; the third communication device processing the first AI model based on the second model information to obtain a first global model; the third communication device receiving third model information, which indicates information about an intermediate model obtained after the first AI model undergoes local data processing by a fourth communication device; the third communication device processing the first global model based on the third model information to obtain a second global model; wherein the second global model is used to determine the second AI model. Specifically, the third communication device may perform two or more model processing steps based on intermediate models obtained from multiple communication devices (e.g., the first communication device, the fourth communication device, etc.) to obtain a global fine-tuned model. In this case, the intermediate models obtained from multiple communication devices may be processed through different processing steps of the third communication device, and the model processing steps of the third communication device can be understood as asynchronous federated learning (AFL) processes. In this way, during the implementation of AFL, since the third communication device does not need to wait for all communication devices to send intermediate models, the processing latency can be effectively reduced, thereby improving the model processing efficiency and solving the above-mentioned problem 1.
[0240] Optionally, during the process of the third communication device receiving the second model information, the second model information may come from the first communication device or from a DSF entity / DSF network element used to back up the model information of the first communication device (such as the second communication device described above). Similarly, during the process of the third communication device receiving the third model information, the second model information may come from the fourth communication device or from a DSF entity / DSF network element used to back up the model information of the fourth communication device.
[0241] Figure 4 As an example of AFL implementation, in the process of large model fine-tuning, to solve the problem of excessive latency caused by asynchronous training progress due to device training interruptions, different capabilities, and possible communication interruptions, AFL can be used for large model fine-tuning. The following will use the first communication device as the terminal, the second communication device as a DSF entity / DSF network element, and the third communication device as a server as an example, and the specific implementation steps include:
[0242] Step 1: The DO / DC establishes a data bearer. The edge cloud or base station acts as the AFL server. Several terminals can initiate AFLs. Based on the capabilities reported by the terminals, the DO / DC assigns temporary DSF addresses to them. These temporary DSF addresses can include the DSF's geographical location, DSID, and terminal identifier. The specific format of this temporary DSF address is: DSF geographical location + DSID + terminal identifier. This temporary DSF address is valid during the temporary address fine-tuning process and is released after the fine-tuning is completed. In addition, the DO / DC notifies the DSF entity / DSF network element to reserve storage space and informs the DSF entity / DSF network element of the corresponding terminal identifier.
[0243] Step 2: The server distributes the global universal model to each terminal.
[0244] Step 3: Each terminal performs fine-tuning locally based on local data. Each gradient calculation using a mini-batch of local data samples is considered an update. Every few updates, a checkpoint is triggered, and the backed-up model parameters are uploaded to the DSF entity / DSF network element until all local data has been learned. When training is interrupted at a terminal (e.g., terminal 1), a model parameter recovery request is initiated. The DSF entity / DSF network element then sends the backed-up model parameters from the checkpoint for terminal 1 to terminal 1.
[0245] It is understandable that step 3 is an example of implementing steps S301 and S302 mentioned above.
[0246] Optionally, DSF entities / DSF network elements may not have unlimited storage. Therefore, storage rules can be determined through pre-configuration or server configuration. For example, a storage parameter limit can be pre-configured for each DSF entity / DSF network element. When the stored backup parameters reach the limit, and new parameters are added, assuming the different backup parameters have similar or identical values, the oldest parameter will be discarded. Then, for each new parameter, an older parameter will be discarded. If a discard (or deletion) rule is configured for excess parameters on the DSF entity / DSF network element via DO / DC or other network elements, the priority of this configured rule is higher than the pre-configured rule.
[0247] Step 4: After each terminal has learned all the local data, it asynchronously uploads the local large model to the server.
[0248] Step 5: The server updates the global model once a large model is received. For example, as mentioned earlier, the server can update "epoch+1" locally. After multiple epochs of asynchronous federated learning, the learning ends when the required number of epochs is reached, and the large model for the vertical domain is obtained.
[0249] Step 6: The server distributes the large vertical domain model to each terminal, and each terminal generates its own large terminal model through local fine-tuning and other methods.
[0250] Optionally, any first communication device can be connected to one or more subordinate nodes, and the first communication device can be the central node of the one or more subordinate nodes. The first communication device and the one or more subordinate nodes can perform model processing using either SFL or AFL; no limitation is made here.
[0251] Furthermore, when model processing between the first communication device and the one or more subordinate nodes can be performed using SFL (Small Federation), the relationship between the first communication device and the one or more subordinate nodes can be referred to as a small federation (or a cluster), and the relationship between the third communication device and multiple first communication devices can be referred to as a large federation. In this implementation, the aggregation model (cluster model) of each small federation can learn data from both strong and weak devices equally, enhancing the availability of weak device data within the small federation. Simultaneously, using AFL (Advanced Framework) between different clusters avoids excessively long training times and low training efficiency. Thus, it balances improving the utilization rate of weak device data with system training efficiency.
[0252] For example, in the implementation of AFL and SFL combined fine-tuning, consider one or more lower-level nodes connected to the first communication device, including some weak devices and some strong devices. Several weak devices and several strong devices form a cluster. SFL is executed within each cluster, and the first communication device can perform model aggregation for each cluster, aggregating them into a cluster model. AFL is executed between multiple clusters, and the third communication device updates the global model. The availability of the cluster model is determined by the weak devices, and the "greedy strategy" scheme described earlier and the "breakpoint direct transmission" scheme described later can be used. Unlike before, the lower-level nodes may also be allocated DSF, which uploads the lower-level node's model to the first communication device, which has the model aggregation function of SFL. Optionally, the "weak device near-term direct transmission" scheme described later can also be used between the first communication device and one or more lower-level nodes.
[0253] Scenario 2: The scenario where the first AI model fine-tunes the model itself.
[0254] In scenario two, the first communication device can perform one or more model processing steps on its own deployed first AI model to obtain a fine-tuned second AI model.
[0255] Optionally, in scenario two, the first communication device can process the first AI model to obtain the second AI model through interaction with the second communication device. The first communication device can also communicate with other communication devices (e.g., Figure 2g The interaction between neighboring nodes (as shown) enables the collection of local data.
[0256] Optionally, the second communication device can be an entity / network element / device used to back up model information. For example, the second communication device can be a data storage function (DSF) entity / network element, or other names, which are not limited here.
[0257] Depend on Figure 3 As shown in the diagram, the second communication device can store the model information backed up by the first communication device in step S302. Subsequently, the second communication device, acting as a backup node, can provide the model information backed up by the second communication device in various ways, which will be described below in conjunction with the following implementation methods.
[0258] Implementation Method 1: The second communication device sends the backup model information to the first communication device, such as... Figure 3 In step S303, the second communication device can send the first model information to the first communication device.
[0259] In implementation method one, Figure 3 The method shown may also include:
[0260] Step S303. The first communication device sends second information to the second communication device, the second information being used to request first model information, which is one of the N model information; the first communication device receives the first model information from the second communication device. In other words, the first communication device can send second information to the second communication device to request the first model information, causing the second communication device to send the first model information to the first communication device. Wherein, in the event of a model processing interruption, the first communication device can obtain intermediate model information through the second communication device, enabling the first communication device to locally recover the intermediate model, thereby improving the processing efficiency of the AI model.
[0261] Optionally, the N model information pieces correspond to N time information pieces, and the first model information piece is the latest model information piece corresponding to the time information of the N model information pieces; or, the second information piece includes the index of the first model information piece.
[0262] Optionally, implementation example one is applied to Figure 4 In the scenario shown, the second communication device can Figure 4 After step 3 and before step 4, the backup model information is sent to the first communication device.
[0263] Method 2: The second communication device sends the backup model information to the third communication device.
[0264] In implementation method two, Figure 3 The method shown may also include:
[0265] Step S304. The second communication device sends second model information to the third communication device, whereby the second model information is one of the N model information. Specifically, the second communication device can provide the third communication device with the intermediate model processed by the first communication device, enabling the third communication device to obtain the information of the backup intermediate model and further process the intermediate model to improve the processing efficiency of the AI model.
[0266] Optionally, the second model information may be included in a certain information, which may include other information besides the second model information. For example, it may include one or more of the following: the index / identifier of the second intermediate model, the indication information of the number of model processing times corresponding to the second intermediate model, the indication information of the period corresponding to the second intermediate model, the identifier of the second communication device, the address of the second communication device, the address of the first communication device, and the identifier of the first communication device, so that the third communication device can process the second model based on the one or more of the information.
[0267] Furthermore, in implementation method two, the second communication device can trigger the transmission of backup model information to the third communication device in various ways. Some implementation examples will be provided below.
[0268] In one implementation example of the second method (denoted as Example A), the second communication device can send corresponding backup model information to the third communication device based on the instruction or request of the first communication device.
[0269] In Example A, Figure 3 The method further includes: the first communication device sending third information to the second communication device, the third information requesting the second communication device to send second model information to the third communication device, the second model information being one of the N model information sets. In other words, the first communication device can send third information to the second communication device requesting the second communication device to send the second model information to the third communication device, causing the second communication device to send the second model information to the third communication device. In the event of a model processing interruption, the first communication device can provide the intermediate model it processed to the third communication device through the second communication device, enabling the third communication device to obtain the information of the backup intermediate model and further process it to improve the processing efficiency of the AI model.
[0270] Optionally, the second model information sent by the second communication device to the third communication device based on the third information may be contained within a certain information. This information, in addition to including the second model information (i.e., the model indicated by the second model information is model A), may also include other information. For example, the other information may include one or more of the following: the index of model A, the identifier of model A, indication information indicating the number of model processing times corresponding to model A, and indication information indicating the period corresponding to model A. Accordingly, the third information can be used not only to request the second communication device to send the second model information to the third communication device, but also to request the second communication device to send the other information to the third communication device.
[0271] Optionally, the N model information pieces correspond to N time information pieces respectively, and the second model information is the latest model information corresponding to the time information of the N model information pieces (i.e., the first model information and the second model information can be the same); or, the third information includes the index of the second model information.
[0272] In one possible implementation of Example A, the process of the first communication device sending the third information may include: the first communication device sending the third information when a first condition is met, wherein the first condition includes any one of the following:
[0273] The model processing of the intermediate model corresponding to the first AI model (such as any intermediate model obtained by the first communication device, the current intermediate model, the most recently processed intermediate model, etc.) is interrupted;
[0274] Determine the current intermediate model of the first AI model to be nearing its expiration date (or about to expire, about to become invalid, etc.).
[0275] Specifically, when the first communication device determines that the model processing of the intermediate model corresponding to the first AI model is interrupted, the first communication device can request the second communication device to provide the intermediate model processed by the first communication device to the third communication device via a third information request. This allows the third communication device to perform model processing based on the information of the intermediate model sent by the second communication device. This solves the latency and model expiration problems caused by the first communication device's fine-tuning interruption and retraining, and improves the utilization rate of local data.
[0276] And / or, if the first communication device determines that the current intermediate model of the first AI model is nearing expiration, the first communication device can request the second communication device to provide the intermediate model processed by the first communication device to the third communication device via a third information request. This allows the third communication device to perform model processing based on the intermediate model information sent by the second communication device. Therefore, the potential problem of model expiration can be solved, effectively preventing model expiration and improving the utilization rate of the local data of the first communication device.
[0277] Optionally, the first communication device determines the current intermediate model nearing its expiration date for the first AI model by: receiving indication information indicating the global model processing count; and determining the current intermediate model nearing its expiration date for the first AI model when the difference between the local model processing count and the global model processing count meets a second condition. In other words, when the difference between the local model processing count and the global model processing count meets the second condition, the first communication device can determine that it is a weaker device, meaning its AI model processing speed is lower than that of other communication devices at the same level. Therefore, while other communication devices have uploaded multiple models and performed multiple global model updates, the local model update of the first communication device has not yet ended. To this end, the first communication device can determine the current intermediate model nearing its expiration date for the first AI model and, through the triggering of third information, enable the third communication device to obtain the intermediate model information of the first communication device, thus maximizing the use of local data from each device in the global model processing.
[0278] Optionally, the indication information used to indicate the number of global model processing times can be periodically sent information or information triggered based on a request from the first communication device; no limitation is made here.
[0279] For example, the third information includes any of the following:
[0280] The first indication information is used to indicate that the third information is triggered by a model processing interruption based on the intermediate model corresponding to the first AI model;
[0281] The second instruction information is used to indicate that the third information is triggered based on the current intermediate model of the first AI model.
[0282] Therefore, the third information sent by the first communication device may include any of the aforementioned indication information, enabling the recipient of the third information to determine the reason why the first communication device sent the third information based on the third information.
[0283] Optionally, in implementation method two, if the second communication device determines that the current intermediate model of the first AI model is nearing its end, the third communication device will perform model processing based on the intermediate model information sent by the second communication device. To this end, the second communication device can send an instruction message to the first communication device, so that the first communication device stops the model processing of the current intermediate model corresponding to the first AI model based on the instruction message, so as to avoid unnecessary processing overhead.
[0284] In another implementation example of implementation method two (denoted as Example B), the second communication device can send the backup model information to the third communication device locally based on the judgment results of certain parameters.
[0285] In Example B, Figure 3 The method further includes: a second communication device receiving indication information indicating the global number of model processing iterations; and the second communication device sending the second model information when the difference between the number of model processing iterations corresponding to the N model information and the global number of model processing iterations meets a second condition. Specifically, if the second communication device determines that the current intermediate model of the first AI model is nearing expiration, the second communication device can provide the intermediate model processed by the first communication device to the third communication device, enabling the third communication device to perform model processing based on the intermediate model information sent by the second communication device. This solves the potential problem of model expiration, effectively avoids model expiration, and improves the utilization rate of the local data of the first communication device.
[0286] Optionally, in Example B, the second communication device may also send an instruction to the first communication device to instruct the cessation of model processing for the current intermediate model corresponding to the first AI model. Specifically, if the second communication device determines that the current intermediate model of the first AI model is nearing its end, the third communication device will perform model processing based on the intermediate model information sent by the second communication device. To this end, the second communication device may send an instruction to the first communication device, causing the first communication device to cessate model processing for the current intermediate model corresponding to the first AI model based on the instruction, thereby avoiding unnecessary processing overhead.
[0287] As can be seen from Example B above, under certain conditions, the first or second communication device will be triggered to send the intermediate model of the first communication device to the third communication device. This can effectively utilize the local data of each first communication device to improve the effect of federated fine-tuning and solve the aforementioned problem 2. The following explanation will take the first communication device as the terminal and the second communication device as a DSF entity / DSF network element as an example.
[0288] As an example, in the model processing, based on a "greedy strategy" approach, learning interruptions and retraining cause latency, leading to model obsolescence: multiple interruptions and retrainings during training increase time and computational costs. After multiple interruptions and retrainings, by the time the terminal has learned all the local data samples, the local model may have become outdated, reducing its usability. In other words, although the local model has learned this data, it has not been learned by the global model.
[0289] To address this issue, the solution proposed in Example B can be considered a "breakpoint direct transmission" solution. When the terminal reaches a certain criterion, "breakpoint direct transmission" is triggered, meaning that training does not need to be restarted. Instead, the DSF entity / DSF network element is requested to upload the latest fine-tuning parameters, effectively avoiding the latency caused by interrupted retraining. After the server distributes the global model, it then learns from the untrained data based on this model.
[0290] For example, assuming there are K = 100 checkpoints, if the process is interrupted when k = 99, it means that most of the data has been learned, and the "breakpoint direct transmission" option can be selected; if the process is interrupted when k = 1, it means that only a small amount of data has been learned, and the "greedy strategy" option can be selected. Under this assumption, the response rule after "fine-tuning interruption" can be set to k >= (or <) K / 2.
[0291] Therefore, the "breakpoint direct transmission" solution can balance avoiding model expiration and maximizing the utilization of local data, allowing terminal data to be effectively learned by the global model, improving the utilization rate of terminal data, and this solution can be adopted for both strong and weak devices.
[0292] As another example, regarding the issue of insufficient data utilization from weak devices, even without interrupting retraining, there are differences in training speed between devices due to their varying capabilities. If high-capability terminals always update faster, the global model will learn more from the data of high-capability terminals, and its performance will favor high-capability terminals. However, weak-capability terminals may possess high-value data but rarely learn from it. This makes it difficult for the trained global model to adapt to weak-capability terminals (because it has little knowledge of high-capability terminals), and it also hinders federated fine-tuning to obtain a good model.
[0293] To address this issue, the solution proposed in Example B can be viewed as a "near-expiration direct transmission for weak devices" solution, preventing model expiration for weak devices from a temporal perspective. The weak device periodically retrieves the global model processing counter `t`, and its local model processing counter is denoted by `τ`. If `t - τ` equals the threshold (an integer), it indicates that the local model of the weak device is nearing expiration (expiration is defined as having minimal impact on the new global model; the threshold is determined by the cloud-side function αt = s(t - τ)). The device then notifies the DSF entity / DSF network element to send the model parameters from the previous checkpoint to the weak device. The global processing counter `t` can also be distributed to the DSF entity / DSF network element corresponding to the weak device. The DSF entity / DSF network element maintains the local model processing counter `τ`, which can be obtained from the information carried when the weak device uploads a backup model. When `t - τ` equals the threshold, the DSF entity / DSF network element sends the latest model parameters to the weak device.
[0294] Therefore, the "direct transmission to weak devices nearing expiration" solution can effectively avoid model expiration and improve the utilization rate of data from weak devices. Regardless of how much data is used when the model is nearing expiration, even if it's only a small portion, this data can still contribute to the global model, avoiding a situation where, although the data has been learned, its usability for the global model is greatly reduced due to model expiration. In some embodiments, since the update of the global processing counter t is led by the strong device, t cannot be updated if the strong device model is not uploaded, thus eliminating the expiration problem. Therefore, this method can disregard the expiration of models on strong devices.
[0295] Please see Figure 5This application provides a communication device 500, which can realize the functions of the first communication device (or second communication device) in the above method embodiments, and thus also achieve the beneficial effects of the above method embodiments. In this application embodiment, the communication device 500 can be the first communication device (or the second communication device), or it can be an integrated circuit or component inside the first communication device (or the second communication device), such as a chip, baseband chip, modem chip, SoC chip containing a modem core, system-in-package (SIP) chip, communication module, chip system, processor, etc.
[0296] It should be noted that the transceiver unit 502 may include a transmitting unit and a receiving unit, which are used to perform transmitting and receiving respectively.
[0297] In one possible implementation, when the device 500 is for performing Figure 3 When the method executed by the first communication device in the relevant embodiments is performed, the device 500 includes a processing unit 501 and a transceiver unit 502; the processing unit 501 is used to determine first information, the first information including N model information, the N model information being used to indicate the information of N intermediate models obtained after the first AI model has undergone first local data processing, where N is a positive integer; the N intermediate model information is used to determine a second AI model; the transceiver unit 502 is used to send the first information to a second communication device, and the second communication device is used to back up the first information.
[0298] In one possible implementation, when the device 500 is for performing Figure 3 When the method executed by the second communication device in the relevant embodiments is used, the device 500 includes a processing unit 501 and a transceiver unit 502; the transceiver unit 502 is used to receive first information from the first communication device, the first information includes N model information, the N model information is used to indicate the information of N intermediate models obtained after the first AI model has undergone the first local data processing, where N is a positive integer; the N intermediate model information is used to determine the second AI model; the processing unit 501 is used to store the first information.
[0299] In one possible implementation, when the device 500 is for performing Figure 3When the method executed by the third communication device in the relevant embodiments is implemented, the device 500 includes a processing unit 501 and a transceiver unit 502; the transceiver unit 502 is used to receive second model information, which is used to indicate the information of the intermediate model obtained after the first AI model has undergone local data processing by the first communication device; the processing unit 501 is used to process the first AI model based on the second model information to obtain a first global model; the transceiver unit 502 is also used to receive third model information, which is used to indicate the information of the intermediate model obtained after the first AI model has undergone local data processing by the fourth communication device; the processing unit 501 is also used to process the first global model based on the third model information to obtain a second global model; wherein, the second global model is used to determine the second AI model.
[0300] In one possible design, when the communication device 500 is a terminal device or a communication module within a terminal, the function of the processing unit 501 can be implemented by one or more processors. Specifically, the processor may include a modem chip, or a SoC chip or SIP chip containing a modem core. The function of the transceiver unit 502 can be implemented by transceiver circuitry.
[0301] In one possible design, when the communication device 500 is a circuit or chip in a terminal responsible for communication functions, such as a modem chip or a SoC chip or SIP chip containing a modem core, the function of the processing unit 501 can be implemented by a circuit system in the aforementioned chip that includes one or more processors or processor cores. The function of the transceiver unit 502 can be implemented by the interface circuit or data transceiver circuit on the aforementioned chip.
[0302] It should be noted that the information execution process of the unit of the above-mentioned communication device 500 can be specifically described in the method embodiment shown above in this application, and will not be repeated here.
[0303] Please see Figure 6 This is another schematic structural diagram of the communication device 600 provided in this application. The communication device 600 includes a logic circuit 601 and an input / output interface 602. The communication device 600 can be a chip or an integrated circuit.
[0304] in, Figure 5 The transceiver unit 502 shown can be a communication interface, which can be... Figure 6 The input / output interface 602 may include an input interface and an output interface. Alternatively, the communication interface may also be a transceiver circuit, which may include an input interface circuit and an output interface circuit.
[0305] In one possible implementation, when the device 600 is for performing Figure 3 When the method executed by the first communication device in the relevant embodiments is performed, the logic circuit 601 is used to determine the first information, which includes N model information, which are respectively used to indicate the information of N intermediate models obtained after the first AI model has undergone the first local data processing, where N is a positive integer; the N intermediate model information is used to determine the second AI model; the input / output interface 602 is used to send the first information to the second communication device, and the second communication device is used to back up the first information.
[0306] In one possible implementation, when the device 600 is for performing Figure 3 When the second communication device in the related embodiments executes the method, the input / output interface 602 is used to receive first information from the first communication device. The first information includes N model information, which are used to indicate the information of N intermediate models obtained after the first AI model has undergone the first local data processing, where N is a positive integer. The N intermediate model information is used to determine the second AI model. The logic circuit 601 is used to store the first information.
[0307] In one possible implementation, when the device 600 is for performing Figure 3 When the method executed by the third communication device in the relevant embodiments is performed, the input / output interface 602 is used to receive second model information, which is used to indicate the information of the intermediate model obtained after the first AI model has undergone local data processing by the first communication device; the logic circuit 601 is used to process the first AI model based on the second model information to obtain a first global model; the input / output interface 602 is also used to receive third model information, which is used to indicate the information of the intermediate model obtained after the first AI model has undergone local data processing by the fourth communication device; the logic circuit 601 is also used to process the first global model based on the third model information to obtain a second global model; wherein, the second global model is used to determine the second AI model.
[0308] The logic circuit 601 and the input / output interface 602 can also perform other steps performed by the first or second communication device in any embodiment and achieve corresponding beneficial effects, which will not be elaborated here.
[0309] In one possible implementation, Figure 5 The processing unit 501 shown can be Figure 6 The logic circuit 601 in the middle.
[0310] Optionally, the logic circuit 601 can be a processing device, the functions of which can be partially or entirely implemented in software.
[0311] Optionally, the processing apparatus may include a memory and a processor, wherein the memory is used to store a computer program, and the processor reads and executes the computer program stored in the memory to perform the corresponding processing and / or steps in any of the method embodiments.
[0312] Optionally, the processing device may consist of only a processor. A memory for storing computer programs is located outside the processing device, and the processor is connected to the memory via circuitry / wires to read and execute the computer programs stored in the memory. The memory and processor may be integrated together or physically independent of each other.
[0313] Optionally, the processing device may be one or more chips, or one or more integrated circuits. For example, the processing device may be one or more field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), system on-chips (SoCs), central processing units (CPUs), network processors (NPs), digital signal processors (DSPs), microcontroller units (MCUs), programmable logic controllers (PLDs), or other integrated chips, or any combination of the above chips or processors.
[0314] Please see Figure 7 The communication device 700 provided in the above embodiments of this application can specifically be the communication device that serves as a terminal device in the above embodiments. Figure 7 The example shown illustrates how a terminal device can be implemented through a terminal device (or a component within a terminal device).
[0315] The present invention provides a possible logical structure diagram of the communication device 700, which may include, but is not limited to, at least one processor 701 and a communication port 702.
[0316] in, Figure 5 The transceiver unit 502 shown can be a communication interface, which can be... Figure 7 The communication port 702 in the diagram may include an input interface and an output interface. Alternatively, the communication port 702 may also be a transceiver circuit, which may include an input interface circuit and an output interface circuit.
[0317] Further optionally, the device may also include at least one of a memory 703 and a bus 704. In the embodiments of this application, the at least one processor 701 is used to control the operation of the communication device 700.
[0318] Furthermore, the processor 701 can be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a digital signal processor and a microprocessor, etc. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0319] It should be noted that, Figure 7 The communication device 700 shown can be used to implement the steps implemented by the terminal device in the aforementioned method embodiments, and to achieve the corresponding technical effects of the terminal device. Figure 7 The specific implementation of the communication device shown can be referred to the description in the foregoing method embodiments, and will not be repeated here.
[0320] Please see Figure 8 The above-described embodiment of the communication device 800 is a schematic diagram illustrating the structure of the communication device 800 provided in the embodiments of this application. Specifically, the communication device 800 can be a communication device serving as a network device as described in the above-described embodiment. Figure 8 The example shown illustrates a network device implemented through a network device (or a component within a network device). The structure of this communication device can be referenced. Figure 8 The structure shown.
[0321] The communication device 800 includes at least one processor 811 and at least one network interface 814. Optionally, the communication device further includes at least one memory 812, at least one transceiver 813, and one or more antennas 815. The processor 811, memory 812, transceiver 813, and network interface 814 are connected, for example, via a bus. In this embodiment, the connection may include various interfaces, transmission lines, or buses, etc., and this embodiment is not limited thereto. The antenna 815 is connected to the transceiver 813. The network interface 814 enables the communication device to communicate with other communication devices through a communication link. For example, the network interface 814 may include a network interface between the communication device and core network equipment, such as an S1 interface, or a network interface between the communication device and other communication devices (e.g., other network devices or core network equipment), such as an X2 or Xn interface.
[0322] in, Figure 5 The transceiver unit 502 shown can be a communication interface, which can be... Figure 8 The network interface 814 may include an input interface and an output interface. Alternatively, the network interface 814 may also be a transceiver circuit, which may include input interface circuitry and output interface circuitry.
[0323] The processor 811 is mainly used to process communication protocols and communication data, control the entire communication device, execute software programs, and process data from the software programs, for example, to support the communication device in performing the actions described in the embodiments. The communication device may include a baseband processor and a central processing unit. The baseband processor is mainly used to process communication protocols and communication data, while the central processing unit is mainly used to control the entire terminal device, execute software programs, and process data from the software programs. Figure 8 The processor 811 can integrate the functions of a baseband processor and a central processing unit. Those skilled in the art will understand that the baseband processor and the central processing unit can also be independent processors interconnected via technologies such as buses. Those skilled in the art will understand that a terminal device can include multiple baseband processors to adapt to different network standards, and a terminal device can include multiple central processing units to enhance its processing capabilities. The various components of the terminal device can be connected via various buses. The baseband processor can also be described as a baseband processing circuit or a baseband processing chip. The central processing unit can also be described as a central processing circuit or a central processing chip. The function of processing communication protocols and communication data can be built into the processor or stored in memory as a software program, with the processor executing the software program to implement the baseband processing function.
[0324] The memory is primarily used to store software programs and data. The memory 812 can exist independently or be connected to the processor 811. Optionally, the memory 812 can be integrated with the processor 811, for example, integrated into a single chip. The memory 812 can store program code that executes the technical solutions of the embodiments of this application, and its execution is controlled by the processor 811. The various types of computer program code being executed can also be considered as drivers for the processor 811.
[0325] Figure 8 Only one memory and one processor are shown. In actual terminal devices, there may be multiple processors and multiple memories. Memory can also be called storage medium or storage device, etc. Memory can be a storage element on the same chip as the processor, i.e., an on-chip storage element, or it can be a separate storage element; this application does not limit this.
[0326] Transceiver 813 can be used to support the reception or transmission of radio frequency (RF) signals between a communication device and a terminal. Transceiver 813 can be connected to antenna 815. Transceiver 813 includes a transmitter Tx and a receiver Rx. Specifically, one or more antennas 815 can receive RF signals. The receiver Rx of transceiver 813 receives the RF signals from the antennas, converts the RF signals into digital baseband signals or digital intermediate frequency (IF) signals, and provides the digital baseband signals or IF signals to processor 811 so that processor 811 can perform further processing on the digital baseband signals or IF signals, such as demodulation and decoding. Furthermore, the transmitter Tx in transceiver 813 is also used to receive modulated digital baseband signals or IF signals from processor 811, convert the modulated digital baseband signals or IF signals into RF signals, and transmit the RF signals through one or more antennas 815. Specifically, the receiver Rx can selectively perform one or more stages of downmixing and analog-to-digital conversion on the radio frequency signal to obtain a digital baseband signal or a digital intermediate frequency (IF) signal. The order of these downmixing and IF conversion processes is adjustable. The transmitter Tx can selectively perform one or more stages of upmixing and digital-to-analog conversion on the modulated digital baseband signal or digital IF signal to obtain a radio frequency signal. The order of these upmixing and IF conversion processes is also adjustable. The digital baseband signal and the digital IF signal can be collectively referred to as digital signals.
[0327] The transceiver 813 can also be called a transceiver unit, transceiver, transceiver device, etc. Optionally, the device in the transceiver unit that performs the receiving function can be regarded as the receiving unit, and the device in the transceiver unit that performs the transmitting function can be regarded as the transmitting unit. That is, the transceiver unit includes a receiving unit and a transmitting unit. The receiving unit can also be called a receiver, input port, receiving circuit, etc., and the transmitting unit can be called a transmitter, transmitter, or transmitting circuit, etc.
[0328] It should be noted that, Figure 8 The communication device 800 shown can be used to implement the steps implemented by the network device in the aforementioned method embodiments, and to achieve the corresponding technical effects of the network device. Figure 8 The specific implementation of the communication device 800 shown can be referred to the description in the foregoing method embodiments, and will not be repeated here.
[0329] Please see Figure 9 The above-described embodiments of the communication device provided in this application are schematic diagrams of the structure of the communication device.
[0330] It is understood that the communication device 900 includes, for example, modules, units, elements, circuits, or interfaces, which are appropriately configured together to execute the technical solutions provided in this application. The communication device 900 may be the terminal device or network device described above, or a component (e.g., a chip) within these devices, used to implement the methods described in the following method embodiments. The communication device 900 includes one or more processors 901. The processor 901 may be a general-purpose processor or a dedicated processor, for example, a baseband processor or a central processing unit. The baseband processor can be used to process communication protocols and communication data, and the central processing unit can be used to control the communication device (e.g., a RAN node, terminal, or chip), execute software programs, and process data from the software programs.
[0331] Optionally, in one design, processor 901 may include program 903 (sometimes also referred to as code or instructions), which can be executed on processor 901 to cause communication device 900 to perform the methods described in the embodiments below. In yet another possible design, communication device 900 includes circuitry (…). Figure 9 (Not shown).
[0332] Optionally, the communication device 900 may include one or more memories 902 storing a program 904 (sometimes referred to as code or instructions), which can be run on the processor 901 to cause the communication device 900 to perform the methods described in the above method embodiments.
[0333] Optionally, the processor 901 and / or memory 902 may include AI modules 907 and 908, which are used to implement AI-related functions. The AI modules can be implemented through software, hardware, or a combination of both. For example, the AI module may include a radio intelligence control (RIC) module. For example, the AI module may be a near real-time RIC or a non-real-time RIC.
[0334] Optionally, the processor 901 and / or memory 902 may also store data. The processor and memory may be configured separately or integrated together.
[0335] Optionally, the communication device 900 may further include a transceiver 905 and / or an antenna 906. The processor 901, sometimes referred to as a processing unit, controls the communication device (e.g., a RAN node or terminal). The transceiver 905, sometimes referred to as a transceiver unit, transceiver, transceiver circuit, or transceiver, is used to implement the transmission and reception functions of the communication device via the antenna 906.
[0336] in, Figure 5 The processing unit 501 shown may be a processor 901. Figure 5 The transceiver unit 502 shown can be a communication interface, which can be... Figure 9 The transceiver 905 in the diagram may include an input interface and an output interface. Alternatively, the transceiver 905 may also be a transceiver circuit, which may include an input interface circuit and an output interface circuit.
[0337] This application also provides a computer-readable storage medium for storing one or more computer-executable instructions. When the computer-executable instructions are executed by a processor, the processor performs the method described in the possible implementations of the first or second communication device in the foregoing embodiments.
[0338] This application also provides a computer program product (or computer program) that, when executed by a processor, executes the method described above for the possible implementation of the first or second communication device.
[0339] This application also provides a chip system including at least one processor for supporting a communication device in implementing the functions involved in the possible implementations of the communication device described above. Optionally, the chip system further includes an interface circuit that provides program instructions and / or data to the at least one processor. In one possible design, the chip system may also include a memory for storing the program instructions and data necessary for the communication device. The chip system may be composed of chips or may include chips and other discrete devices, wherein the communication device may specifically be the first communication device or the second communication device in the aforementioned method embodiments.
[0340] This application also provides a communication system, the network system architecture of which includes the first communication device and / or the second communication device in any of the above embodiments.
[0341] In the embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between devices or units, and may be electrical, mechanical, or other forms. Whether a function is implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0342] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0343] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A communication method, characterized in that, include: First information is determined, which includes N model information, wherein the N model information is used to indicate the information of N intermediate models obtained after the first AI model undergoes first local data processing, where N is a positive integer; the N intermediate model information is used to determine the second AI model. The first information is sent to a second communication device, which is used to back up the first information.
2. The method according to claim 1, characterized in that, The method further includes: Send a second message to the second communication device, the second message being used to request first model information, the first model information being one of the N model information; Receive the first model information from the second communication device.
3. The method according to claim 2, characterized in that, The N model information pieces correspond to N time information pieces respectively, and the first model information piece is the model information piece with the latest time information corresponding to the N model information pieces; or, The second information includes the index of the first model information.
4. The method according to claim 1, characterized in that, The method further includes: Send a third message to the second communication device, the third message being used to request the second communication device to send a second model message to the third communication device, the second model message being one of the N model messages.
5. The method according to claim 4, characterized in that, The N model information pieces correspond to N time information pieces, and the second model information is the model information with the latest time information corresponding to the N model information pieces; or, The third information includes the index of the second model information.
6. The method according to claim 4 or 5, characterized in that, The sending of the third information includes: When the first condition is met, the third information is sent, wherein the first condition includes any one of the following: The model processing of the intermediate model corresponding to the first AI model is interrupted; Determine the current intermediate stage of the first AI model.
7. The method according to claim 6, characterized in that, Determining the current intermediate model phase of the first AI model includes: Receive indication information indicating the total number of model processing operations globally; When the difference between the number of local model processing times and the number of global model processing times meets the second condition, the current intermediate model period of the first AI model is determined.
8. The method according to any one of claims 4 to 7, characterized in that, The third information includes any of the following: The first indication information is used to indicate that the third information is triggered by a model processing interruption based on the intermediate model corresponding to the first AI model; The second indication information is used to indicate that the third information is triggered based on the current intermediate model of the first AI model.
9. The method according to any one of claims 1 to 8, characterized in that, The method further includes: Receive indication information for indicating the address of the second communication device.
10. The method according to claim 9, characterized in that, The address of the second communication device includes at least one of the following fields: The first field is used to indicate the geographical location of the second communication device; The second field is used to indicate the identifier of the bearer transmitting the first information; The third field is used to indicate the identifier of the first communication device.
11. The method according to any one of claims 1 to 10, characterized in that, The first communication device is a terminal device, and the second communication device is a Data Storage Function (DSF) entity.
12. A communication method, characterized in that, include: The system receives first information from a first communication device. The first information includes N model information, which are used to indicate information of N intermediate models obtained after the first AI model undergoes first local data processing, where N is a positive integer. The N intermediate model information is used to determine the second AI model. Store the first information.
13. The method according to claim 12, characterized in that, The method further includes: Receive second information from the first communication device, the second information being used to request first model information, the first model information being one of the N model information; Send the first model information.
14. The method according to claim 13, The N model information pieces correspond to N time information pieces respectively, and the first model information piece is the model information piece with the latest time information corresponding to the N model information pieces; or, The second information includes the index of the first model information.
15. The method according to claim 12, characterized in that, The method further includes: Send second model information to a third communication device, wherein the second model information is one of the N model information.
16. The method according to claim 15, characterized in that, Before sending the second model information to the third communication device, the method further includes: The system receives third information from the first communication device, the third information being used to request the second communication device to send first model information to the third communication device.
17. The method according to claim 15, characterized in that, Sending the second model information includes: Receive indication information indicating the total number of model processing operations globally; When the difference between the number of model processing times corresponding to the N model information and the number of model processing times globally satisfies the second condition, the second model information is sent.
18. The method according to any one of claims 15 to 17, characterized in that, The method further includes: Send an instruction to the first communication device to indicate that model processing of the current intermediate model corresponding to the first AI model should be stopped.
19. A communication method, characterized in that, include: Receive second model information, which is used to indicate the information of the intermediate model obtained after the first AI model has undergone local data processing by the first communication device; The first AI model is processed based on the second model information to obtain the first global model; Receive third model information, which is used to indicate the information of the intermediate model obtained after the first AI model has undergone local data processing by the fourth communication device; The first global model is processed based on the third model information to obtain a second global model; wherein the second global model is used to determine the second AI model.
20. A communication device, characterized in that, Includes a module for performing the method as described in any one of claims 1 to 19.
21. A communication device, characterized in that, It includes at least one processor, said at least one processor being used to perform the method as described in any one of claims 1 to 19.
22. The communication device according to claim 21, characterized in that, The communication device is a chip or chip system.
23. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program or instructions that, when executed by a communication device, implement the method as described in any one of claims 1 to 19.
24. A computer program product, characterized in that, It includes a computer program or instructions that, when executed by a computer, implement the method as described in any one of claims 1 to 19.