Communication method and related apparatus
By selecting a communication device with strong connectivity or large transmission capacity as the server in the distributed learning system, the problem of improving the performance of distributed learning is solved, and more efficient training accuracy and speed are achieved.
Patent Information
- Application Number
- PCT/CN2025/102425
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-29
- Filing Date
- 2025-06-20
- Publication Date
- 2026-01-02
AI Technical Summary
How to improve the performance of distributed learning systems, especially in AI systems involving multiple communication devices, and how to optimize the connectivity and transmission capacity of distributed training to improve training accuracy and speed.
The first communication device receives distributed training capability information from other communication devices, determines the distributed training server, and selects a communication device with strong connectivity or large transmission capacity as the server to improve the performance of distributed learning.
It improves the training accuracy and speed of distributed learning, enhances the transmission of training data and the utilization of computing resources, and improves the overall learning performance.
Smart Images

Figure CN2025102425_02012026_PF_FP_ABST
Abstract
Description
Communication method and related apparatus
[0001] This application claims priority to the Chinese Patent Application No. 202410875277.2, filed on June 29, 2024, and entitled "A communication method and related apparatus", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0002] The present application relates to the field of communication, and in particular to a communication method and related apparatus. BACKGROUND
[0003] With the development of communication technology, in a communication system, in addition to traditional communication services, the services performed by communication devices can also include other new services, such as artificial intelligence (AI) services.
[0004] Currently, a communication system capable of processing AI services can also be referred to as an AI system. One or more communication devices can act as participating nodes of the AI system, and the computing power and / or data of the one or more communication devices are applied to a certain link of the AI system. Generally speaking, an AI system in which two or more communication devices participate can be referred to as a distributed learning system.
[0005] However, in the distributed learning system, how to improve the performance of distributed learning is a technical problem to be solved. SUMMARY
[0006] The present application provides a communication method and related apparatus for improving the performance of distributed learning.
[0007] The first aspect of the present application provides a communication method, which is performed by a first communication device. The first communication device can be a communication device (such as a terminal device, a network device (for example, an access network device), or a central node for managing (or serving, or controlling) one or more network devices), or the first communication device can be a part of a communication device (for example, a circuit or a chip responsible for communication functions (such as a Modem chip, also known as a baseband chip, or a system on chip (SoC) chip or a system in package (SIP) chip containing a modem core)), or the first communication device can also be a logic module or software that can implement all or part of the functions of the communication device. In the method, the first communication device receives N pieces of information, the i-th piece of information in the N pieces of information indicates the distributed training capability information of the i-th communication device, N is a positive integer, and i takes a value from 1 to N; wherein the distributed training capability information of the i-th communication device is used to indicate the connectivity information and / or transmission capacity information of one or more communication devices connected to the i-th communication device; and the first communication device determines a distributed training server based on the N pieces of information.
[0008] Based on the above scheme, the N pieces of information received by the first communication device correspond to N communication devices respectively, that is, the i-th piece of information in the N pieces of information indicates the distributed training capability information of the i-th communication device in the N communication devices, and thereafter, the first communication device can determine a distributed training server based on the N pieces of information. Wherein the distributed training capability information of the i-th communication device is used to indicate the connectivity information and / or transmission capacity information of one or more communication devices connected to the i-th communication device. In this way, the first communication device can take the connectivity information and / or transmission capacity information of other communication devices connected by the N communication devices as the basis for determining the distributed training server, so that the distributed training server determined by the first communication device can obtain stronger connectivity and / or larger transmission capacity gain in the distributed learning process, to improve the performance of distributed learning.
[0009] For example, in the case where the distributed training capability information indicates the connectivity information, the first communication device can take a certain communication device (or a communication device corresponding to the certain communication device) with stronger connectivity as the distributed training server, so that the distributed training server can obtain more training data and more computing power resources in the distributed learning process, to improve the training accuracy and improve the training speed, and thus improve the performance of distributed learning.
[0010] For example, in the case that the distributed training capability information indicates the transmission capacity information, the first communication apparatus can select a certain communication apparatus (or a communication device corresponding to the certain communication apparatus) with a larger transmission capacity as a distributed training server, so that the distributed training server can transmit more data in the same time (and / or transmit the same amount of data in a shorter time) in the distributed learning process, so as to obtain more training data (and / or reduce the transmission delay), thereby improving the performance of the distributed learning.
[0011] In this application, the terms of distributed training, distributed learning, federated learning, federated training, etc. can be replaced with each other. Similarly, the terms of distributed training capability information, distributed learning capability information, federated learning capability information, federated training capability information, etc. can be replaced with each other.
[0012] In this application, the distributed training server can be a certain communication apparatus (or a communication device corresponding to a certain communication apparatus), which can be referred to as a node. Correspondingly, the server can also be referred to as a server node, a master node, a coordinator node, etc.
[0013] In addition, the distributed training client referred to hereinafter can be a certain communication apparatus (or a communication device corresponding to a certain communication apparatus), which can be referred to as a node. Correspondingly, the client can also be referred to as a slave node, a participant node, etc.
[0014] It should be understood that the first communication apparatus can receive the N pieces of information in various ways.
[0015] For example, the first communication apparatus can receive the N pieces of information from the N communication apparatuses through the communication links (or communication interfaces) between the first communication apparatus and the N communication apparatuses, respectively.
[0016] For another example, the first communication apparatus can communicate through one or more relay nodes, and the one or more relay nodes can be connected to part or all of the N communication apparatuses. In addition, part or all of the N communication apparatuses can send part or all of the N pieces of information to the first communication apparatus through the one or more relay nodes.
[0017] For another example, the first communication apparatus can receive the N pieces of information (which can be carried in one or more messages / signaling / information, etc.) through a communication link (or a communication interface) between the first communication apparatus and a center node. The center node is used to manage, control, or serve a plurality of communication apparatuses (including the N communication apparatuses).
[0018] Optionally, the connectivity information of the one or more communication devices connected to the i-th communication device comprises at least one of: a number of the one or more communication devices, identification information of the one or more communication devices, or computing power information of the one or more communication devices.
[0019] In a possible implementation of the first aspect, the distributed training server is one of the N communication devices, which is different from the first communication device.
[0020] Based on the above scheme, the first communication device can determine one of the N communication devices as the distributed training server based on the distributed training capability information of the N communication devices indicated by the N information respectively, so that the first communication device can assist the distributed learning task indicated (or required) by the first communication device through the communication device with stronger connectivity and / or larger transmission capacity as the distributed training server, to improve the performance of the distributed learning.
[0021] In a possible implementation of the first aspect, the first communication device determines the distributed training server based on the N information, comprising: the first communication device determines the distributed training server based on the N information and the distributed training capability information of the first communication device; wherein the distributed training server is one of the N communication devices and the first communication device.
[0022] It should be understood that the distributed training capability information of the first communication device can indicate the connectivity information and / or the transmission capacity information of the one or more communication devices connected to the first communication device. Similarly, the connectivity information of the one or more communication devices connected to the first communication device comprises at least one of: a number of the one or more communication devices, identification information of the one or more communication devices, or computing power information of the one or more communication devices.
[0023] Based on the above scheme, the first communication device can determine one of the N communication devices and the first communication device as the distributed training server based on the distributed training capability information of the N communication devices indicated by the N information respectively and the distributed training capability information of the first communication device itself, so that the first communication device can assist the distributed learning task indicated (or required) by the first communication device through the communication device with stronger connectivity and / or larger transmission capacity as the distributed training server, to improve the performance of the distributed learning.
[0024] Optionally, in the process that the first communication device determines the distributed training server based on the N information and the distributed training capability information of the first communication device, the first communication device can determine a communication device with stronger connectivity and / or larger transmission capacity among the N+1 communication devices (i.e., the N communication devices and the first communication device).
[0025] Alternatively, in the process that the first communication device determines the distributed training server based on the N information and the distributed training capability information of the first communication device, the first communication device can determine itself as the distributed training server in a case that the connectivity of the N communication devices are all lower than or equal to a threshold (and / or the transmission capacity of the N communication devices are all lower than or equal to a threshold).
[0026] In a possible implementation of the first aspect, before the first communication device receives the N information, the method further includes: the first communication device sends N request information, the N request information being used to request the N information.
[0027] Based on the above scheme, the first communication device can send the N request information to trigger the sending of the N information through the N request information, so that the first communication device can receive the N information to determine the distributed training server based on the N information.
[0028] It should be understood that the first communication device can send the N request information in various ways.
[0029] For example, the first communication device can send the N request information to the N communication devices respectively through the communication links (or communication interfaces) between the first communication device and the N communication devices.
[0030] For another example, the first communication device can connect one or more relay nodes for communication, and the one or more relay nodes can connect part or all of the N communication devices; and the first communication device can send part or all of the N request information to part or all of the N communication devices through the one or more relay nodes.
[0031] For another example, the first communication device can send the N information (which can be carried in one or more messages / signaling / information, etc.) through a communication link (or a communication interface) between the first communication device and a center node, and the center node can subsequently send the N information to the N communication devices. The center node is used to manage, control, or serve multiple communication devices (including the N communication devices).
[0032] Optionally, the N request information can be replaced by N indication information, the N indication information being used to indicate the reporting, querying, or providing of the N information.
[0033] In a possible implementation of the first aspect, before the first communication device sends the N request information, the method further includes: the first communication device sends M request information, the M request information are respectively used to request M communication devices to feed back whether the M communication devices support the capability of being a distributed training server, the N communication devices are included in the M communication devices, and M is greater than or equal to N; and the first communication device receives N capability information, the N capability information is used to indicate that the N communication devices (devices or devices corresponding to chips) support the capability of being a distributed training server.
[0034] Based on the above scheme, the first communication device can send M request information to trigger the sending of the N capability information through the M request information, so that the first communication device can receive the N capability information, so as to explicitly determine the N communication devices supporting the capability of being a distributed training server in the M communication devices.
[0035] In addition, the first communication device can obtain distributed training capability information corresponding to the N communication devices supporting the capability of being a distributed training server based on the N capability information. In the case of M greater than N, the first communication device can not need to obtain the distributed training capability information of the other M-N communication devices in the M communication devices except the N communication devices, so as to reduce the overhead.
[0036] It should be understood that the first communication device can send the M request information in various ways.
[0037] For example, the first communication device can send the M request information to the M communication devices respectively through a communication link (or a communication interface) between the first communication device and the M communication devices.
[0038] For another example, the first communication device can be connected to one or more relay nodes for communication, and the one or more relay nodes can be connected to part or all of the M communication devices. The first communication device can send part or all of the M request information to part or all of the M communication devices through the one or more relay nodes.
[0039] For another example, the first communication device can send the M information (the N information can be carried in one or more messages / signaling / information, etc.) through a communication link (or a communication interface) between the first communication device and a center node. Subsequently, the center node can send the M information to the M communication devices. The center node is used to manage, control, or serve a plurality of communication devices (including the M communication devices).
[0040] Optionally, the M pieces of request information can be replaced by M pieces of indication information, the M pieces of indication information being used to indicate whether the capability of serving as a distributed training server is supported or not.
[0041] In a possible implementation of the first aspect, part or all of the N communication devices support the capability of serving as a distributed training server.
[0042] Based on the above scheme, in the N communication devices, the first communication device can determine one of the communication devices as a distributed training server from part or all of the communication devices supporting the capability of serving as a distributed training server. Thus, the first communication device can quickly determine the distributed training server without considering the communication devices not supporting the capability of serving as a distributed training server, and the efficiency of determining the distributed training server can be improved.
[0043] In a possible implementation of the first aspect, the distributed training capability information of the ith communication device is further used to indicate at least one of the following: whether the ith communication device supports the capability of serving as a distributed training server, whether the ith communication device supports the capability of serving as a distributed training client, or the computing power information of the ith communication device.
[0044] Based on the above scheme, the distributed training capability information of the ith communication device is further used to indicate the at least one, so that the first communication device can determine the distributed training server based on more information, and the performance of the distributed training server in distributed learning can be improved.
[0045] Similarly, the distributed training capability information of the first communication device described in the foregoing is further used to indicate at least one of the following: whether the first communication device supports the capability of serving as a distributed training server, whether the first communication device supports the capability of serving as a distributed training client, or the computing power information of the first communication device.
[0046] In a possible implementation of the first aspect, the distributed training capability information of the ith communication device is associated with one or more distributed training use cases; or the distributed training capability information of the ith communication device is associated with any distributed training use case of the ith communication device; or the distributed training capability information of the ith communication device is associated with any distributed training use case supported by the ith communication device.
[0047] Based on the above scheme, the distributed training capability information of the ith communication device can be implemented in any of the above manners, so that the flexibility of implementing the scheme can be improved.
[0048] Similarly, the distributed training capability information of the first communication device described in the foregoing is associated with one or more distributed training use cases; or, the distributed training capability information of the first communication device is associated with any distributed training use case of the first communication device; or, the distributed training capability information of the first communication device is associated with any distributed training use case supported by the first communication device.
[0049] In a possible implementation of the first aspect, the first communication device determines the distributed training service end based on the N pieces of information, including that the first communication device determines the distributed training service end based on the N pieces of information in the case that any of the following conditions is met, including:
[0050] The first communication device determines to perform model training in a distributed training manner.
[0051] The first communication device receives first information, where the first information is used to indicate (request) that model training is performed in a distributed training manner; and the fourth information includes one or more of use case information of model training, quantity information of distributed training clients, training sample quantity information of distributed training, training time information of distributed training, iteration quantity information of distributed training, or target performance information of distributed training.
[0052] Based on the foregoing scheme, the first communication device can trigger determination of the distributed training service end based on the N pieces of information in the foregoing multiple manners, so as to improve flexibility of implementation of the scheme.
[0053] Optionally, the first communication device can determine to perform model training in a distributed training manner in multiple manners. For example, the first communication device determines that a training period of model A has arrived (for example, a model training period of the first communication device is x (x is a positive number) days / hours / minutes / seconds, so that the first communication device triggers model training every x days / hours / minutes / seconds), or the first communication device determines to perform model training in a distributed training manner because of occurrence of a certain specific event (for example, the first communication device detects that model performance of model A has decreased to a predetermined threshold, and then triggers model retraining). The model A can be a model processed by the first communication device, or can be a model processed by another communication device managed / supervised / scheduled by the first communication device, which is not limited herein.
[0054] In a possible implementation of the first aspect, the method further includes that the first communication device sends second information, where the second information is used to indicate the distributed training service end.
[0055] Based on the above scheme, after determining the distributed training service end, the first communication device can further indicate the distributed training service end through the sent second information, so that the receiving end of the second information can perform distributed learning through the distributed training service end based on the second information.
[0056] In a possible implementation manner of the first aspect, the one or more communication devices connected by the i-th communication device support (or have, or are capable of, or can) the capability of participating in distributed training as a distributed training client.
[0057] Based on the above scheme, the one or more communication devices connected by the i-th communication device support the capability of participating in distributed training as a distributed training client, so that in the process of determining the distributed training service end, the first communication device does not need to consider the connectivity information and / or transmission capacity information of the communication device that does not support the capability of participating in distributed training as a distributed training client, to quickly determine the distributed training service end, thereby improving the efficiency of determining the distributed training service end.
[0058] The second aspect of the present application provides a communication method, which is performed by a second communication device. The second communication device can be a communication device (such as a terminal device or a network device), or the second communication device can be a part of the communication device (for example, a circuit or a chip responsible for communication function (such as a Modem chip, also known as a baseband chip, or a SoC chip or a SIP chip containing a modem core, etc.), or the second communication device can also be a logic module or software capable of realizing all or part of the communication device function. In the method, the second communication device determines distributed training capability information, the distributed training capability information indicating connectivity information and / or transmission capacity information of one or more communication devices connected by the second communication device, and the distributed training capability information is used to determine a distributed training service end; and the second communication device sends third information, the third information indicating the distributed training capability information.
[0059] Based on the above scheme, the third information sent by the second communication device is used to indicate the connectivity information and / or transmission capacity information of the one or more communication devices connected by the second communication device, and thereafter, the receiving party (for example, the first communication device) of the third information can determine the distributed training service end based on the third information. In this way, the first communication device can take the connectivity information and / or transmission capacity information of the other communication devices connected by the one or more communication devices as the basis for determining the distributed training service end, so that the distributed training service end determined by the first communication device can obtain stronger connectivity and / or larger transmission capacity gain in the distributed learning process, thereby improving the performance of distributed learning.
[0060] For example, in the case that the distributed training capability information indicates connectivity information, the first communication apparatus can select a certain communication apparatus (or a communication device corresponding to the certain communication apparatus) with stronger connectivity as a distributed training server, so that the distributed training server can obtain more training data and more computing resources in the distributed learning process, to improve training accuracy and training speed, and thus improve the performance of the distributed learning.
[0061] For example, in the case that the distributed training capability information indicates connectivity information, the first communication apparatus can select a certain communication apparatus (or a communication device corresponding to the certain communication apparatus) with stronger connectivity as a distributed training server, so that the distributed training server can obtain more training data and more computing resources in the distributed learning process, to improve training accuracy and training speed, and thus improve the performance of the distributed learning.
[0062] It should be understood that the third information is one of the N information in the first aspect described above. In other words, the number of the second communication apparatuses can be N, and the N second communication apparatuses are the N communication apparatuses described above, or the second communication apparatus is one of the N communication apparatuses described above.
[0063] Optionally, the connectivity information includes at least one of the following: the number of the one or more communication apparatuses, the identification information of the one or more communication apparatuses, or the computing power information of the one or more communication apparatuses.
[0064] In a possible implementation of the second aspect, before the second communication apparatus sends the third information, the method further includes: the second communication apparatus receives first request information, the first request information being used to request the third information.
[0065] Based on the above scheme, the second communication apparatus can receive the first request information from the first communication apparatus, so as to trigger the sending of the third information through the first request information, so that the first communication apparatus can receive the N information, so as to determine the distributed training server based on the N information.
[0066] It should be understood that the first request information is one of the N request information in the first aspect described above.
[0067] Optionally, the request information can be replaced by indication information, the indication information being used to indicate reporting, querying, or providing the third information.
[0068] In a possible implementation of the second aspect, before the second communication device receives the first request information, the method further includes: receiving, by the second communication device, second request information, the second request information being used to request the second communication device to feed back whether the second communication device supports the capability of being a distributed training server; and sending, by the second communication device, first capability information, the first capability information being used to indicate that the second communication device (or a device corresponding to the chip) supports the capability of being a distributed training server.
[0069] Based on the above scheme, the second communication device can receive the second request information from the first communication device, so as to trigger the sending of the first capability information through the second request information, so that the first communication device can receive the first capability information, so as to explicitly determine the second communication device supporting the capability of being a distributed training server in the one or more communication devices.
[0070] It should be understood that the second request information is one of the M request information in the first aspect.
[0071] Optionally, the request information can be replaced by indication information, the indication information being used to indicate whether the second communication device supports the capability of being a distributed training server.
[0072] In a possible implementation of the second aspect, the distributed training capability information of the second communication device is further used to indicate at least one of the following: whether the second communication device supports the capability of being a distributed training server, whether the second communication device supports the capability of being a distributed training client, or computing power information of the second communication device.
[0073] Based on the above scheme, the distributed training capability information of the second communication device is further used to indicate the at least one, so that the first communication device can determine the distributed training server based on more information, so as to improve the performance of the distributed training server in the distributed learning.
[0074] In a possible implementation of the second aspect, the distributed training capability information of the second communication device is associated with one or more distributed training use cases; or the distributed training capability information of the second communication device is associated with any distributed training use case of the second communication device; or the distributed training capability information of the second communication device is associated with any distributed training use case supported by the second communication device.
[0075] Based on the above scheme, the distributed training capability information of the second communication device can be implemented in any of the above manners, so as to improve the flexibility of the implementation of the scheme.
[0076] In a possible implementation manner of the second aspect, the method further includes: receiving, by the second communication device, second information, the second information being used to indicate the distributed training server.
[0077] Based on the above scheme, after determining the distributed training server, the first communication device can further indicate the distributed training server to the second communication device through the transmitted second information, so that the second communication device can perform distributed learning based on the second information through the distributed training server.
[0078] In a possible implementation manner of the second aspect, the one or more communication devices connected to the second communication device support (have, are capable of, can) the capability of participating in distributed training as a distributed training client.
[0079] Based on the above scheme, the one or more communication devices connected to the second communication device support the capability of participating in distributed training as a distributed training client, so that the first communication device does not need to consider the connectivity information and / or the transmission capacity information of the communication device that does not support the capability of participating in distributed training as a distributed training client in the process of determining the distributed training server, thereby quickly determining the distributed training server and improving the efficiency of determining the distributed training server.
[0080] The third aspect of the present application provides a communication device, which is a first communication device, and the device includes a transceiver unit and a processing unit; the transceiver unit is configured to receive N pieces of information, the i th piece of information in the N pieces of information indicating distributed training capability information of an i th communication device, N being a positive integer, and i being an integer from 1 to N; wherein the distributed training capability information of the i th communication device is used to indicate connectivity information and / or transmission capacity information of one or more communication devices connected to the i th communication device; and the processing unit is configured to determine a distributed training server based on the N pieces of information.
[0081] In the third aspect of the present application, the component modules of the communication device can also be configured to perform the steps performed in the various possible implementation manners of the first aspect and achieve the corresponding technical effects, which can be referred to the first aspect for details and will not be described herein.
[0082] The fourth aspect of the present application provides a communication device, which is a second communication device, and the device includes a transceiver unit and a processing unit; the processing unit is configured to determine distributed training capability information, the distributed training capability information indicating connectivity information and / or transmission capacity information of one or more communication devices connected to the second communication device, the distributed training capability information being used to determine a distributed training server; and the transceiver unit is configured to transmit third information, the third information indicating the distributed training capability information.
[0083] In the fourth aspect of the present application, the constituent module of the communication device can also be used to perform the steps performed in the various possible implementation manners of the second aspect and achieve the corresponding technical effects. For details, please refer to the second aspect, which will not be described here again.
[0084] The fifth aspect of the present application provides a communication device, comprising at least one processor coupled with a memory; the memory is used to store programs or instructions; the at least one processor is used to execute the programs or instructions, so that the device implements the method of any one of the possible implementation manners of the first aspect to the second aspect. Optionally, the communication device can include the memory.
[0085] The sixth aspect of the present application provides a communication device, comprising at least one logic circuit and an input / output interface; the logic circuit is used to execute the method of any one of the possible implementation manners of the first aspect to the second aspect.
[0086] The seventh aspect of the present application provides a communication system, comprising the first communication device and the second communication device.
[0087] The eighth aspect of the present application provides a computer readable storage medium, which is used to store one or more computer execution instructions; when the computer execution instructions are executed by a processor, the processor executes the method of any one of the possible implementation manners of the first aspect to the second aspect.
[0088] The ninth aspect of the present application provides a computer program product (or computer program), when the computer program in the computer program product is executed by the processor, the processor executes the method of any one of the possible implementation manners of the first aspect to the second aspect.
[0089] The tenth aspect of the present application provides a chip or chip system, comprising at least one processor, which is used to support the communication device to implement the method of any one of the possible implementation manners of the first aspect to the second aspect. For example, the chip can be a baseband chip, a modem chip, a SoC chip (such as a SoC chip containing a modem core), a SIP chip, or a communication module, etc.
[0090] In a possible design, the chip or chip system can also include a memory, which is used to save the necessary program instructions and data of the communication device. The chip system can be composed of a chip, or can include a chip and other discrete devices. Optionally, the chip system also includes an interface circuit, which provides program instructions and / or data for the at least one processor.
[0091] The technical effects brought by any one of the designs of the third aspect to the tenth aspect can refer to the technical effects brought by the different designs of the first aspect to the second aspect, which will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS
[0092] FIGS. 1a to 1c are schematic diagrams of a communication system provided by the present application;
[0093] FIGS. 2a to 2h are schematic diagrams of an AI processing process related to the present application;
[0094] FIG. 3 is an interaction schematic diagram of a communication method provided by the present application;
[0095] FIGS. 4a to 4j are some schematic diagrams of application scenarios of the communication method provided by the present application;
[0096] FIGS. 5 to 9 are schematic diagrams of communication devices provided by the present application. DETAILED DESCRIPTION
[0097] First, some terms in the embodiments of the present application are explained to facilitate understanding by those skilled in the art.
[0098] (1) Terminal device: can be a wireless terminal device capable of receiving network device scheduling and instruction information, the wireless terminal device can be a device that provides voice and / or data connectivity to a user, or a handheld device with wireless connection function, or other processing devices connected to a wireless modem.
[0099] A terminal device can communicate with one or more core networks or the Internet via a radio access network (RAN), and the terminal device can be a mobile terminal device, such as a mobile phone (or called "cellular" phone, mobile phone), a computer, and a data card, for example, which can be a portable, pocket, hand-held, computer- built-in, or vehicle-mounted mobile device that exchanges voice and / or data with a radio access network. For example, a personal communication service (PCS) phone, a cordless phone, a session initiation protocol (SIP) phone, a wireless local loop (WLL) station, a personal digital assistant (PDA), a tablet, a computer with wireless transceiver, and the like. The wireless terminal device can also be referred to as a system, a subscriber unit, a subscriber station, a mobile station (MS), a remote station, an access point (AP), a remote terminal, an access terminal, a user terminal, a user agent, a subscriber station (SS), customer premises equipment (CPE), a terminal, user equipment (UE), a mobile terminal (MT), and the like.
[0100] By way of example and not limitation, in embodiments of the present application, the terminal device can also be a wearable device. The wearable device can also be referred to as a smart wearable device or a smart wearable device, etc., which is a general term for devices that can be designed and developed by applying wearable technology to daily wear, such as glasses, gloves, watches, clothing, and shoes, etc. The wearable device is a portable device that can be directly worn on the body or integrated into the user's clothes or accessories. The wearable device is not only a hardware device, but also a powerful function realized through software support and data interaction, cloud interaction. The general wearable smart device includes a full function, large size, and can realize complete or partial functions without relying on a smart phone, such as smart watches or smart glasses, etc., and focuses on a certain application function, and needs to cooperate with other devices such as a smart phone, such as various smart wristbands, smart helmets, smart jewelry, etc. for monitoring vital signs.
[0101] The terminal can also be a drone, a robot, a terminal in device-to-device (D2D) communication, a terminal in vehicle to everything (V2X), a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a wireless terminal in industrial control, a wireless terminal in self driving, a wireless terminal in telemedicine or telehealth services, a wireless terminal in smart grid, a wireless terminal in transportation safety, a wireless terminal in smart city, a wireless terminal in smart home, etc.
[0102] In addition, the terminal device can also be a terminal device in a communication system evolved after the 5th generation (5G) communication system (such as 5G Advanced or 6th generation (6G) communication system, etc.) or a terminal device in a future evolved public land mobile network (PLMN), etc. For example, 5G Advanced or 6G network can further expand the form and function of 5G communication terminal, and 6G terminal includes but is not limited to vehicle, cellular network terminal (integrating satellite terminal function), drone, internet of things (IoT) device.
[0103] In the embodiments of the present application, the terminal device can also obtain an artificial intelligence (AI) service provided by the network device. Optionally, the terminal device can also have AI processing capability.
[0104] (2) Network device: can be a device in a wireless network, for example, the network device can be a RAN node (or device) for accessing the terminal device to the wireless network, which can also be referred to as a base station. At present, some examples of RAN devices are: base station (base station, BS), evolved NodeB (eNodeB), base station gNB (gNodeB) in 5G communication system, transmission reception point (transmission reception point, TRP), evolved Node B (eNB), radio network controller (radio network controller, RNC), Node B (Node B, NB), home base station (for example, home evolved Node B, or home Node B, HNB), baseband unit (baseband unit, BBU) or wireless fidelity (wireless fidelity, Wi-Fi) access point (AP) and the like. In addition, in one network structure, the network device can include a central unit (central unit, CU) node, or a distributed unit (distributed unit, DU) node, or a RAN device including a CU node and a DU node.
[0105] Optionally, the RAN node can also be a macro base station, a micro base station or an indoor station, a relay node or a donor node, or a wireless controller in a cloud radio access network (cloud radio access network, CRAN) scenario. The RAN node can also be a server, a wearable device, a vehicle or a vehicle-mounted device, etc. For example, the access network device in vehicle external connection (V2X) technology can be a road side unit (road side unit, RSU).
[0106] In another possible scenario, multiple RAN nodes cooperate to assist a terminal to implement wireless access, and different RAN nodes respectively implement part of functions of a base station. For example, a RAN node can be a CU, a DU, a CU-control plane (CP), a CU-user plane (UP), or a radio unit (RU), etc. The CU and the DU can be separately configured, or can also be included in the same network element, for example, in a baseband unit (BBU). The RU can be included in a radio frequency device or a radio frequency unit, for example, included in a remote radio unit (RRU), an active antenna unit (AAU), a radio head (RH), or a remote radio head (RRH).
[0107] In different systems, the CU (or CU-CP and CU-UP), the DU, or the RU can also have different names, but those skilled in the art can understand their meanings. For example, in an open RAN (O-RAN or ORAN) system, the CU can also be referred to as an O-CU (open CU), the DU can also be referred to as an O-DU, the CU-CP can also be referred to as an O-CU-CP, the CU-UP can also be referred to as an O-CU-UP, and the RU can also be referred to as an O-RU. For the convenience of description, the CU, the CU-CP, the CU-UP, the DU, and the RU are taken as examples for description in this application. Any one of the CU (or the CU-CP, the CU-UP), the DU, and the RU in this application can be implemented by a software module, a hardware module, or a combination of a software module and a hardware module.
[0108] The communication between the access network device and the terminal device follows a certain protocol layer structure. The protocol layer can include a control plane protocol layer and a user plane protocol layer. The control plane protocol layer can include at least one of the following: a radio resource control (RRC) layer, a packet data convergence protocol (PDCP) layer, a radio link control (RLC) layer, a media access control (MAC) layer, or a physical (PHY) layer, etc. The user plane protocol layer can include at least one of the following: a service data adaptation protocol (SDAP) layer, a PDCP layer, an RLC layer, a MAC layer, or a physical layer, etc.
[0109] For the correspondence between the network elements in the ORAN system and the protocol layer functions that can be implemented by the network elements, refer to Table 1 below.
[0110] Table 1
[0111] The network device can be another device that provides a wireless communication function for the terminal device. Embodiments of the present application do not limit the specific technology and specific device form adopted by the network device. For the convenience of description, embodiments of the present application do not limit.
[0112] The network device can also include a core network device, for example, a mobility management entity (MME) in a fourth generation (4G) network, a home subscriber server (HSS), a serving gateway (S-GW), a policy and charging rules function (PCRF), a public data network gateway (PDN gateway or P-GW), a network element such as an access and mobility management function (AMF) in a 5G network, a user plane function (UPF), or a session management function (SMF). In addition, the core network device can also include other core network devices in the 5G network and the next generation network of the 5G network.
[0113] In the embodiments of the present application, the network device described above can also be an AI-capable network node, which can provide AI services for terminals or other network devices, for example, AI nodes, computing power nodes, AI-capable RAN nodes, AI-capable core network elements, etc. on the network side (access network or core network).
[0114] In the embodiments of the present application, the device for implementing the function of the network device can be a network device or a device capable of supporting the network device to implement the function, such as a chip system, which can be arranged in the network device. In the technical solutions provided in the embodiments of the present application, the device for implementing the function of the network device is taken as an example to describe the technical solutions provided in the embodiments of the present application.
[0115] (3) Configuration and pre-configuration: in this application, both configuration and pre-configuration will be used. Among them, configuration refers to that the network device / server sends some parameter configuration information or parameter values to the terminal through messages or signaling, so that the terminal determines the communication parameters or resource in transmission according to the values or information. Pre-configuration is similar to configuration, which can be parameter information or parameter values agreed by the network device / server and the terminal device in advance, or parameter information or parameter values adopted by the base station / network device or the terminal device according to the standard protocol, or parameter information or parameter values pre-stored in the base station / server or the terminal device. This application does not limit this.
[0116] Further, these values and parameters can be changed or updated.
[0117] (4) The terms "system" and "network" in the embodiments of the present application can be used interchangeably. "Multiple" means two or more. "And / or" describes the association between the associated objects, which means that there can be three relationships, for example, A and / or B, which means that A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects. "At least one of the following" or similar expressions means any combination of these items, including any combination of single item or multiple items. For example, "at least one of A, B and C" includes A, B, C, AB, AC, BC or ABC. In addition, unless otherwise specified, the ordinal numbers "first", "second", etc. mentioned in the embodiments of the present application are used to distinguish multiple objects, and are not used to limit the order, time sequence, priority or importance of the multiple objects.
[0118] (5) In the embodiments of the present application, "sending" and "receiving" represent the direction of signal transmission. For example, "sending information to XX" can be understood as that the destination of the information is XX, which can include direct sending through the air interface, or indirect sending through the air interface by other units or modules. "Receiving information from YY" can be understood as that the source of the information is YY, which can include direct receiving from YY through the air interface, or indirect receiving from YY through the air interface by other units or modules. "Sending" can also be understood as the "output" of the chip interface, and "receiving" can also be understood as the "input" of the chip interface.
[0119] In other words, sending and receiving can be carried out between devices, such as between network devices and terminal devices, or within devices, such as between components, modules, chips, software modules or hardware modules within a device through buses, wires or interfaces.
[0120] It can be understood that the information can be processed, such as encoding and modulation, between the source end and the destination end of the information transmission, but the destination end can understand the effective information from the source end. Similar expressions in this application can be similarly understood, and will not be repeated here.
[0121] (6) In the embodiments of the present application, “indication” can include direct indication and indirect indication, and can also include explicit indication and implicit indication. The information indicated by certain information (indication information described below) is referred to as to-be-indicated information. In the implementation process, there are many ways to indicate the to-be-indicated information, for example, but not limited to, the to-be-indicated information can be directly indicated, such as the to-be-indicated information itself or the index of the to-be-indicated information. The to-be-indicated information can also be indirectly indicated by indicating other information, where the other information and the to-be-indicated information have an association relationship. The to-be-indicated information can also be indicated only by a part, and the other part of the to-be-indicated information is known or agreed in advance. For example, the indication of a specific information can be achieved by means of the arrangement order of each information agreed in advance (for example, protocol predefined), thereby reducing the indication overhead to a certain extent. The present application does not limit the specific manner of indication. It can be understood that for the sender of the indication information, the indication information can be used to indicate the to-be-indicated information, and for the receiver of the indication information, the indication information can be used to determine the to-be-indicated information.
[0122] In the present application, the same or similar parts of each embodiment can be mutually referenced unless otherwise specified. In the present application, the terms and / or descriptions of different embodiments and the methods / designs / implementation manners in each embodiment are consistent and can be mutually referenced if not specially specified and there is no logical conflict. The technical features of different embodiments and the methods / designs / implementation manners in each embodiment can be combined to form new embodiments, methods, or implementation manners according to their inherent logical relationship. The implementation manners of the present application described below do not constitute a limitation on the protection scope of the present application.
[0123] The present application can be applied to a long term evolution (LTE) system, a new radio (NR) system, or a communication system evolved after 5G (such as 6G, etc.). The communication system includes at least one network device and / or at least one terminal device.
[0124] Please refer to FIG. 1a, which is a schematic diagram of a communication system in the present application. In FIG. 1a, a network device and six terminal devices are exemplarily shown, which are terminal device 1, terminal device 2, terminal device 3, terminal device 4, terminal device 5 and terminal device 6. In the example shown in FIG. 1a, the terminal device 1 is exemplarily taken as a smart tea cup, the terminal device 2 is exemplarily taken as a smart air conditioner, the terminal device 3 is exemplarily taken as a smart gas station, the terminal device 4 is exemplarily taken as a vehicle, the terminal device 5 is exemplarily taken as a mobile phone, and the terminal device 6 is exemplarily taken as a printer.
[0125] As shown in FIG. 1a, the sending entity of the AI configuration information can be the network device. The receiving entity of the AI configuration information can be the terminal devices 1-6. In this case, the network device and the terminal devices 1-6 form a communication system, in which the terminal devices 1-6 can send data to the network device, and the network device receives the data sent by the terminal devices 1-6. The network device can send configuration information to the terminal devices 1-6.
[0126] Exemplarily, in FIG. 1a, the terminal devices 4-6 can also form a communication system. Among them, the terminal device 5 acts as a network device, i.e., the sending entity of the AI configuration information; the terminal devices 4 and 6 act as terminal devices, i.e., the receiving entity of the AI configuration information. For example, in a vehicle-to-everything system, the terminal device 5 sends AI configuration information to the terminal devices 4 and 6, and receives data sent by the terminal devices 4 and 6; correspondingly, the terminal devices 4 and 6 receive the AI configuration information sent by the terminal device 5, and send data to the terminal device 5.
[0127] Taking the communication system shown in FIG. 1a as an example, different devices (including network devices, network devices and terminal devices, and / or terminal devices) can perform AI-related services in addition to performing communication-related services.
[0128] As shown in FIG. 1b, taking a network device as a base station as an example, the base station can perform communication-related services and AI-related services with one or more terminal devices, and different terminal devices can also perform communication-related services and AI-related services.
[0129] As shown in FIG. 1c, taking a terminal device including a television and a mobile phone as an example, the television and the mobile phone can also perform communication-related services and AI-related services.
[0130] The technical solutions provided in the present application can be applied to a wireless communication system (for example, the system shown in FIG. 1a, FIG. 1b or FIG. 1c), for example, an AI network element can be introduced in the communication system provided in the present application to implement part or all of the AI related operations. The AI network element can also be referred to as an AI node, an AI device, an AI entity, an AI module, an AI model, or an AI unit, etc. The AI network element can be built-in in a network element of the communication system. For example, the AI network element can be an AI module built-in in an access network device, a core network device, a cloud server, or an operation, administration and maintenance (OAM) to implement AI related functions. The OAM can be a network management of the core network device and / or a network management of the access network device. Alternatively, the AI network element can also be a network element independently arranged in the communication system. Optionally, an AI entity can also be included in a terminal or a chip built-in in the terminal to implement AI related functions.
[0131] Optionally, in the communication system, the AI application cases can include but are not limited to: channel status information (CSI) feedback enhancement, beam management enhancement, positioning accuracy enhancement, network energy saving, load balancing, and mobility optimization. The following will be described respectively.
[0132] 1. CSI feedback enhancement
[0133] CSI is the channel property of the communication link, and is the channel quality information reported by the terminal device to the network device. The terminal device reports the channel quality information to the network device, so as to select a suitable modulation and coding scheme (MCS) for the terminal device, so that the wireless channel can be adapted to the change. For example, the terminal device performs channel estimation according to the received channel state information-reference signal (CSI-RS), and then feeds back the channel quality information to the network device. The information is used as the input of the model of the network device, so that the network device can implement AI model training. By applying AI to CSI feedback enhancement, the overhead can be reduced, the accuracy can be improved, and prediction can be realized.
[0134] CSI-RS feedback enhancement can include at least one sub-function, such as: CSI compression, CSI prediction, and CSI-RS configuration signaling reduction, respectively. CSI compression can be further divided into CSI compression in at least one of spatial, time, and frequency domains.
[0135] 2. Beam management enhancement
[0136] BM is mainly to find the strongest transmit / receive beam pair. Based on AI sparse beam prediction, accuracy can be improved. According to AI training and inference, it can be divided into network side AI sparse beam prediction and terminal device side AI sparse beam prediction. Taking the terminal device side AI sparse beam prediction as an example, the pre-trained AI model of the terminal device side can be delivered by the network side or pre-stored by the terminal device side. In the training phase, the network device scans all possible beams, and then the network reports the transmit beam pattern to the terminal device. When the model training is completed, the network only needs to scan a small part of the beam, and then the terminal device feeds back the inference result to the network. Based on AI beam management, beam prediction in time and / or spatial domain can be realized to reduce overhead and delay and improve beam selection accuracy.
[0137] Beam management enhancement can include at least one sub-function, such as: beam scanning matrix prediction, optimal beam prediction, respectively.
[0138] 3. Positioning enhancement
[0139] In line of sight (LOS) or non-line of sight (NLOS) scenarios, AI-based positioning can improve positioning accuracy with a smaller number of TRP antennas. Positioning enhancement can include at least one sub-function, such as: access network device-based positioning enhancement, positioning management function network element-based positioning enhancement, and terminal device-based positioning enhancement, respectively.
[0140] 4. Network energy saving
[0141] Network energy saving can be achieved through cell activation / deactivation, load reduction, improved coverage, or other RAN setting adjustments. AI technology can be used to optimize energy saving decisions by utilizing data collected in the RAN network. AI algorithms can predict the energy efficiency and load status of the next period, which can be used to assist in decision-making for cell activation / deactivation to save energy. Based on the predicted load, the system can dynamically configure energy saving strategies to maintain a balance between system performance and energy efficiency and reduce energy consumption.
[0142] 5. Load balancing
[0143] Load balancing can make the load evenly distributed among cells and among areas within a cell, or divert part of the traffic from congested cells, or split users among cells, carriers or access technologies to improve network performance. AI model based load balancing can provide higher quality user experience and improve system capacity.
[0144] 6. Mobility management
[0145] Mobility management is a solution to ensure service continuity during terminal device movement by minimizing dropped calls, radio link failure (RLF), unnecessary handover and ping-pong effect. AI can enhance mobility management, such as reducing the probability of unexpected events, predicting terminal device location / mobility / performance, and traffic steering, etc.
[0146] It should be understood that the definitions of the above technical terms are only examples. For example, as technology continues to evolve, the scope of the above definitions can also change, and the embodiments of the present application are not limited.
[0147] For example, an AI function can include multiple AI sub-functions.
[0148] Optionally, the AI application case is also referred to as an AI application scenario or an AI function.
[0149] As described above, AI can be widely used in CSI feedback enhancement, beam management, positioning accuracy enhancement, energy saving, mobility enhancement, load balancing and other aspects to improve network performance. AI models can be deployed on the network side and / or the terminal device side, and the training of AI models depends on the collection of training data, which can come from terminal device measurement and feedback.
[0150] The concepts that can be involved in the present application will be briefly introduced below.
[0151] AI can give machines human intelligence, for example, machines can use computer hardware and software to simulate some intelligent behaviors of humans. To achieve artificial intelligence, machine learning methods can be used. In machine learning methods, machines learn (or train) models using training data. The model represents the mapping between input and output. The learned model can be used for inference (or prediction), i.e., the model can be used to predict the output corresponding to a given input. The output can also be referred to as the inference result (or prediction result).
[0152] Machine learning can include supervised learning, unsupervised learning, and reinforcement learning. Among them, unsupervised learning can also be referred to as non-supervised learning.
[0153] Supervised learning learns the mapping relationship from sample values to sample labels according to the collected sample values and sample labels, and uses an AI model to express the learned mapping relationship. The process of training a machine learning model is the process of learning such a mapping relationship. In the training process, the sample values are input into the model to obtain the predicted values of the model, and the model parameters are optimized by calculating the error between the predicted values of the model and the sample labels (ideal values). After the mapping relationship is learned, the learned mapping can be used to predict new sample labels. The learned mapping relationship of supervised learning can include linear mapping or nonlinear mapping. According to the type of label, the learned task can be divided into classification task and regression task.
[0154] Unsupervised learning uses algorithms to discover the internal patterns of samples according to the collected sample values. In unsupervised learning, a class of algorithms uses the sample itself as a supervision signal, that is, the model learns the mapping relationship from the sample to the sample, which is called self-supervised learning. In training, the model parameters are optimized by calculating the error between the predicted values of the model and the sample itself. Self-supervised learning can be used for signal compression and decompression recovery applications. Common algorithms include autoencoders and generative adversarial networks.
[0155] Reinforcement learning is different from supervised learning, and is a class of algorithms that learn strategies to solve problems by interacting with the environment. Unlike supervised and unsupervised learning, reinforcement learning problems do not have clear "correct" action label data. The algorithm needs to interact with the environment to obtain the reward signal of the environment feedback, and then adjust the decision action to obtain a larger reward signal value. In the following power control, the reinforcement learning model adjusts the downlink transmission power of each user according to the system total throughput rate feedback by the wireless network, and then expects to obtain a higher system throughput rate. The goal of reinforcement learning is also to learn the mapping relationship between the environment state and the optimal (for example, the optimal) decision action. However, because the "correct action" label cannot be obtained in advance, the network cannot be optimized by calculating the error between the action and the "correct action". Reinforcement learning training is achieved through iterative interaction with the environment.
[0156] Neural network (NN) is a specific model in machine learning technology. According to the universal approximation theorem, neural network can theoretically approximate any continuous function, so that neural network has the ability to learn any mapping. Traditional communication systems need to use rich expert knowledge to design communication modules, while deep learning communication systems based on neural network can automatically discover the implicit pattern structure from a large amount of data set, establish the mapping relationship between data, and obtain better performance than traditional modeling methods.
[0157] The idea behind neural networks comes from the neuronal structure of the brain. For example, each neuron performs a weighted summation of its input values and outputs the result through an activation function.
[0158] Figure 2a shows a schematic diagram of a neuron structure. Assume the input to the neuron is x = [x0, x1, ..., x...]. n The weights corresponding to each input are w = [w0, w1, ..., w] n ], where n is a positive integer, w i and x i It can be any possible type, such as a decimal, an integer (e.g., 0, a positive integer, or a negative integer), or a complex number. i As x i The weights are used to assign weights to x. i Weighting is applied. The bias for the weighted sum of the input values is, for example, b. Activation functions can take many forms. Suppose the activation function of a neuron is: y = f(z) = max(0, z), then the output of that neuron is: For example, if the activation function of a neuron is y = f(z) = z, then the output of that neuron is: Here, b can be any possible type, such as a decimal, an integer (e.g., 0, a positive integer, or a negative integer), or a complex number. The activation functions of different neurons in a neural network can be the same or different.
[0159] Furthermore, neural networks generally consist of multiple layers, each of which may include one or more neurons. Increasing the depth and / or width of a neural network can improve its expressive power, providing more powerful information extraction and abstract modeling capabilities for complex systems. The depth of a neural network can refer to the number of layers it includes, and the number of neurons in each layer can be called the width of that layer. In one implementation, a neural network includes an input layer and an output layer. The input layer processes the received input information through neurons and passes the processing result to the output layer, which then obtains the output of the neural network. In another implementation, a neural network includes an input layer, hidden layers, and an output layer. The input layer processes the received input information through neurons and passes the processing result to the hidden layer. The hidden layer calculates the received processing result and passes the calculation result to the output layer or the next adjacent hidden layer, ultimately obtaining the output of the neural network. A neural network may include one hidden layer or multiple sequentially connected hidden layers, without limitation.
[0160] The neural network is, for example, a deep neural network (DNN). According to the construction manner of the network, the DNN can include a feedforward neural network (FNN), a convolutional neural network (CNN) and a recurrent neural network (RNN).
[0161] Fig. 2b is a schematic diagram of a FNN network. The FNN network is characterized by that the neurons in adjacent layers are fully connected to each other. This feature makes the FNN usually need a large amount of storage space and lead to a high computational complexity.
[0162] The CNN is a neural network specially designed to process data with a similar grid structure. For example, time series data (e.g. time axis discrete sampling) and image data (e.g. two-dimensional discrete sampling) can be considered as data with a similar grid structure. The CNN does not use all the input information for operation at one time, but uses a fixed size window to extract part of the information for convolution operation, which greatly reduces the calculation amount of model parameters. In addition, according to the different types of information extracted by the window (such as people and objects in the same image are different types of information), each window can use different convolution kernel operations, which makes the CNN better extract the features of the input data.
[0163] The RNN is a kind of neural network that uses feedback time series information. The input of the RNN includes the new input value at the current time and the output value of itself at the previous time. The RNN is suitable for obtaining sequence features with temporal correlation, such as speech recognition, channel coding and decoding applications.
[0164] In the above model training process of machine learning, a loss function can be defined. The loss function describes the gap or difference between the output value of the model and the ideal target value. The loss function can be embodied in various forms, and the specific form of the loss function is not limited. The model training process can be regarded as the following process: by adjusting part or all of the parameters of the model, the value of the loss function is less than the threshold value or meets the target demand.
[0165] The model can also be referred to as an AI model, a rule, or other names, etc. The AI model can be considered as a specific method to implement an AI function. The AI model represents a mapping relationship or a function between the input and the output of the model. The AI function can include one or more of the following: data collection, model training (or model learning), model information publishing, model inference (or model reasoning, reasoning, or prediction, etc.), model monitoring or model verification, or inference result publishing, etc. The AI function can also be referred to as an AI (related) operation, or an AI-related function.
[0166] The implementation process of the neural network will be described below with reference to the accompanying drawings.
[0167] 1. Fully connected neural network, also known as multilayer perceptron (MLP).
[0168] As shown in FIG. 2c, an MLP includes an input layer (left side), an output layer (right side), and multiple hidden layers (middle). Each layer of the MLP includes a number of nodes, referred to as neurons. The neurons of adjacent two layers are connected to each other.
[0169] Optionally, considering the neurons of adjacent two layers, the output h of the neuron of the next layer is the weighted sum of all the neurons x of the previous layer connected to it and is processed by an activation function, which can be represented as: h = f(wx + b).
[0170] where w is a weight matrix, b is a bias vector, and f is an activation function.
[0171] Further optionally, the output of the neural network can be recursively expressed as: y = f z (w z f z-1 (…)+b z ).
[0172] where z is the index of the layer of the neural network, z is greater than or equal to 1, and z is less than or equal to Z, where Z is the total number of layers of the neural network.
[0173] In other words, the neural network can be understood as a mapping relationship from a set of input data to a set of output data. Usually, the neural network is randomly initialized, and the process of obtaining this mapping relationship from the random w and b with the existing data is called training of the neural network.
[0174] Optionally, the specific way of training is to evaluate the output result of the neural network by using a loss function.
[0175] As shown in Figure 2d, the error can be backpropagated, and the neural network parameters (including w and b) can be iteratively optimized using gradient descent until the output of the loss function reaches its minimum value, which is the "better point (e.g., the optimal point)" in Figure 2d. It can be understood that the neural network parameters corresponding to the "better point (e.g., the optimal point)" in Figure 2d can be used as the neural network parameters in the trained AI model information.
[0176] Alternatively, the gradient descent process can be represented as:
[0177] Where θ represents the parameters to be optimized (including w and b), L is the loss function, and η is the learning rate, controlling the step size of gradient descent. This represents the differentiation operation. This indicates taking the derivative of θ with respect to L.
[0178] Alternatively, the backpropagation process can utilize the chain rule for partial derivatives.
[0179] As shown in Figure 2e, the gradient of the parameters in the previous layer can be recursively calculated from the gradient of the parameters in the next layer, and can be expressed as:
[0180] Among them, w ij Let s be the weight of the connection between node j and node i. i The weighted sum of the inputs at node i.
[0181] 2. Distributed learning, also known as distributed training, federated learning (FL), federated training, etc.
[0182] The concept of distributed learning effectively addresses the current challenges in the development of artificial intelligence. While fully protecting user data privacy and security, it enables various edge devices and central servers to collaborate efficiently to complete the model's learning task.
[0183] As an implementation example, as shown in Figure 2f, the FL architecture is a widely used training architecture in the current field of distributed learning. The FedAvg algorithm is the basic algorithm of FL, and the algorithm flow of FedAvg is roughly as follows:
[0184] (1) Initialize the model to be trained at the center end. And broadcast it to all clients.
[0185] (2) In the t∈[1,T] round, the client k∈[1,K] is based on the local dataset. For the received global model Perform E epochs of training to obtain the local training results. The local training results sent by the distributed nodes n, k, and m are respectively denoted as G n , G k , and G m .
[0186] (3) The central node collects the local training results from all (or part) of the clients, assuming that the set of clients uploading the local model in the t-th round is The central end will obtain a new global model by weighted averaging with the sample number of the corresponding client as the weight, and the specific updating rule is Then the central end broadcasts the latest version of the global model to all clients for a new round of training.
[0187] (4) Repeat steps (2) and (3) until the model converges or the training round reaches the upper limit.
[0188] Optionally, in addition to reporting the local model , the client can also report the trained local gradient , and the central node averages all the local gradients reported by the clients and updates the global model according to the average gradient.
[0189] As can be seen, in the FL framework, the data set exists in the distributed node (such as the client), that is, the distributed node collects the local data set and performs local training, and reports the local result (model or gradient) obtained by training to the central node. The central node itself can not have a data set, and can be responsible for fusing the training results of the distributed nodes to obtain a global model and issuing it to the distributed nodes.
[0190] The technical scheme provided in the present application can be applied in a communication system (such as the system shown in FIG. 1a or FIG. 1b or FIG. 1c), in which the communication nodes generally have signal transceiving capability and computing capability. Taking a network device with computing capability as an example, the computing capability of the network device is mainly to provide computing power support for the signal transceiving capability (for example: to perform sending processing and receiving processing on the signal), so as to realize the communication task of the network device and other communication nodes.
[0191] With the development of communication technology, in a communication system, in addition to the traditional communication service, the service performed by the communication device can also include other new services, such as artificial intelligence (AI) service. At present, the communication system capable of processing AI service can also be called AI system. Among them, one or more communication devices can act as a participant node of the AI system, and the computing power and / or data of the one or more communication devices are applied to a certain link of the AI system. Generally speaking, the AI system participated by two or more communication devices can be called a distributed learning system.
[0192] However, in the distributed learning system, how to improve the performance of distributed learning is a technical problem to be solved.
[0193] In a possible implementation, one way to improve the performance of distributed learning can be to control or schedule a plurality of client nodes by a server node of the distributed learning, so as to realize the cooperation between a plurality of distributed nodes. By utilizing the cooperation of the computing power and / or data of a plurality of nodes, the performance of distributed learning can be improved.
[0194] As shown in the example of FIG. 2g, the participants of the distributed learning include one server node and a plurality of client nodes, which can all be communication devices, such as terminal devices and / or network devices. Among them, the process of the distributed learning is as follows:
[0195] Step 1, initial model distribution: the server node distributes the initial model parameters to each client node, which can be represented as the initial global parameters in the figure.
[0196] Step 2, parallel training: each client node trains the model based on local data to obtain local model parameters (for example, gradient, step length), and reports the local model parameters to the server node.
[0197] Step 3, parameter update: the server node calculates new global model parameters (for example, gradient, step length) according to the local model parameters reported by each node, and sends the new global model parameters to each client node.
[0198] Step 4, iterative optimization: steps 2 and 3 are repeated until the model converges or reaches a predetermined number of iterations.
[0199] As shown in the example of FIG. 2h, taking the case where both the server node and the client node are base stations in a wireless network as an example, the base stations can be used to participate in training of the "user location prediction model". In the wireless network, distributed learning can be performed among multiple base stations. For example, multiple client base stations (i.e., base stations acting as distributed learning clients) each collect location information of users under the base station, and then perform distributed learning under coordination of a server base station (i.e., a base station acting as a distributed learning server) to jointly complete training of the "user location prediction model". During training, each client base station performs local model training based on local data (i.e., the location information of users under the base station collected by the client base station itself).
[0200] Generally, before performing distributed learning, the server base station and the client base stations need to be determined, i.e., "server base station selection" and "client base station selection" need to be performed. Taking the scenario shown in FIG. 2h as an example, during server base station selection, the algorithm capability of the base station is mainly considered, i.e., whether the base station has a relevant algorithm (e.g., a global model update algorithm) for acting as a server. If a base station has the above algorithm capability, the base station can be selected as a server base station. However, if the server base station is selected only based on the algorithm capability, the selected server base station can not meet the requirements of distributed learning. For example, the selected server base station can not be able to connect to enough client base stations, resulting in too little data available for training and affecting training accuracy. For another example, the interaction delay between the server base station and the client base stations is large, resulting in too long overall time of the training process. Therefore, how to select a suitable server base station to improve the performance of distributed learning is a technical problem to be solved.
[0201] To solve the above problem, the present application provides a communication method and related apparatus, which will be described in detail below with reference to the accompanying drawings.
[0202] Please refer to FIG. 3, which is an implementation schematic diagram of the communication method provided by the present application. The method includes the following steps.
[0203] It should be noted that in the following, the first communication apparatus and other communication apparatuses (e.g., N communication apparatuses) in FIG. 3 are taken as an example to illustrate the execution subject of the interaction, but the present application does not limit the execution subject of the interaction. For example, the communication apparatus can be a communication device (e.g., a terminal device or a network device), or a chip, a baseband chip, a modem chip, an SoC chip (such as an SoC chip containing a modem core), a SIP chip, a communication module, a chip system, a processor, a logic module or software in the communication device, etc.
[0204] As an example, the first communication apparatus and the N communication apparatuses can each be a network device (e.g., an access network device).
[0205] As another example, the first communication device can be a terminal device, and part or all of the N communication devices can be network devices. Optionally, if there are other parts of the N communication devices that are not network devices, the other parts can be terminal devices.
[0206] As another example, the first communication device can be a network device, and part or all of the N communication devices can be terminal devices. Optionally, if there are other parts of the N communication devices that are not terminal devices, the other parts can be network devices.
[0207] As another example, the first communication device and the N communication devices are all terminal devices, i.e., the scheme shown in FIG. 3 can be applied to a sidelink communication scenario.
[0208] Optionally, in any of the above examples, the first communication device can be a center node in distributed learning, where the center node is used to manage, control, or serve a plurality of communication devices (including the N communication devices shown in FIG. 3). For example, the center node can be an access network device (such as a master base station), a core network device, a network management device, a radio intelligent unit (RIU), or a radio access network intelligent controller (RIC), etc.
[0209] Optionally, in any of the above examples, the network device can be a device in an O-RAN architecture. For example, referring to FIG. 4a, an O-RAN network diagram is shown, including one or more network elements: a Non-Real-Time Radio access network Intelligent Controller (Non-RT RIC), a Near-Real-Time Radio access network Intelligent Controller (Near-RT RIC), an O-RAN Central Unit Control Plane (O-CU-CP), an O-RAN Central Unit User Plane (O-CU-UP), an O-RAN Distributed Unit (O-DU), and an O-RAN evolved NodeB (O-eNB). In this case, the network device can be the O-CU-CP, the O-DU, or the O-eNB, and the central node can be a network management device, a core network device, a Non-RT RIC, a Near-RT RIC, etc.
[0210] S301. The N communication devices send N pieces of information, and the first communication device receives the N pieces of information. The i-th piece of information in the N pieces of information indicates the distributed training capability information of the i-th communication device, N is a positive integer, and i takes a value from 1 to N. The distributed training capability information of the i-th communication device is used to indicate the connectivity information and / or the transmission capacity information of one or more communication devices connected to the i-th communication device.
[0211] S302. The first communication device determines a distributed training server based on the N pieces of information.
[0212] In this application, the terms distributed training, distributed learning, federated learning, and federated training can be replaced with each other. Similarly, the terms distributed training capability information, distributed learning capability information, federated learning capability information, and federated training capability information can be replaced with each other.
[0213] In this application, the distributed training server can be a certain communication device (or a communication device corresponding to a certain communication device), which can be referred to as a node. Correspondingly, the server can also be referred to as a server node, a master node, a coordinator node, etc.
[0214] Further, the distributed training client to be referred to later can be a certain communication device (or a communication device corresponding to a certain communication device), which can be referred to as a node. Correspondingly, the client can also be referred to as a slave node, a participant node, etc.
[0215] It should be understood that the first communication device can receive the N information in various ways.
[0216] For example, the first communication device can receive the N information from the N communication devices respectively through the communication links (or communication interfaces) between the first communication device and the N communication devices.
[0217] For another example, the first communication device can communicate with one or more relay nodes, and the one or more relay nodes can be connected to part or all of the N communication devices; and part or all of the N communication devices can send part or all of the N information to the first communication device through the one or more relay nodes.
[0218] For another example, the first communication device can receive the N information (which can be carried in one or more messages / signals / information, etc.) through a communication link (or a communication interface) between the first communication device and a center node. The center node is used to manage, control, or serve a plurality of communication devices (including the N communication devices).
[0219] Optionally, the connectivity information of the one or more communication devices connected to the i-th communication device includes at least one of the following: the number of the one or more communication devices, the identification information of the one or more communication devices, or the computing power information of the one or more communication devices.
[0220] In a possible implementation, before the first communication device receives the N information in step S301, the method further includes: the first communication device sends N request information, and the N request information is used to request the N information. Specifically, the first communication device can send the N request information to trigger the sending of the N information through the N request information, so that the first communication device can receive the N information, so as to determine the distributed training server based on the N information.
[0221] It should be understood that the first communication device can send the N request information in various ways.
[0222] For example, the first communication device can send the N request information to the N communication devices respectively through the communication links (or communication interfaces) between the first communication device and the N communication devices.
[0223] For example, the first communication device can send the N request information to the N communication devices respectively through the communication links (or communication interfaces) between the first communication device and the N communication devices.For another example, the first communication device can connect one or more relay nodes for communication, the one or more relay nodes can connect part or all of the N communication devices; and the first communication device can send part or all of the N request information to part or all of the N communication devices through the one or more relay nodes.
[0224] For another example, the first communication device can send the N information (which can be carried in one or more messages / signaling / information, etc.) through a communication link (or communication interface) with a center node, and the center node can subsequently send the N information to the N communication devices. The center node is used to manage, control, or serve multiple communication devices (including the N communication devices).
[0225] Optionally, the N request information can be replaced by N indication information, which is used to indicate the reporting, querying, or providing of the N information.
[0226] Based on the scheme shown in FIG. 3, the N information received by the first communication device in step S301 respectively corresponds to the N communication devices, i.e., the i-th information in the N information indicates the distributed training capability information of the i-th communication device in the N communication devices. Thereafter, the first communication device can determine a distributed training server based on the N information in step S302. The distributed training capability information of the i-th communication device is used to indicate the connectivity information and / or transmission capacity information of one or more communication devices connected by the i-th communication device. In this way, the first communication device can use the connectivity information and / or transmission capacity information of other communication devices connected by the N communication devices as a basis for determining the distributed training server, so that the distributed training server determined by the first communication device can obtain stronger connectivity and / or greater transmission capacity gain in the distributed learning process, thereby improving the performance of distributed learning.
[0227] For example, in the case where the distributed training capability information indicates the connectivity information, the first communication device can use a certain communication device with stronger connectivity (or a communication device corresponding to the certain communication device) as the distributed training server, so that the distributed training server can obtain more training data and more computing power resources in the distributed learning process, thereby improving the training accuracy and the training speed, and further improving the performance of distributed learning.
[0228] For example, in the case where the distributed training capability information indicates the transmission capacity information, the first communication device can select a certain communication device (or a communication device corresponding to the certain communication device) with a larger transmission capacity as a distributed training server, so that the distributed training server can transmit more data in the same time (and / or transmit the same amount of data in a shorter time) in the distributed learning process, so as to obtain more training data (and / or reduce the transmission delay), thereby improving the performance of the distributed learning.
[0229] As an application example, taking the scenario shown in FIG. 4b as an example, the N communication devices can include four BSs, i.e., BS1, BS2, BS3 and BS4 in FIG. 4b, and the first communication device is a node responsible for server base station selection, which can be one of the four BSs or a central node. Based on the method shown in FIG. 3, in step S301, the first communication device can obtain the connectivity information (for example, the number of potential client base stations connected by the base station) and the transmission capacity information (for example, the transmission capacity between the base station and the potential client base station) of each base station, and in step S302, select a suitable server base station from the above base stations according to the information, so that the available training data is as much as possible and the information interaction delay is as small as possible, thereby improving the performance of the distributed learning (improving the model accuracy and reducing the model training time).
[0230] As shown in the example of FIG. 4c, it is assumed that BS1 and BS2 can serve as server base stations (with server-related algorithms). Next, according to the connectivity information and the transmission capacity information, it is determined which of BS1 and BS2 is more suitable as a server base station.
[0231] For the connectivity information: BS1 is connected to 3 potential client base stations (BS2, BS3 and BS4), and BS2 is connected to 2 potential client base stations (BS1 and BS3). Therefore, if BS1 is selected as a server base station, more client base stations will join the distributed learning, and more client base stations mean more training data, thereby improving the accuracy of model training.
[0232] For the transmission capacity information: the transmission load between BS1 and each potential client base station is low, which means that BS1 and its client base stations can complete the model information interaction in a shorter time; while the transmission load between BS2 and some client base stations (BS3) is high, which means that BS2 and these client base stations have a large delay in information interaction, resulting in a longer training time.
[0233] From the above two points, it can be seen that BS1 is more suitable as a server base station, which can improve the performance of the distributed learning as much as possible.
[0234] In a possible implementation, the distributed training server determined by the first communication device in step S302 can be implemented in various ways, which will be described below in connection with Example One and Example Two.
[0235] Example One, the distributed training server is one of the N communication devices, which is different from the first communication device.
[0236] In Example One, the first communication device can determine one of the N communication devices as the distributed training server based on the distributed training capability information of the N communication devices respectively indicated by the N information, so that the first communication device can use the communication device with stronger connectivity and / or larger transmission capacity as the distributed training server to assist the distributed learning task indicated (or required) by the first communication device, thereby improving the performance of the distributed learning.
[0237] Example Two, the first communication device determines the distributed training server based on the N information, including: the first communication device determines the distributed training server based on the N information and the distributed training capability information of the first communication device; wherein the distributed training server is one of the N communication devices and the first communication device.
[0238] In Example Two, the first communication device can determine one of the N communication devices and the first communication device as the distributed training server based on the distributed training capability information of the N communication devices respectively indicated by the N information and the distributed training capability information of the first communication device, so that the first communication device can use the communication device with stronger connectivity and / or larger transmission capacity as the distributed training server to assist the distributed learning task indicated (or required) by the first communication device, thereby improving the performance of the distributed learning.
[0239] It should be understood that the distributed training capability information of the first communication device can indicate the connectivity information and / or transmission capacity information of one or more communication devices connected to the first communication device. Similarly, the connectivity information of the one or more communication devices connected to the first communication device includes at least one of: the number of the one or more communication devices, the identification information of the one or more communication devices, or the computing power information of the one or more communication devices.
[0240] Optionally, in Example Two, in the process of determining the distributed training server by the first communication device based on the N information and the distributed training capability information of the first communication device, the first communication device can determine the communication device with stronger connectivity and / or larger transmission capacity among the N+1 communication devices (i.e., the N communication devices and the first communication device).
[0241] Alternatively, in Example II, in a process in which the first communication device determines the distributed training server based on the N information and the distributed training capability information of the first communication device, the first communication device can determine itself as the distributed training server in a case where the connectivity of the N communication devices are all lower than or equal to a threshold value (and / or the transmission capacity of the N communication devices are all lower than or equal to a threshold value).
[0242] In a possible implementation, in step S302, the first communication device can determine one of the N communication devices as the distributed training server based on the N information. In the N communication devices, all the communication devices can have the capability to support as the distributed training server, or some of the communication devices can have the capability to support as the distributed training server. The following describes the two cases in combination with Example A and Example B.
[0243] Example A. All the N communication devices have the capability to support as the distributed training server.
[0244] In Example A, before the first communication device sends the N request information, the method further includes: the first communication device sends M request information, the M request information are respectively used to request M communication devices to feed back whether the M communication devices have the capability to support as the distributed training server, the N communication devices are included in the M communication devices, and M is greater than or equal to N; and the first communication device receives N capability information, the N capability information is used to indicate that the N communication devices (devices or devices corresponding to chips) have the capability to support as the distributed training server. Thus, the first communication device can send the M request information to trigger the sending of the N capability information through the M request information, so that the first communication device can receive the N capability information, so as to explicitly determine the N communication devices in the M communication devices that have the capability to support as the distributed training server.
[0245] In addition, the first communication device can obtain the distributed training capability information corresponding to the N communication devices that have the capability to support as the distributed training server based on the N capability information. In a case where M is greater than N, the first communication device can not need to obtain the distributed training capability information of the other M-N communication devices in the M communication devices except for the N communication devices, so as to reduce the overhead.
[0246] It should be understood that the first communication device can send the M request information in various ways.
[0247] For example, the first communication device can send the M request information to the M communication devices respectively through the communication links (or communication interfaces) between the first communication device and the M communication devices.
[0248] For example, the first communication device can connect one or more relay nodes for communication, and the one or more relay nodes can connect part or all of the M communication devices; and the first communication device can send part or all of the M request information to part or all of the M communication devices through the one or more relay nodes.
[0249] For example, the first communication device can send the M information (the N information can be carried in one or more messages / signaling / information, etc.) through the communication link (or communication interface) between the first communication device and the center node, and the center node can send the M information to the M communication devices. The center node is used to manage, control, or serve multiple communication devices (including the M communication devices).
[0250] Optionally, the M request information can be replaced by M indication information, and the M indication information is used to indicate whether the support for the capability of serving as a distributed training server is supported.
[0251] Example B. Part or all of the N communication devices have the capability of supporting the distributed training server.
[0252] In example B, among the N communication devices, the first communication device can determine one of the communication devices having the capability of supporting the distributed training server as the distributed training server. Thus, the first communication device can quickly determine the distributed training server without considering the communication devices not having the capability of supporting the distributed training server, and the efficiency of determining the distributed training server can be improved.
[0253] In a possible implementation, as described in step S301, the distributed training capability information of the ith communication device is used to indicate the connectivity information and / or the transmission capacity information of one or more communication devices connected by the ith communication device, and the distributed training capability information of the ith communication device is also used to indicate at least one of the following: whether the ith communication device supports the capability of serving as a distributed training server, whether the ith communication device supports the capability of serving as a distributed training client, or the computing power information of the ith communication device. Specifically, the distributed training capability information of the ith communication device is also used to indicate the at least one, so that the first communication device can determine the distributed training server based on more information, and the performance of the distributed training server in the distributed learning can be improved.
[0254] Similarly, the distributed training capability information of the first communication device described in the foregoing can also be used to indicate at least one of the following: whether the first communication device supports the capability of being a distributed training server, whether the first communication device supports the capability of being a distributed training client, or the computing power information of the first communication device.
[0255] In a possible implementation, the distributed training capability information of the ith communication device is associated with one or more distributed training use cases; or the distributed training capability information of the ith communication device is associated with any distributed training use case of the ith communication device; or the distributed training capability information of the ith communication device is associated with any distributed training use case supported by the ith communication device. Specifically, the distributed training capability information of the ith communication device can be implemented in any of the above manners to improve the flexibility of the scheme implementation.
[0256] Similarly, the distributed training capability information of the first communication device described in the foregoing is associated with one or more distributed training use cases; or the distributed training capability information of the first communication device is associated with any distributed training use case of the first communication device; or the distributed training capability information of the first communication device is associated with any distributed training use case supported by the first communication device.
[0257] For example, the distributed training use cases described above can be used for CSI feedback enhancement, beam management, positioning accuracy enhancement, energy saving, mobility enhancement, load balancing, and the like to improve network performance. For another example, the distributed training use cases described above can be used for other applications, such as transmission and / or enhancement of audio / video, data processing related to large language models, and the like.
[0258] In a possible implementation, the method further includes: the first communication device sending second information, the second information being used to indicate the distributed training server. Specifically, after determining the distributed training server, the first communication device can also indicate the distributed training server through the sent second information, so that the receiving end of the second information can perform distributed learning through the distributed training server based on the second information.
[0259] In a possible implementation, the one or more communication devices connected to the i-th communication device support (or have, or are capable of, or can) the capability of participating in the distributed training as a distributed training client. Specifically, the one or more communication devices connected to the i-th communication device support the capability of participating in the distributed training as a distributed training client, so that the first communication device does not need to consider the connectivity information and / or the transmission capacity information of the communication devices that do not support the capability of participating in the distributed training as a distributed training client in determining the distributed training server, thereby improving the efficiency of determining the distributed training server.
[0260] In a possible implementation, the process in which the first communication device determines the distributed training server based on the N pieces of information in step S302 includes: when any of the following conditions is met, the first communication device determines the distributed training server based on the N pieces of information, including:
[0261] The first communication device determines to perform model training in a distributed training manner.
[0262] The first communication device receives first information, where the first information is used to indicate (request) that the model training is performed in a distributed training manner; and the fourth information includes one or more of the following: use case information of the model training, quantity information of the distributed training clients, training sample quantity information of the distributed training, training time information of the distributed training, iteration quantity information of the distributed training, or target performance information of the distributed training.
[0263] Therefore, the first communication device can trigger the determination of the distributed training server based on the N pieces of information in the above-mentioned multiple manners, thereby improving the flexibility of the implementation of the scheme.
[0264] Optionally, the first communication device can determine to perform model training in a distributed training manner in multiple manners. For example, the first communication device determines that the training period of model A has arrived (for example, the model training period of the first communication device is x (x is a positive number) days / hours / minutes / seconds, and the first communication device triggers model training every x days / hours / minutes / seconds), or the model training is triggered because of a certain specific event (for example, the first communication device detects that the model performance of model A has decreased to a predetermined threshold, and then triggers model retraining). The model A can be a model processed by the first communication device, or a model processed by other communication devices managed / supervised / scheduled by the first communication device, which is not limited here.
[0265] From the above process, it can be known that the first communication device can trigger the process of determining the distributed training service end based on the N pieces of information in multiple ways, and the process of distributed training can also be triggered in multiple ways. The following will be described in combination with some embodiments based on the scenario shown in FIG. 4b.
[0266] Embodiment one: whether to perform distributed learning is determined by BS1, and the base station selected as the distributed training service end is also selected by BS1.
[0267] As shown in FIG. 4d, the process of embodiment one is as follows. It should be understood that in the implementation shown in FIG. 4d, BS1 is an implementation example of the first communication device shown in the foregoing, BS2, BS3 and BS4 are an implementation example of the N (N=3) communication devices shown in the foregoing. In addition, the following step 4 is an implementation example of the foregoing step S301, and the following step 5 is an implementation example of the foregoing step S302.
[0268] Step 1: BS1 determines to perform distributed training.
[0269] For example, the distributed training can be for a specific use case, such as distributed training for a “user location prediction” use case, distributed training for “downlink power allocation”, and the like. BS1 determines to perform model training, which can be because the model training period has arrived (for example, the model training period of BS1 is 12 hours, so BS1 triggers model training once every 12 hours), or because a certain specific event occurs (for example, BS1 detects that the model performance has dropped to a predetermined threshold, and then triggers model retraining). When BS1 determines to perform model training, it needs to further determine whether to use centralized training or distributed training. As an example, BS1 can determine to use distributed training in the following cases:
[0270] Case 1: BS1 lacks local data and needs to train in combination with the data of other base stations, but at this time the transmission rate of other base stations is limited, which causes a large amount of data to be unable to be transmitted to BS1, so the distributed training method is used.
[0271] Case 2: BS1 lacks local data and needs to train in combination with the data of other base stations, but the data of other base stations contains privacy information, which causes the data to be unable to be transmitted to BS1, so the distributed training method is used.
[0272] Step 2: BS1 sends a distributed training capability information request to other base stations (for example, BS2, BS3, BS4), to request other base stations to feed back their respective distributed training capability information. Here, the “other base stations” should be base stations that have an interface with BS1. The request contains at least one of the following information:
[0273] Information of use case, for indicating which use case or which use cases the distributed training is for. This information can indicate the name of the use case (e.g. "user location prediction"), or the identity of the use case (e.g. use case 10), or the identity of the model the use case is associated with (e.g. model 3), etc.
[0274] Type of capability information, for indicating which capability information about the distributed training the feedback is requested for. For example, the following at least one capability information can be contained:
[0275] 1. Whether the base station can serve as a server of the distributed training.
[0276] 2. Whether the base station can serve as a client of the distributed training.
[0277] 3. The computing power information of the base station.
[0278] 4. The connectivity information of the base station.
[0279] 5. The transmission capacity information of the base station.
[0280] Each of the above capability information can or can not correspond to a specific use case. For example, capability information 1, 2, 4, 5 can correspond to a specific use case (the use case specified by the "information of use case" or the use case supported by the base station), and capability information 3 can not correspond to a specific use case. For example, Table 2 gives an example of the distributed training capability information request.
[0281] Table 2
[0282] Note: The above 5 kinds of capability information can correspond to a specific use case (i.e. use cases 1 and 2), or can not correspond to a specific use case. Whether each capability information corresponds to a specific use case can be specified in the communication protocol or pre-configured by the network.
[0283] Step 3, other base stations (e.g. BS2, BS3, BS4) determine their respective distributed training capability information. For details, see the "determining distributed training capability information" sub-process in the subsequent paragraphs.
[0284] Step 4, other base stations (e.g. BS2, BS3, BS4) send the distributed training capability information to BS1. The information contains the following at least one:
[0285] Indication information 1, for indicating whether the base station can serve as a server of the distributed training. Example: the indication information is of Boolean type, and when the value is "True", it means that the base station can serve as a server of the distributed training, and when the value is "False" or the indication information is not contained, it means that the base station cannot serve as a server of the distributed training.
[0286] An indication information 2, used to indicate whether the base station can be a client of distributed training. Example: the indication information is of Boolean type, and when the value is "True", it means that the base station can be a client of distributed training, and when the value is "False" or the indication information is not included, it means that the base station cannot be a client of distributed training.
[0287] Computing power information, used to indicate the computing power of the base station. The information can include at least one of the following: the total computing power of the base station, the computing power already used by the base station, the computing power not used by the base station, the proportion of computing power already used by the base station, and the proportion of computing power not used by the base station. The computing power can be the number of floating point operations per second (FLOPS) or the number of instructions per second (IPS), and the proportion of computing power can be a percentage value or an interval of percentage values.
[0288] Connectivity information, used to indicate the base stations connected to the base station and capable of being clients of distributed training. The information includes at least one of the following: 1) the number of base stations connected to the base station and capable of being clients of distributed training; 2) the identity of each base station connected to the base station and capable of being a client of distributed training; 3) the computing power information of each base station connected to the base station and capable of being a client of distributed training. The computing power information can include at least one of the following: the total computing power of the base station, the computing power already used by the base station, the computing power not used by the base station, the proportion of computing power already used by the base station, and the proportion of computing power not used by the base station; and 4) whether each base station connected to the base station and capable of being a client of distributed training is currently performing distributed training.
[0289] Transmission capacity information, used to indicate the transmission capacity between the base station and the base stations connected to the base station and capable of being clients of distributed training. The information includes at least one of the following: 1) the average, maximum, and / or minimum transmission capacity between the base station and the base stations connected to the base station and capable of being clients of distributed training; and 2) the transmission capacity between the base station and each base station connected to the base station and capable of being a client of distributed training.
[0290] Note: The "transmission capacity" between base station A and base station B can include at least one of the following: the total uplink (or downlink, or uplink plus downlink) transmission capacity size, the already used uplink (or downlink, or uplink plus downlink) transmission capacity size, the unused uplink (or downlink, or uplink plus downlink) transmission capacity size, the proportion of already used uplink (or downlink, or uplink plus downlink) transmission capacity, and the proportion of unused uplink (or downlink, or uplink plus downlink) transmission capacity. The uplink refers to transmission from base station A to base station B, and the downlink refers to transmission from base station B to base station A. The capacity size can be the number of bits per second (bps) or the number of bytes per second (Bps). The capacity proportion can be a percentage value or a percentage value interval.
[0291] Each of the above information (indication information 1, indication information 2, computing power information, connectivity information, and transmission capacity information) can correspond to a certain use case (a certain use case among all use cases specified by the "use case information" or a certain use case among all use cases supported by the base station), all use cases (all use cases specified by the "use case information" or all use cases supported by the base station), or no specific use case. If a certain information corresponds to a certain use case, the "distributed training capability information" can include the information of the use case to which the certain information corresponds (for example, the name or identifier of the use case, etc.).
[0292] Table 3 gives an example of distributed training capability information.
[0293] Table 3
[0294] For steps 1-4, the following is explained: 1. Step 2 is an optional step, i.e., step 4 can not depend on step 2 (i.e., base stations 2, 3, and 4 can actively report their distributed training capability information to BS1). 2. Steps 2-4 can also occur before step 1.
[0295] Step 5: BS1 selects a suitable base station as the server of distributed training according to the distributed training capability information of other base stations (e.g., BS2, BS3, and BS4). For example: BS1 first excludes base stations that cannot be used as distributed training servers (i.e., base stations with indication information 1 taking the value "False"), then excludes base stations with transmission capacity less than a predetermined threshold, and finally selects the base station with the best connectivity among the remaining base stations (e.g., the base station with the most connected client base stations).
[0296] Step 6, BS1 sends a distributed training instruction to the selected base station (e.g. BS2) as the distributed training server, which contains the distributed training parameters, such as: the identification of the use case, the training completion time, the maximum number of iterations in training, the minimum number of data samples used in training, etc.
[0297] Step 7, the base station (e.g. BS2) as the distributed training server performs distributed training, including steps such as client base station selection, model distribution, model parameter aggregation, etc.
[0298] As an example, the flow shown in Figure 4e is a sub-flow of step 3 "determine distributed training capability information" shown in Figure 4d. For other base stations (e.g. BS2), the sub-flow contains the following steps.
[0299] Step 1, the surrounding base stations (e.g. BS x, BS y, BS z) send capability and state information to BS2. The information contains at least one of the following:
[0300] Indication information for indicating whether the base station can be a client of distributed training. Example: the indication information is of Boolean type, and when the value is "True", it means that the base station can be a client of distributed training, and when the value is "False" or the indication information is not included, it means that the base station cannot be a client of distributed training.
[0301] Computing power information for indicating the computing power information of the base station. The information can contain at least one of the following: the total computing power size of the base station, the computing power size already used by the base station, the computing power size not used by the base station, the proportion of computing power already used by the base station, the proportion of computing power not used by the base station. Wherein, the computing power size can be the number of floating point operations per second (FLOPS) or the number of instructions per second (IPS), etc., and the computing power proportion can be a percentage value or a percentage value interval, etc.
[0302] Transmission capacity information for indicating the transmission capacity information of the base station to BS2. The information contains at least one of the following: total transmission capacity size, already used transmission capacity size, unused transmission capacity size, already used transmission capacity proportion, unused transmission capacity proportion. Wherein, the capacity size can be the number of bits per second (bps) or the number of bytes per second (Bps), etc.; the capacity proportion can be a percentage value or a percentage value interval, etc.
[0303] Each of the above information (indication information, computing power information, transmission capacity information) can correspond to a certain use case, can correspond to all use cases, or can not correspond to a specific use case. If a certain item of information corresponds to a certain use case, the information of the use case to which the item of information corresponds (for example, the name or identifier of the use case, etc.) can be included in the "distributed training capability information".
[0304] Step 2, BS2 determines its own distributed training capability information according to the above information. For example, the number of base stations that can be used as clients for distributed training in the surrounding base stations is counted as the "connectivity information" of itself; the capacity transmitted by these client base stations to itself is counted, and the capacity transmitted by itself to these client base stations is added as the "transmission capacity information" of itself.
[0305] Optionally, in the embodiment I shown in FIG. 4d, other base stations (for example, BS2, BS3, BS4) send distributed training capability information to BS1, which includes 5 items of information: indication information 1, indication information 2, computing power information, connectivity information, and transmission capacity information (see step 4). In the embodiment I shown in FIG. 4d, the 5 items of information are "together" sent to BS1. In fact, the 5 items of information can also be "separately" or "in batches" sent to BS1. Here is an example: other base stations first send "indication information 1 (i.e., whether the base station can be used as a server for distributed training)" to BS1, and BS1 can select a number of base stations as candidate server base stations according to the information; then BS1 obtains the other 4 items of information from these candidate server base stations, and determines the final server base stations according to the information.
[0306] As an example, as shown in the flowchart of FIG. 4f, the flow of this example is as follows. It should be understood that in the implementation shown in FIG. 4f, BS1 is an implementation example of the first communication device shown in the foregoing, BS2, BS3, and BS4 are an implementation example of N (N = 3) communication devices shown in the foregoing. In addition, the following step 7 is an implementation example of the foregoing step S301, and the following step 8 is an implementation example of the foregoing step S302.
[0307] Step 1, same as step 1 of the embodiment shown in FIG. 4d.
[0308] Step 2, BS1 requests other base stations (for example, BS2, BS3, BS4) to feedback their respective distributed training capability information A, which is used to indicate whether the base station can be used as a server for distributed training. Optionally, the request can include information of a use case, which is used to indicate which use case or which use cases the above distributed training is for.
[0309] Step 3, other base stations (e.g. BS2, BS3, BS4) send distributed training capability information A to BS1. The information contains indication information 1, which indicates whether the base station can serve as a server of distributed training. Example: the indication information is of Boolean type, and takes the value "True" when it can serve as a server of distributed training, and takes the value "False" or does not contain the indication information when it cannot serve as a server of distributed training. The indication information can correspond to a certain use case, or to all use cases, or not to a specific use case. If the indication information corresponds to a certain use case, the "distributed training capability information A" can contain information of the use case (e.g. the name or identifier of the use case, etc.).
[0310] Note: Step 2 is an optional step, i.e. step 3 can not depend on step 2 (i.e. base stations 2, 3, 4 can actively report their distributed training capability information A to BS1). In addition, steps 2-3 can also occur before step 1.
[0311] Step 4, BS1 determines one or more base stations (e.g. BS2, BS3) as candidate servers of distributed training according to the above information. Specifically, BS1 can take all base stations whose indication information 1 takes the value "True" as candidate servers of distributed training.
[0312] Step 5, BS1 sends a distributed training capability information B request to the above candidate base stations (e.g. BS2, BS3).
[0313] Step 6, the candidate base stations (e.g. BS2, BS3) determine their respective distributed training capability information B.
[0314] Step 7, the candidate base stations (e.g. BS2, BS3) send the distributed training capability information B to BS1.
[0315] Note: The above steps 5-7 are the same as steps 2-4 of the first embodiment shown in FIG. 4d, with the difference that the distributed training capability information B does not contain "indication information 1".
[0316] Step 8, the same as step 5 of the embodiment shown in FIG. 4d.
[0317] Step 9, the same as step 6 of the embodiment shown in FIG. 4d.
[0318] Step 10, the same as step 7 of the embodiment shown in FIG. 4d.
[0319] From the above implementation process, it can be known that in steps 4 and 5 shown in FIG. 4d (or steps 7 and 8 shown in FIG. 4f): other base stations send distributed training capability information to BS1, and the information at least contains connectivity information and / or transmission capacity information, so that BS1 selects a suitable base station as a server of distributed training according to the above information. Therefore, through the information, BS1 can more accurately select a base station as a server of distributed training, thereby improving the performance of distributed training. For example, selecting a base station with good connectivity as a server of distributed training can have more client base stations participate in distributed training, so as to provide more training data (thereby improving training accuracy) and more computing power resources (thereby improving training speed). For another example, selecting a base station with large transmission capacity as a server of distributed training can complete various information interactions between the server base station and the client base station more quickly, thereby improving the training speed.
[0320] Embodiment two: the center node decides whether to perform distributed learning, and the center node selects a base station as a server of distributed training.
[0321] As shown in FIG. 4g, the flow of embodiment two is as follows. It should be understood that in the implementation shown in FIG. 4g, BS1 is an implementation example of the first communication device shown in the foregoing, BS2, BS3 and BS4 are an implementation example of N (N = 3) communication devices shown in the foregoing. In addition, step 4 described below is an implementation example of step S301 described in the foregoing, and step 5 described below is an implementation example of step S302 described in the foregoing.
[0322] It should be noted that steps 1-7 of the embodiment shown in FIG. 4g can refer to steps 1-7 of the embodiment shown in FIG. 4d, and the difference is that “BS1” in the embodiment shown in FIG. 4d is replaced by “center node” in the embodiment shown in FIG. 4g. Steps 0a, 0b, 0c and 0d of the embodiment shown in FIG. 4g give a method of triggering the center node to decide whether to perform model training, that is, the center node receives model performance information reported by a base station (BS1) managed by the center node, and decides whether to retrain a model for the base station according to the model performance information.
[0323] Step 0a, BS1 deploys a model, and BS1 runs the model (that is, performs inference based on the model).
[0324] Step 0b, the center node sends model performance reporting configuration information to BS1. The configuration information contains the following contents:
[0325] Type of model reporting. The type can be periodic reporting and / or event-triggered reporting.
[0326] Model reporting parameters. If the model reporting type is periodic reporting, this parameter contains the reporting period, e.g. 100 ms or 100 radio frames. If the model reporting type is event triggered reporting, this parameter contains the event information, e.g. "the cosine similarity of the model is less than 0.80" or "the user throughput is less than 5 Mbps".
[0327] Step 0c, BS1 monitors the model performance and records.
[0328] Step Od, BS1 reports the model performance to the central node. The reporting is in compliance with the model performance reporting configuration information, e.g. reporting according to the configured period, or reporting when a configured event is found.
[0329] Optionally, a variant of the embodiment shown in Fig. 4g is shown in Fig. 4h. Steps 1-10 shown in Fig. 4h can refer to steps 1-10 of the embodiment shown in Fig. 4f, except that "BS1" is replaced by "central node" in steps 1-10. Steps 0a, 0b, 0c and Od shown in Fig. 4h can refer to steps 0a, 0b, 0c and Od of the embodiment shown in Fig. 4g.
[0330] Embodiment three, BS1 decides whether to perform distributed learning, and after deciding to perform distributed learning, BS1 requests the central node to select a base station as the distributed training server, and the central node selects a base station as the distributed training server.
[0331] As shown in Fig. 4i, the flow of embodiment three is as follows. It should be understood that in the implementation shown in Fig. 4i, BS1 is an implementation example of the first communication device shown in the foregoing, BS2, BS3 and BS4 are an implementation example of the N (N = 3) communication devices shown in the foregoing. In addition, step 4 below is an implementation example of step S301 in the foregoing, and step 5 below is an implementation example of step S302 in the foregoing.
[0332] Step 0, BS1 decides to perform distributed training. It can refer to step 1 of the embodiment shown in Fig. 4d.
[0333] Step 1, BS1 sends a distributed training server selection request to the central node, for requesting the central node to select a base station as the distributed training server. The request contains at least one of the following information:
[0334] The information of the use case, for indicating which use case or which use cases the distributed training is for. The information can be the name of the use case (e.g. "user location prediction"), or the identification of the use case (e.g. use case 10), or the identification of the model associated with the use case (e.g. model 3), etc.
[0335] information of the number of distributed training client base stations. For example, the maximum and / or minimum of the number of distributed training client base stations.
[0336] information of the number of data samples used for distributed training. For example, the maximum and / or minimum of the number of data samples used for distributed training.
[0337] information of the time used for distributed training. For example, the maximum of the time used for distributed training, or the latest time when distributed training is completed.
[0338] information of the number of rounds iterated for distributed training. For example, the maximum of the number of rounds iterated for distributed training.
[0339] information of the performance achieved by distributed training. For example, the maximum of the Mean Squared Error (MSE) of the model achieved by distributed training.
[0340] The above information is used to help the center node determine which distributed training capabilities of other base stations to request (step 2), and to determine which base stations to select as distributed training servers (step 5).
[0341] Steps 2-7 shown in FIG. 4i can refer to steps 2-7 shown in FIG. 4d, with the difference that “BS1” shown in FIG. 4d is replaced by “center node” shown in FIG. 4i.
[0342] Optionally, a variant of the embodiment shown in FIG. 4i is shown in FIG. 4j. Steps 0-1 shown in FIG. 4j can refer to steps 0-1 of the embodiment shown in FIG. 4i, with the difference that “BS1” therein is replaced by “center node”. Steps 2-10 shown in FIG. 4j can refer to steps 2-10 of the embodiment shown in FIG. 4f.
[0343] Referring to FIG. 5, an embodiment of the present application provides a communication apparatus 500, which can implement the functions of the first communication apparatus (or the second communication apparatus) in the above method embodiments, and thus can also implement the beneficial effects possessed by the above method embodiments. In an embodiment of the present application, the communication apparatus 500 can be the first communication apparatus (or the second communication apparatus), or an integrated circuit or element etc. inside the first communication apparatus (or the second communication apparatus), such as a chip, a baseband chip, a modem chip, an SoC chip (such as an SoC chip containing a modem core), a SIP chip, a communication module, a chip system, a processor, etc.
[0344] It should be noted that the transceiver unit 502 can include a transmitting unit and a receiving unit, which are respectively used for performing transmission and reception.
[0345] In a possible implementation, when the apparatus 500 is configured to perform the method performed by the first communication device in FIG. 3 and related embodiments, the apparatus 500 includes a processing unit 501 and a transceiver unit 502; the transceiver unit 502 is configured to receive N pieces of information, the i th piece of information in the N pieces of information indicating distributed training capability information of an i th communication device, N is a positive integer, and i is an integer from 1 to N; wherein the distributed training capability information of the i th communication device is used to indicate connectivity information and / or transmission capacity information of one or more communication devices connected to the i th communication device; and the processing unit 501 is configured to determine a distributed training server based on the N pieces of information.
[0346] In a possible implementation, when the apparatus 500 is configured to perform the method performed by the third communication device in FIG. 3 and related embodiments, the apparatus 500 includes a processing unit 501 and a transceiver unit 502; the processing unit 501 is configured to determine distributed training capability information, the distributed training capability information indicating connectivity information and / or transmission capacity information of one or more communication devices connected to a second communication device, and the distributed training capability information being used to determine a distributed training server; and the transceiver unit 502 is configured to send third information, the third information indicating the distributed training capability information.
[0347] In a possible design, when the communication apparatus 500 is a communication module in a terminal device or a terminal, the function of the processing unit 501 can be implemented by one or more processors. Specifically, the processor can include a modem chip, a SoC chip (such as a SoC chip including a modem core), or a SIP chip. The function of the transceiver unit 502 can be implemented by a transceiver circuit.
[0348] In a possible design, when the communication apparatus 500 is a circuit or chip responsible for communication functions in a terminal, such as a modem chip or a SoC chip or a SoC chip including a modem core or a SIP chip, the function of the processing unit 501 can be implemented by a circuit system including one or more processors or processor cores in the above-mentioned chip. The function of the transceiver unit 502 can be implemented by an interface circuit or a data transceiver circuit on the above-mentioned chip.
[0349] It should be noted that the information execution process and the like of the units of the communication apparatus 500 are specifically described in the foregoing method embodiments of the present application, and will not be described here.
[0350] Please refer to FIG. 6, which is another schematic structural diagram of a communication apparatus 600 provided by the present application. The communication apparatus 600 includes a logic circuit 601 and an input / output interface 602. The communication apparatus 600 can be a chip or an integrated circuit.
[0351] The transceiving unit 502 shown in FIG. 5 can be a communication interface, which can be an input and output interface 602 in FIG. 6, and the input and output interface 602 can include an input interface and an output interface. Alternatively, the communication interface can also be a transceiving circuit, which can include an input interface circuit and an output interface circuit.
[0352] In a possible implementation, when the apparatus 600 is configured to perform the method performed by the first communication device in FIG. 3 and related embodiments, the input and output interface 602 is configured to receive N pieces of information, the i th piece of information in the N pieces of information indicating distributed training capability information of an i th communication device, N is a positive integer, and i is an integer from 1 to N; wherein the distributed training capability information of the i th communication device is configured to indicate connectivity information and / or transmission capacity information of one or more communication devices connected to the i th communication device; and the logic circuit 601 is configured to determine a distributed training server based on the N pieces of information.
[0353] In a possible implementation, when the apparatus 600 is configured to perform the method performed by the second communication device in FIG. 3 and related embodiments, the logic circuit 601 is configured to determine distributed training capability information, the distributed training capability information indicating connectivity information and / or transmission capacity information of one or more communication devices connected to the second communication device, and the distributed training capability information is configured to determine a distributed training server; and the input and output interface 602 is configured to send third information, the third information indicating the distributed training capability information.
[0354] The logic circuit 601 and the input and output interface 602 can also perform other steps performed by the first communication device or the second communication device in any of the embodiments and achieve corresponding beneficial effects, which are not described herein again.
[0355] In a possible implementation, the processing unit 501 shown in FIG. 5 can be the logic circuit 601 in FIG. 6.
[0356] Optionally, the logic circuit 601 can be a processing apparatus, and the functions of the processing apparatus can be partially or entirely implemented through software.
[0357] Optionally, the processing apparatus can include a memory and a processor, wherein the memory is configured to store a computer program, and the processor is configured to read and execute the computer program stored in the memory to perform corresponding processing and / or steps in any one of the method embodiments.
[0358] Optionally, the processing device can only include a processor. The memory for storing the computer program is located outside the processing device, and the processor is connected with the memory through the circuit / wire to read and execute the computer program stored in the memory. Among them, the memory and the processor can be integrated together, or they can also be physically independent of each other.
[0359] Optionally, the processing device can be one or more chips, or one or more integrated circuits. For example, the processing device can be one or more field-programmable gate arrays (FPGA), application specific integrated circuits (ASIC), system on chips (SoC), central processing units (CPU), network processors (NP), digital signal processors (DSP), micro controller units (MCU), programmable logic devices (PLD) or other integrated chips, or any combination of the above chips or processors, etc.
[0360] Please refer to FIG. 7, the communication device 700 involved in the above-mentioned embodiments provided by the embodiments of the present application, which can be specifically the communication device as the terminal device in the above-mentioned embodiments, and the example shown in FIG. 7 is realized by the terminal device (or components in the terminal device).
[0361] Among them, a possible logical structure diagram of the communication device 700, the communication device 700 can include but not limited to at least one processor 701 and a communication port 702.
[0362] Among them, the transceiver unit 502 shown in FIG. 5 can be a communication interface, which can be a communication port 702 in FIG. 7, and the communication port 702 can include an input interface and an output interface. Alternatively, the communication port 702 can also be a transceiver circuit, which can include an input interface circuit and an output interface circuit.
[0363] Further optionally, the device can also include at least one of a memory 703, a bus 704, and in the embodiments of the present application, the at least one processor 701 is used to control and process the actions of the communication device 700.
[0364] Further, the processor 701 can be a central processing unit, a general purpose processor, a digital signal processor, an application specific integrated circuit, a field programmable gate array, or other programmable logic device, transistor logic, hardware components, or any combination thereof. It can implement or execute various example logical blocks, modules, and circuits described in connection with the disclosure. The processor can also be a combination of computing functions, such as a combination of one or more microprocessors, a combination of a digital signal processor and a microprocessor, and the like. For the convenience and brevity of description, the specific working processes of the above-described system, device, and unit can be referred to the corresponding processes in the foregoing method embodiments, which will not be described herein.
[0365] It should be noted that the communication device 700 shown in FIG. 7 can be specifically used to implement the steps implemented by the terminal device in the foregoing method embodiments, and achieve the corresponding technical effects of the terminal device. The specific implementation of the communication device shown in FIG. 7 can be referred to the description in the foregoing method embodiments, which will not be described herein.
[0366] Please refer to FIG. 8, which is a structural schematic diagram of a communication device 800 involved in the foregoing embodiments provided by the embodiments of the present application. The communication device 800 can be specifically the communication device as the network device in the foregoing embodiments, and the example shown in FIG. 8 is implemented by the network device (or components in the network device). The structure of the communication device can be referred to the structure shown in FIG. 8.
[0367] The communication device 800 includes at least one processor 811 and at least one network interface 814. Further optionally, the communication device further includes at least one memory 812, at least one transceiver 813, and one or more antennas 815. The processor 811, the memory 812, the transceiver 813, and the network interface 814 are connected, for example, through a bus. In the embodiments of the present application, the connection can include various interfaces, transmission lines, or buses, etc., which are not limited in the embodiments. The antenna 815 is connected to the transceiver 813. The network interface 814 is used to enable the communication device to communicate with other communication devices through a communication link. For example, the network interface 814 can include a network interface between the communication device and the core network device, such as an S1 interface. The network interface can include a network interface between the communication device and other communication devices (such as other network devices or core network devices), such as an X2 or Xn interface.
[0368] The transceiver unit 502 shown in FIG. 5 can be a communication interface, which can be the network interface 814 in FIG. 8. The network interface 814 can include an input interface and an output interface. Alternatively, the network interface 814 can be a transceiver circuit, which can include an input interface circuit and an output interface circuit.
[0369] The processor 811 is mainly used for processing communication protocols and communication data, and controlling the whole communication device, executing software programs, processing data of the software programs, for example, for supporting the communication device to perform the actions described in the embodiments. The communication device can include a baseband processor mainly used for processing communication protocols and communication data, and a central processor mainly used for controlling the whole terminal device, executing software programs, and processing data of the software programs. The processor 811 in FIG. 8 can integrate the functions of the baseband processor and the central processor, and those skilled in the art can understand that the baseband processor and the central processor can also be independent processors interconnected by a bus or the like. Those skilled in the art can understand that the terminal device can include multiple baseband processors to adapt to different network modes, and the terminal device can include multiple central processors to enhance its processing capability, and various components of the terminal device can be connected by various buses. The baseband processor can also be referred to as a baseband processing circuit or a baseband processing chip. The central processor can also be referred to as a central processing circuit or a central processing chip. The function of processing communication protocols and communication data can be built into the processor, or stored in the memory in the form of a software program, and the processor executes the software program to realize the baseband processing function.
[0370] The memory is mainly used for storing software programs and data. The memory 812 can exist independently and be connected with the processor 811. Alternatively, the memory 812 can be integrated with the processor 811, for example, integrated in a chip. The memory 812 can store program codes for executing the technical solutions of the embodiments of the present application, and the processor 811 controls the execution. Various computer programs executed can also be regarded as a driver of the processor 811.
[0371] FIG. 8 only shows one memory and one processor. In an actual terminal device, multiple processors and multiple memories can exist. The memory can also be referred to as a storage medium or a storage device, etc. The memory can be a storage element on the same chip as the processor, that is, an on-chip storage element, or an independent storage element, and the embodiments of the present application do not limit this.
[0372] The transceiver 813 can be configured to support the receiving or transmitting of radio frequency signals between the communication device and a terminal. The transceiver 813 can be connected to the antenna 815. The transceiver 813 includes a transmitter Tx and a receiver Rx. Specifically, the one or more antennas 815 can receive radio frequency signals, the receiver Rx of the transceiver 813 is configured to receive the radio frequency signals from the antenna and convert the radio frequency signals into digital baseband signals or digital intermediate frequency signals, and provide the digital baseband signals or digital intermediate frequency signals to the processor 811 for further processing, such as demodulation processing and decoding processing, by the processor 811. In addition, the transmitter Tx in the transceiver 813 is also configured to receive modulated digital baseband signals or digital intermediate frequency signals from the processor 811, and convert the modulated digital baseband signals or digital intermediate frequency signals into radio frequency signals, and transmit the radio frequency signals through the one or more antennas 815. Specifically, the receiver Rx can selectively perform one or more levels of down-mixing processing and analog-to-digital conversion processing on the radio frequency signals to obtain the digital baseband signals or digital intermediate frequency signals, and the order of the down-mixing processing and the analog-to-digital conversion processing can be adjustable. The transmitter Tx can selectively perform one or more levels of up-mixing processing and digital-to-analog conversion processing on the modulated digital baseband signals or digital intermediate frequency signals to obtain the radio frequency signals, and the order of the up-mixing processing and the digital-to-analog conversion processing can be adjustable. The digital baseband signals and the digital intermediate frequency signals can be collectively referred to as digital signals.
[0373] The transceiver 813 can also be referred to as a transceiving unit, a transceiver, a transceiving device, etc. Optionally, the devices in the transceiving unit for implementing the receiving function can be regarded as a receiving unit, and the devices in the transceiving unit for implementing the transmitting function can be regarded as a transmitting unit, i.e., the transceiving unit includes the receiving unit and the transmitting unit, the receiving unit can also be referred to as a receiver, an input port, a receiving circuit, etc., and the transmitting unit can be referred to as a transmitter, a transmitter, or a transmitting circuit, etc.
[0374] It should be noted that the communication device 800 shown in FIG. 8 can be specifically configured to implement the steps implemented by the network device in the foregoing method embodiments, and achieve the corresponding technical effects of the network device. The specific implementation mode of the communication device 800 shown in FIG. 8 can be referred to the description in the foregoing method embodiments, which will not be described here one by one.
[0375] Please refer to FIG. 9, which is a structural schematic diagram of a communication device involved in the above embodiments provided by the embodiments of the present application.
[0376] It can be understood that the communication apparatus 900 includes, for example, modules, units, elements, circuits, or interfaces, and the like, which are appropriately configured together to perform the technical solutions provided in the present application. The communication apparatus 900 can be a terminal device or a network device described above, or can be a component (for example, a chip) of the devices, to implement the methods described in the following method embodiments. The communication apparatus 900 includes one or more processors 901. The processor 901 can be a general processor or a special-purpose processor, and the like. For example, it can be a baseband processor or a central processing unit. The baseband processor can be used to process communication protocols and communication data, and the central processing unit can be used to control the communication apparatus (such as a RAN node, a terminal, or a chip, and the like), execute software programs, and process data of the software programs.
[0377] Optionally, in one design, the processor 901 can include a program 903 (which can also be referred to as code or instructions at times) that can be run on the processor 901, so that the communication apparatus 900 performs the methods described in the following embodiments. In yet another possible design, the communication apparatus 900 includes a circuit (not shown in FIG. 9).
[0378] Optionally, the communication apparatus 900 can include one or more memories 902 having a program 904 (which can also be referred to as code or instructions at times) stored thereon, which can be run on the processor 901, so that the communication apparatus 900 performs the methods described in the above method embodiments.
[0379] Optionally, the processor 901 and / or the memory 902 can include an AI module 907, 908, which is used to implement AI-related functions. The AI module can be implemented in software, hardware, or a combination of software and hardware. For example, the AI module can include a radio intelligence control (RIC) module. For example, the AI module can be a near-real-time RIC or a non-real-time RIC.
[0380] Optionally, the processor 901 and / or the memory 902 can also store data. The processor and the memory can be separately arranged, or can be integrated together.
[0381] Optionally, the communication apparatus 900 can also include a transceiver 905 and / or an antenna 906. The processor 901 can also be referred to as a processing unit, which controls the communication apparatus (such as a RAN node or a terminal). The transceiver 905 can also be referred to as a transceiving unit, a transceiver, a transceiving circuit, or a transceiver, and the like, which is used to realize the transceiving function of the communication apparatus through the antenna 906.
[0382] The processing unit 501 shown in FIG. 5 can be the processor 901. The transceiving unit 502 shown in FIG. 5 can be a communication interface, which can be the transceiver 905 in FIG. 9, and the transceiver 905 can include an input interface and an output interface. Alternatively, the transceiver 905 can also be a transceiving circuit, which can include an input interface circuit and an output interface circuit.
[0383] The embodiments of the present application further provide a computer readable storage medium for storing one or more computer-executable instructions, which, when executed by a processor, cause the processor to perform the method described in the possible implementation manners of the first communication device or the second communication device.
[0384] The embodiments of the present application further provide a computer program product (or computer program), which, when executed by a processor, causes the processor to perform the method described in the possible implementation manners of the first communication device or the second communication device.
[0385] The embodiments of the present application further provide a chip system, which includes at least one processor for supporting the communication device to implement the functions involved in the possible implementation manners of the communication device. Optionally, the chip system further includes an interface circuit for providing program instructions and / or data for the at least one processor. In a possible design, the chip system can further include a memory for storing necessary program instructions and data of the communication device. The chip system can be composed of a chip, or can include a chip and other discrete devices. The communication device can be the first communication device or the second communication device in the method embodiments.
[0386] The embodiments of the present application further provide a communication system, which includes the first communication device and / or the second communication device in any of the above embodiments.
[0387] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other manners. For example, the described device embodiments are merely schematic. For example, the division of the units is only a logical function division. There can be another division manner for the actual implementation, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between different units, can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.
[0388] The units described as separated components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purposes of the embodiments of the present application.
[0389] In addition, each functional unit in the embodiments of the present application can be integrated in one processing unit, or each unit can exist physically as a separate unit, or two or more units can be integrated in one unit. The integrated unit can be implemented in the form of hardware, or in the form of software functional units. When the integrated unit is implemented in the form of software functional units and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or substantially, or all or part of the technical solutions, can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the embodiments of the present application. The foregoing storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, and various other media that can store program codes.
Claims
1. A communication method, characterized in that, include: Receive N pieces of information, wherein the i-th piece of information in the N pieces of information indicates the distributed training capability information of the i-th communication device, where N is a positive integer and i takes the value from 1 to N; wherein the distributed training capability information of the i-th communication device is used to indicate the connectivity information and / or transmission capacity information of one or more communication devices connected to the i-th communication device; The distributed training server is determined based on the N pieces of information.
2. The method according to claim 1, characterized in that, The method is applied to a first communication device; The distributed training server is one of N communication devices, and the N communication devices are different from the first communication device.
3. The method according to claim 1, characterized in that, The method is applied to a first communication device; The process of determining the distributed training server based on the N pieces of information includes: The distributed training server is determined based on the N pieces of information and the distributed training capability information of the first communication device; wherein, the distributed training server is one of the N communication devices and the first communication device.
4. The method according to any one of claims 1 to 3, characterized in that, Before receiving N pieces of information, the method further includes: Send N request messages, wherein the N request messages are used to request the N messages.
5. The method according to claim 4, characterized in that, Before sending N request messages, the method further includes: Send M request messages, each of which requests the M communication devices to provide feedback on whether the M communication devices support the capability of acting as a distributed training server. The N communication devices are included in the M communication devices, where M is greater than or equal to N. Receive N capability information, which are used to indicate that the N communication devices support the capability of acting as a distributed training server.
6. The method according to claim 4, characterized in that, Some or all of the N communication devices support the capability to serve as distributed training servers.
7. The method according to any one of claims 1 to 6, characterized in that, The connectivity information includes at least one of the following: The number of the one or more communication devices; Identification information of the one or more communication devices; or, The computing power information of the one or more communication devices.
8. The method according to any one of claims 1 to 7, characterized in that, The distributed training capability information of the i-th communication device is also used to indicate at least one of the following: Does the i-th communication device support the capability to act as a distributed training server? Does the i-th communication device support the ability to act as a distributed training client? The computing power information of the i-th communication device.
9. The method according to any one of claims 1 to 8, characterized in that, The distributed training capability information of the i-th communication device is associated with one or more distributed training use cases; or, The distributed training capability information of the i-th communication device is associated with any distributed training use case of the i-th communication device; or, The distributed training capability information of the i-th communication device is associated with any distributed training use case supported by the i-th communication device.
10. The method according to any one of claims 1 to 9, characterized in that, The process of determining the distributed training server based on the N pieces of information includes: The distributed training server is determined based on the N pieces of information if any of the following conditions are met: Determine to train the model using a distributed training approach; Receive first information, which is used to instruct model training based on distributed training; wherein, the fourth information includes one or more of the following: use case information for model training, number of distributed training clients, number of training samples for distributed training, training time for distributed training, number of iterations for distributed training, or target performance information for distributed training.
11. The method according to any one of claims 1 to 10, characterized in that, The method further includes: Send a second message, which is used to instruct the distributed training server.
12. The method according to any one of claims 1 to 11, characterized in that, The i-th communication device is connected to one or more communication devices that support the ability to participate in distributed training as a distributed training client.
13. The method according to any one of claims 1 to 12, characterized in that, The method is applied to terminal devices, network devices, or central nodes used to manage one or more network devices.
14. A communication method, characterized in that, include: Determine distributed training capability information, which indicates the connectivity information and / or transmission capacity information of one or more communication devices connected to the second communication device, and the distributed training capability information is used to determine the distributed training server. Send a third message, which indicates the distributed training capability information.
15. The method according to claim 14, characterized in that, Before sending the third information, the method further includes: Receive a first request message, which is used to request the third information.
16. The method according to claim 15, characterized in that, Before receiving the first request information, the method further includes: Receive a second request message, the second request message being used to request the second communication device to provide feedback on whether the second communication device supports the capability of acting as a distributed training server; Send first capability information, which indicates that the second communication device supports the capability of acting as a distributed training server.
17. The method according to any one of claims 14 to 16, characterized in that, The connectivity information includes at least one of the following: The number of the one or more communication devices; Identification information of the one or more communication devices; or, The computing power information of the one or more communication devices.
18. The method according to any one of claims 14 to 17, characterized in that, The distributed training capability information of the second communication device is also used to indicate at least one of the following: Does the second communication device support the capability to act as a distributed training server? Does the second communication device support the capability to act as a distributed training client? The computing power information of the second communication device.
19. The method according to any one of claims 14 to 18, characterized in that, The distributed training capability information of the second communication device is associated with one or more distributed training use cases; or, The distributed training capability information of the second communication device is associated with any distributed training use case of the second communication device; or, The distributed training capability information of the second communication device is associated with any distributed training use case supported by the second communication device.
20. The method according to any one of claims 14 to 19, characterized in that, The method further includes: Receive second information, which is used to instruct the distributed training server.
21. The method according to any one of claims 14 to 20, characterized in that, The second communication device is connected to one or more communication devices that support the ability to participate in distributed training as a distributed training client.
22. The method according to any one of claims 14 to 21, characterized in that, The method is applied to terminal devices or network devices.
23. A communication device, characterized in that, Includes a module for performing the method as described in any one of claims 1 to 22.
24. A communication device, characterized in that, It includes at least one processor, said at least one processor being used to perform the method as described in any one of claims 1 to 22.
25. The communication device according to claim 24, characterized in that, The communication device is a chip or chip system.
26. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program or instructions that, when executed by a communication device, implement the method as described in any one of claims 1 to 22.
27. A computer program product, characterized in that, It includes a computer program or instructions that, when executed by a computer, implement the method as described in any one of claims 1 to 22.
28. A communication device, characterized in that, It includes at least one logic circuit and an input / output interface; the input / output interface is used to input and / or output signals, and the logic circuit is used to perform the method as described in any one of claims 1 to 22.
29. A communication system, characterized in that, It includes a first communication device and a second communication device, wherein the first communication device is used to perform the method as described in any one of claims 1 to 13; and the second communication device is used to perform the method as described in any one of claims 14 to 22.
Citation Information
Patent Citations
Distributed model training device and method, and computer program product
CN112561079A
Model training method and device, electronic equipment, server and storage medium
CN117632451A
Communication method, system and related equipment
CN117676716A
Deterministic control method and device for distributed training
CN117792890A