Communication method and communication device

By receiving and shaping a reference distribution to train a second AI model, the problem of AI models trained independently on different devices having difficulty working together is solved, thus improving the model's learning performance and collaborative working ability.

CN121241353APending Publication Date: 2025-12-30HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380098838.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-06-13
Filing Date
2023-10-17
Publication Date
2025-12-30

AI Technical Summary

Technical Problem

AI models trained independently on different devices have difficulty working together, resulting in poor learning performance.

Method used

The second AI model is obtained by receiving instructions on Q reference distributions corresponding to the Q layers of the first AI model, and then shaping it according to these reference distributions to minimize layer distribution differences and improve model training performance.

Benefits of technology

It enables interconnection and interoperability between multiple AI models, improving the training performance and collaborative working capabilities of AI models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121241353A_ABST
    Figure CN121241353A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a communication method and a communication device. The communication method includes: receiving first information indicating Q reference distributions corresponding to Q layers of a first AI model; and obtaining a second AI model according to q reference distributions in the Q reference distributions. According to the technical scheme, the learning performance of the AI model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Cross Reference to Related Applications

[0001] This application is related to U.S. Provisional Patent Application No. 63 / 507,767, filed on June 13, 2023, entitled “AI MODEL TRAINING WITH REFERENCE DISTRIBUTION,” and claims priority to the U.S. Provisional Patent Application.

[0002] The disclosure of the above application is incorporated herein by reference in its entirety. TECHNICAL FIELD

[0003] Embodiments of the present application relate to the field of communication, and more particularly, to a communication method and a communication apparatus. BACKGROUND

[0004] Artificial intelligence (AI) based algorithms have been introduced in wireless communications to solve some wireless problems, such as channel estimation, scheduling, channel state information (CSI) compression, positioning, beam management, etc. AI algorithms are a data-driven approach that adjusts some pre-defined architecture through a set of data samples called training dataset.

[0005] The learning performance of an AI model is crucial to its application. For example, multiple AI models deployed on different devices can need to work collaboratively. However, the AI models can be independently trained by different providers, and it is difficult to guarantee the training quality, which can result in the AI models failing to work collaboratively.

[0006] Therefore, how to improve the learning performance of an AI model is a technical problem to be solved. SUMMARY

[0007] Embodiments of the present application provide a communication method and a communication apparatus. The technical solution can improve the learning performance of an AI model.

[0008] According to a first aspect, embodiments of the present application provide a communication method, comprising: receiving first information indicating Q reference distributions corresponding to Q layers of a first AI model, wherein Q is a positive integer; and obtaining a second AI model according to q reference distributions in the Q reference distributions, wherein q is a positive integer and q≤Q.

[0009] According to the above technical solution, the second AI model is obtained according to the q reference distributions, which is conducive to shaping the Q layers of the second AI model according to the q reference distributions. In this way, the reference distributions can be set as needed to obtain the required AI model, which is conducive to improving the training performance of the AI model.

[0010] Optionally, the reference distribution can be in the form of a parametric standard distribution.

[0011] Optionally, the reference distribution can be in the form of a combination of multiple parametric standard distributions.

[0012] Optionally, the reference distribution can be in the form of multiple reference data samples.

[0013] In a possible design, obtaining the second AI model according to the q reference distributions of the Q reference distributions includes: obtaining the second AI model with one or more regularizations, the one or more regularizations being used to minimize differences between the q reference distributions and distributions of the corresponding q layers of the Q layers.

[0014] According to the technical solution described above, the q layers of the second AI model can be shaped by the q reference distributions, which is beneficial to realizing interconnection and intercommunication of multiple AI models. For example, the q reference distributions can be consistent with outputs of the q layers in another AI model that needs to work together with the second AI model. According to the technical solution described above, the distributions of the q layers in the second AI model can be as close as possible to the distributions of the q layers in the other AI model, which is beneficial to realizing interconnection and intercommunication between the two AI models.

[0015] In a possible design, the method further includes: sending second information indicating an optimization result of the one or more regularizations.

[0016] In a possible design, the Q layers include one or more latent layers of the first AI model.

[0017] In a possible design, the q layers can include one or more latent layers of the first AI model.

[0018] Optionally, the q layers can be q latent layers of the first AI model.

[0019] In addition, the outputs of the q layers can include outputs of at least one latent layer, and in this case, the outputs of the at least one latent layer can be shaped according to the reference distribution.

[0020] In a possible design, the method further includes: receiving third information indicating the Q layers.

[0021] In a possible design, the method further includes: receiving fourth information indicating Q scoring functions, the Q scoring functions being used to measure differences between the Q reference distributions and the distributions of the Q layers.

[0022] According to a second aspect, embodiments of the present application provide a communication device, comprising: obtaining Q reference distributions corresponding to Q layers of a first AI model, wherein the Q reference distributions are used to obtain a second AI model, and Q is a positive integer; and sending first information indicating the Q reference distributions.

[0023] In a possible design, the Q reference distributions are used to establish one or more regularizations, and the one or more regularizations are used to minimize differences between the Q reference distributions and distributions of the Q layers.

[0024] In a possible design, the method further includes: receiving second information indicating an optimization result of the one or more regularizations.

[0025] In a possible design, the Q layers include one or more latent layers of the first AI model.

[0026] In a possible design, the method further includes: receiving third information indicating the Q layers.

[0027] In a possible design, the method further includes: receiving fourth information indicating Q scoring functions, and the Q scoring functions are used to measure the differences between the Q reference distributions and the distributions of the Q layers.

[0028] According to a third aspect, a communication apparatus is provided. The communication apparatus includes functions or units for performing the method according to the first aspect or any possible design of the first aspect.

[0029] For example, the communication apparatus can be a network device or a chip in a network device. For another example, the communication apparatus can be a terminal device or a chip in a terminal device.

[0030] According to a fourth aspect, a communication apparatus is provided. The communication apparatus includes functions or units for performing the method according to the second aspect or any possible design of the second aspect.

[0031] For example, the communication apparatus can be a terminal device or a chip in a terminal device. For another example, the communication apparatus can be a network device or a chip in a network device.

[0032] According to a fifth aspect, a system is provided. The system includes the communication apparatus according to the third aspect and the communication apparatus according to the fourth aspect.

[0033] According to a sixth aspect, a communication apparatus is provided. The communication apparatus includes at least one processor coupled to at least one memory. The at least one memory is configured to store a computer program or one or more instructions. The at least one processor is configured to invoke the computer program or the one or more instructions from the at least one memory and run the computer program or the one or more instructions, so that the communication apparatus performs the method in the first aspect or any possible design of the first aspect, or so that the communication apparatus performs the method in the second aspect or any possible design of the second aspect.

[0034] For example, the communication apparatus can be a network device or a component (e.g., a chip or an integrated circuit) installed in a network device. For another example, the communication apparatus can be a terminal device or a component (e.g., a chip or an integrated circuit) installed in a terminal device.

[0035] According to a seventh aspect, a communication apparatus is provided. The communication apparatus includes a processor and a communication interface. The processor is connected to the communication interface. The processor is configured to execute the one or more instructions, and the communication interface is configured to communicate with other network elements under control of the processor. The processor is configured to perform the method according to the first aspect or any possible design of the first aspect, or the second aspect or any possible design of the second aspect.

[0036] According to an eighth aspect, a computer storage medium is provided. The computer storage medium stores program codes for executing one or more instructions of the method according to the first aspect or any possible design of the first aspect, or the second aspect or any possible design of the second aspect.

[0037] According to a ninth aspect, the present application provides a computer program product including one or more instructions, wherein when the computer program product runs on a computer, the computer executes the method according to the first aspect or any possible design of the first aspect, or the second aspect or any possible design of the second aspect. BRIEF DESCRIPTION OF DRAWINGS

[0038] Figure 1 is a schematic diagram of an application scenario provided by the present application; Figure 2 An exemplary communication system 100 is shown; Figure 3 An exemplary device in a communication system is shown; Figure 4 is a schematic diagram of a device in two cycles provided by an embodiment of the present application; Figure 5An exemplary local data of a device provided by embodiments of the application is shown; Figure 6 is a schematic diagram of an exemplary scenario; Figure 7 is a flowchart of a communication method provided by embodiments of the application; Figure 8 is a schematic diagram of an exemplary regularization of a distribution of outputs of a latent layer provided by embodiments of the application; Figure 9 is a schematic diagram of an exemplary training process of an AE provided by embodiments of the application; Figure 10 is a schematic diagram of another exemplary training process of an AE provided by embodiments of the application; Figures 11 to 15 is a schematic block diagram of a possible device provided by embodiments of the application. DETAILED DESCRIPTION

[0039] The technical solutions of the present application are described below with reference to the accompanying drawings.

[0040] Embodiments of the application can be applied to a communication system of the next generation (e.g., sixth generation (6G) or higher), fifth generation (5th Generation, 5G), new radio (NR), long term evolution (long term evolution, LTE), etc.

[0041] Figure 1 is a schematic diagram of the structure of an exemplary communication system.

[0042] Reference Figure 1 , as a non-limiting illustrative example, a simplified schematic diagram of a communication system is provided. The communication system 100 includes a wireless access network 120. The wireless access network 120 can be a next generation (e.g., 6G or higher) wireless access network, or a conventional (e.g., 5G, 4G, 3G or 2G) wireless access network. One or more communication electronic devices (electric device, ED) 110a to 120j (generally referred to as 110) can be interconnected to each other, or connected to one or more network nodes (170a, 170b, generally referred to as 170) in the wireless access network 120. The core network 130 can be part of the communication system, and can depend on or be independent of the wireless access technology used in the communication system 100. In addition, the communication system 100 includes a public switched telephone network (public switched telephone network, PSTN) 140, the Internet 150 and other networks 160.

[0043] Figure 2is a block diagram of the structure of another example communication system.

[0044] Generally, the communication system 100 enables multiple wireless or wireline elements to communicate data and other content. The communication system 100 can aim to provide voice, data, video and / or text content through broadcast, multicast, and unicast, among other approaches. The communication system 100 can operate through sharing of resources in a carrier frequency spectrum band, among other approaches. The communication system 100 can include terrestrial communication systems and / or non-terrestrial communication systems. The communication system 100 can provide a wide range of communication services and applications (e.g., earth monitoring, remote sensing, passive sensing and positioning, navigation and tracking, autonomous delivery and mobility, among others). The communication system 100 can provide high availability and robustness through joint operation of terrestrial communication systems and non-terrestrial communication systems. For example, integration of non-terrestrial communication systems (or components thereof) into terrestrial communication systems can result in a heterogeneous network comprising multiple tiers. The heterogeneous network can achieve better overall performance compared to traditional communication networks through efficient multi-link joint operation, more flexible function sharing, and faster physical layer link switching between terrestrial and non-terrestrial networks.

[0045] The terrestrial communication systems and non-terrestrial communication systems can be considered as subsystems of a communication system. In the illustrated example, the communication system 100 includes electronic devices (EDs) 110a-110d (generally referred to as EDs 110), radio access networks (RANs) 120a and 120b, a non-terrestrial communication network 120c, a core network 130, a public switched telephone network (PSTN) 140, the Internet 150, and other networks 160. The RANs 120a and 120b include respective base stations (BSs) 170a and 170b, which can be generally referred to as terrestrial transmit and receive points (T-TRPs) 170a and 170b. The non-terrestrial communication network 120c includes an access node 120c, which can be generally referred to as a non-terrestrial transmit and receive point (NT-TRP) 172.

[0046] Any of the EDs 110 can alternatively or additionally be used to connect, interface, or communicate with any of the other T-TRPs 170a and 170b, the NT-TRP 172, the Internet 150, the core network 130, the PSTN 140, the other networks 160, or any combination of the above. In some examples, the ED 110a can perform uplink and / or downlink transmissions with the T-TRP 170a over the interface 190a. In some examples, the EDs 110a, 110b, and 110d can also communicate directly with each other over one or more sidelink air interfaces 190b. In some examples, the ED 110d can perform uplink and / or downlink transmissions with the NT-TRP 172 over the interface 190c.

[0047] The air interfaces 190a and 190b can use similar communication techniques, such as any suitable wireless access technique. For example, the communication system 100 can implement one or more channel access methods in the air interfaces 190a and 190b, such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), or single-carrier FDMA (SC-FDMA). The air interfaces 190a and 190b can utilize other higher-dimensional signal spaces, which can involve combinations of orthogonal and non-orthogonal dimensions.

[0048] The air interface 190c can enable communication between the ED 110d and the one or more NT-TRPs 172 through a wireless link or simply a link. For some examples, the link is a dedicated connection for unicast transmissions, a connection for broadcast transmissions, or a connection between a group of EDs and one or more NT-TRPs for groupcast transmissions.

[0049] The RANs 120a and 120b are in communication with the core network 130 to provide the EDs 110a, 110b, and 110c with access to various services, such as voice, data, and other services. The RANs 120a and 120b and / or the core network 130 can be in direct or indirect communication with one or more other RANs (not shown) that can or can not be directly served by the core network 130 and can or can not utilize the same radio access technology as the RANs 120a and / or 120b. The core network 130 can also serve as a gateway for the RANs 120a and 120b, or EDs 110a, 110b, and 110c, or both, to access other networks (such as PSTN 140, the Internet 150, and other networks 160) by providing means for converting between the air interface protocols of the RANs 120a and 120b and the protocols used by the other networks. Further, some or all of the EDs 110a, 110b, and 110c can include functionality for communicating over different wireless links with different wireless networks using different wireless technologies and / or protocols. In place of (or in addition to) wireless communication, the EDs 110a, 110b, and 110c can communicate with service providers or switches (not shown) and with the Internet 150 over wired communication channels. The PSTN 140 can include a circuit-switched telephone network for providing plain old telephone service (POTS). The Internet 150 can include networks of computers and subnetworks (intranets) or both, and incorporates protocols such as Internet Protocol (IP), transmission control protocol (TCP), and user datagram protocol (UDP). The EDs 110a, 110b, and 110c can be multi-mode devices capable of operating according to multiple wireless access technologies and include multiple transceivers needed to support these technologies.

[0050] The ED 110 can be widely used in various scenarios, such as cellular communication, device-to-device (D2D), vehicle-to-everything (V2X), peer-to-peer (P2P), machine-to-machine (M2M), machine-type communication (MTC), Internet of Things (IoT), virtual reality (VR), augmented reality (AR), industrial control, autonomous driving, telemedicine, smart grids, smart furniture, smart offices, smart wearables, smart transportation, smart cities, drones, robots, remote sensing, passive sensing, positioning, navigation and tracking, autonomous delivery and mobility, etc.

[0051] Each ED 110 represents any suitable wirelessly operating end-user equipment and may include (or be referred to as): user equipment / device (UE), wireless transmit / receive unit (WTRU), mobile station, fixed or mobile subscriber unit, cellular phone, station (STA), machine type communication (MTC) device, personal digital assistant (PDA), personal communications service (PCS) phone, session initiation protocol phone, wireless local loop (WLL) station, smartphone, laptop, computer, tablet, wireless sensor, consumer electronics device, smartbook, vehicle, automobile, truck, bus, train, or IoT device, industrial equipment, or devices within the aforementioned equipment (e.g., communication modules, modems, or chips), etc. Future generations of ED 110 may be referred to using other terms. Base stations 170a and 170b are T-TRPs, hereinafter referred to as T-TRP 170. NT-TRPs will be hereinafter referred to as NT-TRP 172. Each ED 110 connected to T-TRP 170 and / or NT-TRP 172 can be dynamically or semi-statically turned on (i.e., established, activated, or enabled), turned off (i.e., released, deactivated, or disabled), and / or configured in response to one or more of connectivity availability and connectivity necessity.

[0052] In some implementations, T-TRP 170 can be referred to by other names, such as base station, base transceiver station (BTS), wireless base station, network node, network equipment, network-side equipment, transmit / receive node, NodeB, evolved NodeB (eNodeB or eNB), home eNodeB, generation NodeB (gNB), transmission point (TP), site controller, access point (AP) or wireless router, relay station, remote radio head, ground node, ground network equipment or ground base station, base band unit (BBU), remote radio unit (RRU), active antenna unit (AAU), remote radio head (RRH), central unit (CU), distributed unit (DU), positioning node, etc. T-TRP 170 can be a macro BS, micro BS, relay node, host node, or a combination thereof. T-TRP 170 may refer to the aforementioned equipment or a device within the aforementioned equipment (e.g., a communication module, modem, or chip).

[0053] In some embodiments, the various parts of T-TRP 170 can be distributed. For example, some modules of T-TRP 170 may be located remotely from the device housing the antenna of T-TRP 170 and may be coupled to the device housing the antenna via a communication link (not shown), sometimes referred to as the fronthaul, such as the Common Public Radio Interface (CPRI). Therefore, in some embodiments, the term T-TRP 170 may also refer to modules on the network side that perform processing operations such as determining the location of ED 110, resource allocation (scheduling), message generation, and encoding / decoding, and are not necessarily part of the device housing the antenna of T-TRP 170. These modules may also be coupled to other T-TRPs. In some embodiments, T-TRP 170 may actually be multiple T-TRPs operating together to serve ED 110 through cooperative multicast or similar methods.

[0054] In some implementations, NT-TRP 172 may be referred to by other names, such as non-terrestrial node, non-terrestrial network device, or non-terrestrial base station.

[0055] Artificial intelligence (AI) technology can be applied to communications, including AI / ML-based communications at the physical layer and AI / ML-based communications at higher levels such as the medium access control (MAC) layer. For example, at the physical layer, AI / ML-based communications can aim to optimize component design and / or improve algorithm performance. AI / ML can be used to achieve, for instance, channel coding, channel modeling, channel estimation, channel decoding, modulation, demodulation, multiple-input multiple-output (MIMO), waveform generation, multiple access, physical layer element parameter optimization and updating, beamforming, tracking, sensing and / or localization, and so on. At the MAC layer, AI / ML-based communications can aim to leverage AI / ML capabilities for learning, prediction, and / or decision-making to solve complex optimization problems using better strategies and / or optimal solutions, such as optimizing functionality within the MAC layer. For example, AI / ML can be applied to achieve: intelligent transmission and reception point (TRP) management, intelligent beam management, intelligent channel resource allocation, intelligent power control, intelligent spectrum utilization, intelligent modulation and coding scheme (MCS), hybrid automatic repeat request (HARQ) strategy, intelligent transmit / receive (Tx / Rx) mode adaptation, etc.

[0056] To facilitate understanding of the embodiments of this application, the AI / ML related terms that may be involved in the embodiments of this application are described below.

[0057] (1) Data collection Data is a crucial component of AI / ML technology. Data collection is the process by which network nodes, management entities, or user experiences gather data for AI / ML model training, data analysis, and inference.

[0058] (2) AI / ML model training AI / ML model training is the process of learning input / output relationships in a data-driven manner to train an AI / ML model and then using the trained AI / ML model for inference.

[0059] (3) AI / ML model inference The process of using a trained AI / ML model to produce a set of outputs from a set of inputs.

[0060] (4) AI / ML model validation As a sub-process of training, validation is used to evaluate the quality of the AI / ML model using a different dataset than the one used for model training. Validation can help in selecting model parameters that generalize well beyond the dataset used for model training. The trained model parameters can be further tuned through the validation process.

[0061] (5) AI / ML model testing Similar to validation, testing is also a sub-process of training. It is used to evaluate the performance of the final AI / ML model using a different dataset than that used for model training and validation. Unlike AI / ML model validation, testing does not assume subsequent model tuning.

[0062] (6) Online training Online training refers to the AI / ML training process in which the model used for inference is typically trained continuously (nearly in real time) as new training samples arrive.

[0063] (7) Offline training: Offline training is an AI / ML training process in which a model is trained on a collected dataset, and the trained model is then used for inference or transmitted for inference.

[0064] (8) AI / ML model delivery / transmission AI / ML model transfer / transmission is a general term referring to the transfer of an AI / ML model from one entity to another in any way. Transferring an AI / ML model over an air interface includes transferring parameters of the model structure known to the receiving end, as well as transferring a new model with parameters. The transfer can include a complete model or a partial model.

[0065] (9) Life cycle management (LCM) When training and / or inferring AI / ML models on a device, the entire AI / ML process needs to be monitored and managed to ensure the performance gains achieved through AI / ML technology. For example, due to the randomness of wireless channels and the mobility of UEs, the propagation environment of wireless signals changes frequently. However, it is difficult for AI / ML models to maintain optimal performance in all scenarios, and performance may even degrade sharply in some scenarios. Therefore, lifecycle management (LCM) of AI / ML models is crucial for the sustainable operation of AI / ML in the NR air interface.

[0066] Lifecycle management encompasses the entire process of applying AI / ML technologies across one or more nodes. Specifically, lifecycle management includes at least one of the following sub-processes: data collection, model training, model identification, model registration, model deployment, model configuration, model inference, model selection, model activation, deactivation, model switching, model rollback, model monitoring, model update, model transmission / transmission, and UE capability reporting.

[0067] Model monitoring can be based on inference accuracy, including metrics related to key performance indicators (KPIs), or on system performance, including metrics related to system performance KPIs, such as accuracy and relevance, overhead, complexity (computational and memory costs), latency (timeliness of monitoring results, from model failure to recovery), and power consumption. Furthermore, due to environmental changes, data distribution may change after deployment; therefore, models based on input or output data distribution should also be considered.

[0068] (10) Supervised learning The goal of supervised learning algorithms is to train a model that maps feature vectors (inputs) to labels (outputs) based on training data that includes exemplary feature-label pairs. Supervised learning analyzes the training data and produces an inference function that can be used to map inference data.

[0069] (11) Federated learning (FL) Federated learning is a machine learning technique used to train AI / ML models through a central node (e.g., a server) and multiple distributed edge nodes (e.g., UEs, next-generation base stations, "gNBs"). The central node can also be called a central device. Edge nodes can also be called worker nodes or worker devices. The central device connects to the worker devices.

[0070] According to wireless FL technology, the central node can provide edge nodes with a set of model parameters (e.g., weights, biases, gradients) describing the global AI / ML model. Edge nodes can use the received global AI / ML model parameters to initialize their local AI / ML models. Then, the edge nodes can use local data samples to train their local AI / ML models, resulting in trained local AI / ML models. Subsequently, the edge nodes can provide the central node with a set of AI / ML model parameters describing their local AI / ML models.

[0071] Upon receiving multiple sets of AI / ML model parameters describing the corresponding local AI / ML models at multiple edge nodes, the central node can aggregate the local AI / ML model parameters reported from the multiple edge nodes and update the global AI / ML model based on this aggregation. Subsequent iterations proceed very similarly to the first iteration. The central node can then send the aggregated global model to multiple edge nodes. This process iterates multiple times until the global AI / ML model is considered finalized, for example, when the AI / ML model converges or the training stopping condition is met.

[0072] Wireless FL technology does not involve the exchange of local data samples. In fact, local data samples are retained at the respective edge nodes.

[0073] AI-based algorithms have been introduced into wireless communication to solve many wireless problems, such as channel estimation, scheduling, CSI compression (from UE to BS), MIMO beamforming, and localization. AI algorithms are a data-driven approach that tunes some predefined architecture using a set of data samples called a training dataset.

[0074] Neural networks are a typical way to implement AI algorithms. Taking deep neural networks (DNNs) as an example, a training dataset can be used to train a DNN to obtain a model for inference. Recent AI approaches use the stochastic gradient descent (SGD) algorithm to build neurons and train DNN architectures. Examples of DNNs include CNNs, RNNs, and Transformers.

[0075] A communication system comprises multiple connected devices. For example, a device can be a BS (Browser) or UE (User Equipment). For example, a communication system can be... Figure 1 or Figure 2 The communication system 100 in the middle, the equipment can be Figure 1 or Figure 2 The network element shown.

[0076] Figure 3 This is a schematic diagram of the device provided in an embodiment of this application. For example... Figure 3 As shown, the device may include at least one of a sensing module, a communication module, or an AI module. The sensing module can be used to sense and collect signals and / or data. The communication module can be used to send and receive signals and / or data. The AI ​​module can be used to train and / or infer AI implementations.

[0077] To facilitate understanding of the embodiments of this application, the following uses DNN as an example to illustrate the AI ​​implementation method in the embodiments of this application.

[0078] An exemplary AI implementation is based on DNN and consists of two cycles: a training cycle and an inference cycle. The training cycle can also be called the learning cycle. The inference cycle can also be called the reasoning loop.

[0079] Figure 4 This is a schematic diagram of the device in two cycles provided in the embodiments of this application.

[0080] For example, during an inference cycle, the device's AI module can perform one or more inferences using one or more DNNs to complete one or more tasks. The device's perception module can generate signals and / or data, and the device's communication module can receive signals and / or data from one or more other devices. For instance, the input to one or more DNNs can be signals and / or data generated by the device's perception module, and / or signals and / or data received by the device's communication module. After the device's AI module completes the inference, the device's communication module can send the inference results to one or more other devices.

[0081] As another example, during a training cycle, the device's AI module can train one or more DNNs, where the device's perception module can generate signals and / or data, and the device's communication module can receive signals and / or data from one or more other devices. For example, the training data for one or more DNNs can be signals and / or data generated by the device's perception module, and / or signals and / or data received by the device's communication module. During and / or after the AI ​​module completes training, the device's communication module can send the training results to one or more other devices.

[0082] AI implementations can switch between two cycles, or remain in both cycles simultaneously.

[0083] For example, the device's AI module can train a DNN during a training cycle. At the end of the training cycle, the AI ​​implementation switches to an inference cycle, meaning the AI ​​module performs inference based on the trained DNN. At the end of the inference cycle, the AI ​​implementation switches back to a training cycle, and so on.

[0084] For example, the device's AI module can train a second DNN, but still perform inference based on the first DNN.

[0085] The device described above is merely an example. Figure 3 and Figure 4The way the modules are divided and the number of modules do not constitute any limitation on the embodiments of this application. For example, the communication module can be replaced by two modules, namely, a transmitting module and a receiving module. The transmitting module can be used to transmit signals and / or data, and the receiving module can be used to receive signals and / or data. As another example, the sensing module and the communication module can be integrated into one module. Furthermore, the device may also include a processing module. The processing module can be used to process signals and / or data. As yet another example, the device may not include an AI module. Moreover, the AI ​​module may only be used for inferring AI implementation methods, or the AI ​​module may only remain in the inference cycle.

[0086] Wireless systems can enable AI to generalize and interconnect during both the learning and inference cycles.

[0087] Figure 5 Exemplary local data of the device is shown. The device's local data may include at least one of the following: local sensing data provided by the device's sensing module, local channel data provided by the device's communication module, local AI model data provided by the device's AI module, or local potential output data provided by the device's AI module. Local channel data is based on channel measurement results. Local channel data can also be considered as sensing results. Therefore, local channel data can be considered as being provided by the communication module or the sensing module.

[0088] For example, such as Figure 5 As shown, local sensing data may include at least one of RGB data, LiDAR data, temperature, air pressure, or power outage.

[0089] For example, such as Figure 5 As shown, local channel data may include at least one of channel state information (CSI), received signal strength indication (RSSI), or delay.

[0090] Local AI model data can also be called neuron data. For example, such as Figure 5 As shown, local AI model data may include at least one of the following: some or all neurons in a local AI model deployed on the device, or some or all gradients of a local AI model deployed on the device. A neuron can be considered as a function including weights.

[0091] For example, such as Figure 5 As shown, local potential output data can include one or more potential outputs of a local AI model deployed on the device.

[0092] The device can receive local data from one or more other devices. As an example, the data received by the device's communication module may include at least one of the following: perception data from one or more other devices, channel data from one or more other devices, AI model data from one or more other devices, or potential output data from one or more other devices.

[0093] For example, the data received by the communication module of device #A may include channel data from devices #B and #C, as well as AI model data from device #C. The channel data from devices #B and #C refers to the local channel data of device #B and the local channel data of device #C, respectively. The AI ​​model data from device #C refers to the local AI model data of device #C. Devices #A, #B, and #C are different devices.

[0094] For example, the sensing data received by the communication module may include at least one of RGB data, LiDAR data, temperature, air pressure, or power outage.

[0095] For example, the channel data received by the communication module may include at least one of CSI, RSSI, and delay.

[0096] For example, the AI ​​model data received by the communication module may include at least one of the following: some or all of the neurons in the AI ​​model, or some or all of the gradients of the AI ​​model.

[0097] For example, the potential output data received by the communication module may include one or more potential outputs of the AI ​​model.

[0098] During the training period, the device's AI module can operate in single-user mode or collaborative mode.

[0099] In single-user mode, the device's AI module can use the device's local data to train one or more local AI models.

[0100] In collaborative mode, the device's AI module can use data received from the device's communication module to train one or more local AI models.

[0101] For example, data received from the device's communication module can be used by the AI ​​module to train a local AI model in the following ways.

[0102] Alternative solution #1: The perception data received by the device's communication module can be accumulated into a training dataset for training a local AI model.

[0103] Alternative solution #2: The channel data received by the device's communication module can be accumulated into a training dataset for training a local AI model.

[0104] Alternative Solution #3: Some or all neurons in the local AI model can be configured based on the AI ​​model data received by the device's communication module. For example, in federated learning mode, the neurons of an AI model on one device can be configured based on the neurons or gradients of AI models on other devices. Alternatively, the neurons in the local AI model can be updated using the gradients received by the device's communication module.

[0105] Alternative Solution #4: The potential output received by the device's communication module can be input into the device's local AI model. For example, when devices #A and #B work together to train a DNN, device #A trains the first part of the DNN, and device #B trains the second part. Device #A's communication module sends the potential output of the first part of the DNN to device #B. Device #B receives the potential output of the first part and inputs it into the second part of the DNN.

[0106] In addition, the device's local data and the data received by the device's communication module can be used together to train a local AI model.

[0107] For example, the device's local data and the data received by the device's communication module can be used by the AI ​​module to train a local AI model in the following ways.

[0108] Alternative Solution #1: The local perception data provided by the device's perception module and the perception data received by the device's communication module can be combined into a training dataset for training a local AI model.

[0109] Alternative solution #2: The local channel data provided by the device's perception module and the channel data received by the device's communication module can be combined into a training dataset for training a local AI model.

[0110] Alternative Solution #3: The neurons in the updated local AI model can be averaged between some or all of the neurons in the device's AI module's local AI model and the corresponding neurons received by the device's communication module. Alternatively, the neurons in the local AI model can be updated using some or all of the gradients from the device's AI module's local AI model and the corresponding gradients received by the device's communication module.

[0111] Alternative solution #4: The potential outputs of the device’s AI module and the potential outputs received by the device’s communication module can be averaged and then fed into the device’s DNN.

[0112] The training performance of AI models is crucial to their applications. For example, in some scenarios, multiple AI models deployed on different devices may need to work together. However, AI models may be trained independently by different providers, making it difficult to guarantee training quality and potentially causing the AI ​​models to fail to work together.

[0113] Figure 6 This is a schematic diagram of an exemplary scenario.

[0114] like Figure 6 As shown, the encoder deployed on the UE and the decoder deployed on the BS need to work together. However, the encoder and decoder may be from different providers (e.g., Figure 6 Providers #1 and #2 in the code are trained independently, which may affect the interoperability of the encoder and decoder.

[0115] This application provides a communication method in which the difference between a reference distribution and local data can be applied to improve training performance.

[0116] Figure 7 This is a flowchart illustrating the communication method provided in an embodiment of this application.

[0117] like Figure 7 As shown, method 700 includes the following steps.

[0118] 710: The first network element receives information #1 (an example of the first information) indicating Q reference distributions from the second network element. Q is a positive integer.

[0119] The Q reference distributions can each correspond to one or more layers of an AI model. One reference distribution corresponds to one layer, which can be understood as the reference distribution corresponding to the output of that layer.

[0120] The Q reference distributions can also be referred to as the reference distributions of the Q layers.

[0121] For ease of description, this application embodiment uses Q layers belonging to an AI model (first AI model) as an example for illustration.

[0122] 720: The first network element obtains the second AI model based on q out of Q reference distributions. q is a positive integer. q≤Q.

[0123] For example, the first network element could be Figure 3 The device in the network element. The communication module of the first network element can receive information #1. The AI ​​module of the first network element can execute step 720.

[0124] For example, the first network element can be a terminal device or a network device.

[0125] For example, the second network element could beFigure 3 The device in the network. The communication module of the second network element can send information #1.

[0126] For example, the second network element can be a network device or a terminal device.

[0127] In step 720, the first network element can train the first AI model to obtain the second AI model.

[0128] The first AI model and the second AI model are models with the same structure. The parameters of the first AI model and the second AI model are different. The first AI model and the second AI model can be understood as AI models at different stages of the training process. For example, the second AI model can be considered as a first AI model that has already been trained. That is, the second AI model can be considered as the training result of the first AI model.

[0129] Both AI models have the same structure. For ease of description, the first AI model and the second AI model will be collectively referred to as AI models below. The AI ​​model during the training cycle or the AI ​​model to be trained below can be regarded as the first AI model, and the AI ​​model trained using method 700 can be regarded as the second AI model.

[0130] During the training cycle, the device's AI module can operate in either single-user or collaborative mode. In both modes, the device can train the AI ​​model using one or more reference distributions.

[0131] Taking a reference distribution as an example, some exemplary representations of the reference distribution are described.

[0132] In some embodiments, the reference distribution may be a parameterized standard distribution. The distribution takes the form of a normal distribution, such as the Poisson distribution, Rayleigh distribution, etc. R represents the reference distribution. This represents the parameterized standard distribution. and This represents the statistical parameters used to describe the parameterized standard distribution.

[0133] In some embodiments, the reference distribution may be a combination of multiple parameterized standard distributions.

[0134] For example, the reference distribution can be a linear combination of multiple Gaussian distributions.

[0135] In some embodiments, the reference distribution can be in the form of multiple reference data samples, for example... . This is the first reference sample used to represent the reference distribution. This is the second reference sample used to represent the reference distribution, and so on. M is the number of reference data samples used to represent the reference distribution.

[0136] The reference data sample can be the original dimension of the reference data or the dimension after compression of the reference data. In other words, the reference data sample can be the original data sample or the compressed data sample. The original dimension can be the dimension of the output data of the layer corresponding to the reference distribution.

[0137] In one possible implementation, the compressed data sample can be obtained by compressing the original data sample according to the first transformation matrix. The first transformation matrix can be a unitary matrix or an orthonormal matrix.

[0138] In some embodiments, each basis vector of the first transformation matrix can be a standard basis such as a Fourier basis, a DCT basis, a wavelet basis, etc.

[0139] In some embodiments, the basis vectors of the first transformation matrix can be constructed as needed. For example, the basis vectors of the first transformation matrix can be constructed based on the reference distribution.

[0140] For example, an original data sample x can be expressed as sample, where n is an integer greater than 1. The first transformation matrix U can be expressed as matrix, where r is a positive integer less than n. The matrix U can be a unitary matrix. In this case, , . c is the compressed data sample. The reference data sample can be x or c.

[0141] In one possible implementation, the compressed data sample can be obtained by compressing the sampling result of the original data sample according to the second transformation matrix. The sampling result of the original data sample is obtained by sampling some values of the original data sample with a sampling matrix.

[0142] Optionally, the sampling matrix can be a random matrix or a pseudo-random matrix. The first transformation matrix can be sampled into a compact matrix smaller than the first transformation matrix through the sampling matrix.

[0143] For example, the sampling matrix P can be as follows: .

[0144] The number of rows in the sampling matrix is the number of positions sampled in the original data sample.

[0145] For example, the sampling matrix P corresponding to the original data sample x can be applied to the first transformation matrix U. P can be denoted as matrix, where m < n and m is a positive integer. Further, m << n. P can be used to "compress" U into the second transformation matrix , which is a matrix of matrix. The compressed data samples can be obtained through... get. It consists of the values ​​sampled from x through the sampling matrix P. sample. This is the left inverse of the second transformation matrix.

[0146] The above are just some examples. Compressed data samples can also be obtained using other compression methods. The embodiments in this application are not limited to these.

[0147] A reference distribution represented by one or more standard distributions saves more radio resources than a reference distribution represented by multiple reference data samples.

[0148] When Q>1, the representations of the Q reference distributions can be the same or different.

[0149] For example, the second network element can send information #1 via broadcast, multicast, or unicast.

[0150] The following describes some examples of information #1.

[0151] For example, information #1 can indicate the statistical parameters of Q reference distributions.

[0152] As mentioned earlier, the reference distribution can be a single parameterized standard distribution or a combination of multiple parameterized standard distributions. Information #1 can indicate some statistical parameters describing the reference distribution.

[0153] For example, information #1 can indicate a reference data sample used to represent Q reference distributions.

[0154] For example, information #1 can indicate the index of Q reference distributions.

[0155] For example, there may be multiple candidate reference distributions in the first network element. Information #1 may include the indices of Q reference distributions among the multiple candidates.

[0156] Information #1 can also be in other forms, as long as the form can indicate Q reference distributions.

[0157] Step 710 is optional. The first network element can determine the Q reference distributions in other ways. For example, the Q reference distributions can be predefined. Alternatively, the Q reference distributions can be determined by the first network element itself.

[0158] Each of the Q layers can include at least one potential layer of one or more AI models.

[0159] For example, the Q layers include one or more potential layers of the first AI model.

[0160] In addition, the Q layers may include one or more input layers and / or output layers of one or more AI models.

[0161] In some embodiments, the Q layers can be determined by a second network element.

[0162] Method 700 may further include: the second network element may send information #2 (an example of third information) indicating Q layers to the first network element.

[0163] Information #2 is used to indicate which layer each of the Q reference distributions corresponds to.

[0164] Taking Q=2 as an example. The Q reference distributions can include reference distribution #1 and reference distribution #2. Information #2 is used to indicate which layer reference distribution #1 corresponds to and which layer reference distribution #2 corresponds to.

[0165] For example, information #2 may include Q indications that indicate Q layers respectively.

[0166] For example, these Q indicators could be indices for Q layers.

[0167] Information #2 can also be in other forms, as long as the form can indicate which reference distribution corresponds to which layer.

[0168] In some embodiments, the Q layers can be determined by a first network element. Before step 710, the first network element can send information #3 indicating the Q layers to a second network element. The second network element can determine the Q reference distributions based on the information #3.

[0169] The format of Message #3 can be referenced from Message #2, and will not be repeated here.

[0170] In some embodiments, the correspondence between the Q layers and the Q reference distributions can be predefined.

[0171] In some embodiments, the first network element may select q reference distributions from Q reference distributions to train the AI ​​model.

[0172] The q layers corresponding to the q reference distributions may include at least one latent layer of the AI ​​model (the first AI model).

[0173] Optionally, the q layers corresponding to the q reference distributions can be the q latent layers of the AI ​​model (the first AI model).

[0174] In addition, the q layers may also include the output layer of the AI ​​model (the first AI model).

[0175] Step 720 may include: the first network element trains the AI ​​model based on the distance between the q reference distributions and the distributions of the corresponding q layers of the AI ​​model (the first AI model).

[0176] In the embodiments of this application, the distance between the two can also be understood as the difference between them. For example, the distance between q reference distributions and the distributions of the corresponding q layers can also be understood as the difference between the q reference distributions and the distributions of the corresponding q layers.

[0177] The training dataset can be found in the previous text, for example... Figure 5 The relevant content will not be repeated in this article.

[0178] Each of the q reference distributions corresponds to one of the q layers. Each of the q layers belongs to a further Q layers. The distance between the reference distribution of each of the q layers and the distribution of that layer can be measured.

[0179] Layer distribution refers to the distribution of the output of that layer.

[0180] The distance between the reference distribution of a layer and the distribution of that layer can also be regarded as the distance between the reference distribution of that layer and the layer output during the training period.

[0181] For example, the q reference distributions can include reference distribution #1 and reference distribution #2. Reference distribution #1 can correspond to the output of layer #1, and reference distribution #2 can correspond to the output of layer #2.

[0182] The distance between the reference distribution of each of the q layers and the distribution of that layer can be measured using a scoring function.

[0183] The scoring function is mathematically differentiable.

[0184] For example, the scoring function can be based on one of the following: Kullback-Leibler divergence (KL divergence), graph edit distance, Wasserstein distance, or Jensen-Shanon distance (JSD distance).

[0185] When q>1, the scoring function can be the same or different for different reference distributions.

[0186] Taking q=2 as an example. The Q reference distributions can include reference distribution #1 and reference distribution #2 Scoring function #1 Used for measurement and reference distribution #1 The distance. This can represent the reference distribution #1 The corresponding latent layer output. Scoring function #2 Used for measurement and reference distribution #2 The distance. This can represent the reference distribution #2 The corresponding latent layer output. and They can be the same or different.

[0187] In some embodiments, the q scoring functions used to measure the distance between the q reference distributions and the distributions of the q layers can be determined by the second network element.

[0188] Method 700 may further include: the first network element may receive information #4 (an example of fourth information) indicating Q scoring functions from the second network element, the Q scoring functions being used to measure the distance between the Q reference distributions and the distributions of the Q layers respectively.

[0189] Taking Q=2 as an example. The Q reference distributions can include reference distribution #1 and reference distribution #2. Information #4 indicates: the scoring function #1 used to measure the distance to reference distribution #1, and the scoring function #2 used to measure the distance to reference distribution #2.

[0190] For example, information #4 may include Q scoring functions.

[0191] For example, information #4 could include indices of Q scoring functions.

[0192] In some embodiments, the q scoring functions can be determined by the first network element.

[0193] The first network element can also obtain q scoring functions through other methods. For example, the correspondence between the Q scoring functions and the Q layers can be predefined. The first network element can obtain q scoring functions based on the q layers.

[0194] In step 720, the first network element can establish one or more regularizations using q reference distributions. These one or more regularizations are used to minimize the distance between the q reference distributions and the distributions of the corresponding q layers.

[0195] The first network element can establish q regularizations for the distributions of q layers based on q reference distributions.

[0196] The first network element can acquire a second AI model with one or more regularizations, which are used to minimize the difference between the distributions of q reference distributions and the distributions of the corresponding q layers in the first AI model.

[0197] In other words, the first network element can be trained on an AI model based on an objective function that includes one or more regularizations, which are used to minimize the distance between the q reference distributions and the distributions of the corresponding q layers.

[0198] During the training period, the first network element can minimize the scoring function as one or more training-optimal regularizations. These one or more regularizations can be viewed as additional constraints in the learning process.

[0199] In this way, the first network element can shape the outputs of q layers using q reference distributions. Furthermore, the outputs of the q layers can include the output of at least one latent layer, in which case the output of that at least one latent layer can be shaped according to the reference distributions.

[0200] A reference distribution can be set as needed to obtain the required AI model, which helps improve the training performance of the AI ​​model.

[0201] The above technical solution can facilitate the interconnection and interoperability of multiple AI models.

[0202] The q reference distributions can be consistent with the outputs of q layers in another AI model that needs to work with the second AI model. According to the above technical solution, the distribution of the q layers in the second AI model can be as close as possible to the distribution of the q layers in the other AI model, which facilitates interoperability between the two AI models.

[0203] In some scenarios, multiple AI models may need to work collaboratively. The output of the latent layer of one model (e.g., model #A) can be the input of another model (e.g., model #B), or the input of the latent layer of another model. According to the technical solution of this application embodiment, training model #A based on a reference distribution is beneficial in making the distribution of the latent layer output of model #A as close as possible to the reference distribution. The reference distribution can be consistent with the input of model #B. In this case, the distribution of the latent layer output of model #A is as close as possible to the distribution of the input of model #B, which facilitates the interconnection and interoperability between model #A and model #B. A relevant example can be found in Exemplary Scenario-2 below.

[0204] The above technical solutions can help improve the training performance of AI models.

[0205] According to embodiments of this application, adding one or more regularizations to the intermediate part of the AI ​​model is equivalent to introducing additional loss. Regularizing the output of the latent layer with one or more regularizations facilitates more thorough training of the latent layer, improves the transparency and directness of the latent layer learning process, and enhances convergence efficiency.

[0206] Figure 8 This is a schematic diagram illustrating an exemplary regularization of the distribution of the output of a potential layer.

[0207] like Figure 8As shown, the first network element trains a DNN-based autoencoder (AE) using the training dataset. The AE includes the encoder. and decoder The input to the encoder (AE) is the input to the encoder. The output of the encoder is the input to the decoder. The output of the encoder (AE) is the output of the decoder. The relationship between the encoder's input and output can be expressed as follows: . These represent the parameters of the encoder. The relationship between the input and output of the decoder can be expressed as: . This represents the parameters of the decoder. The training objective of the AE is to minimize the input. With output The differences between them. For example, the loss function can be based on the mean squared error (MSE). The training objective of AE is to minimize the loss function, for example... .

[0208] The encoder's output can be viewed as the latent layer output of the AE. Regularization of the latent layer output distribution is used to minimize the scoring function, for example... Regularization can be viewed as an additional constraint in the learning process. This represents the ratio of the reference distribution R to the output of the latent layer. A scoring function for the differences between them.

[0209] The above are merely exemplary scenarios and do not constitute a limitation on the technical solutions of the embodiments of this application. For example, the AI ​​model can also be other models. Furthermore, q can also be other values.

[0210] Taking q=2 as an example, the first network element can establish regularization #1, which minimizes the reference distribution #1. With the corresponding latent layer output #1 Distance between: ,in this case, . Indicates the input of the AI ​​model With the corresponding latent layer output #1 The mapping relationship between them. This represents the parameters of the AI ​​model. This represents the parameters of the AI ​​model after the parameter update. The first network element can establish regularization #2, which minimizes the distance between the reference distribution #2 and the corresponding latent layer output #2: ,in this case, . This represents the mapping relationship between the input of the AI ​​model and the corresponding latent layer output #2. The two regularizations mentioned above can be regarded as additional constraints in the learning process.

[0211] If the reference distribution is in the form of multiple reference data samples, the first network element can input them into the corresponding regularization.

[0212] Optionally, the reference data sample can be shuffled by the first network element before being input into regularization.

[0213] This allows for data augmentation, which helps improve the training performance of AI models.

[0214] Taking the two regularization methods mentioned above as examples: If reference distribution #1 is in the form of multiple reference data samples, the first network element can shuffle the reference data samples and input them into regularization #1. If reference distribution #2 is in the form of multiple reference data samples, the first network element can shuffle the reference data samples and input them into regularization #2.

[0215] When the reference distribution is in the form of multiple reference data samples, and the reference data samples are compressed data samples, the distance between the reference distribution and the distribution of the corresponding layer can be the distance in the original dimension or the distance in the compressed dimension.

[0216] Optionally, when the reference data sample is a compressed data sample, the distance between the reference distribution and the distribution of the corresponding layer can be based on the compression result of the reference data sample and the output of the corresponding layer. In this case, the distance between the reference distribution and the distribution of the corresponding layer is the distance of the compressed dimension.

[0217] For example, when the reference data sample is a compressed data sample, the first network element can shuffle the reference data sample and input it into the corresponding regularization. The first network element can compress the output of the corresponding layer and input it into the corresponding regularization.

[0218] The compression method can be found in the previous text, and will not be repeated here.

[0219] Optionally, when the reference data sample is a compressed data sample, the distance between the reference distribution and the distribution of the corresponding layer can be based on the decompression result of the reference data sample and the output of the corresponding layer. In this case, the distance between the reference distribution and the distribution of the corresponding layer is the distance of the original dimension.

[0220] For example, when the reference data sample is a compressed data sample, the first network element can decompress the reference data sample, shuffle the decompression results, and input them into the corresponding regularization. The first network element can also input the output of the corresponding layer into the corresponding regularization.

[0221] The decompression method can refer to the previous compression method, which involves reversing the compression process, and will not be elaborated on in this article.

[0222] If the reference distribution is in the form of a parameterized standard distribution, the first network element can generate multiple reference data samples according to the parameterized standard distribution and input these reference data samples into the corresponding regularization.

[0223] For example, the first network element can generate multiple reference data samples by randomly parameterizing a standard distribution.

[0224] Taking the two regularization methods mentioned above as examples: If reference distribution #1 is in the form of a parameterized standard distribution, then the first network element can generate multiple reference data samples using a randomized parameterized standard distribution and input these reference data samples into regularization #1. If reference distribution #2 is in the form of a parameterized standard distribution, then the first network element can generate multiple reference data samples using a randomized parameterized standard distribution and input these reference data samples into regularization #2.

[0225] When q>1, the q reference distributions can be the same or different.

[0226] The first network element can measure the result scores of q rules. The result scores of q rules are related to the distances between the q reference distributions and the corresponding layer outputs of the current AI model.

[0227] There is a positive correlation between the resulting scores of the q rules and the distances between the q reference distributions and the corresponding layer outputs of the current AI model.

[0228] For example, the resulting scores of q rules can be the distances between q reference distributions and the corresponding layer outputs of the current AI model.

[0229] The positive correlation described above is merely an example; other relationships, such as a negative correlation, can also exist between the two. The technical solutions in this application embodiment are illustrated using only a positive correlation as an example. When other relationships exist between the two, the corresponding descriptions can be adjusted appropriately.

[0230] For example, in each round, the first network element can measure the lowest score it can achieve through optimization rules.

[0231] In each round, the minimum score can be the result score of q rules measured after a round ends.

[0232] In other words, after a round, the first network element can measure the result scores of q rules.

[0233] For example, the first network element can measure the result scores of q rules after training is complete.

[0234] Taking the two regularization methods mentioned above as examples, the score for regularization #1 is... ,in this case, The score obtained after regularization #2 is... ,in this case, .

[0235] Optionally, the first network element can memorize the result scores of q rules.

[0236] Alternatively, method 700 may also include step 730.

[0237] 730: The first network element sends information #5 (an example of the second information) indicating the optimization results of one or more regularizations to one or more other devices, such as the second network element.

[0238] For example, the optimization result of one or more regularizations can be the result scores of q regularizations.

[0239] For example, step 730 can be performed by the communication module of the first network element.

[0240] The first network element can send the result score via broadcast, multicast, or unicast.

[0241] Message #5 can be represented in various forms.

[0242] For example, message #5 may include q regularized result scores. In other words, the first network element sends q regularized result scores.

[0243] For example, multiple ranges can exist. Each range corresponds to a level. Message #5 can indicate the q levels corresponding to the ranges to which the q result scores belong.

[0244] The following describes an exemplary description of the timing for sending message #5.

[0245] In some embodiments, once the result scores of q rules have been measured, the first network element can send information #5.

[0246] For example, after a round ends, the first network element can send message #5.

[0247] For example, after a batch ends, the first network element can send message #5.

[0248] For example, after training is complete, the first network element can send message #5.

[0249] In some embodiments, the first network element may send information #5 in response to a request for a result score sent by the second network element or other network elements.

[0250] In some embodiments, if the resulting score is less than or equal to one or more thresholds, the first network element may send information #5.

[0251] If the resulting score is less than or equal to one or more thresholds, the AI ​​model is considered to meet the requirements. For example, the second network can consider the AI ​​model on the first network element as a candidate that meets the usage conditions.

[0252] For example, if the resulting score is consistently below the corresponding threshold, the AI ​​model meets the requirements. In the case of q > 1, the q thresholds can be the same or different. The thresholds can be predefined. Alternatively, the thresholds can be received by the device. Or, the thresholds can be determined by the device itself.

[0253] Take the interconnection of AEs on two devices as an example. The output of the encoder on device #1 is the input of the decoder on device #2. The distribution of the input of the decoder on device #2 can be used as a reference distribution. If, after training, the distance between the output of the encoder on device #1 and the reference distribution is less than or equal to a threshold, then the AE on device #1 can be used as a candidate for interconnection with the AE on device #2.

[0254] In some embodiments, if the result score is greater than or equal to one or more thresholds, the first network element may send information #5.

[0255] The method 700 of this application embodiment is illustrated below based on two examples (Exemplary Scenario-1 and Exemplary Scenario-2).

[0256] Exemplary Scenario-1 Optionally, method 700 can be applied to federated learning.

[0257] One type of communication system includes a central device and multiple working devices. For example, the working devices may include... Figure 3 The modules shown include a perception module for collecting local data, an AI module for training local AI models such as DNNs, and a communication module for receiving signals and / or data from and transmitting signals and / or data to the central device. The central device may include at least... Figure 3 The communication module and AI module are shown.

[0258] For example, the central device can be the second network element in method 700, and the working device can be the first network element in method 700.

[0259] The following describes an example of one possible implementation of federated learning.

[0260] The central device and worker devices can collaborate round-by-round using federated learning. Specifically, the communication module of a worker device sends all or a portion of its local neurons to the central device. The central device's communication module receives these neurons from multiple worker devices, its AI module aggregates these neurons and updates its AI model accordingly, then the communication module broadcasts or multicasts the updated neurons to the worker devices. For example, the central device's AI module averages these neurons, and then its communication module sends the averaged neurons to the worker devices. The worker device's communication module receives the updated neurons, and its AI module sets them into its local DNN. Then, the worker device's AI module trains the updated local DNN. This process is repeated round-by-round and batch-by-batch until both the central device and worker devices have completed DNN training. In federated learning, all DNNs trained on the involved worker devices must have the same architecture.

[0261] Based on the aforementioned traditional federated learning, the following describes an example of how the technical solution of this application is applied to federated learning.

[0262] The communication module of the central device can send information #1, which indicates Q reference distributions (e.g., Q=2), to the working device via broadcast, multicast, or unicast.

[0263] In addition, the central device can also instruct one or more of the following: the scoring functions corresponding to the two reference distributions, and the layers of the AI ​​model corresponding to the two reference distributions.

[0264] For example, the communication module of the central device can send messages to the working devices via broadcast, multicast, or unicast to notify the working devices which layer, which scoring function, and which reference distribution to use for regularizing the training scores.

[0265] The above process corresponds to step 710. For a detailed description, please refer to step 710. This article will not repeat it here.

[0266] The working device's AI module regularizes its undertrained AI model using a reference distribution. The working device trains its AI model in one round.

[0267] In addition, the AI ​​model of the working device can remember the regularized result score after the round ends.

[0268] The above process corresponds to step 720. For a detailed description, please refer to step 720. This article will not repeat it here.

[0269] In addition, the communication module of the working device can send the regularized result score to the central device.

[0270] The above process corresponds to step 730. For a detailed description, please refer to step 730. This article will not repeat it here.

[0271] The above is merely an exemplary process of applying the technical solutions in this application to federated learning. When the technical solutions in this application are applied to federated learning, they can also be implemented in other ways; relevant descriptions can be found in method 700, and will not be repeated here.

[0272] The regularized result score can reflect the training status of the AI ​​model on the work device. For example, an AI model with a lower result score may have higher training quality and / or faster training speed.

[0273] In addition, the central equipment can schedule working equipment based on the result score of the regularization.

[0274] Exemplary Scenario-2 In some scenarios, multiple AI models deployed on different devices may need to work together. These AI models may be trained independently by different providers.

[0275] For example, encoders and decoders deployed on different devices may need to work together.

[0276] Optionally, method 700 can be applied to train autoencoders on different devices. After training, the encoder can be deployed at the transmitting end, and the decoder can be deployed at the receiving end. The transmitting end is the encoding device. The receiving end is the decoding device. The encoder of the encoding device can output to the decoder of the decoding device.

[0277] The following example uses a DNN-based autoencoder. The encoder can be an encoding DNN, and the decoder can be a decoding DNN.

[0278] There are two devices, device #1 and device #2, used for training the Advanced Effect (AE). For example, device #1 may include... Figure 3 The illustrated modules include a perception module for collecting local data, an AI module for training a DNN-based autoencoder #1 using its local data, and a communication module for receiving and transmitting signals and / or data. Device #2 may include... Figure 3 The modules shown include a perception module for collecting local data, an AI module for training a DNN-based autoencoder #2 using its local data, and a communication module for receiving and / or sending signals and / or data.

[0279] Device #1 can be the first network element.

[0280] Taking Q=1 as an example, the reference distribution can correspond to the output of the encoder in AE.

[0281] The AI ​​model of device #1 is regularized on its undertrained DNN-based autoencoder #1 using a reference distribution. Device #1 trains its DNN-based autoencoder #1 in one round.

[0282] In addition, the AI ​​model of device #1 can remember the regularized result score after the round ends.

[0283] The above process corresponds to step 720. For a detailed description, please refer to step 720. This article will not repeat it here.

[0284] In some implementations, the reference distribution can be the distribution of encoder outputs in the DNN-based autoencoder #2 on device #2.

[0285] For example, after a round, device #2 can send the distribution of encoder output in the DNN-based autoencoder #2 to device #1. In this case, device #2 can be regarded as a second network element.

[0286] The above process corresponds to step 710. For a detailed description, please refer to step 710. This article will not repeat it here.

[0287] Figure 9 This is a schematic diagram of an exemplary training process for AE.

[0288] The AI ​​module of device #2 trains autoencoder #2 using its local data. The training objective could be to minimize the input of autoencoder #2 (e.g., Figure 9 In ) and the output of the automatic encoder #2 (e.g., Figure 9 In The difference between ). and The difference between them can be measured by their mean squared error, which is, for example, Figure 9 mse( , ). This indicates the encoder for automatic encoder #2. Indicates encoder The parameters. This indicates the decoder for autoencoder #2. Indicate decoder The parameters are as follows. The encoder's output is the decoder's input.

[0289] For example, after a round is completed, device #2 can send the distribution of the encoder output in the automatic encoder #2 as a reference distribution R to device #1. The reference distribution R can be... This refers to the output of the encoder in autoencoder #2. Device #1 establishes regularization on its undertrained autoencoder #1 using the reference distribution R. The reference distribution R corresponds to the latent layer output. That is, the output of the encoder. This is a scoring function used to measure the distance between two distributions.

[0290] The AI ​​module of device #1 trains autoencoder #1 using its local data. The training objective could be to minimize the input of autoencoder #1 (e.g., Figure 9 In ) and the output of the automatic encoder #1 (e.g., Figure 9 In The difference between ). and The difference between them can be measured by their mean squared error, which is, for example, Figure 9 mse( , ). This indicates the encoder for automatic encoder #1. Indicates encoder The parameters. This indicates the decoder for autoencoder #1. Indicate decoder The parameters are as follows. The encoder's output is the decoder's input. Furthermore, during the training cycle, the AI ​​module of device #1 can also minimize the scoring function. As the optimal regularization for training.

[0291] In addition, the communication module of device #1 can send the regularized result score to device #2.

[0292] For example, the communication system of device #2 can send a message to device #1 to request device #1 to measure and return the result score of the regularization.

[0293] The above process corresponds to step 730. For a detailed description, please refer to step 730. This article will not repeat it here.

[0294] In some implementations, the reference distribution can be a common reference distribution used for both the DNN-based autoencoder #1 and the DNN-based autoencoder #2.

[0295] In this scenario, the AI ​​model of device #1 also builds regularization on its undertrained DNN-based autoencoder #2 using a reference distribution. Device #2 trains its DNN-based autoencoder #2 in one round.

[0296] In addition, the AI ​​model of device #2 can remember the regularized result score after the round ends.

[0297] The above process corresponds to step 720. For a detailed description, please refer to step 720. This article will not repeat it here.

[0298] For example, a reference distribution can be sent from device #2 to device #1. Alternatively, a reference distribution can be sent from device #3 to both device #1 and device #2.

[0299] The device that transmits the reference distribution can be considered as a second network element.

[0300] The above process corresponds to step 710. For a detailed description, please refer to step 710. This article will not repeat it here.

[0301] Figure 10 This is a schematic diagram of an exemplary training process for AE.

[0302] The AI ​​module of device #1 establishes regularization on its undertrained autoencoder #1 using a reference distribution R. The reference distribution R corresponds to the latent layer output. This refers to the encoder's output. The AI ​​module of device #2 establishes regularization on its undertrained autoencoder #2 using the reference distribution R. The reference distribution R corresponds to the latent layer output. That is, the output of the encoder.

[0303] This is a scoring function used to measure the distance between two distributions.

[0304] The AI ​​module of device #1 trains autoencoder #1 using its local data. The training objective can be to minimize the difference between the input and output of autoencoder #1. and The difference between them can be measured by their mean squared error, which is, for example, Figure 10 mse( , Furthermore, during the training cycle, the AI ​​module of device #1 can also minimize the scoring function. As the optimal regularization for training.

[0305] The AI ​​module of device #2 trains autoencoder #2 using its local data. The training objective can be to minimize the difference between the input and output of autoencoder #2. and The difference between them can be measured by their mean squared error, which is, for example, Figure 10 mse( , Furthermore, during the training cycle, the AI ​​module of device #2 can also minimize the scoring function. This serves as the optimal regularization for training.

[0306] According to the technical solution of this application embodiment, training autoencoder #1 and autoencoder #2 based on the same reference distribution R is beneficial in making the distribution of the latent layer outputs of autoencoder #1 and autoencoder #2 as close as possible to the reference distribution. In this way, the distributions of the latent layer outputs of autoencoder #1 and autoencoder #2 can be consistent with each other, which is beneficial for achieving interconnection between model #A and model #B.

[0307] In addition, the communication module of device #1 can send the regularized result score to device #2 and / or device #3.

[0308] For example, the communication system of device #2 can send a message to device #1 to request device #1 to measure and return the result score of the regularization.

[0309] The above process corresponds to step 730. For a detailed description, please refer to step 730. This article will not repeat it here.

[0310] For example, such as Figure 9 or Figure 10 As shown, the encoding DNN of autoencoder #1 trained on device #1 can be deployed on the encoding device, and the decoding DNN of autoencoder #2 trained on device #2 can be deployed on the decoding device.

[0311] Alternatively, the decoding DNN of the autoencoder #1 trained on device #1 can be deployed on the decoding device, and the encoding DNN of the autoencoder #2 trained on device #2 can be deployed on the encoding device.

[0312] The above is merely an exemplary process of applying the technical solutions in this application to AE training. When the technical solutions in this application are applied to AE training, they can also be implemented in other ways; relevant descriptions can be found in method 700, and will not be repeated here.

[0313] The transmission processes in Exemplary Scenario-1 and Exemplary Scenario-2 are merely examples. For other implementation methods, please refer to Method 700.

[0314] The communication method provided in the embodiments of this application has been described in detail above. The following will refer to Figures 11 to 15 The communication device provided in the embodiments of this application is described in detail.

[0315] Figure 11 This is a schematic block diagram of the communication device 10 provided in an embodiment of this application. Figure 11 As shown, the communication device 10 may include: Transceiver module 11 is used to receive first information indicating Q reference distributions corresponding to Q layers of the first AI model, where Q is a positive integer; Processing module 12 is used to obtain a second AI model based on q of the Q reference distributions, where q is a positive integer and q≤Q.

[0316] The communication device 10 in this embodiment can correspond to the first network element in the communication method described above. The management operations and / or functions of each module of the communication device 10, as well as other management operations and / or functions, are intended to implement the corresponding steps of the above method. For the sake of brevity, these will not be elaborated further here.

[0317] The transceiver module 11 in this embodiment can be implemented by a transceiver. The processing module 12 in this embodiment can be implemented by a processor.

[0318] like Figure 12 As shown, the communication device 20 may include a transceiver 21. Optionally, the communication device 20 may also include a processor 22 and / or a memory 23. The memory 23 may be used to store instruction information, or to store code, instructions, etc., to be executed by the processor 22.

[0319] Figure 13 This is a schematic block diagram of the communication device 30 provided in an embodiment of this application. Figure 13 As shown, the communication device 30 may include: Processing module 31 is used to obtain Q reference distributions corresponding to Q layers of the first AI model, wherein the Q reference distributions are used to obtain the second AI model, and Q is a positive integer; The transceiver module 32 is used to send first information indicating the Q reference distributions.

[0320] The communication device 30 in this embodiment can correspond to the second network element in the communication method described above. The management operations and / or functions of each module of the communication device 30, and other management operations and / or functions, are designed to implement the corresponding steps of the above method. For the sake of brevity, further details are omitted here.

[0321] The processing module 31 in this embodiment can be implemented by a processor. The transceiver module 32 in this embodiment can be implemented by a transceiver.

[0322] like Figure 14 As shown, the communication device 40 may include a transceiver 41. Optionally, the communication device 40 may also include a processor 42 and / or a memory 43. The memory 43 may be used to store instruction information, or to store code, instructions, etc., to be executed by the processor 42.

[0323] Processor 22 or processor 42 can be an integrated circuit chip with signal processing capabilities. In implementation, each step in the above method embodiments can be implemented using hardware integrated logic circuits in the processor or by using software instructions. Processor 22 or processor 42 can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. All methods, steps, and logic block diagrams disclosed in the embodiments of this application can be implemented or executed. A general-purpose processor can be a microprocessor, or the processor can be any conventional processor, etc. The steps of the methods disclosed in the embodiments of this invention can be directly executed and completed by a hardware decoding processor, or executed and completed using a combination of hardware and software modules in the decoding processor. The software modules can reside in storage media known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. The storage medium is located in memory, and the processor reads information from the memory and combines the processor's hardware to complete the steps of the above methods.

[0324] It is understood that the memory 23 or memory 43 in the embodiments of the present invention can be volatile memory or non-volatile memory, and may include both volatile and non-volatile memory. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM) used as an external cache. By way of example rather than limitation, many forms of RAM may be used, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus dynamic random access memory (DR RAM). The storage of the systems and methods described in this specification is intended to include, but is not limited to, these and any other suitable storage.

[0325] One embodiment of this application also provides a system. For example... Figure 15 As shown, system 50 includes: The communication device 10 and the communication device 20 provided in the embodiments of this application are described in this application.

[0326] An embodiment of this application also provides a computer storage medium that can store one or more instructions for performing any of the above methods.

[0327] Alternatively, the storage medium may specifically be memory 23 or 43.

[0328] Those skilled in the art will recognize that, in conjunction with the examples described in the embodiments disclosed in this specification, each unit and algorithm step can be implemented using electronic hardware or a combination of computer software and electronic hardware. Whether the function is performed by hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but should not consider that this embodiment is beyond the scope of this application.

[0329] Those skilled in the art will understand that, for convenience and brevity, the detailed working process of the above-described systems, devices, and units can be referred to the corresponding process in the above-described method embodiments, and will not be repeated here.

[0330] In the several embodiments provided in this application, the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the described apparatus embodiments are merely examples. For example, unit partitioning is a logical functional partitioning, and other partitioning methods can be used in actual embodiments. For example, multiple units or components can be merged or integrated into another system, or some features can be ignored or not performed. Furthermore, the mutual coupling or direct coupling or communication connection shown or described can be implemented through various communication interfaces. Indirect coupling or communication connection between devices or units can be implemented electronically, mechanically, or otherwise.

[0331] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, these components may be located in one unit or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0332] Furthermore, the functional units in the embodiments of this application can be integrated into a processing unit. Each of these units can exist physically independently, or two or more units can be integrated into one unit.

[0333] When these functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. The technical solution of this application can be implemented as a software product. The software product is stored in a storage medium and includes several instructions to instruct a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the embodiments of this application. The aforementioned storage medium includes any medium capable of storing program code, such as a USB flash drive, removable hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk, optical disk, etc.

[0334] The above description is merely a specific embodiment of this application and is not intended to limit the scope of protection of this application. Any variations or substitutions that are readily conceived by those skilled in the art within the scope of the technology disclosed in this application should fall within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A communication method characterized by comprising: comprising: receiving first information indicating Q reference distributions corresponding to Q layers of a first AI model, wherein Q is a positive integer; obtaining a second AI model according to q reference distributions of the Q reference distributions, wherein q is a positive integer, q≤Q.

2. The communication method according to claim 1, characterized by, The obtaining of the second AI model according to q reference distributions of the Q reference distributions comprises: obtaining the second AI model with one or more regularizations for minimizing differences between the q reference distributions and distributions of corresponding q layers of the Q layers.

3. The communication method according to claim 2, wherein, Further comprising: sending second information indicating an optimization result of the one or more regularizations.

4. The communication method according to any one of claims 1 to 3, characterized by, The Q layers comprise one or more latent layers of the first AI model.

5. The communication method according to any one of claims 1 to 4, characterized by, Further comprising: receiving third information indicating the Q layers.

6. The communication method according to any one of claims 1 to 5, characterized by, Further comprising: receiving fourth information indicating Q scoring functions for measuring differences between the Q reference distributions and distributions of the Q layers.

7. A communication method characterized by comprising: comprising: obtaining Q reference distributions corresponding to Q layers of a first AI model, wherein the Q reference distributions are used for obtaining a second AI model, Q being a positive integer; sending first information indicating the Q reference distributions.

8. The communication method according to claim 7, wherein, The Q reference distributions are used for establishing one or more regularizations for minimizing differences between the Q reference distributions and distributions of the Q layers.

9. The communication method according to claim 8, wherein, Further comprising: receiving second information indicating an optimization result of the one or more regularizations.

10. The communication method according to any one of claims 7 to 9, characterized by, The Q layers comprise one or more latent layers of the first AI model.

11. The communication method according to any one of claims 7 to 10, characterized by, Further comprising: receiving third information indicating the Q layers.

12. The communication method according to any one of claims 7 to 11, characterized by, Further comprising: receiving fourth information indicating Q scoring functions for measuring differences between the Q reference distributions and distributions of the Q layers.

13. An apparatus, comprising: The apparatus comprises a processor and a memory, the memory storing one or more instructions executable on the processor, when the one or more instructions are executed, the apparatus is capable of performing the method according to any one of claims 1 to 6 or performing the method according to any one of claims 7 to 12.

14. An apparatus, comprising: The apparatus comprises units for performing the method according to any one of claims 1 to 6 or performing the method according to any one of claims 7 to 12.

15. A communication system, characterized by comprising a first communication apparatus and a second communication apparatus, wherein the first communication apparatus performs the method according to any one of claims 1 to 6, and the second communication apparatus performs the method according to any one of claims 7 to 12.

16. A computer-readable storage medium, characterized in that, comprising one or more instructions, when the one or more instructions are executed on a computer, the computer performs the method according to any one of claims 1 to 6 or the method according to any one of claims 7 to 12.