Model training method and communication device

By referring to the local data sets of other devices in the wireless communication system for model training, the problem of long model alignment time and large overhead caused by independent model training of sending and receiving devices is solved, and the effect of shortening model alignment time and reducing complexity and overhead is achieved.

CN120166437APending Publication Date: 2025-06-17HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311737963.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-15
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

In wireless communication systems, the model training process of the transmitting and receiving devices is independent, resulting in a long model alignment time and a large overhead.

Method used

By referring to the local data set of other devices for model training, the connection between the model of the first device and the model of the second device is established, thereby simplifying the subsequent model alignment process.

Benefits of technology

Through this method, the time of model alignment can be shortened, and the complexity and overhead of model alignment can be reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120166437A_ABST
    Figure CN120166437A_ABST
Patent Text Reader

Abstract

The invention provides a model training method and a communication device, and the method comprises the steps: a first device receives data feature information of a second local data set from a second device, and the second local data set is a data set of the second device; based on the data feature information of the second local data set and the first local data set, a first training data set is determined, the first local data set is a data set of first equipment, and the data feature information of the first training data set comprises the data feature information of the second local data set and the data feature information of the first local data set; and obtaining a first model based on the first training data set, wherein the first model is a model deployed in first equipment. Through the method, the first device can refer to the local data sets of other devices to carry out model training, subsequent model alignment with other devices is facilitated, and the time of model alignment is shortened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technologies, and in particular, to a model training method and a communication device. Background Art

[0002] Combining a wireless communication system with a neural network, that is, training a neural network model in the wireless communication system through data-driven, is beneficial to improving the performance of the wireless communication system. The neural network model can be applied to various devices in the wireless communication system (including receiving-end devices or transmitting-end devices, etc.). For example, applying the neural network model to the channel decoding module in the receiving-end device, through data-driven training, is beneficial to improving the channel decoding performance of the receiving-end device.

[0003] However, the transmitting-end device and the receiving-end device in the communication process may belong to different manufacturers, that is, the model deployed on the transmitting-end device (abbreviated as the transmitting model) and the model deployed on the receiving-end device (abbreviated as the receiving model) correspond to different training processes, which may cause the transmitting-end device and the receiving-end device to be unable to communicate. To avoid this situation, before the transmitting-end device and the receiving-end device communicate, the transmitting model and the receiving model need to be aligned (also referred to as adaptation or joint training) operation.

[0004] Generally, the alignment process of the transmitting model and the receiving model can be understood as a process of aligning and training the transmitting model and the receiving model through the data in a training dataset. However, since the training processes of the transmitting model and the receiving model are independent, the alignment training process takes a long time and has a large overhead. Summary of the Invention

[0005] This application provides a model training method and a communication device, which can refer to the local dataset of other devices for model training, which is beneficial to subsequent model alignment with other devices, thereby helping to shorten the time for model alignment.

[0006] In a first aspect, the present application provides a model training method, which is applied to a first device or a module in the first device (such as a chip or a chip system, etc.). Taking the application to the first device as an example, the method includes: the first device receives data feature information of a second local dataset from a second device, and the second local dataset is a dataset of the second device; further, based on the data feature information of the second local dataset and a first local dataset, a first training dataset is determined, where the first local dataset is a dataset of the first device, and the data feature information of the first training dataset includes the data feature information of the second local dataset and the data feature information of the first local dataset; furthermore, the first device obtains a first model based on the first training dataset, and the first model is a model deployed on the first device. Optionally, the data feature information of the second local dataset is indicated by a first indication information, that is, the first device receives the first indication information from the second device, and the first indication information is used to indicate the data feature information of the second local dataset.

[0007] In a possible implementation method, the first device receives a second local dataset from the second device, and the second local dataset is a dataset of the second device; further, the first device determines a first training dataset based on the second local dataset and a first local dataset, where the first local dataset is a dataset of the first device; furthermore, the first device obtains a first model based on the first training dataset, and the first model is a model deployed on the first device.

[0008] Based on the method described in the first aspect, before obtaining the first model deployed on the first device, the first device obtains the first training dataset by referring to the data feature information of the second local dataset of the second device, and further, obtains the first model based on the first training dataset. By such a method, in terms of feature representation, the first model can have similarity with the model (such as the second model) obtained based on the second local dataset, which is beneficial to establishing the connection between the model of the first device (i.e., the first model) and the model of the second device (i.e., the second model), and further beneficial to reducing the complexity of subsequent model alignment between the first model and the second model and shortening the time of model alignment.

[0009] In a possible implementation manner, the data feature information includes one or more of the following information: data identification information, data distribution information, or data classification information; wherein, the data identification information is used to indicate the data included in the dataset, the data distribution information is used to indicate the distribution of the data in the dataset, and the data classification information is used to indicate the classification of the data in the dataset.

[0010] In a possible implementation, the data distribution information includes one or more of clustering distribution information, probability distribution information, or model parameter information of a data generation model; the data classification information includes one or more of geographical region information, signal-related information, device configuration information, quality-related information, or classification center information.

[0011] In a possible implementation, the first device determines the data feature information of the second dataset difference set from the data feature information of the second local dataset based on the data feature information of the first local dataset. The data feature information of the second dataset difference set is the data feature information not included in the first local dataset. Further, the first device sends the data feature information of the second dataset difference set to the second device; and receives the second dataset difference set from the second device; the first training dataset is the union of the second dataset difference set and the first local dataset. Optionally, the data feature information of the second dataset difference set is indicated by the second indication information, that is, the first device sends the second indication information to the second device, and the second indication information is used to indicate the data feature information of the second dataset difference set.

[0012] In a possible implementation, the first device aligns the first model and the second model. The second model is a model deployed on the second device and is obtained based on the second training dataset. The data feature information of the first training dataset includes the data feature information of the second training dataset. By implementing this possible implementation, the data feature information of the first training dataset includes the data feature information of the second training dataset. Based on this, the first model and the second model obtained will have a certain similarity in feature representation. Further, in the process of model alignment, compared with aligning models without similarity, this application aligns models with a certain similarity, which is beneficial to improving the speed of model alignment.

[0013] Optionally, the second training dataset includes the data feature information of the second local dataset, or the second training dataset includes the data feature information of the second local dataset and the data feature information of the first local dataset. In the case where the second training dataset includes the data feature information of the second local dataset and the data feature information of the first local dataset, the data feature information of the second training dataset is the same as the data feature information of the first training dataset.

[0014] In a possible implementation, the first device determines the data feature information of the first dataset difference set from the first local dataset according to the data feature information of the second local dataset, where the data feature information of the first dataset difference set is the data feature information not included in the second local dataset; further, the first device sends the first dataset difference set to the second device. Optionally, the first dataset difference set is indicated by the third indication information, that is, the first device sends the third indication information to the second device, and the third indication information is used to indicate the first dataset difference set.

[0015] In a possible implementation, the first device determines a reference model based on the size of the first training dataset and the size of the second training dataset, where the reference model is the first model or the second model; further, the first device aligns the first model and the second model based on the reference model.

[0016] In a possible implementation, the data feature information included in the first training dataset further includes the data feature information of the third training dataset, and the third training dataset is used to train the third model, where the third model is a model deployed on the third device.

[0017] In a second aspect, the present application provides a data transmission method, which is applied to the second device or a module in the second device (such as a chip or a chip system, etc.). Taking the application to the second device as an example, the method includes: the second device sends the data feature information of the second local dataset to the first device, where the second local dataset is the dataset of the second device. Optionally, the data feature information of the second local dataset is indicated by the first indication information, that is, the second device sends the first indication information to the first device, and the first indication information is used to indicate the data feature information of the second local dataset.

[0018] In a possible implementation method, the second device sends the second local dataset to the first device.

[0019] Based on the method described in the second aspect, the second device can provide its own local dataset (i.e., the second local dataset) for reference by other devices (such as the first device), which is beneficial to establishing a connection between the model of the first device (i.e., the first model) and the model of the second device (i.e., the second model), and further beneficial to reducing the complexity of subsequent model alignment between the first model and the second model and shortening the time of model alignment. For the beneficial effects of other implementation manners described in the second aspect, reference may be made to the beneficial effects of the implementation manners described in the first aspect, and details will not be described hereinafter.

[0020] In a possible implementation, the data feature information includes one or more of the following information: data identification information, data distribution information, or data classification information; wherein, the data identification information is used to indicate the data included in the data set, the data distribution information is used to indicate the distribution of the data in the data set, and the data classification information is used to indicate the classification of the data in the data set.

[0021] In a possible implementation, the data distribution information includes one or more of clustering distribution information, probability distribution information, or model parameter information of a data generation model; the data classification information includes one or more of geographical region information, signal-related information, device configuration information, quality-related information, or classification center information.

[0022] In a possible implementation, the second device receives the data feature information of the second data set difference set from the first device, and the data feature information of the second data set difference set is the data feature information included in the second local data set and not included in the first local data set, and the first local data set is the data set of the first device; further, the second device sends the second data set difference set to the first device based on the data feature information of the second data set difference set. Optionally, the data feature information of the second data set difference set is indicated by a second indication information, that is, the second device receives the second indication information from the first device, and the second indication information is used to indicate the data feature information of the second data set difference set.

[0023] In a possible implementation, the second device aligns the first model and the second model, the second model is a model deployed on the second device, and the second model is obtained based on a second training data set, and the data feature information of the first training data set includes the data feature information of the second training data set.

[0024] Optionally, the second training data set includes the data feature information of the second local data set, or the second training data set includes the data feature information of the second local data set and the data feature information of the first local data set. In the case where the second training data set includes the data feature information of the second local data set and the data feature information of the first local data set, the data feature information of the second training data set is the same as the data feature information of the first training data set.

[0025] In a possible implementation, the second device receives the first data set difference set from the first device. The data feature information of the first data set difference set is included in the first local data set and not included in the data feature information of the second local data set. The second training data set is the union of the first data set difference set and the second local data set. Optionally, the first data set difference set is indicated by the third indication information. That is, the second device receives the third indication information from the first device, and the third indication information is used to indicate the first data set difference set.

[0026] In a possible implementation, the second device determines a reference model based on the size of the first training data set and the size of the second training data set. The reference model is the first model or the second model. Further, the second device aligns the first model and the second model based on the reference model.

[0027] In a possible implementation, the data feature information included in the first training data set further includes the data feature information of the third training data set, and the third training data set is used to train the third model, and the third model is a model deployed on the third device.

[0028] In a third aspect, the present application provides a communication device, which may be the first device, or a device in the first device, or a device that can be used in combination with the first device. Among them, the communication device may also be a chip system. The communication device can execute the method described in the first aspect. The functions of the communication device can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more units or modules corresponding to the above functions. The unit or module may be software and / or hardware. The operations and beneficial effects performed by the communication device can refer to the method and beneficial effects described in the first aspect above.

[0029] In a fourth aspect, the present application provides a communication device, which may be the second device, or a device in the second device, or a device that can be used in combination with the second device. Among them, the communication device may also be a chip system. The communication device can execute the method described in the second aspect. The functions of the communication device can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more units or modules corresponding to the above functions. The unit or module may be software and / or hardware. The operations and beneficial effects performed by the communication device can refer to the method and beneficial effects described in the second aspect above.

[0030] Fifth aspect, the present application provides a communication device, which includes a processor and an interface circuit. The interface circuit is configured to receive signals from other communication devices outside the communication device and transmit them to the processor, or send signals from the processor to other communication devices outside the communication device. The processor is configured to implement the method described in the first aspect through logic circuits or by executing code instructions, or the processor is configured to implement the method described in the second aspect through logic circuits or by executing code instructions.

[0031] Sixth aspect, the present application provides a communication device, which includes a processor connected to a memory, and is configured to call a program stored in the memory to execute the method described in the first aspect or the second aspect above. The memory may be located inside the first device or the second device, or may be located outside the first device or the second device. And the processor includes one or more.

[0032] Seventh aspect, the present application provides a computer-readable storage medium, in which a computer program or instruction is stored. When the computer program or instruction is executed by a communication device, the method described in the first aspect is implemented, or the method described in the second aspect is implemented.

[0033] Eighth aspect, the present application provides a computer program product including instructions. When a communication device reads and executes the instructions, the communication device is caused to execute the method described in the first aspect, or the communication device is caused to execute the method described in the second aspect.

[0034] Ninth aspect, the present application provides a communication system, which includes a communication device for executing the method described in the first aspect above and a communication device for executing the method described in the second aspect above. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1a is a schematic diagram of the architecture of a communication system provided by an embodiment of the present application;

[0036] Figure 1b is another schematic diagram of a wireless communication system applicable to an embodiment of the present application;

[0037] Figure 2 is a schematic diagram of a wireless communication system based on a neural network transmitter and receiver provided by an embodiment of the present application;

[0038] Figure 3 is a schematic diagram of several neural network model optimized transceivers in different modules provided by an embodiment of the present application;

[0039] Figure 4 is a schematic flowchart of a model training method provided by an embodiment of the present application;

[0040] Figure 5 Classification schematic diagram of a second local data set provided by an embodiment of the present application;

[0041] Figure 6 Schematic diagram of the relationship between a first local data set and a second local data set provided by an embodiment of the present application;

[0042] Figure 7 Flow schematic diagram of a model alignment method provided by an embodiment of the present application;

[0043] Figure 8 Schematic diagram of model alignment provided by an embodiment of the present application;

[0044] Figure 9 Another schematic diagram of model alignment provided by an embodiment of the present application;

[0045] Figure 10 Yet another schematic diagram of model alignment provided by an embodiment of the present application;

[0046] Figure 11 Schematic diagram of the structure of a communication device provided by an embodiment of the present application;

[0047] Figure 12 Another schematic diagram of the structure of a communication device provided by an embodiment of the present application. Detailed implementation manners

[0048] For the convenience of understanding the embodiments of the present application, the system architecture involved in the embodiments of the present application will be introduced first below.

[0049] Figure 1a It is a schematic diagram of the architecture of a communication system 1000 to which the embodiments of the present application are applied. As Figure 1a shown, the communication system includes a radio access network (RAN) 100 and a core network 200. Optionally, the communication system 1000 may further include the Internet 300. Among them, the RAN 100 includes at least one RAN node (such as Figure 1a 110a and 110b in Figure 1a , collectively referred to as 110), and may further include at least one terminal (such as Figure 1a(not shown). The terminal 120 is connected to the RAN node 110 wirelessly, and the RAN node 110 is connected to the core network 200 wirelessly or wired. The core network devices in the core network 200 and the RAN nodes 110 in the RAN 100 can be independent different physical devices, or the same physical device integrating the logical functions of the core network devices and the logical functions of the RAN nodes. Terminals can be connected to each other, and RAN nodes can be connected to each other, either wirelessly or wired.

[0050] The RAN 100 can be an evolved universal terrestrial radio access (E-UTRA) system, a new radio (NR) system, and a future radio access system defined in the 3rd generation partnership project (3GPP). The RAN 100 can also include two or more different radio access systems mentioned above. The RAN 100 can also be an open RAN (O-RAN).

[0051] The RAN node, also known as a radio access network device, a RAN entity, or an access node, can also be referred to as a network device hereinafter, and is used to help terminals access the communication system wirelessly. In one application scenario, the RAN node can be a base station, an evolved NodeB (eNodeB), a transmission reception point (TRP), a next generation NodeB (gNB) in the 5th generation (5G) mobile communication system, a next generation NodeB in the 6th generation (6G) mobile communication system, or a base station in a future mobile communication system. The RAN node can be a macro base station (such as Figure 1a 110a in Figure 1a ), or a micro base station or an indoor station (such as

[0052] In another application scenario, the wireless access of a terminal can be assisted through the cooperation of multiple RAN nodes, and different RAN nodes respectively implement some functions of a base station. For example, the RAN node can be a central unit (CU), a distributed unit (DU), or a radio unit (RU). Here, the CU completes the functions of the radio resource control protocol and the packet data convergence protocol (PDCP) of the base station, and can also complete the function of the service data adaptation protocol (SDAP); the DU completes the functions of the radio link control layer and the medium access control (MAC) layer of the base station, and can also complete some or all of the functions of the physical layer. For the specific descriptions of the above various protocol layers, reference can be made to the relevant technical specifications of 3GPP. The RU can be used to implement the functions of transmitting and receiving radio frequency signals. The CU and the DU can be two independent RAN nodes, or can be integrated in the same RAN node, for example, integrated in the baseband unit (BBU). The RU can be included in the radio frequency device, for example, included in the remote radio unit (RRU) or the active antenna unit (AAU). The CU can be further divided into two types of RAN nodes: CU-control plane and CU-user plane.

[0053] The RAN node can support one or more types of fronthaul interfaces. Different fronthaul interfaces respectively correspond to DUs and RUs with different functions. If the fronthaul interface between the DU and the RU is the common public radio interface (CPRI), the DU is configured to implement one or more of the baseband functions, and the RU is configured to implement one or more of the radio frequency functions. If the fronthaul interface between the DU and the RU is the enhanced common public radio interface (eCPRI), compared with the CPRI, some of the downlink and / or uplink baseband functions are moved from the DU to the RU for implementation. Different splitting methods between the DU and the RU correspond to different categories (Cat) of eCPRI, such as eCPRI Cat A, B, C, D, E, F.

[0054] Taking eCPRI Cat A as an example, for downlink transmission, with layer mapping as the segmentation, the DU is configured to implement layer mapping and one or more functions before it (i.e., one or more of encoding, rate matching, scrambling, modulation, layer mapping), while other functions after layer mapping (e.g., one or more of resource element (RE) mapping, digital beamforming (BF), or inverse fast Fourier transform (IFFT) / adding cyclic prefix (CP)) are moved to the RU for implementation. For uplink transmission, with de-RE mapping as the segmentation, the DU is configured to implement de-mapping and one or more functions before it (i.e., one or more of decoding, de-rate matching, descrambling, demodulation, inverse discrete Fourier transform (IDFT), channel equalization, de-RE mapping), while other functions after de-mapping (e.g., one or more of digital BF or fast Fourier transform (FFT) / removing CP) are moved to the RU for implementation. It can be understood that for the function descriptions of the DU and RU corresponding to various types of eCPRI, reference can be made to the eCPRI protocol, which will not be elaborated here.

[0055] In a possible design, the processing unit in the BBU for implementing baseband functions is called the baseband high (BBH) unit, and the processing unit in the RRU / AAU / RRH for implementing baseband functions is called the baseband low (BBL) unit.

[0056] In different systems, RAN nodes may have different names. For example, in the O-RAN system, the CU can be called the open CU (O-CU), the DU can be called the open DU (O-DU), and the RU can be called the open RU (O-RU). The RAN nodes in the embodiments of this application can be implemented in the form of software modules, hardware modules, or a combination of software modules and hardware modules. For example, the RAN node can be a server loaded with the corresponding software module. The embodiments of this application do not limit the specific technologies and specific device forms adopted by the RAN nodes. For the convenience of description, the base station is taken as an example of the RAN node in the following description.

[0057] A terminal is a device with wireless transceiver capabilities that can send signals to a base station or receive signals from a base station. A terminal can also be referred to as a terminal device, user equipment (UE), mobile station, mobile terminal, etc. Terminals can be widely applied in various scenarios, such as device-to-device (D2D), vehicle to everything (V2X) communication, machine-type communication (MTC), Internet of Things (IoT), virtual reality, augmented reality, industrial control, autonomous driving, telemedicine, smart grid, smart home, smart office, smart wearables, smart transportation, smart city, etc. A terminal can be a mobile phone, tablet computer, computer with wireless transceiver capabilities, wearable device, vehicle, aircraft, ship, robot, robotic arm, smart home device, etc. Embodiments of the present application do not limit the specific technologies and specific device forms adopted by the terminal.

[0058] The base station and the terminal can be fixed in position or movable. The base station and the terminal can be deployed on land, including indoor or outdoor, handheld or vehicle-mounted; they can also be deployed on water; they can also be deployed on aircraft, balloons, and artificial satellites. Embodiments of the present application do not limit the application scenarios of the base station and the terminal.

[0059] The roles of the base station and the terminal can be relative. For example, Figure 1a the helicopter or drone 120i in [figure] can be configured as a mobile base station. For the terminals 120j that access the radio access network 100 through 120i, the terminal 120i is a base station; but for the base station 110a, 120i is a terminal, that is, the communication between 110a and 120i is through the radio air interface protocol. Of course, the communication between 110a and 120i can also be through the interface protocol between base stations. In this case, relative to 110a, 120i is also a base station. Therefore, both the base station and the terminal can be uniformly referred to as communication devices. Figure 1a The 110a and 110b in [figure] can be referred to as communication devices with base station functions. Figure 1a The 120a - 120j in [figure] can be referred to as communication devices with terminal functions.

[0060] Communication can be carried out between a base station and a terminal, between base stations, and between terminals through licensed spectrum, unlicensed spectrum, or both simultaneously; communication can be carried out through spectrum below 6 gigahertz (GHz), through spectrum above 6 GHz, or by using both spectrum below 6 GHz and spectrum above 6 GHz simultaneously. Embodiments of this application do not limit the spectrum resources used for wireless communication.

[0061] In embodiments of this application, the functions of a base station can also be performed by a module (such as a chip) in the base station or by a control subsystem with base station functions. The control subsystem with base station functions here can be a control center in the above application scenarios such as smart grid, industrial control, intelligent transportation, and smart city. The functions of a terminal can also be performed by a module (such as a chip or a modem) in the terminal or by a device with terminal functions.

[0062] Please refer to Figure 1b , Figure 1b which is another schematic diagram of a wireless communication system applicable to embodiments of this application.

[0063] As Figure 1b shown, a radio access network (RAN) intelligent controller (RIC) is included in the wireless communication system. As an example, the RIC can be used to implement functions related to artificial intelligence (AI). As an example, the RIC includes a near-real time RIC (near-RT RIC) and a non-real time RIC (Non-RT RIC). Among them, the non-real time RIC mainly processes non-real time information, such as data that is not sensitive to latency, and the latency of this data can be in seconds. The real time RIC mainly processes near-real time information, such as data that is relatively sensitive to latency, and the latency of this data is in tens of milliseconds.

[0064] The near-RT RIC is used for model training and inference. For example, it is used to train an AI model and perform inference using this AI model. The near-RT RIC can obtain information on the network side and / or the terminal side from RAN nodes (such as CU, CU-CP, CU-UP, DU, and / or RU) and / or terminals. This information can be used as training data or inference data. Optionally, the near-RT RIC can submit the inference result to RAN nodes and / or terminals. Optionally, the inference result can be exchanged between the CU and the DU, and / or between the DU and the RU. For example, the near-RT RIC submits the inference result to the DU, and the DU sends it to the RU.

[0065] The non-real-time RIC is also used for model training and inference. For example, it is used to train an AI model and perform inference using the model. The non-real-time RIC can obtain network-side and / or terminal-side information from RAN nodes (such as CU, CU-CP, CU-UP, DU, and / or RU) and / or terminals. This information can be used as training data or inference data, and the inference result can be delivered to RAN nodes and / or terminals. Optionally, the inference result can be interacted between the CU and the DU, and / or between the DU and the RU. For example, the non-real-time RIC delivers the inference result to the DU, and the DU sends it to the RU.

[0066] The near-real-time RIC and the non-real-time RIC can also be separately set as a network element. Optionally, the near-real-time RIC and the non-real-time RIC can also be part of other devices. For example, the near-real-time RIC is set in a RAN node (such as in a CU or a DU), and the non-real-time RIC is set in a network management (operation, administration and maintenance, OAM), a cloud server, a core network device, or other network devices.

[0067] In practical applications, the wireless communication system can include multiple network devices (also referred to as access network devices) at the same time, and can also include multiple terminal devices without limitation. One network device can serve one or more terminal devices at the same time. One terminal device can also access one or more network devices at the same time. The embodiments of the present application do not limit the number of terminal devices and network devices included in the wireless communication system.

[0068] To facilitate the understanding of the relevant content of the embodiments of the present application, some terms involved in the embodiments of the present application are further explained below. This part is only for easy understanding and cannot be regarded as a disclosure or specific limitation of the technical solution of the present application.

[0069] 1. Neural network

[0070] A neural network can be composed of neural units. A neural unit can refer to an operation unit with x s as the input, and the output of the operation unit can be shown in formula (1).

[0071]

[0072] where s = 1, 2,..., n, and n is a natural number greater than 1, and W s is x swhere \(w\) is the weight of the neuron, \(b\) is the bias of the neuron, and \(f\) is the activation function of the neuron, which is used to introduce non - linear characteristics into the neural network to convert the input signal in the neuron into an output signal. The output signal of this activation function can be used as the input of the next convolutional layer. The activation function can be the sigmoid function. A neural network is a network formed by connecting many such single neurons together, that is, the output of one neuron can be the input of another neuron. The input of each neuron can be connected to the local receptive field of the previous layer to extract the features of the local receptive field, and the local receptive field can be a region composed of several neurons.

[0073] It should be noted that the neural network model mentioned in this application can be one or more of the network models of neural networks, the network models of deep neural networks (DNN), the network models of convolutional neural networks (CNN), the network models of recurrent neural networks (RNN), the network models of generative adversarial networks, or a deformation (or combination) of their combinations. This application does not make specific limitations on this.

[0074] 2. Intelligent air interface technology

[0075] Intelligent air interface technology can be understood as a technology that applies a neural network model in a wireless communication system to optimize the performance of the wireless communication system. Please refer to Figure 2 as shown. Figure 2 This is a wireless communication system provided by this application based on a neural network transmitter and receiver (collectively referred to as transceiver) hereinafter. In Figure 2 the system shown, the transceiver applying the neural network model can optimize the performance of the transceiver signals through data - driven training.

[0076] It can be understood that optimizing the performance of the transceiver by applying the neural network model is to optimize the channel coding, modulation, waveform, pilot and other modules used for signal processing in the transceiver through the neural network model. Exemplarily, please refer to Figure 3 as shown. Figure 3 This is a schematic diagram of several neural network models provided by the embodiments of this application for optimizing different modules in the transceiver. Among them, Figure 3 Figure 3a is a schematic diagram of optimizing the modulation module and waveform through a neural network model; Figure 3 Figure 3b is a schematic diagram of optimizing the coding module and modulation module through a neural network model; Figure 3 Figure 3c is a schematic diagram of optimizing the reference signal through a neural network model.

[0077] In Figure 3 3a, after the bits to be transmitted are encoded by the forward error correction (FEC) of the transmitting device to obtain encoded bits, the encoded bits are modulated by the modulation neural network model (i.e., the NN-Mod module in 3a of Figure 3 ), mapped to physical resources, transformed into a time-domain signal through an inverse fast Fourier transform (IFFT), and then sent after being processed by the time-signal neural network model (i.e., the T-NN module in 3a of Figure 3 ); the received signal is first processed by the time-signal neural network model of the receiving device, transformed into a frequency-domain signal through a fast Fourier transform (FFT), and then demodulated by the demodulation neural network model (i.e., the NN-DeMod module in 3a of Figure 3 ), and then sent to the FEC decoder.

[0078] In Figure 3 3b, after the bits to be transmitted are processed by the encoding and modulation neural network model of the transmitting device (i.e., the NN-CoMo module in 3b of Figure 3 ), modulation symbols are obtained, and then sent after being processed by modules such as IFFT; the receiving device directly sends the signal processed by modules such as FFT to the demodulation and decoding neural network model (i.e., the NN-DeCoMo module in 3b of Figure 3 ) for processing to obtain the estimated transmitted bits.

[0079] In Figure 3 3c, the transmitting device generates a reference signal based on the reference signal neural network model (i.e., the NN-RS module in 3c of Figure 3 ), and the receiving device processes the received signal based on the reference signal and the corresponding channel estimation neural network model (i.e., the NN-CE module in 3c of Figure 3 ) to obtain the estimated channel.

[0080] In the application process of intelligent air interface technology, there is a situation where the models deployed on the sending device (abbreviated as the sending model) and the models deployed on the receiving device (abbreviated as the receiving model) correspond to different training processes. For example, the sending device and the receiving device belong to different manufacturers. In this case, if the sending model and the receiving model are not aligned before the sending device and the receiving device communicate, it may occur that the receiving model cannot parse the data processed by the sending model, resulting in the inability to communicate between the sending device and the receiving device. Generally, the process of aligning the sending model and the receiving model is a process of aligning (or understood as joint training, adaptation training, etc.) the sending model and the receiving model through the data in a training dataset. The time and cost of the model alignment process are relatively long and large.

[0081] This application provides a model training method, which can improve the correlation between the training dataset corresponding to the sending model and the training dataset corresponding to the receiving model, thereby facilitating the improvement of the correlation between the sending model and the receiving model and facilitating subsequent model alignment. The model training method and communication device provided by this application are further introduced below with reference to the accompanying drawings:

[0082] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of a model training method provided by an embodiment of this application. As Figure 4 shown, the model training method includes the following S401 to S403. Figure 4 The execution subject of the method shown can be the first device and the second device, or Figure 4 the execution subject of the method shown can be the module in the first device and the module in the second device, or Figure 4 the execution subject of the method shown can be the chip of the first device and the chip of the second device. Figure 4 Taking the first device and the second device as the execution subjects of the method as an example for illustration. It should be noted that the first device (or the second device) mentioned in this application can be Figure 1a the network device shown, or Figure 1a the terminal device shown. This application does not make specific limitations. It should also be noted that the first device in this application can be the sending device or the receiving device in the communication process, and the second device in this application can also be the sending device or the receiving device; when the first device is the sending device, the second device is the receiving device; when the first device is the receiving device, the second device is the sending device. Among them:

[0083] S401. The second device sends the data feature information of the second local dataset to the first device, and the second local dataset is the dataset of the second device.

[0084] It can be understood that in the present application, one or more of the data sets collected by the first device, the data sets generated by the data generation model, or the stored data sets are denoted as the first local data set, and this first local data set is the data set of the first device; one or more of the data sets collected by the second device, the data sets generated by the data generation model, or the stored data sets are denoted as the second local data set, and this second local data set is the data set of the second device. After the second device obtains the data feature information of the second local data set, the second device indicates the data feature information of the second local data set to the first device. In the following text, the present application takes the example of the second device indicating the data feature information of the second local data set to the first device through the first indication information for illustration.

[0085] Among them, the data feature information is used to indicate the features of the data included in the second local data set. The present application does not specifically limit the specific manifestation form of this data feature information. For the convenience of understanding, the present application provides several types of data feature information for exemplary illustration, which should not be regarded as a specific limitation of the present application. In a possible implementation method, the data feature information includes one or more of data identification information, data distribution information, or data classification information. Among them:

[0086] ①. The data identification information is used to indicate the data included in the data set. Among them, the data identification information can be the identity document (ID) of the data in the data set. In this case, the first indication information includes the IDs of all the data in the second local data set.

[0087] ②. The data distribution information (or referred to as data statistical information) is used to indicate the distribution of the data in the data set. Among them, the data distribution information can be clustering distribution information (such as the number of clustering clusters corresponding to the second local data set, the clustering centers of each clustering cluster, the cluster size of each clustering cluster, or the cluster density of each clustering cluster, etc.), probability distribution information (used to indicate the parameters of the probability distribution satisfied by the data in the data set. For example, if the data in the second local data set satisfies a Gaussian distribution, then the probability distribution information is the mean and variance of the Gaussian distribution), or model parameter information of the data generation model (used to indicate the model parameters for generating the data in the case where the data in the data set is generated by the data generation model).

[0088] ③. The data classification information is used to indicate the classification of data in the data set. Among them, the data classification information includes information about the categories included in the data set, the size of each category's corresponding subset (i.e., the number of data included in each subset), or one or more of the data included in each category's corresponding subset. It should be noted that this application does not specifically limit the category information included in the data set. This category information can be geographical area information (such as cell ID, site ID, longitude and latitude information where the data collection device is located, etc.), signal-related information (such as Doppler spread information, delay spread information, or antenna correlation coefficient corresponding to the signal when collecting data, etc.), device configuration information (such as the moving speed of the device for collecting data or antenna configuration information, etc.), quality-related information (such as data processing algorithm-related information when obtaining data, channel estimation accuracy information when collecting data, etc.), classification center information (such as the clustering distribution information corresponding to the data set, etc.), or one or more of them.

[0089] For example, taking the Doppler spread information and delay spread information as an example, the multiple signals included in the second local data set are classified to obtain 4 subsets as shown in Figure 5 Subset 1 to Subset 4. Among them, Subset 1 corresponds to Doppler spread range 1 and delay spread range 1, Subset 2 corresponds to Doppler spread range 1 and delay spread range 2, Subset 3 corresponds to Doppler spread range 2 and delay spread range 1, and Subset 4 corresponds to Doppler spread range 2 and delay spread range 2. Further, the second device sends first indication information to the first device. The first indication information indicates the Doppler spread range and delay spread range corresponding to each subset among Subset 1 to Subset 4 included in the second local data set, or the first indication information indicates the number or proportion of the data included in each subset of the second local data set.

[0090] It should be noted that this application does not specifically limit the acquisition method for the device to obtain the data feature information of its local data set, that is, this application does not limit the method for the second device to obtain the data feature information of the second local data set, and similarly does not limit the method for the first device to obtain the data feature information of the first local data set. For example, the second device analyzes the second local data set through a data feature information analysis model (such as a clustering algorithm, etc.); or the second device does not need to analyze the second local data set through a data feature information analysis model and can obtain the data feature information of the second local data set from the attribute information of the second local data set (which can be understood as including information for describing the local data set), etc.

[0091] It should also be noted that the devices in this application have a consensus on the data feature information (i.e., the understanding is the same). For example, when the first indication information indicates that the data feature information of the second local data set is data identification information, the data feature information of the first local data set of the first device is also data identification information; when the first indication information indicates that the data feature information of the second local data set is data distribution information, and the data distribution information is clustering distribution information, the data feature information of the first local data set of the first device is also clustering distribution information; when the first indication information indicates that the data feature information of the second local data set is data classification information, and the second local data set is classified according to Doppler spread range 1 and delay spread range 1, the data feature information of the first local data set of the first device will also be classified according to Doppler spread range 1, other Doppler spread ranges (ranges other than Doppler spread range 1), delay spread range 1, and other delay spread ranges (i.e., ranges other than delay spread range 1).

[0092] In a possible implementation, before S401 is executed, the first device and the second device can perform data set alignment negotiation (or be understood as starting the data set alignment process, that is, triggering the execution of S401 to S402).

[0093] For example, in a possible implementation, the first device sends a data set alignment request message to the second device; the second device sends a data set alignment confirmation message to the first device according to the data set alignment request message; further, the second device sends the first indication information to the first device. Among them, the data set alignment confirmation message can be the same message as the message carrying the first indication information in S401, or can be a different message from the message carrying the first indication information in S401.

[0094] For another example, in another possible implementation, the second device sends a data set alignment request message to the first device; the first device sends a data set alignment confirmation message to the second device according to the data set alignment request message; further, the second device sends the first indication information to the first device.

[0095] S402. The first device determines a first training data set based on the data feature information of the second local data set and the first local data set. The first local data set is the data set of the first device, and the data feature information of the first training data set includes the data feature information of the second local data set and the data feature information of the first local data set.

[0096] That is to say, the first device combines the data feature information of the first local data set and the data feature information of the second local data set to obtain a first training data set. The first training data set can be the union of the first local data set and the second local data set, or the data feature information of the first training data set includes the data feature information of the first local data set and the data feature information of the second local data set.

[0097] In a possible implementation manner of S402, the first device generates a data set (denoted as the first generated data set) that is the same as the data feature information of the second local data set according to the data feature information of the second local data set. Further, the first device determines the first training data set according to the first generated data set and the first local data set, and the first training data set is the union of the first generated data set and the first local data.

[0098] For example, the data feature information of the second local data set is the model parameter information of the data generation model 1. In this case, the first device constructs the data generation model 1 according to the model parameter information, and generates the first generated data set through the data generation model 1. Further, the first device determines the union of the first generated data set and the first local data set as the first training data set. That is, it can be considered that the data feature information of the first training data set includes the data feature information of the first local data set and the data feature information of the second local data set.

[0099] In another possible implementation manner of S402, the first device determines the data feature information of the second data set difference set from the second local data set according to the data feature information of the first local data set, and the data feature information of the second data set difference set is the data feature information not included in the first local data set; further, the first device sends the data feature information of the second data set difference set to the second device, and receives the second data set difference set from the second device, and determines the first training data set according to the second data set difference set and the first local data set. The first training data set is the union of the first local data set and the second data set difference set. In the following, this application takes the first device indicating the data feature information of the second data set difference set to the second device through the second indication information as an example for description.

[0100] For ease of understanding, please refer to Figure 6As shown, there is an intersection between the data feature information of the first local data set and the data feature information of the second local data set. This intersection can be an empty set or a non-empty set, and this application does not make specific limitations on this. In this application, the difference set between the data feature information of the first local data set and this intersection is called the data feature information of the first data set difference set; the difference set between the data feature information of the second local data set and this intersection is called the data feature information of the second data set difference set. After the first device receives the data feature information of the second local data set, it compares the data feature information of the first local data set and the data feature information of the second local data set, and determines the data feature information different from the data feature information of the first local data set (that is, the data feature information of the second data set difference set) from the data feature information of the second local data set. Further, the first device sends second indication information to the second device, and this second indication information is used to indicate the data feature information of the second data set difference set. The second device sends the second data set difference set to the first device according to the data feature information indicated by the second indication information. Further, the first device determines the union of the second data set difference set and the first local data set as the first training data set.

[0101] S403. The first device obtains a first model based on the first training data set, and this first model is a model deployed on the first device.

[0102] That is to say, the first device trains a neural network model based on the first training data set to obtain the first model. This first model can be a model deployed in the transmitter of the first device or a model deployed in the receiver of the first device.

[0103] In summary, the first device obtains the first training data set by referring to the data feature information of the second local data set of the second device, and obtains the first model based on this first training data set. It can be understood that the first model and the model obtained based on the second local data set (such as the second model) have similarities in feature representation, which is conducive to establishing a connection between the model of the first device (that is, the first model) and the model of the second device (that is, the second model), and further conducive to reducing the complexity of subsequent model alignment between the first model and the second model and shortening the time of model alignment.

[0104] Optionally, this application also provides a Figure 4 specific application scenario to which it applies. In this application scenario, the first device can be an access network device, and the second device is a terminal device served by the first device. Or, the first device is a terminal device, and the second device is an access network device that provides services for the first device.

[0105] To facilitate understanding of the model alignment process mentioned in this application, please refer to Figure 7 ,Figure 7 is a schematic flowchart of a model alignment method provided by an embodiment of the present application. As Figure 7 shown, the model alignment method includes the following S701 - S703. Figure 7 The execution subject of the method shown can be the first device and the second device, or Figure 7 the execution subject of the method shown can be a module in the first device and a module in the second device, or Figure 7 the execution subject of the method shown can be a chip of the first device and a chip of the second device. Figure 7 Taking the first device and the second device as the execution subjects of the method as an example for illustration. It should be noted that the descriptions of the first device and the second device can be referred to in the foregoing Figure 4 and will not be elaborated here. Among them:

[0106] S701. The first device obtains a first model based on a first training dataset.

[0107] Among them, the description of S701 can be referred to the specific implementation manners of the foregoing S401 - S403 and will not be described here.

[0108] S702. The second device obtains a second model based on a second training dataset.

[0109] That is to say, the second device trains a neural network model based on the second training dataset to obtain a second model. The second model is a model deployed on the second device. For example, the second model is a model deployed in a transmitter of the second device or can also be a model deployed in a receiver of the second device.

[0110] Among them, the data feature information of the first training dataset includes the data feature information of the second training dataset. It should be noted that the descriptions of the data feature information and the first training dataset can be referred to the relevant descriptions of the data feature information and the first training dataset in the foregoing S401 - S403 and will not be elaborated here.

[0111] For ease of understanding, the relationship between the first training dataset and the second training dataset will be described in the following two cases.

[0112] Case 1. The second training dataset is a second local dataset.

[0113] That is to say, the second training dataset includes the data feature information of the second local dataset. In this case, the data feature information of the first training dataset includes, in addition to the data feature information of the second training dataset, other data feature information. For example, as Figure 6As shown, the data feature information of the first training data set further includes the data feature information of the first data set difference set.

[0114] In a possible implementation manner, the data feature information of the first training data set includes, in addition to the data feature information of the second training data set, the data feature information of a third training data set. The third training data set is used to train a third model, and the third model is a model deployed on a third device.

[0115] Exemplarily, the central node maintains a first local data set, the distributed node 1 maintains a second local data set, and the distributed node 2 maintains a third local data set. The central node sends the data feature information of the first local data set to the distributed node 1 and the distributed node 2. Based on the data feature information of the first local data set and the data feature information of the second local data set, the distributed node 1 sends an indication message 1 to the central node. The indication message 1 is used to indicate a second data set difference set, and the second data set difference set is composed of data corresponding to the data feature information that is included in the second local data set but not included in the first local data set. Based on the data feature information of the first local data set and the data feature information of the third local data set, the distributed node 2 sends an indication message 2 to the central node. The indication message 2 is used to indicate a third data set difference set, and the third data set difference set is composed of data corresponding to the data feature information that is included in the third local data set but not included in the first local data set. Further, the central node determines the union of the first local data set, the second data set difference set, and the third data set difference set as the first training data set for obtaining a first model deployed on the central node. The distributed node 1 determines the second local data set as the second training data set for obtaining a second model deployed on the distributed node 1. The distributed node 2 determines the third local data set as the third training data set for obtaining a third model deployed on the distributed node 2. That is, it can be understood that the data feature information of the first training data set includes the feature information of the second training data set and the feature information of the third training data set. Optionally, in an application scenario of this example, the central node may be an access network device or a core network device, and the multiple distributed nodes may be different terminal devices respectively.

[0116] Case 2: The second training data set is the union of the second local data set and the first data set difference set.

[0117] In this case, the data feature information of the first training data set is the same as that of the second training data set. That is to say, the second training data set includes the data feature information of the second local data set and the data feature information of the first local data set. In a possible implementation manner, after the first device obtains the data feature information of the second local data set, according to the data feature information of the second local data set, the first device determines the data feature information of the first data set difference set from the first local data set. The data feature information of the first data set difference set is the data feature information that is included in the first local data set but not included in the second local data set. Further, the first device sends the first data set difference set to the second device. The second device determines the union of the first data set difference set and the second local data set as the second training data set. Optionally, the first device indicates the first data set difference set to the second device through the third indication information.

[0118] For ease of understanding, refer to, for example Figure 6 As shown, the data feature information of the first training data set is the union of the data feature information of the first local data set and the data feature information of the second data set difference set, and the data feature information of the second training data set is the union of the data feature information of the second local data set and the data feature information of the first data set difference set. That is to say, the data feature information of the first training data set is the same as that of the second training data set.

[0119] S703. The first device and the second device align the first model and the second model.

[0120] That is to say, the first device and the second device align the first model and the second model (which can be understood as training data for model alignment) according to the alignment data. It can be understood that the goal of aligning the first model and the second model is to enable the first device and the second device to communicate normally.

[0121] For example, in an example 1, the first model is a model deployed on the transmitter of the first device, and the second model is a model deployed on the receiver of the second device. After aligning the first model and the second model, the first model processes data #1 to obtain data #2; the first device sends data #2 to the second device; the second model processes data #2 to obtain data #3; where data #3 is relatively close to data #1, or it can be understood that the error between data #3 and data #1 is small (for example, the error is less than the first threshold, and the specific value of the first threshold in this application is not limited).

[0122] For example, in an example 2, the first model is a model deployed in the receiver of the first device, and the second model is a model deployed in the transmitter of the second device. After aligning the first model and the second model, the second model processes data #1 to obtain data #2; the second device sends data #2 to the first device; the first model processes data #2 to obtain data #3; wherein, the data #3 is relatively close to the data #1, or it can be understood that the error between the data #3 and the data #1 is small.

[0123] In a possible implementation manner, aligning the first model and the second model includes but is not limited to the following three methods:

[0124] Method 1: Jointly train the first model and the second model.

[0125] For the process of this joint training, please refer to Figure 8 as shown. Among them, when the first model is a model of a transmitter deployed in the first device (i.e., the sending-end device), and the second model is a model of a receiver deployed in the second device (i.e., the receiving-end device), the process of this joint training is as shown in Figure 8 8a in. When the first model is a model of a receiver deployed in the first device (i.e., the receiving-end device), and the second model is a model of a transmitter deployed in the second device (i.e., the sending-end device), the process of this joint training is as shown in Figure 8 8b in.

[0126] Next, taking Figure 8 8a in as an example, the process of this joint training will be described exemplarily. The first device performs forward inference on the alignment data through the first model to obtain a forward inference result, and sends the forward inference result to the second device; after the second device receives the forward inference result from the first device, it calculates the forward inference result through the second model (for example, including at least one of forward inference, loss function calculation, or reverse gradient calculation) to obtain a feedback result, and feeds back the feedback result to the first device. Among them, the feedback result can be the loss result of the second model's inference calculation, or the gradient obtained by the second model's reverse gradient calculation. This application does not make specific limitations on this. The first device updates the model parameters of the first model based on the feedback result and the forward inference result. Optionally, the second device can also update the model parameters of the second model based on the forward inference result.

[0127] Method 2: Perform adaptation training on the first model and the second model according to the third model in the adaptation node. Wherein, the adaptation node is the first device or the second device.

[0128] It should be understood that when the adaptation node is the first device, if the first model is the transmitter model in the first device and the second model is the receiver model in the second device, then the third model is the receiver model in the first device; if the first model is the receiver model in the first device and the second model is the transmitter model in the second device, then the third model is the transmitter model in the first device. When the adaptation node is the second device, if the first model is the transmitter model in the first device and the second model is the receiver model in the second device, then the third model is the transmitter model in the second device; if the first model is the receiver model in the first device and the second model is the transmitter model in the second device, then the third model is the receiver model in the second device.

[0129] It should also be noted that the present application does not specifically limit the method for determining the adaptation node. For example, the first device and the second device determine the adaptation node based on the size of the first training dataset and the size of the second training dataset. The adaptation node can be the device with the larger training dataset or the device with the smaller training dataset, and the present application does not make specific limitations. For the sake of understanding, hereinafter, an example in which the adaptation node is the device with the larger training dataset (i.e., the first device) will be used for illustration.

[0130] In a possible implementation manner of Method 2, the adaptation node constructs a fourth training dataset according to the third model, and the adaptation node sends the fourth training dataset to other devices (i.e., the devices other than the adaptation node in the first device and the second device). It can be understood that the first device and the second device align the first model and the second model based on the fourth training dataset.

[0131] For example, the first device is the adaptation node, and the first device trains the transmitter model and the receiver model in the first device based on the first training dataset; for the process of this adaptation training, please refer to Figure 9 as shown. Among them, when the first model is the transmitter model deployed in the first device (i.e., the sending-end device), the third model is the receiver model deployed in the first device, and the second model is the receiver model deployed in the second device (i.e., the receiving-end device), the process of this adaptation training is as shown in Figure 9 9a in. When the first model is the receiver model deployed in the first device (i.e., the receiving-end device), the third model is the transmitter model deployed in the first device, and the second model is the transmitter model deployed in the second device (i.e., the sending-end device), the process of this adaptation training is as shown in Figure 9 9b in. Taking Figure 9Taking 9a in [the relevant context] as an example, the process of this adaptation training will be described exemplarily. The first device constructs a fourth training data set based on the input of the third model (i.e., the output of the first model) and the output of the third model, and sends the fourth training data set to the second device; the second device trains the second model based on the fourth training data set, so that when the input of the second model is the same as the input of the third model, the output of the second model is as close as possible to the output of the third model (or understood as having a smaller error).

[0132] In another possible implementation manner of the second method, the adaptation node constructs a training reference model according to the third model, and the adaptation node sends the model parameters of the training reference model to other devices (i.e., the devices among the first device and the second device that are not the adaptation node). It can be understood that the first device and the second device align the first model and the second model based on the training reference model.

[0133] For example, the first device is the adaptation node; the first device is the sending device, and the first device trains the model of the transmitter in the first device (i.e., the first model mentioned in this application) and the model of the receiver (i.e., the third model mentioned in this application) based on the first training data set; the second device is the receiving device, and the second device trains the model of the receiver in the second device (i.e., the second model mentioned in this application) based on the second training data set. In this case, the first device constructs a receiver training reference model according to the third model. When the input of the receiver training reference model is the same as the input of the third model, the output of the receiver training reference model is the same as or similar to the output of the third model. Further, the first device sends the model parameters of the receiver training reference model to the second device; the second device constructs a fifth training data set based on the model parameters of the receiver training reference model, and trains the second model based on the fifth training data set, so that when the input of the second model is the same as the input of the receiver training reference model, the output of the second model is as close as possible to the output of the receiver training reference model.

[0134] Method 3: Align the first model and the second model based on a reference model. Among them, the reference model is the first model or the second model, and the reference model is determined by the first device and the second device based on the size of the first training data set and the size of the second training data set.

[0135] It should be noted that the reference model can be the model obtained based on the larger training data set (i.e., the training data set corresponding to the larger value between the first data volume and the second data volume) among the first model and the second model; or, the reference model can also be the model obtained based on the smaller training data set (i.e., the training data set corresponding to the smaller value between the first data volume and the second data volume) among the first model and the second model; this application does not make specific limitations on this.

[0136] That is, before aligning the first model and the second model, the first device and the second device interact with the size of the first training data set (denoted as the first data volume) and / or the size of the second training data set (denoted as the second data volume); furthermore, the first device and / or the second device can determine a reference model according to the first data volume and the second data volume.

[0137] For example, in an example 3, the first device sends the first data volume to the second device. After the second device obtains the first data volume, it determines that the value of the first data volume is greater than the value of the second data volume. Further, the second device determines the first model as the reference model. The second device sends the fourth indication information to the first device, and the fourth indication information is used to indicate that the reference model is the first model. Further, the second device and the first device align the second model with the first model based on the first model. In an example 4, the second device sends the second data volume to the first device. After the first device obtains the second data volume, it determines that the value of the first data volume is greater than the value of the second data volume. Further, the first device determines the first model as the reference model. The first device sends the fifth indication information to the second device, and the fifth indication information is used to indicate that the reference model is the first model. Further, the second device and the first device align the second model with the first model based on the first model. In an example 5, the first device sends the first data volume to the second device, and the second device sends the second data volume to the first device. The first device and the second device determine that the value of the first data volume is greater than the value of the second data volume. Further, the first device and the second device determine the first model as the reference model. Further, the second device and the first device align the second model with the first model based on the first model.

[0138] In a possible implementation manner of Mode 3, when the reference model is the first model, the second device trains the second model and the second adaptation layer based on the first model. For example, when the first device is a sending device and the second device is a receiving device, and the first model is a model deployed on the transmitter of the first device and the second model is a model deployed on the receiver of the second device, the process of the adaptation training is as Figure 10 shown in Figure 10a. The first device processes Data #1 through the first model to obtain Data #2; the first device sends Data #2 to the second device; the second device processes Data #2 through the second adaptation layer to obtain Data #3, and processes Data #3 through the second model to obtain Data #4. The training objective of the adaptation training can be understood as making the error between Data #4 and Data #1 smaller. Or it can be understood as making the output after the joint processing of the second adaptation layer and the second model as close as possible to the output of the model of the receiver of the first device under the same input.

[0139] In another possible implementation manner of Mode 3, when the reference model is the second model, the first device trains the first model and the first adaptation layer based on the second model. For example, when the first device is a sending device and the second device is a receiving device, the first model is a model deployed on the transmitter of the first device, and the second model is a model deployed on the receiver of the second device, the process of the adaptation training is as follows Figure 10 shown in 10b. The first device processes Data #1 through the first model to obtain Data #2, and processes Data #2 through the first adaptation layer to obtain Data #3; the first device sends Data #3 to the second device; the second device processes Data #3 through the second model to obtain Data #4. The training objective of this adaptation training can be understood as minimizing the error between Data #4 and Data #1. Or it can be understood as making the output after the combined processing of the first model and the first adaptation layer as close as possible to the output of the model of the transmitter of the second device under the same input.

[0140] It can be seen that by implementing the model alignment method described in this application Figure 7 the data feature information of the first training data set includes the data feature information of the second training data set, and the first model and the second model obtained based on this will have a certain similarity in terms of feature representation. Further, in the process of model alignment, compared with aligning models without similarity, this application aligns models with a certain similarity, which is beneficial to improving the speed of model alignment.

[0141] It can be understood that in order to implement the above functions, the above device includes the corresponding hardware structure and / or software module for executing each function. Those skilled in the art should easily realize that, combining the units and algorithm steps of each example described in the embodiments disclosed in this article, this application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0142] The embodiments of this application can divide the first device or the second device into functional modules according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. It should be noted that the division of modules in the embodiments of this application is illustrative, only a logical function division, and there may be other division methods in actual implementation.

[0143] Please refer to Figure 11 , Figure 11 which shows a schematic structural diagram of a communication device 1100 according to an embodiment of the present application. Figure 11 The shown communication device may be the first device, a device in the first device, or a device that can be used in matching with the first device. Figure 11 The shown communication device may include a communication unit 1101 and a processing unit 1102; Figure 11 The shown communication device may be the second device, a device in the second device, or a device that can be used in matching with the second device. Figure 11 The shown communication device may include a communication unit 1101 and a processing unit 1102. Specifically, the processing unit 1102 is used to process data, which may be the data received by the communication unit 1101, and the processed data may also be sent by the communication unit 1101; the communication unit 1101 may be understood as a transceiver unit, including a receiving module and / or a sending module, and the receiving module is used to execute Figure 4 or Figure 7 any one of the receiving actions of the device (i.e., the first device or the second device) in the embodiments of Figure 4 or Figure 7 any one of the sending actions of the device (i.e., the first device or the second device) in the embodiments of.

[0144] In one implementation, when the communication device 1100 is the first device, a device in the first device (such as a chip or a chip system in the first device), or a device that can be used in matching with the first device, wherein:

[0145] The communication unit 1101 is used to receive the data feature information of the second local data set from the second device, and the second local data set is the data set of the second device; the processing unit 1102 is used to determine a first training data set based on the data feature information of the second local data set and the first local data set, and the first local data set is the data set of the first device, and the data feature information of the first training data set includes the data feature information of the second local data set and the data feature information of the first local data set; the processing unit 1102 is further used to obtain a first model based on the first training data set, and the first model is a model deployed on the first device.

[0146] In a possible implementation, the data feature information includes one or more of the following information: data identification information, data distribution information, or data classification information; wherein, the data identification information is used to indicate the data included in the data set, the data distribution information is used to indicate the distribution of the data in the data set, and the data classification information is used to indicate the classification of the data in the data set.

[0147] In a possible implementation, the data distribution information includes one or more of clustering distribution information, probability distribution information, or model parameter information of a data generation model; the data classification information includes one or more of geographical region information, signal-related information, device configuration information, quality-related information, or classification center information, Doppler information, or delay spread information.

[0148] In a possible implementation, the processing unit 1102 is further configured to determine, based on the data feature information of the first local data set, the data feature information of the second data set difference set from the data feature information of the second local data set, where the data feature information of the second data set difference set is the data feature information not included in the first local data set; the communication unit 1101 is further configured to send the data feature information of the second data set difference set to the second device; the communication unit 1101 is further configured to receive the second data set difference set from the second device; the first training data set is the union of the second data set difference set and the first local data set.

[0149] In a possible implementation, the processing unit 1102 is further configured to determine, based on the data feature information of the second local data set, the data feature information of the first data set difference set from the first local data set, where the data feature information of the first data set difference set is the data feature information not included in the second local data set; the communication unit 1101 is further configured to send the first data set difference set to the second device.

[0150] In a possible implementation, the processing unit 1102 is further configured to align the first model and the second model, where the second model is a model deployed on the second device and is obtained based on the second training data set, and the data feature information of the first training data set includes the data feature information of the second training data set.

[0151] In a possible implementation, the data feature information of the second training data set is the same as the data feature information of the first training data set.

[0152] In a possible implementation, the processing unit 1102 is further configured to determine, based on the data feature information of the second local data set, the data feature information of the first data set difference set from the first local data set, where the data feature information of the first data set difference set is the data feature information not included in the second local data set; the communication unit 1101 is further configured to send the first data set difference set to the second device.

[0153] In a possible implementation, the processing unit 1102 is further configured to determine a reference model based on the size of the first training data set and the size of the second training data set, where the reference model is the first model or the second model; the processing unit 1102 is further configured to align the first model and the second model based on the reference model.

[0154] In a possible implementation, the data feature information included in the first training data set further includes the data feature information of the third training data set, and the third training data set is used to train the third model, and the third model is a model deployed on the third device.

[0155] For a more detailed description of the above communication unit 1101 and processing unit 1102, reference can be made to Figure 4 or Figure 7 the relevant description of the first device in the method embodiment shown.

[0156] In one implementation, when the communication device 1100 is a device in the second device, a device in the second device, or a device that can be used in matching with the second device, where:

[0157] The communication unit 1101 is configured to send the data feature information of the second local data set to the first device, and the second local data set is a data set of the second device.

[0158] In a possible implementation, the data feature information includes one or more of the following information: data identification information, data distribution information, or data classification information; wherein, the data identification information is used to indicate the data included in the data set, the data distribution information is used to indicate the distribution of the data in the data set, and the data classification information is used to indicate the classification of the data in the data set.

[0159] In a possible implementation, the data distribution information includes one or more of clustering distribution information, probability distribution information, or model parameter information of a data generation model; the data classification information includes one or more of geographical region information, signal-related information, device configuration information, quality-related information, or classification center information.

[0160] In a possible implementation, the communication unit 1101 is further configured to receive the data feature information of the difference set of the second data set from the first device, and the data feature information of the difference set of the second data set is the data feature information included in the second local data set and not included in the first local data set, and the first local data set is a data set of the first device; the communication unit 1101 is further configured to send the difference set of the second data set to the first device based on the data feature information of the difference set of the second data set.

[0161] In a possible implementation, the processing unit 1102 is further configured to align the first model and the second model, and the second model is a model deployed on the second device, and the second model is obtained based on the second training data set, and the data feature information of the first training data set includes the data feature information of the second training data set.

[0162] In a possible implementation, the data feature information of the second training data set is the same as that of the first training data set.

[0163] In a possible implementation, the communication unit 1101 is further configured to receive a first data set difference set from the first device, where the data feature information of the first data set difference set is included in the first local data set and not included in the data feature information of the second local data set; the second training data set is the union of the first data set difference set and the second local data set.

[0164] In a possible implementation, the processing unit 1102 is further configured to determine a reference model based on the size of the first training data set and the size of the second training data set, where the reference model is the first model or the second model; the processing unit 1102 is further configured to align the first model and the second model based on the reference model.

[0165] In a possible implementation, the data feature information included in the first training data set further includes the data feature information of a third training data set, where the third training data set is used to train a third model, and the third model is a model deployed on a third device.

[0166] For a more detailed description of the above communication unit 1101 and processing unit 1102, reference can be made to Figure 4 or Figure 7 the relevant description of the second device in the method embodiment shown.

[0167] In a possible implementation manner, when the communication device 1100 is a chip, the communication unit 1101 may be a communication interface, a pin, a circuit, etc. The communication interface can be used to input data to be processed into the processor and can output the processing result of the processor outward. In a specific implementation, the communication interface can be a general purpose input output (GPIO) interface, and can be connected to multiple peripheral devices (such as a display (LCD), a camera, a radio frequency (RF) module, an antenna, etc.). The communication interface is connected to the processor through a bus.

[0168] The processing unit 1102 may be a processor, and the processor can execute the computer execution instructions stored in the storage module to enable the chip to execute Figure 4 or Figure 7The method involved in any of the illustrated embodiments. Further, the processor may include a controller, an arithmetic unit, and registers. Exemplarily, the controller is mainly responsible for instruction decoding and issuing control signals for the operations corresponding to the instructions. The arithmetic unit is mainly responsible for performing fixed-point or floating-point arithmetic operations, shift operations, and logical operations, etc., and can also perform address operations and conversions. The registers are mainly responsible for storing register operands and intermediate operation results temporarily stored during the execution of instructions, etc. In a specific implementation, the hardware architecture of the processor may be an application-specific integrated circuit (ASIC) architecture, a microprocessor without interlocked piped stages architecture (MIPS) architecture, an advanced RISC machines (ARM) architecture, or a network processor (NP) architecture, etc. The processor can be single-core or multi-core. The storage module may be a storage module within the chip, such as registers, caches, etc. The storage module may also be a storage module located outside the chip, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.

[0169] It should be noted that the functions corresponding to the processor and the interface can be implemented through hardware design, software design, or a combination of software and hardware, and there is no limitation here.

[0170] Figure 12 This is a schematic structural diagram of another communication device provided by an embodiment of the present application. It can be understood that the communication device 1200 includes means in the necessary forms such as modules, units, components, circuits, or interfaces, which are appropriately configured together to execute the present solution. The communication device 1200 may be the above-mentioned first device or second device, or a component (such as a chip) in these devices, for implementing the method described in the above method embodiments.

[0171] In a possible design, as Figure 12 shown, the communication device 1200 includes a processor 1210 and an interface circuit 1220. The processor 1210 and the interface circuit 1220 are coupled to each other.

[0172] Optionally, the communication device 1200 may include one or more processors 1210. The processor 1210 may be a general-purpose processor or a dedicated processor, etc. For example, it may be a baseband processor or a central processing unit. The baseband processor may be used to process communication protocols and communication data, and the central processing unit may be used to control the communication device (such as a terminal device, a network device, or a chip, etc.), execute software programs, and process the data of software programs.

[0173] It can be understood that the interface circuit 1220 may be a transceiver or an input / output interface. When the communication device 1200 is the first device or the second device, the interface circuit 1220 is a transceiver, including a transmitter and / or a receiver. Among them, the transmitter may be referred to as a sending unit, a transmitter, or a sending circuit, etc., for implementing the sending function, and the receiver may be referred to as a receiving unit, a receiver, or a receiving circuit, etc., for implementing the receiving function. When the communication device 1200 is a chip in the first device or the second device, the interface circuit 1220 is the input / output interface of the chip. Optionally, the communication device 1200 may further include an antenna (not shown in the figure), and the interface circuit 1220 may sometimes also be referred to as a transceiver unit, a transceiver, a transceiver circuit, or a transceiver, etc., for implementing the transceiver function of the communication device through the antenna.

[0174] Optionally, the communication device 1200 may further include a memory 1230, which is used to store instructions executed by the processor 1210 or input data required for the processor 1210 to run instructions or data generated after the processor 1210 runs instructions. Optionally, the processor 1210 and the memory 1230 may be provided separately or integrated together.

[0175] When the communication device 1200 is used to implement Figure 4 or Figure 7 the method shown, the processor 1210 is used to implement the functions of the above-mentioned processing unit 1102, and the interface circuit 1220 is used to implement the functions of the above-mentioned communication unit 1101.

[0176] When the above-mentioned communication device is a chip applied to the first device, the terminal chip implements the functions of the first device in the above-mentioned method embodiment. The first device chip receives information from the second device. It can be understood that the information is first received by other modules (such as a radio frequency module or an antenna) in the first device, and then sent to the first device chip by these modules. The first device chip sends information to the second device. It can be understood that the information is first sent to other modules (such as a radio frequency module or an antenna) in the first device, and then sent to the second device by these modules.

[0177] When the above communication device is a chip applied to a second device, the second device chip implements the functions of the second device in the above method embodiments. The second device chip receives information from the first device, which can be understood as the information is first received by other modules (such as a radio frequency module or an antenna) in the second device, and then sent by these modules to the second device chip. The second device chip sends information to the first device, which can be understood as the information is sent to other modules (such as a radio frequency module or an antenna) in the second device, and then sent by these modules to the first device.

[0178] An embodiment of the present application further provides a computer-readable storage medium. The computer-readable storage medium stores computer instructions, and when the computer instructions are executed, the computer is caused to execute the method as described in any item of any embodiment in Figure 4 or Figure 7 any one of the above.

[0179] An embodiment of the present application further provides a computer program product. The computer program product includes: computer program code. When the computer program code is run on a computer, the computer is caused to execute the method as described in any item of any embodiment in Figure 4 or Figure 7 any one of the above.

[0180] In the present application, when entity A sends information to entity B, it can be that A directly sends to B, or A indirectly sends to B through other entities. Similarly, when entity B receives information from entity A, it can be that entity B directly receives the information sent by entity A, or entity B indirectly receives the information sent by entity A through other entities. Here, entity A and B can be RAN nodes or terminals, or modules inside RAN nodes or terminals. The sending and receiving of information can be information interaction between a RAN node and a terminal, for example, information interaction between a base station and a terminal; the sending and receiving of information can also be information interaction between two RAN nodes, for example, information interaction between a CU and a DU; the sending and receiving of information can also be information interaction between different modules inside a device, for example, information interaction between a terminal chip and other modules of the terminal, or information interaction between a base station chip and other modules in the base station.

[0181] It can be understood that the processor in the embodiments of the present application may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.

[0182] The method steps in the embodiments of the present application may be implemented in hardware or in software instructions executable by a processor. The software instructions may be composed of corresponding software modules, and the software modules may be stored in a random access memory, flash memory, read-only memory, programmable read-only memory, erasable programmable read-only memory, electrically erasable programmable read-only memory, registers, hard disk, removable hard disk, CD-ROM, or any other form of storage medium well-known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. The storage medium may also be a component of the processor. The processor and the storage medium may be located in an ASIC. Additionally, the ASIC may be located in a base station or a terminal. The processor and the storage medium may also exist as discrete components in a base station or a terminal.

[0183] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are executed in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user device, or other programmable devices. The computer program or instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer program or instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center integrating one or more available media. The available medium can be a magnetic medium, such as a floppy disk, a hard disk, or a magnetic tape; it can also be an optical medium, such as a digital video disc; or it can be a semiconductor medium, such as a solid-state drive. The computer-readable storage medium can be a volatile or non-volatile storage medium, or can include both volatile and non-volatile types of storage media.

[0184] Reference to "embodiment" herein means that a particular feature, structure, or characteristic described in connection with an embodiment can be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0185] In various embodiments of the present application, if there is no special explanation and logical conflict, the terms and / or descriptions between different embodiments are consistent and can be cross-referenced, and the technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationships.

[0186] In this application, "at least one" means one or more, and "a plurality of" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B may be singular or plural. In the written description of this application, the character " / " generally represents an "or" relationship between the associated objects before and after; in the formulas of this application, the character " / " represents a "division" relationship between the associated objects before and after. "Including at least one of A, B, and C" may represent: including A; including B; including C; including A and B; including A and C; including B and C; including A, B, and C.

[0187] In the description, claims, and drawings of this application, terms such as "first" and "second" are used to distinguish different objects rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of operations or units is not limited to the listed operations or units, but may optionally further include operations or units not listed, or may optionally further include other operations or units inherent to these processes, methods, products, or devices.

[0188] In this application, "send" and "receive" represent the direction of signal transmission. For example, "sending information to XX" can be understood as the destination of the information being XX, which may include directly sending through the air interface or indirectly sending through other units or modules via the air interface. "Receiving information from YY" can be understood as the source of the information being YY, which may include directly receiving from YY through the air interface or indirectly receiving from YY through other units or modules via the air interface. "Send" can also be understood as the "output" of the chip interface, and "receive" can also be understood as the "input" of the chip interface. In other words, sending and receiving can be carried out between devices, for example, between a network device and a terminal device, or can be carried out within a device, for example, sending or receiving between components, modules, chips, software modules, or hardware modules within a device through a bus, trace, or interface. It can be understood that necessary processing may be performed on the information between the source and destination of the information transmission, such as encoding, modulation, etc., but the destination can understand the valid information from the source. Similar expressions in this application can be understood similarly and will not be elaborated further.

[0189] The "indication" in this application may include direct indication and indirect indication, and may also include explicit indication and implicit indication. If the information indicated by a certain piece of information (such as the indication information described below) is called the information to be indicated, then in the specific implementation process, there are many ways to indicate the information to be indicated. For example, but not limited to, the information to be indicated can be directly indicated, such as the information to be indicated itself or the index of the information to be indicated, etc. It is also possible to indirectly indicate the information to be indicated by indicating other information, where there is an association relationship between the other information and the information to be indicated; it is also possible to only indicate a part of the information to be indicated, while the other parts of the information to be indicated are known or pre-agreed. For example, the arrangement order of each piece of information pre-agreed (such as protocol pre-definition) can be used to achieve the indication of specific information, thereby reducing the indication overhead to a certain extent. This application does not limit the specific manner of indication. It can be understood that for the sender of the indication information, the indication information can be used to indicate the information to be indicated, and for the receiver of the indication information, the indication information can be used to determine the information to be indicated.

[0190] It can be understood that the various numerical numbers involved in the embodiments of this application are only for the convenience of description and are not used to limit the scope of the embodiments of this application. The magnitudes of the serial numbers of the above processes do not mean the sequence of execution, and the execution sequence of each process should be determined by its function and internal logic.

Claims

1. A model training method, characterized in that, The method includes: Receiving data feature information of a second local data set from a second device, where the second local data set is a data set of the second device; Determining a first training data set based on the data feature information of the second local data set and a first local data set, where the first local data set is a data set of a first device, and the data feature information of the first training data set includes the data feature information of the second local data set and the data feature information of the first local data set; Obtaining a first model based on the first training data set, where the first model is a model deployed on the first device.

2. The method according to claim 1, characterized in that, The determining the first training data set based on the data feature information of the second local data set and the first local data set includes: Determining data feature information of a second data set difference set from the data feature information of the second local data set based on the data feature information of the first local data set, where the data feature information of the second data set difference set is data feature information not included in the first local data set; Sending the data feature information of the second data set difference set to the second device; Receiving the second data set difference set from the second device; the first training data set is the union of the second data set difference set and the first local data set.

3. The method according to claim 1 or 2, characterized in that, The method further includes: Aligning the first model and a second model, where the second model is a model deployed on the second device, the second model is obtained based on a second training data set, and the data feature information of the first training data set includes the data feature information of the second training data set; Wherein, the second training data set includes the data feature information of the second local data set, or the second training data set includes the data feature information of the second local data set and the data feature information of the first local data set.

4. The method according to claim 3, characterized in that, Before aligning the first model and the second model, the method further includes: Determining data feature information of a first data set difference set from the first local data set according to the data feature information of the second local data set, where the data feature information of the first data set difference set is data feature information not included in the second local data set; Sending the first data set difference set to the second device, and the second training data set is the union of the first data set difference set and the second local data set.

5. A data transmission method, characterized in that, The method includes: Sending data feature information of a second local data set to a first device, where the second local data set is a data set of a second device.

6. The method according to claim 5, characterized in that, The method further includes: Receiving data feature information of a second data set difference set from the first device, where the data feature information of the second data set difference set is data feature information included in the second local data set and not included in a first local data set, and the first local data set is a data set of the first device; Sending the second data set difference set to the first device based on the data feature information of the second data set difference set.

7. The method according to claim 5 or 6, characterized in that, The method further includes: Align the first model and the second model, where the second model is a model deployed on the second device and is obtained based on a second training dataset, and the data feature information of the first training dataset includes the data feature information of the second training dataset; Among them, the second training dataset includes the data feature information of the second local dataset, or the second training dataset includes the data feature information of the second local dataset and the data feature information of the first local dataset.

8. The method according to claim 7, characterized in that, Before aligning the first model and the second model, the method further includes: Receiving a first dataset difference set from the first device, where the data feature information of the first dataset difference set is included in the first local dataset and not included in the data feature information of the second local dataset; the second training dataset is the union of the first dataset difference set and the second local dataset.

9. The method according to claim 3, 4, 7 or 8, characterized in that, Aligning the first model and the second model includes: Determining a reference model based on the size of the first training dataset and the size of the second training dataset, where the reference model is the first model or the second model; Aligning the first model and the second model based on the reference model.

10. The method according to claim 3, 4, 7, 8 or 9, characterized in that, The data feature information included in the first training dataset further includes the data feature information of a third training dataset, where the third training dataset is used to train a third model, and the third model is a model deployed on a third device.

11. The method according to any one of claims 1 - 10, characterized in that, The data feature information includes one or more of the following information: data identification information, data distribution information, or data classification information; Among them, the data identification information is used to indicate the data included in the dataset, the data distribution information is used to indicate the distribution of data in the dataset, and the data classification information is used to indicate the classification of data in the dataset.

12. According to the method described in claim 11, wherein, The data distribution information includes one or more of clustering distribution information, probability distribution information, or model parameter information of a data generation model; The data classification information includes one or more of geographical region information, signal-related information, device configuration information, quality-related information, or classification center information.

13. A communication device, wherein, Includes a module for executing the method according to any one of claims 1-12.

14. A communication device, wherein, Includes a processor and an interface circuit, where the interface circuit is used to receive signals from other communication devices outside the communication device and transmit them to the processor or send signals from the processor to other communication devices outside the communication device, and the processor is used to implement the method according to any one of claims 1-12 through logic circuits or by executing code instructions.

15. According to the device described in claim 14, wherein, The communication device is a chip or a chip system.

16. A computer-readable storage medium, wherein, A computer program or instruction is stored in the storage medium, and when the computer program or instruction is executed by the communication device, the method according to any one of claims 1-15 is implemented.