Model training method, terminal device, and network device

By selecting a suitable training mode from a variety of training modes and utilizing network device resources to assist terminal devices with weak computing power, the problem of low training efficiency caused by the heterogeneity of terminal device computing power is solved, and the model training efficiency is improved.

WO2025194348A1PCT designated stage Publication Date: 2025-09-25QUECTEL WIRELESS SOLUTIONS CO LTD

Patent Information

Application Number
PCT/CN2024/082509
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-19
Publication Date
2025-09-25

AI Technical Summary

Technical Problem

When terminal devices of different types or different architectures jointly train models, the heterogeneous computing power leads to low training efficiency, and existing technologies cannot efficiently perform model training.

Method used

The first terminal device selects a training mode that matches it from a plurality of training modes based on the received information, and uses the computing resources of the network device to assist the terminal device with weaker computing power to perform model training, thereby avoiding training delays.

Benefits of technology

It improves the efficiency of jointly training models on multiple terminal devices and reduces the training delay caused by terminal devices with weak computing power.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024082509_25092025_PF_FP_ABST
    Figure CN2024082509_25092025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a model training method, a terminal device, and a network device, helping to improve the model training efficiency. The method comprises: a first terminal device receives a first model and first information that are sent by a network device; and the first terminal device selects a first training mode from among multiple training modes on the basis of the first information, wherein the multiple training modes are together used for training the first model, and the first training mode is used for the first terminal device to perform uplink transmission and / or model training associated with the first model.
Need to check novelty before this filing date? Find Prior Art

Description

Model training method, terminal equipment and network equipment Technical Field

[0001] The present application relates to the field of machine learning technology, and more specifically, to a model training method, terminal device and network device. Background Art

[0002] With the development of communication technologies, some services, supported by intelligent models, can significantly reduce the data transmission burden in wireless networks. For example, task-oriented semantic communication combined with a deep learning-based semantic codec model can support intelligent connectivity for a variety of terminal devices.

[0003] However, different types of devices or different architectures have significant differences in capabilities, and the time required to train the same model also varies. Therefore, when multiple devices are jointly training a model, how to efficiently train the model is an urgent problem that needs to be solved.

[0004] Summary of the Invention

[0005] The present application provides a model training method, a terminal device, and a network device. The following describes various aspects of the embodiments of the present application.

[0006] In a first aspect, a model training method is provided, including: a first terminal device receives a first model and first information sent by a network device; the first terminal device selects a first training mode from multiple training modes based on the first information; wherein the multiple training modes are commonly used to train the first model, and the first training mode is used by the first terminal device to perform uplink transmission and / or model training related to the first model.

[0007] According to a second aspect, a model training method is provided, comprising: a network device sends a first model and first information to a first terminal device; wherein, the first information is used by the first terminal device to select a first training mode from a plurality of training modes, the plurality of training modes are used together to train the first model, and the first training mode is used by the first terminal device to perform uplink transmission and / or model training related to the first model.

[0008] According to a third aspect, a terminal device is provided, which is a first terminal device, and includes: a first receiving unit for receiving a first model and first information sent by a network device; a first processing unit for selecting a first training mode from a plurality of training modes based on the first information; wherein the plurality of training modes are used together to train the first model, and the first training mode is used by the first terminal device to perform uplink transmission and / or model training related to the first model.

[0009] In a fourth aspect, a network device is provided, comprising: a first sending unit for sending a first model and first information; wherein, the first information is used by the first terminal device to select a first training mode from multiple training modes, the multiple training modes are used together to train the first model, and the first training mode is used by the first terminal device to perform uplink transmission and / or model training related to the first model.

[0010] In a fifth aspect, a communication device is provided, comprising a memory and a processor, wherein the memory is used to store a program, and the processor is used to call the program in the memory to execute the method described in the first aspect or the second aspect.

[0011] In a sixth aspect, a device is provided, comprising a processor for calling a program from a memory to execute the method as described in the first aspect or the second aspect.

[0012] In a seventh aspect, a chip is provided, comprising a processor for calling a program from a memory so that a device equipped with the chip executes the method described in the first aspect or the second aspect.

[0013] In an eighth aspect, a computer-readable storage medium is provided, on which a program is stored, wherein the program enables a computer to execute the method as described in the first aspect or the second aspect.

[0014] In a ninth aspect, a computer program product is provided, comprising a program, wherein the program enables a computer to execute the method as described in the first aspect or the second aspect.

[0015] In a tenth aspect, a computer program is provided, which enables a computer to execute the method as described in the first aspect or the second aspect.

[0016] In an embodiment of the present application, the first terminal device determines a corresponding training mode from a plurality of training modes based on the first information. The plurality of training modes are used for collaborative training of the first model by a plurality of terminal devices and a network device. The first terminal device can participate in the training of the first model according to the determined training mode. Thus, the model training method of the embodiment of the present application includes a plurality of training modes, so that the first terminal device can select a training mode that matches it, thereby improving the efficiency of the joint training of the first model by a plurality of terminal devices. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] FIG1 is a wireless communication system used in an embodiment of the present application.

[0018] FIG2 is a schematic diagram of a codec model system suitable for semantic communication.

[0019] FIG3 is a flow chart of a model training method provided in an embodiment of the present application.

[0020] FIG4 is a flow chart of a possible implementation of the method shown in FIG3 .

[0021] FIG5 is a flowchart of another possible implementation of the method shown in FIG3 .

[0022] Figure 6 is a schematic diagram of a training system for a semantic codec model provided in an embodiment of the present application.

[0023] FIG7 is a schematic structural diagram of a terminal device provided in an embodiment of the present application.

[0024] FIG8 is a schematic structural diagram of a control device of the terminal device shown in FIG7 .

[0025] FIG9 is a schematic structural diagram of a network device provided in an embodiment of the present application.

[0026] FIG10 is a schematic structural diagram of a control device of the network device shown in FIG8 .

[0027] FIG11 is a schematic structural diagram of an electronic device provided in an embodiment of the present application.

[0028] FIG12 is a schematic block diagram of a communication device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0029] The following will describe the technical solutions in the embodiments of this application in conjunction with the drawings in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of the embodiments. With respect to the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0030] The embodiments of the present application can be applied to various communication systems. For example, the embodiments of the present application can be applied to global system of mobile communication (GSM) system, code division multiple access (CDMA) system, wideband code division multiple access (WCDMA) system, general packet radio service (GPRS), long term evolution (LTE) system, advanced long term evolution (LTE-A) system, new radio (NR) system, evolution system of NR system, LTE-based access to unlicensed spectrum (LTE-U) system on unlicensed spectrum, NR-based access to unlicensed spectrum (NR-U) system on unlicensed spectrum, NTN system, universal mobile telecommunication system (UMTS), wireless local area networks (WLAN), wireless fidelity (WiFi), and fifth generation communication (5th-generation, 5G) system. The embodiments of the present application can also be applied to other communication systems, such as future communication systems. The future communication system may be, for example, a sixth-generation (6G) mobile communication system or a satellite communication system.

[0031] Traditional communication systems support a limited number of connections and are easy to implement. However, with the development of communication technology, communication systems can not only support traditional cellular communications, but also support one or more other types of communications. For example, a communication system can support one or more of the following communications: device to device (D2D) communication, machine to machine (M2M) communication, machine type communication (MTC), enhanced machine type communication (eMTC), vehicle to vehicle (V2V) communication, and vehicle to everything (V2X) communication, etc. The embodiments of the present application can also be applied to communication systems that support the above-mentioned communication methods.

[0032] The communication system in the embodiment of the present application can be applied to a carrier aggregation (CA) scenario, a dual connectivity (DC) scenario, and a standalone (SA) networking scenario.

[0033] The communication system in the embodiments of the present application can be applied to unlicensed spectrum. The unlicensed spectrum can also be considered a shared spectrum. Alternatively, the communication system in the embodiments of the present application can also be applied to licensed spectrum. The licensed spectrum can also be considered a dedicated spectrum.

[0034] The embodiments of the present application can be applied to NTN systems. As examples, the NTN systems may include 4G-based NTN systems, NR-based NTN systems, Internet of Things (IoT)-based NTN systems, and narrowband Internet of Things (NB-IoT)-based NTN systems.

[0035] A communication system may include one or more terminal devices. The terminal devices mentioned in the embodiments of the present application may also be referred to as user equipment (UE), access terminal, subscriber unit, subscriber station, mobile station, mobile station (MS), mobile terminal (MT), remote station, remote terminal, mobile device, user terminal, terminal, wireless communication device, user agent or user device, etc.

[0036] In some embodiments, the terminal device may be a station (ST) in a WLAN. In some embodiments, the terminal device may be a cellular phone, a cordless phone, a session initiation protocol (SIP) phone, a wireless local loop (WLL) station, a personal digital assistant (PDA) device, a handheld device with wireless communication capabilities, a computing device or other processing device connected to a wireless modem, an in-vehicle device, a wearable device, a terminal device in a next-generation communication system (e.g., a NR system), or a terminal device in a future-evolved public land mobile network (PLMN) network.

[0037] In some embodiments, a terminal device may be a device that provides voice and / or data connectivity to a user. For example, the terminal device may be a handheld device, an in-vehicle device, etc. with wireless connection capabilities. As some specific examples, the terminal device may be a mobile phone, a tablet computer, a laptop computer, a PDA, a mobile internet device (MID), a wearable device, a virtual reality (VR) device, an augmented reality (AR) device, a wireless terminal in industrial control, a wireless terminal in self-driving, a wireless terminal in remote medical surgery, a wireless terminal in a smart grid, a wireless terminal in transportation safety, a wireless terminal in a smart city, a wireless terminal in a smart home, etc.

[0038] In some embodiments, the terminal device can be deployed on land. For example, the terminal device can be deployed indoors or outdoors. In some embodiments, the terminal device can be deployed on the water, such as on a ship. In some embodiments, the terminal device can be deployed in the air, such as on an airplane, a balloon, or a satellite.

[0039] In addition to the terminal device, the communication system may also include one or more network devices. The network device in the embodiment of the present application may be a device for communicating with the terminal device, and the network device may also be referred to as an access network device or a radio access network device. The network device may be, for example, a base station. The network device in the embodiment of the present application may refer to a radio access network (RAN) node (or device) that connects the terminal device to a wireless network. A base station can broadly cover various names as follows, or be replaced with the following names, such as: NodeB, evolved NodeB (eNB), next generation NodeB (gNB), relay station, access point, transmission point (TRP), transmission point (TP), master station MeNB, secondary station SeNB, multi-standard radio (MSR) node, home base station, network controller, access node, wireless node, access point (AP), transmission node, transceiver node, baseband unit (BBU), remote radio unit (RRU), active antenna unit (AAU), remote radio head (RRH), central unit (CU), distributed unit (DU), positioning node, etc. A base station can be a macro base station, a micro base station, a relay node, a donor node or the like, or a combination thereof. A base station can also refer to a communication module, modem or chip used to be set in the aforementioned device or apparatus. The base station can also be a mobile switching center and a device that performs base station functions in D2D, V2X, and M2M communications, a network-side device in a 6G network, or a device that performs base station functions in future communication systems. The base station can support networks with the same or different access technologies. The embodiments of this application do not limit the specific technology and specific device form used by the network equipment.

[0040] Base stations can be fixed or mobile. For example, a helicopter or drone can be configured to act as a mobile base station, and one or more cells can move based on the location of the mobile base station. In other examples, a helicopter or drone can be configured to act as a device that communicates with another base station.

[0041] In some deployments, the network device in the embodiments of the present application may refer to a CU or a DU, or the network device may include a CU and a DU. The gNB may also include an AAU.

[0042] By way of example and not limitation, in embodiments of the present application, a network device may be mobile, for example, a mobile device. In some embodiments of the present application, the network device may be a satellite or balloon station. In some embodiments of the present application, the network device may also be a base station located on land, water, or the like.

[0043] In an embodiment of the present application, the network device can provide services for a cell, and the terminal device communicates with the network device through the transmission resources used by the cell (for example, frequency domain resources, or spectrum resources). The cell can be a cell corresponding to the network device (for example, a base station). The cell can belong to a macro base station or a base station corresponding to a small cell. The small cells here may include: metro cells, micro cells, pico cells, femto cells, etc. These small cells have the characteristics of small coverage and low transmission power, and are suitable for providing high-speed data transmission services.

[0044] In some embodiments, the present application can also be applied to artificial intelligence communication systems. One of the goals of the 3rd Generation Partnership Project (3GPP) version (release) 18 is to enhance the capabilities of 5G and expand its application to new devices, deployments, and industries. As network designs become more complex, which can include a wide range of deployment and usage options, traditional methods will not be able to provide quick solutions. Because manual reconfiguration of cellular communication systems is costly and inefficient, it is necessary to use artificial intelligence (AI) and machine learning (ML) to automate operational processes to reduce costs by automating functions that require human interaction. For example, by using large amounts of data collected from wireless networks, AI and ML can solve complex and unstructured network problems.

[0045] As an example, AI can be used in the core network and RAN to implement intelligent network operations. For example, AI can be used to enhance quality of service (QoS), improve efficiency, simplify deployment, and improve security.

[0046] As an example, on-device AI can benefit the entire communications system. One potential AI-enabled capability is radio awareness. AI can provide valuable knowledge through environmental and contextual awareness, reducing overhead and latency. Through radio awareness, communications systems can support enhanced device experiences, such as intelligent beamforming and power management. Furthermore, AI can help improve system performance, such as reducing interference, achieving better spectrum utilization, and enhancing radio security. For example, AI can help better detect and prevent malicious attacks.

[0047] For example, Figure 1 is a schematic diagram of the architecture of a communication system provided in an embodiment of the present application. As shown in Figure 1, the communication system 100 may include a network device 110, which may be a device that communicates with a terminal device 120 (or also referred to as a communication terminal or terminal). The network device 110 may provide communication coverage for a specific geographic area and may communicate with terminal devices located within the coverage area.

[0048] FIG1 exemplarily shows a network device and two terminal devices. In some embodiments of the present application, the communication system 100 may include multiple network devices and each network device may include another number of terminal devices within its coverage area, which is not limited.

[0049] In an embodiment of the present application, the communication system shown in Figure 1 also includes other network entities such as a mobility management entity (MME) and an access and mobility management function (AMF), but this embodiment of the present application does not limit this.

[0050] It should be understood that in the embodiments of the present application, a device having a communication function in a network / system may be referred to as a communication device. Taking the communication system 100 shown in FIG1 as an example, the communication device may include a network device 110 and a terminal device 120 having a communication function. The network device 110 and the terminal device 120 may be the specific devices described above and will not be described in detail here. The communication device may also include other devices in the communication system 100, such as a network controller, a mobility management entity, and other network entities, which are not limited in the embodiments of the present application.

[0051] To facilitate the detailed explanation of the innovative features of the technical solution, we first introduce some relevant technical knowledge involved in the embodiments of this application. The following related technologies can be combined with the technical solutions of the embodiments of this application as optional solutions, and they all fall within the scope of protection of the embodiments of this application. The embodiments of this application include at least part of the following contents.

[0052] With the continuous advancement of communication technologies, certain services, supported by intelligent models, can reduce the data transmission burden in wireless networks. For example, task-oriented semantic communication, an emerging communication paradigm in 6G networks, only transmits task-related information, significantly reducing the data transmission burden on wireless networks. Furthermore, supported by semantic codec models based on deep learning, semantic communication can achieve intelligent connectivity for everything.

[0053] Task-oriented semantic communication is a task-based, "understand first, transmit later" communication method. A task-oriented semantic communication system typically consists of a transmitter, a receiver, a wireless channel, source data, target data, a semantic encoder model, and a semantic decoder model. The semantic decoder model is also called a semantic decoder model. In standard communication, the transmitter uses the semantic encoder model to encode the source data into a semantic signal; the transmitter then transmits the semantic signal to the receiver; the receiver receives the semantic signal transmitted through the channel and uses the semantic decoder model to decode the semantic signal into the target data.

[0054] For ease of understanding, the following exemplary description of a codec model system applicable to semantic communication is given in conjunction with FIG2 . The model system 200 in FIG2 includes a transmitting end 210 , a receiving end 220 , and a wireless channel 230 .

[0055] 2 , in step S21 , the transmitter 210 inputs the original data X into the semantic encoder model 212 .

[0056] In step S22 , the semantic encoder model 212 encodes the original data X into a semantic signal Z.

[0057] In step S23 , the transmitter 210 transmits the semantic signal Z through the wireless channel 230 .

[0058] In step S24, considering the noise n in the wireless channel 230, the semantic signal Z becomes a wireless semantic signal after passing through the channel

[0059] In step S25, the receiving end 220 uses the semantic decoder model to convert the wireless semantic signal Decoded into target data Y.

[0060] The codec model for semantic communication, described above with reference to Figure 2, requires sufficient data samples for deep learning and training of intelligent models such as the semantic codec model. However, these data samples are distributed and stored across different terminal devices in wireless edge networks. Therefore, efficiently training the semantic codec model becomes a pressing issue.

[0061] Currently, the main approaches for model training in wireless edge networks include federated learning and centralized learning. The following uses the semantic codec model as an example to illustrate federated learning and centralized learning systems.

[0062] A centralized learning system enables centralized training of deep learning-based semantic codec models. A centralized learning system in a wireless edge network also consists of a base station and multiple end devices. First, each end device uploads local data samples to the base station. The base station then uses the collected data samples to train a global semantic codec model until the global semantic codec model meets the preset convergence criteria or reaches the maximum number of preset training rounds, completing the centralized learning process.

[0063] In centralized learning systems, when end devices upload their local data samples to network devices, this can lead to private data leakage. When deploying existing centralized learning systems to train semantic codec models, the local data samples uploaded by each end device contain private information related to the end device. Therefore, uploading local data samples exposes the end device's private data to the base station, posing a risk of private data leakage.

[0064] Federated learning systems enable distributed training of semantic codec models based on deep learning. Typically, a federated learning system in a wireless edge network consists of a base station and multiple terminal devices. The entire training process is divided into multiple rounds. In each training round, the base station broadcasts a global semantic codec model to all terminal devices. Each terminal device uses its local data to train the global semantic codec model, generating a local semantic codec model, and then uploads the local semantic codec model to the base station via a wireless channel. The base station then aggregates the local semantic codec models uploaded by all terminal devices to obtain a global semantic codec model. This process is repeated until the global semantic codec model meets the preset convergence conditions or the maximum preset number of training rounds is reached, thus completing the entire federated learning process.

[0065] While federated learning systems can address the privacy concerns of centralized learning systems, local model training on each terminal device underutilizes base station computing resources. Specifically, when deploying existing federated learning systems to train semantic codec models, training is performed only on each terminal device. The base station is solely responsible for aggregating and broadcasting the semantic codec model, not training it. This significantly wastes base station computing resources.

[0066] The above describes federated learning systems and centralized learning systems. Regardless of the type of learning system, data or models must be transmitted over a communication system. In existing wireless communications, terminal devices often use digital communication. This involves first encoding the original data into a bit stream using source and channel encoding, which is then uploaded to a base station via a wireless channel. The base station then performs channel and source decoding to recover the original data. This process requires encoding the original data into a high-dimensional bit stream, which results in significant latency overhead and reduces the transmission rate. This demonstrates the low efficiency of data transmission based on digital communication. Compared to digital communication, data transmission based on over-the-air computing can significantly improve transmission efficiency.

[0067] Over-the-air computing is a new type of non-orthogonal access technology. In over-the-air computing, multiple terminal devices pre-process the original signal and then leverage the superposition characteristics of the wireless channel to simultaneously transmit information on the same time-frequency resources. The signal received by the base station is the superposition of the signals transmitted by multiple terminal devices. The base station then post-processes the superimposed signal to obtain the target signal, thus enabling signal transmission and computation. Therefore, over-the-air computing requires certain pre-processing and post-processing at both the transmitting and receiving ends. Through these pre-processing and post-processing, over-the-air computing can unify the communication and computation processes during communication, thereby improving transmission efficiency.

[0068] While over-the-air computing can improve transmission efficiency, it requires multiple devices to transmit signals on the same time-frequency resources. However, the devices performing model training may be different. Different types of devices, or even the same type, may have different types, performance, and architectures of computing power, which is known as heterogeneous computing power.

[0069] Due to the heterogeneous nature of terminal devices, the time required to complete the training of the local semantic codec model in the aforementioned federated learning system varies. For example, devices with strong computing power may require less time, while devices with weaker computing power may require more time. As heterogeneity increases, the difference in training time also increases.

[0070] Therefore, terminal devices with weak computing power significantly extend the training latency of federated learning systems. Specifically, when distributed training a semantic codec model based on federated learning is performed, the system training latency is determined by the terminal device that takes the longest to complete local model training. Therefore, when multiple terminal devices are jointly training a model, how to efficiently train the model becomes a pressing issue.

[0071] It should be noted that the above-mentioned problem of long training delay of codec model for semantic communication due to heterogeneity of computing power of different terminal devices is only an example. The embodiments of the present application can be applied to any type of model training scenario where the model training efficiency is low due to heterogeneity of computing power of different terminal devices.

[0072] In response to the above problems, an embodiment of the present application proposes a model training method. Through this method, the first terminal device can select a training mode corresponding to it from a plurality of training modes for training the first model based on the first information. Multiple training modes can be used for terminal devices with different capabilities, thereby avoiding the problem of low training efficiency caused by the heterogeneous capabilities of multiple terminal devices. When the computing power of the first terminal device is weak, sufficient computing resources on the base station side can assist the first terminal device in model training, thereby reducing the training delay caused by the first terminal device. For ease of understanding, the model training method is described in detail below in conjunction with Figure 3.

[0073] Figure 3 illustrates the interaction between a first terminal device and a network device. The first terminal device is one of the aforementioned terminal devices that possesses certain computing and communication capabilities. In some embodiments, the first terminal device can train a first model based on local data samples. In some embodiments, the first terminal device can send the data samples and the trained model to the network device. In some embodiments, the first terminal device can receive the first model broadcasted by the network device.

[0074] In some embodiments, the first terminal device may be any terminal device in a wireless edge network. The first terminal device may store a variety of data samples. The data samples stored by the first terminal device may include local data samples used to train the first model.

[0075] In some embodiments, the first terminal device may use the first model for data processing. Exemplarily, when the first model includes a semantic encoder model, the first terminal device may encode local data based on the encoder model to obtain a semantic signal. Exemplarily, when the first model includes a semantic decoder model, the first terminal device may decode the semantic signal based on the decoder model.

[0076] The first terminal device may be any terminal device among the multiple terminal devices participating in the first model training. For example, the multiple terminal devices may train the first model together with the network device. For example, the multiple terminal devices may train the first model.

[0077] In some embodiments, the first terminal device may store multiple local data samples to participate in the training of the first model. In some embodiments, multiple terminal devices may jointly train the first model based on the local data samples. In some embodiments, multiple terminal devices may provide the network device with local data samples for training the first model.

[0078] In some embodiments, the local data samples may be various data samples used by the first terminal device to train the first model. For example, the local data samples include but are not limited to text, voice, image, video, etc., which are not limited in the embodiments of the present application.

[0079] As an example, the first terminal device may store a local data sample set having multiple samples. For example, when the model training system includes N+K terminal devices (N and K are both integers ≥ 1), the N+K terminal devices may store N+K local data sample sets respectively. The N+K local data sample sets respectively have D1, ..., D N , D N+1 ,…,D N+K samples, which can be expressed as:

[0080] Among them, x m,l is the original data of the lth sample at the mth terminal device, y m,l is the target data of the lth sample; m is a natural number from 1 to N+K, l is a natural number from 1 to D m The natural number D m Indicates the total number of data samples of the mth terminal device.

[0081] In some embodiments, some or all of the terminal devices communicating with the network device participate in model training.

[0082] The network device is any of the aforementioned communication devices that provide services to multiple terminal devices. In some embodiments, the network device is a communication device with powerful computing capabilities. The network device can be the base station that broadcasts the global model to multiple terminal devices based on federated learning, or the base station that trains the global model based on centralized learning, without limitation.

[0083] The network device may communicate with multiple terminal devices, including the first terminal device. For example, the network device may receive data samples or local models sent by the multiple terminal devices. For example, the network device may send a global model of a certain round to the multiple terminal devices. For example, the network device may send a resource allocation policy for model training or data transmission to the multiple terminal devices.

[0084] In some embodiments, the network device can group multiple terminal devices to facilitate the determination of how to train the model. For example, the network device can determine some terminal devices with weaker capabilities based on the capability information of multiple terminal devices and instruct these terminal devices not to perform model training locally, thereby avoiding excessive latency.

[0085] 2 , in step S210 , the first terminal device receives the first model and first information sent by the network device.

[0086] The first model may be a machine learning model that supports multiple communication services. Optionally, the first model may be applied to a wireless edge network.

[0087] In some embodiments, the first model may be a codec model. That is, the first model may include an encoder model and a decoder model. The encoder model is used for encoding at the transmitting end, and the decoder model is used for decoding at the receiving end.

[0088] Optionally, the transmitting end and the receiving end may be the transmitting and receiving ends of a wireless communication link. Exemplarily, when the transmitting end is a network device, the receiving end may be a terminal device, or may be another network device other than the transmitting end. Exemplarily, when the transmitting end is a terminal device, the receiving end may be a network device, or may be any terminal device other than the transmitting end.

[0089] In some embodiments, the first model can be a codec model for semantic communication, which can also be referred to as a semantic codec model or a collaborative semantic codec model. For example, the first model can be a collaborative semantic codec model for semantic communication. The encoder model in the first model is used by the transmitter to semantically encode the data to be sent to obtain a semantic signal, while the decoder model is used by the receiver to receive and decode the semantic signal.

[0090] The first model can be any one or more of a variety of neural network models, which is not limited in the present embodiment. Optionally, the first model includes but is not limited to: a convolutional neural network model, a recurrent neural network model, a variational autoencoder neural model, etc.

[0091] The first model may be a model currently being trained. The training process of the first model may include multiple training cycles. A training cycle may refer to a round of training or learning, also known as a round or learning round. For example, the tth round is training cycle t.

[0092] In some embodiments, the start time of a training cycle can be uniformly indicated by the network device. For example, the network device can send a model training start instruction to all terminal devices participating in the training. Taking the semantic codec model as an example, the model training start instruction can be implemented using a "1" bit signal. When the first terminal device receives a "1" bit, it can start the training of the local semantic codec model or the semantic encoding task of the original data; if it does not receive a "1" bit, it can wait and maintain the current behavior.

[0093] In some embodiments, the duration of a training cycle can be determined based on multiple training modes and / or multiple terminal device groups of the first model. For example, the duration of a training cycle is determined by the longest duration of multiple training modes running in parallel. In another example, the duration of a training cycle is determined by the duration of a group of terminal devices with weaker computing capabilities. This will be described below with reference to multiple terminal device groups as an example.

[0094] In some embodiments, the number of training cycles during the entire training process can be determined by the network device. For example, after the base station determines that the number of training rounds has reached a preset maximum, it broadcasts a training termination instruction to all terminal devices participating in the training. Alternatively, after the base station determines that the training has reached convergence, it broadcasts a training termination instruction to all terminal devices participating in the training.

[0095] As an example, when the training cycle t=T, it is determined that the current training reaches a preset maximum number of rounds of training, where T is the maximum number of rounds of training.

[0096] In some embodiments, during the training of the first model, the first model may be the model applied to any training cycle. That is, the first model may be the model trained in any training cycle. In some embodiments, the first model may be the model determined in the previous training cycle before any training cycle. That is, in any training cycle other than the first training cycle, the first model may be the model determined in the previous training cycle.

[0097] In some embodiments, the first model may be a global model to be trained in any training cycle. For example, the first model may be a global semantic codec model φ broadcasted by the base station to multiple terminal devices in the tth round of training. t ,θ t Among them, φ t Can be a Q-dimensional real vector, used to represent the global semantic encoder model; θ t It can be an R-dimensional real vector, used to represent the global semantic decoder model.

[0098] In some embodiments, the first model may be determined by the network device. For example, the network device may integrate information related to model training in the previous training cycle to determine the first model for the current training cycle. For example, the network device may determine the first model based on the training results of the previous training cycle. For example, the network device may determine the second model for the next training cycle based on the training results of the current training cycle. For the last training cycle, the second model may be the final global model. For example, the network device may determine the second model to be trained in the next training cycle based on the training results of the first model in the current training cycle.

[0099] As an example, the training of the first model may be performed jointly by the network device and a plurality of terminal devices including the first terminal device.

[0100] In some embodiments, the network device may send the first model to multiple terminal devices via broadcasting so that all terminal devices participating in model training receive the first model. In other words, the first terminal device may receive the first model sent by the network device via broadcasting.

[0101] As an example, the first terminal device may receive the global semantic encoder model and the global semantic decoder model via broadcasting.

[0102] In some embodiments, broadcasting the first model to multiple terminal devices by the network device may be part of a current training cycle. For example, the model receiving process may be used to determine the start time of the current training cycle. For example, broadcasting the first model by the network device may indicate the start of the current training cycle. For another example, the network device broadcasts the first model after the current training cycle begins.

[0103] In some embodiments, broadcasting the first model from the network device to multiple terminal devices may not be part of the current training cycle. For example, the current training cycle begins only after all terminal devices receive the first model by executing the model.

[0104] The first information may be sent directly by the network device or instructed by a higher layer, which is not limited here. For example, the network device may send the first information by broadcasting.

[0105] The first information is determined by the network device. Optionally, the network device can determine the first information based on multiple terminal devices participating in the model training.

[0106] In step S320, the first terminal device selects a first training mode from a plurality of training modes according to the first information.

[0107] The training mode can be the way in which one or more devices participating in the training perform model training, which can also be called the training method.

[0108] Multiple training modes are used together to train the first model. That is, the training of the first model is not performed based on a single training mode, but is implemented based on multiple training modes. For example, multiple devices participating in the training of the first model can jointly train the first model based on multiple training modes.

[0109] In some embodiments, the multiple training modes may include the centralized training mode based on the centralized learning system and the distributed training mode based on the federated learning system described above, and may also include other modes for training machine learning models.

[0110] Optionally, the multiple training modes may include training of the first model by some or all terminal devices in the communication system, that is, the multiple training modes may be determined according to the type or number of terminal devices participating in the training.

[0111] Optionally, the multiple training modes may include the network device training some or all of the models in the first model. That is, the multiple training modes may be determined based on the type or number of models in the first model that the network device participates in training. For example, the network device may participate in the training of some or all of the models in the first model.

[0112] Optionally, the multiple training modes may include collaborative training of the first model by multiple terminal devices and network devices.

[0113] In some embodiments, when the first model includes an encoder model and a decoder model, the multiple training modes may include centralized training and distributed training. Exemplarily, distributed training involves training the encoder model and the decoder model separately by multiple terminal devices. Exemplarily, centralized training involves training the decoder model by a network device.

[0114] The first information is used by the first terminal device to determine a first training mode among multiple training modes. In other words, the first terminal device can determine a mode for training the first model based on the first information.

[0115] In some embodiments, the first information is used to determine whether the first training mode includes centralized training of part or all of the first model by the network device. As an example, when the first training mode includes centralized training, the first terminal device does not directly train the first model; when the first training mode does not include centralized training, the first terminal device directly trains the first model. In other words, the first terminal device can determine whether to directly train the first model based on the first information.

[0116] As an example, when the first training mode includes the centralized training, the first terminal device sends first data / first signal for centralized training to the network device; when the first training mode does not include the centralized training, the first terminal device trains the first model.

[0117] Optionally, the first data used for centralized training may be local data of the first terminal device or processed data.

[0118] Optionally, the first signal used for centralized training can be a signal converted by the first terminal device based on local data samples, or a signal converted after processing local data, thereby avoiding privacy leakage during transmission. For example, when the first model is a semantic-oriented codec model, the first terminal device can input local data samples into the encoder to obtain a semantic signal (first signal).

[0119] In some embodiments, the first information may include one or more information for determining the first training mode. As an example, the first information may include a first threshold value related to the capabilities of the terminal device, so that the first terminal device determines the first training mode based on the first information and its capabilities. As an example, the first information may include a first condition. When the first terminal device meets the first condition, the first training mode may instruct the first terminal device to directly train the first model; when the first terminal device does not meet the first condition, the first training mode may instruct the first terminal device not to directly train the first model.

[0120] Optionally, when the first information includes a first threshold, if the capability of the first terminal device is lower than the first threshold, the first training mode includes centralized training of part or all of the first models by the network device.

[0121] Optionally, the first threshold value may be determined according to a capability parameter. For example, when the capability parameter is processor frequency, the first threshold value is a frequency threshold.

[0122] In some embodiments, the first information may include a grouping strategy for multiple terminal device groups. The multiple terminal device groups correspond one-to-one to the multiple training modes. For example, the network device may divide the multiple terminal devices participating in training into multiple terminal device groups based on the multiple training modes. In another example, after receiving the grouping strategy, each terminal device may determine the device group to which it belongs.

[0123] As an example, the grouping strategy in the first information is used by the first terminal device to determine the first training mode according to the terminal device group to which it belongs. For example, the first terminal device can select the terminal device group to which it belongs from multiple terminal device groups according to the grouping strategy.

[0124] In some embodiments, the grouping strategy in the first information may be determined based on capabilities and / or channel coefficients of the plurality of terminal devices. The capabilities of the terminal devices may include computing capabilities and / or transmission capabilities of the terminal devices.

[0125] As an example, the terminal device grouping strategy on the base station side may include but is not limited to: processor frequency of each terminal device, number of floating-point operations performed per second, local data sample size, etc. This application does not impose any limitation on this.

[0126] Exemplarily, the processor of the terminal device includes various processors such as a central processing unit (CPU).

[0127] As an example, a network device can divide multiple terminal devices into a group of high-computing-capability terminal devices and a group of low-computing-capability terminal devices based on their computing capabilities. Examples of high-computing-capability terminal devices include computers or communications devices with similar capabilities, while examples of low-computing-capability terminal devices include mobile phones or communications devices with similar capabilities. For simplicity, the group of high-computing-capability terminal devices will be referred to as the first group of terminal devices, and the group of low-computing-capability terminal devices will be referred to as the second group of terminal devices. The network device will be described using a base station as an example.

[0128] For example, the training system of the first model may include 1 base station and N+K terminal devices. The base station uses the terminal device grouping strategy to group the N+K terminal devices. Among them, N terminal devices are grouped into a first terminal device group; K terminal devices are grouped into a second terminal device group. The first terminal device can be the nth terminal device among the N terminal devices, or the kth terminal device among the K terminal devices, where n is any number from 1 to N and k is any number from 1 to K. Therefore, D n It can represent the local data sample set of the nth terminal device among N terminal devices; D k It can represent the local data sample set of the kth terminal device among K terminal devices.

[0129] As an example, when the first terminal device belongs to the second terminal device group, the first terminal device may not directly train the first model. When the first model is a semantic codec model, the first terminal device can perform semantic codec model training with the assistance of sufficient computing resources on the base station side, thereby reducing the training delay of the semantic codec model. Specifically, the base station groups all terminal devices into a first terminal device group and a second terminal device group, and the two groups of devices execute different training modes respectively. Among them, the terminal devices in the first terminal device group have strong computing power, and the training of the first model will not cause a large delay; the terminal devices in the second terminal device group have weak computing power, and the first model is trained with the assistance of the base station, thereby avoiding large delays and improving training efficiency.

[0130] As can be seen from the foregoing, the duration of a training cycle can be determined based on multiple terminal device groups. Taking the training system of the semantic codec model as an example, the training system may include a base station, a first terminal device group, and a second terminal device group. During a training cycle, the terminal devices in the first terminal device group perform model training and model upload, the terminal devices in the second terminal device group perform signal extraction and upload, and the base station performs centralized training and model aggregation. In this scenario, a training cycle may include a terminal device training cycle for the first terminal device group, a terminal device semantic signal extraction cycle for the second terminal device group, a terminal device model gradient upload cycle for the first terminal device group, a terminal device semantic signal upload cycle for the second terminal device group, a base station centralized semantic decoder model training cycle, and a base station global semantic codec model aggregation cycle.

[0131] Within a training cycle, the terminal device training cycle of the first terminal device group and the terminal device model gradient upload cycle of the first terminal device group are in a serial relationship, forming a first terminal device group cycle. The terminal device semantic signal extraction cycle of the second terminal device group and the terminal device semantic signal upload cycle of the second terminal device group are in a serial relationship, forming a second terminal device group cycle. Furthermore, the first terminal device group cycle and the second terminal device group cycle are in a parallel relationship, forming a terminal device cycle. The terminal device cycle, the base station centralized semantic decoder model training cycle, and the base station global semantic codec model aggregation cycle are in a serial relationship.

[0132] In some embodiments, the computing capabilities or transmission capabilities of the plurality of terminal devices may be determined by capability parameters of each terminal device. The capability parameters of each terminal device may include one or more of the following: processor frequency, maximum processor frequency, number of floating-point operations, local data sample size, maximum uplink transmission power, and maximum energy consumption budget.

[0133] As an example, the capability parameters of the first terminal device may include one or more of the following: the number of floating-point operations of the first terminal device; the local data sample size of the first terminal device; the maximum processor frequency of the first terminal device; the maximum uplink transmission power of the first terminal device; and the maximum energy consumption budget of the first terminal device.

[0134] For example, N+K terminal devices can report the maximum processor frequency separately. Maximum uplink transmit power and the maximum energy consumption budget

[0135] In some embodiments, the grouping strategy can be determined based on the capabilities of multiple terminal devices participating in model training. As an example, the network device can collect capability information from multiple terminal devices to determine the first information. For example, the network device can send a resource data collection instruction or a capability information collection instruction to multiple terminal devices to trigger the multiple terminal devices to transmit capability information. For another example, the local resources that can be reported by the terminal devices may include the maximum processor frequency, the maximum uplink transmission power, and the maximum energy consumption budget.

[0136] As an example, the network device may send a first instruction to the first terminal device to trigger the first terminal device to send capability parameters. The first instruction may be a resource data collection instruction or a capability information collection instruction, or an instruction to achieve similar requirements.

[0137] In some embodiments, the first instruction may be a signal sent by the base station and received by each terminal device. Optionally, the first instruction may be implemented using a "1" bit signal. For example, when a terminal device receives a "1" bit, it may transmit local resources on orthogonal frequency resources and send an orthogonal pilot to the base station. If it does not receive a "1" bit, it may wait and maintain its current behavior.

[0138] In some embodiments, the first information may be determined based on channel coefficients associated with multiple terminal devices. For example, each terminal device may transmit a pilot signal when transmitting capability parameters to facilitate determination of the channel coefficient by the network device. In another example, when receiving resource data, the base station may obtain the channel coefficient of each terminal device based on the received pilot signal.

[0139] Optionally, the pilot signal used to obtain the channel coefficient of each terminal device may be an orthogonal pilot signal or a direct sequence spread spectrum signal.

[0140] Optionally, methods for obtaining channel coefficients include, but are not limited to: zero-forcing channel estimation, least squares channel estimation, etc. This embodiment of the present application does not limit this.

[0141] In some embodiments, the first information may be determined based on the capabilities and channel coefficients of multiple terminal devices. For example, when the first information includes a grouping strategy, the base station may group the terminal devices after receiving resource data of each terminal device and determining the channel coefficient.

[0142] The first training mode is used for the first terminal device to perform uplink transmission and / or model training related to the first model. Exemplarily, the first training mode may include any one or more training modes among multiple training modes, which are not limited here.

[0143] In some embodiments, uplink transmission related to the first model may refer to the first terminal device sending data, signals, or model parameters related to the first model to the network device, without limitation. For example, the first terminal device sends first data / first signal to the network device. In another example, the first terminal device sends a model gradient to the network device.

[0144] In some embodiments, model training related to the first model may refer to the first terminal device locally training part or all of the models in the first model. In other words, the first terminal device directly trains part or all of the models in the first model.

[0145] Optionally, the first training mode may include centralized training of some or all of the first models by the network device. Alternatively, the first training mode may not include distributed training of some or all of the first models by the first terminal device. In these scenarios, the first terminal device may be any terminal device in the second terminal device group.

[0146] Optionally, the first training mode may include distributed training of part or all of the first model by the first terminal device. Alternatively, the first training mode may not include centralized training of part or all of the first model by the network device. In these scenarios, the first terminal device may be any terminal device in the first terminal device group.

[0147] In some embodiments, when the first training mode is distributed training, the first terminal device trains the first model based on local data samples to obtain a first local model gradient. The first local model gradient is aggregated with other local model gradients based on over-the-air calculation.

[0148] As an example, when the first model is a semantic codec model, the first local model gradient is a semantic codec gradient model.

[0149] As an example, multiple terminal devices participating in distributed training upload their local model gradients based on over-the-air computing. For example, multiple terminal devices upload their generated local semantic codec model gradients to the base station on the same time-frequency resources.

[0150] In some embodiments, when the first terminal device belongs to the first terminal device group, the first training mode performed by the first terminal device is distributed training. Exemplarily, the first terminal device can train the first model during a terminal device training cycle and upload the local model gradient obtained from the training during a terminal device model gradient upload cycle.

[0151] For ease of understanding, the following uses the semantic codec model as an example to illustrate the first training mode of the terminal devices in the first terminal device group. Taking the tth round of training (training cycle t) as an example, during the terminal device training cycle, the N terminal devices in the first terminal device group can use the local data sample sets D1, D2, ..., D N Train the global semantic encoder-decoder model φ t ,θ t , get N local semantic encoder-decoder model gradients

[0152] Optionally, during the terminal device training cycle, the training performed by multiple terminal devices may include local semantic signal generation, distortion impact simulation, target data reconstruction, loss function calculation, and related calculations of local model gradients.

[0153] As an example, the nth terminal device in the first terminal device group can use the semantic encoder obtained by broadcasting and the original data x n,l Generate a local semantic signal. For example, the local semantic signal z t,n,l It can be expressed as:

[0154] z t,n,l =S(x n,l ;φ t );

[0155] Among them, z t,n,l is an M-dimensional real signal vector with mean 1 and variance 0, M is an integer ≥ 1, and S(·) is a semantic encoding operation.

[0156] As an example, in order to simulate the distortion effect of the wireless channel on the semantic signal, the nth terminal device in the first terminal device group can apply an influencing factor to the semantic signal generated locally. For example, the terminal device can apply Gaussian white noise to the semantic signal generated locally. The signal after applying Gaussian white noise is It can be expressed as:

[0157] in, Indicates power is Real Gaussian white noise, I M Represents the identity matrix.

[0158] As an example, the nth terminal device in the first terminal device group can use The semantic decoder obtained by broadcasting reconstructs the target data. For example, the reconstructed target data It can be expressed as:

[0159] Among them, S -1 (·) is the semantic decoding operation.

[0160] As an example, the nth terminal device in the first terminal device group can calculate a local loss function and train the received global semantic codec model using a gradient descent method. It can be expressed as:

[0161] Among them, f t,n,l =f(φ t ,θ t ;x n,l ,y n,l ) is the nth terminal device in the first terminal device group about the original data x n,l With the target data y n,l The local loss function.

[0162] Optionally, the local loss function includes but is not limited to: a cross entropy loss function, a mean square error loss function, a hinge loss function, etc. This embodiment of the present application does not limit this.

[0163] As an example, the nth terminal device in the first terminal device group can calculate the local semantic codec model gradient according to the local loss function Exemplarily, the encoder model gradient and the decoder model gradient It can be expressed as:

[0164] in, are the original data x obtained by the nth terminal device in the first terminal device group during training period t. n,l With the target data y n,l gradient; and Both are gradient operators.

[0165] As an example, the nth terminal device in the first terminal device group can calculate the mean and the mean square sum of the local semantic codec model gradients. mean square sum They can be expressed as:

[0166] in, is the qth entry of the local semantic codec model gradient of the nth terminal device located in the first terminal device group; Q and R are both integers greater than or equal to 1.

[0167] As an example, N terminal devices in the first terminal device group can upload the mean and mean square of the local semantic codec model gradient to the base station. The base station can calculate the mean and variance of the global semantic codec model gradient and broadcast the mean and variance of the global semantic codec model gradient. variance They can be expressed as:

[0168] As can be seen from the above, during the terminal device training cycle, the nth terminal device in the first terminal device group can use the local data sample set to train the global semantic codec model. It can be expressed as:

[0169] in, The CPU frequency assigned to the nth terminal device in the first terminal device group in the tth round of training, κ n The number of central processing unit cycles required to train one data sample for the nth terminal device in the first terminal device group.

[0170] Still taking the training cycle t as an example, during the terminal device model gradient upload cycle, the N terminal devices in the first terminal device group upload the local semantic codec model gradient in the same time-frequency resource. Upload to the base station.

[0171] Optionally, during the model gradient upload period, multiple terminal devices respectively perform model gradient normalization processing, model gradient upload, etc.

[0172] As an example, the nth terminal device in the first terminal device group may perform normalization processing on the local semantic codec model gradient. Exemplarily, the normalized local semantic codec model gradient signal s obtained by the normalization processing is t,n It can be expressed as:

[0173] Among them, D P =∑ n∈N D n , D P Indicates the total number of data samples of the first terminal device group.

[0174] As an example, the nth terminal device in the first terminal device group can upload the normalized local semantic codec model gradient signal to the base station. For example, the power v of the uploaded model gradient signal is t,n It can be expressed as:

[0175] v t,n =p t,n s t,n ;

[0176] Among them, p t,n is the uplink transmission power of the nth terminal device in the first terminal device group.

[0177] As an example, N terminal devices in the first terminal device group can upload the local semantic codec model gradient to the base station at the same time and frequency based on over-the-air calculation. For example, the upload period of the nth terminal device is It can be expressed as:

[0178] Among them, S is the number of symbols that can be transmitted on each subchannel in a unit subframe, T subframe is the duration of a unit subframe, and ceil(·) is a rounding function.

[0179] As an example, the base station receives a superimposed gradient signal. t It can be expressed as:

[0180] y t =∑ n∈N h t,n p t,n s t,n +n t ;

[0181] Among them, h t,n is the channel coefficient from the nth terminal device in the first terminal device group to the base station, n t ~CN(0,σ 2 I Q+R ) is the power σ 2 Complex Gaussian white noise, I Q+R Represents the identity matrix.

[0182] As an example, the base station applies a receiving scalar α to the received superimposed gradient signal. t , we can get the federated learning aggregate semantic encoder-decoder model gradient, that is, the aggregate model gradient. Optionally, the aggregate model gradient It can be expressed as:

[0183] In some embodiments, when the first training mode is centralized training, the first terminal device inputs local data samples into the encoder model to obtain encoded first data / first signal. Furthermore, the first terminal device transmits target data and the first data / first signal to the network device. The target data and the first data / first signal are used by the network device to train the decoder model.

[0184] Optionally, the target data and the first data / first signal may be as described above and will not be repeated here.

[0185] In some embodiments, when the first terminal device belongs to the second terminal device group, the first training mode performed by the first terminal device is centralized training. For example, the first terminal device can obtain the semantic signal to be transmitted during the terminal device semantic signal extraction period and upload the semantic signal during the terminal device semantic signal upload period.

[0186] For ease of understanding, the semantic codec model is still used as an example to illustrate the first training mode of the terminal devices in the second terminal device group. Taking the tth round of training (training period t) as an example, during the terminal device semantic signal extraction period of the second terminal device group, the K terminal devices in the second terminal device group use the local data sample set D N+1 , D N+2 ,…,D N+K With the global semantic encoder model φ t , extracting semantic signals from local data samples Among them, z t,N+k,l is the semantic signal of the lth sample in the kth terminal device.

[0187] As an example, the kth terminal device in the second terminal device group uses the semantic encoder obtained by broadcasting and the original data x N+k,l Generate a local semantic signal. For example, the local semantic signal z generated by the kth terminal device t,N+k,l It can be expressed as:

[0188] z t,N+k,l =S(x N+k,l ;φ t );

[0189] Among them, z t,N+k,lis an M-dimensional real signal vector with mean 1 and variance 0, and S(·) is the semantic encoding operation.

[0190] As an example, the kth terminal device in the second terminal device group uses the local data sample set and the global semantic encoder model to extract the semantic signal of the local data sample. It can be expressed as:

[0191] in, The CPU frequency assigned to the kth terminal device in the second terminal device group in the tth round of training, k N+k The number of central processing unit cycles required to extract the semantic signal of a data sample for the k-th terminal device located in the second terminal device group.

[0192] Still taking the training period t as an example, during the terminal device semantic signal uploading period, the K terminal devices in the second terminal device group respectively extract the semantic signals on different frequency resources. With the original target data Upload to the base station. The base station can receive wireless semantic signals through the wireless channel and target data

[0193] As an example, the kth terminal device in the second terminal device group uploads the extracted semantic signal. For example, the power v of the uploaded semantic signal is t,N+k,l It can be expressed as:

[0194] v t,N+k,l =p t,N+k z t,N+k,l ;

[0195] Among them, p t,N+k is the uplink transmission power of the kth terminal device in the second terminal device group.

[0196] As an example, the period of uploading the extracted semantic signal by the kth terminal device in the second terminal device group is It can be expressed as:

[0197] As an example, the original semantic signal y received by the base station t,N+k,l It can be expressed as:

[0198] y t,N+k,l =h t,N+k p t,N+k z t,N+k,l +n t,N+k,l ;

[0199] Among them, n t,N+k,l ~CN(0,σ 2 I M ) is the power σ 2 Complex Gaussian white noise, h t,N+k Represents the channel coefficient of the kth terminal device.

[0200] As an example, the base station obtains the wireless semantic signal using the received original semantic signal. It can be expressed as:

[0201] As can be seen from the foregoing, when the first training mode is centralized training, the base station needs to train at least part of the first model; when the first training mode is distributed training, the base station receives the aggregated model after training of multiple terminal devices.

[0202] In some embodiments, when the first model is a semantic codec model, the base station trains the decoder model to reduce training latency. For example, the base station trains a global semantic decoding model using received wireless semantic signals to obtain a centralized semantic decoder model gradient.

[0203] In some embodiments, the distributed training may determine an aggregated model gradient after aggregating multiple local model gradients in the current training cycle. For example, the base station may receive or determine an aggregated model gradient after aggregating multiple local model gradients in the current training cycle.

[0204] In some embodiments, the centralized training may determine a centralized model gradient after training the decoder model in a current training cycle. Exemplarily, the base station trains the decoder model to determine a centralized model gradient after training the decoder model in the current training cycle.

[0205] In some embodiments, the aggregated model gradient and the centralized model gradient are used together to determine a second model to be trained in a next training cycle. Exemplarily, the base station determines the second model for the next training cycle based on the aggregated model gradient and the centralized model gradient.

[0206] The previous example of a terminal device participating in training described the signal reception and processing performed by the base station. For ease of understanding, the following example uses the semantic codec model as an example to illustrate the model training and model processing performed by the base station in the training system.

[0207] Taking training period t as an example, during the training period of the base station centralized semantic decoder model, the base station can use the received wireless semantic signal Train the global semantic decoder model to obtain the centralized semantic decoder model gradient

[0208] As an example, the base station calculates the global loss function and trains the global semantic decoding model using the gradient descent method. For example, the global loss function F C (φ t ,θ t ) can be expressed as:

[0209] in, When training a centralized semantic decoder model for a base station, the wireless semantic signal With the target data y N+k,l The loss function of D W =∑ k∈κ D N+k , D W Indicates the total number of data samples in the second terminal device group.

[0210] As an example, the global semantic decoder model gradient calculated by the base station according to the global loss function It can be expressed as:

[0211] in, is the gradient of the loss function with respect to the global semantic encoder-decoder model.

[0212] During the aggregation period of the base station global semantic codec model of training period t, the base station can use federated learning to aggregate the semantic codec model gradient and centralized semantic decoder model gradient Update the global semantic encoder-decoder model φ t ,θ t , get the global semantic encoder-decoder model φ for the next round of training t+1 ,θ t+1 That is, the encoder model in the first model is φ t , the decoder model is θ t When the encoder model φ in the second model t+1 and decoder model θ t+1 It can be expressed as:

[0213] Among them, γ represents the training learning rate of the first model in the current training cycle, D P represents the total number of data samples of all terminal devices performing distributed training, D W represents the total number of data samples of all terminal devices performing centralized training, D = D P+D w , represents the encoder model gradient in the aggregated model gradient, represents the decoder model gradient in the aggregated model gradient, represents the centralized model gradient.

[0214] As can be seen from Figure 3, the first terminal device can select the first training mode according to the first information. Multiple terminal devices can respectively select appropriate training modes to avoid delay differences caused by terminal devices of different types or different capabilities performing the same task. Taking the collaborative semantic codec model for semantic communication as an example, the model training system can offload the training tasks of terminal devices with weak computing capabilities to the base station, and use the sufficient computing resources on the base station side to effectively alleviate the impact of terminal devices with weak computing capabilities on the overall semantic codec model training delay. Furthermore, this method can reduce the semantic codec model training delay and improve the global semantic codec model performance. At the same time, each terminal device in the multiple terminal device groups does not transmit original data, which can effectively prevent the privacy leakage of terminal devices.

[0215] However, if terminal devices within the same terminal device group use different resource allocation strategies to execute the first training mode, large time differences may still occur. In order to more efficiently perform model training, the embodiments of the present application propose a resource allocation mechanism for a model training system to achieve collaborative resource allocation among terminal devices and improve computing and transmission efficiency.

[0216] In some embodiments, the first terminal device may receive second information sent by the network device. The second information is used by the first terminal device to determine a resource allocation strategy for executing the first training mode. In other words, the network device may indicate to the first terminal device the relevant resource configuration for participating in the first model training, thereby synchronizing the execution progress of multiple terminal devices.

[0217] In some embodiments, the resource allocation strategy may be used to determine the uplink transmission power, energy consumption budget and / or processor frequency of the first terminal device executing the first training mode.

[0218] Optionally, the resource allocation strategy may include a central processing unit frequency allocation strategy and / or an uplink transmission power allocation strategy.

[0219] In some embodiments, the second information is determined based on capability parameters of a plurality of terminal devices used to train the first model.

[0220] Optionally, the generation basis of the grouping strategy and resource allocation strategy performed by the base station side is the same. For example, the generation basis of the resource allocation strategy includes but is not limited to: the processor frequency of each terminal device, the number of floating-point operations performed per second, the local data sample size, etc.

[0221] In some embodiments, the first information may include the second information. That is, the second information may be sent together with the first information. For example, after the base station generates a terminal device grouping strategy and a central processing unit frequency and uplink transmission power allocation strategy, the base station may broadcast it to each terminal device. This is not limited in the embodiments of the present application.

[0222] As an example, after completing the terminal device grouping strategy and the uplink transmission power and CPU frequency allocation strategy for each terminal device, the base station transmits the grouping information of each terminal device and the CPU frequency and uplink transmission power allocation strategy through broadcast information.

[0223] In some embodiments, the second information may be sent separately. For example, the base station sends the resource allocation strategy after sending the grouping strategy.

[0224] As an example, when the first training mode is distributed training, the first terminal device is a terminal device in the first terminal device group.

[0225] For example, the uplink transmission power p of the nth terminal device in the first terminal device group is t,n It can be expressed as:

[0226] Where i represents any number from 1 to N, which is used to determine the minimum value among N terminal devices; the receiving scalar α t It can be determined as:

[0227] For example, the energy consumed by the uplink transmission of the nth terminal device in the first terminal device group is It can be expressed as:

[0228] For example, the CPU frequency of the nth terminal device in the first terminal device group is It can be expressed as:

[0229] in, is the energy consumption coefficient of the central processing unit of the nth terminal device in the first terminal device group, A maximum energy consumption budget for a central processing unit of the nth terminal device in the first terminal device group.

[0230] As an example, when the first training mode is centralized training, the first terminal device is a terminal device in the second terminal device group. Exemplarily, the uplink transmission power p of the kth terminal device in the second terminal device group is t,N+k It can be expressed as:

[0231] For example, the energy consumed by the uplink transmission of the kth terminal device in the second terminal device group is It can be expressed as:

[0232] For example, the CPU frequency of the kth terminal device in the second terminal device group is It can be expressed as:

[0233] in, is the energy consumption coefficient of the central processing unit of the kth terminal device in the second terminal device group, is the maximum energy consumption budget of the central processor of the kth terminal device in the second terminal device group.

[0234] In some embodiments, after completing resource allocation, the first terminal device may send a resource allocation completion instruction to the base station. Optionally, the resource allocation completion instruction may use a "1" bit signal with the same number of terminal devices and be transmitted on an orthogonal frequency. When the base station receives a "1" bit from the nth terminal device, it determines that the nth terminal device has completed resource allocation; if it does not receive a "1" bit from the nth terminal device, it determines that the nth terminal device has not completed resource allocation.

[0235] As can be seen above, in the resource allocation mechanism for the collaborative semantic codec model training system, the first terminal group allocates uplink transmission power based on channel inversion technology; the second terminal group allocates uplink transmission power based on the wireless channel coefficient to ensure that the semantic signal received by the base station meets the default signal-to-noise ratio. Furthermore, each terminal group allocates CPU frequency to efficiently complete local semantic codec model training or semantic encoding of raw data.

[0236] For ease of understanding, the following flowcharts of Figures 4 and 5 respectively introduce the training mechanism and resource allocation mechanism of a collaborative semantic codec model for semantic communication using an embodiment of the present application.

[0237] As shown in FIG4 , the process of the training mechanism of the collaborative semantic codec model includes steps S410 to S470 .

[0238] In step S410, the current training cycle begins, and the base station broadcasts the global semantic codec model (first model) obtained from the previous round of gradient hybrid aggregation, as well as the grouping and resource allocation strategies for each terminal device. A training cycle includes a terminal device training cycle and model gradient upload cycle for the first terminal device group, a terminal device semantic signal extraction cycle and semantic signal upload cycle for the second terminal device group, a base station centralized semantic decoder model training cycle, and a base station global semantic codec model aggregation cycle.

[0239] In some embodiments, upon entering the current training cycle t, the base station broadcasts the global semantic codec model obtained from the previous round of gradient hybrid aggregation, and broadcasts the grouping strategy, central processing unit frequency, and uplink transmission power allocation strategy of each terminal device to each terminal device. Each terminal device determines the device category based on the received grouping strategy and allocates the central processing unit frequency and uplink transmission power based on the central processing unit frequency and uplink transmission power allocation strategy.

[0240] In step S422, the terminal devices of the first terminal device group train the global semantic codec model using the local data sample set to obtain a local semantic codec model gradient.

[0241] In step S424 , the terminal devices of the second terminal device group use the local data sample set and the global semantic encoder model to extract semantic signals of the local data samples.

[0242] In step S432, the terminal devices of the first terminal device group upload the local semantic codec model gradients to the base station on the same time-frequency resources based on over-the-air computing, and the base station receives the federated learning aggregated semantic codec model gradients.

[0243] In step S434, the terminal devices of the second terminal device group upload the semantic signals of the local data samples to the base station on the orthogonal frequency resources, and the base station receives the wireless semantic signals through the wireless channel.

[0244] In step S440, the base station trains a global semantic decoding model using the received wireless semantic signal to obtain a centralized semantic decoder model gradient.

[0245] In step S450, the base station performs mixed aggregation using the received federated learning aggregated semantic codec model gradient and the trained centralized semantic decoder model gradient to obtain a global semantic codec model.

[0246] In step S460, it is determined whether convergence or a preset maximum number of iterations has been reached. In other words, the base station determines whether the training of the cooperative semantic codec model for semantic communication has reached the preset maximum number of training rounds. If so, step S470 is executed; if not, the process returns to step S410.

[0247] In step S470, the cooperative semantic codec model training for semantic communication is terminated. After determining that the number of rounds of training has reached a preset maximum, the base station broadcasts a training termination instruction to all terminal devices.

[0248] As shown in Figure 4, in the training system for a collaborative semantic codec model for semantic communication, a first group of terminal devices performs local training of the semantic codec model and, based on over-the-air computation, uploads the local semantic codec model gradients to the base station for aggregation. The base station then obtains the federated learning aggregated semantic codec model gradients. A second group of terminal devices performs semantic encoding of local data samples and uploads the target data in the local data samples and the encoded semantic signals to the base station.

[0249] In the example of Figure 4, the training mechanism of the collaborative semantic codec model for semantic communication can increase the amount of local data samples participating in model training in each round of training, realize the hybrid aggregation of the local semantic codec model and the base station-side semantic codec model, and improve the performance of the global model. On the one hand, in each round of training, the base station receives the target data and semantic signals from the second terminal device group, and performs centralized training of the semantic decoder model to obtain a centralized semantic decoder model. On the other hand, the base station uses the received federated learning aggregate semantic encoder model gradient and the centralized semantic decoder model gradient to update the global semantic decoder model. Finally, the base station obtains the global semantic codec model trained in this round.

[0250] As shown in FIG5 , the process of the resource allocation mechanism of the training mechanism of the collaborative semantic codec model includes steps S510 to S550 .

[0251] In step S510, the base station broadcasts a resource data collection instruction to each terminal device, each terminal device reports local resources and sends a pilot signal, and the base station receives the resource data and obtains the channel coefficient of each terminal device based on the received pilot signal.

[0252] In step S520, the base station divides all terminal devices into a first terminal device group and a second terminal device group according to the received resource data and channel coefficients of each terminal device, and generates an uplink transmission power and a central processing unit frequency allocation strategy for each terminal device.

[0253] In step S530, the base station broadcasts the grouping information, central processing unit frequency, and uplink transmission power allocation strategy of each terminal device to each terminal device, and each terminal device receives the grouping information, central processing unit frequency, and uplink transmission power allocation strategy.

[0254] In step S540, each terminal device determines a device category according to the received grouping strategy, and allocates a central processing unit frequency and an uplink transmission power according to the central processing unit frequency and uplink transmission power allocation strategy.

[0255] In step S550, each terminal device completes resource allocation and reports the resource allocation completion instruction to the base station, ending the resource allocation for collaborative semantic codec model training. The base station broadcasts the model training start instruction, and each terminal device performs the corresponding model training and transmission tasks according to the grouping.

[0256] As shown in Figure 5, the resource allocation mechanism for collaborative semantic codec model training in the embodiment of the present application can achieve collaborative resource allocation for terminal devices and improve computing and transmission efficiency. Specifically, the first terminal device group allocates uplink transmission power based on channel inversion technology to achieve precise over-the-air calculation gradient signal transmission and aggregation; the second terminal device group allocates uplink transmission power based on wireless channel coefficients to achieve accurate transmission of semantic signals. Each terminal device group collaboratively allocates central processing unit frequency to achieve efficient execution of local semantic codec model training and semantic encoding tasks for raw data.

[0257] The following describes the embodiment of the present application in more detail with reference to the specific example of Figure 6. It should be noted that the examples of Figures 3 to 5 are merely intended to help those skilled in the art understand the embodiment of the present application, and are not intended to limit the embodiment of the present application to the specific numerical values ​​or specific scenarios illustrated. It is obvious that those skilled in the art can make various equivalent modifications or changes based on the examples of Figures 3 to 5, and such modifications or changes also fall within the scope of the embodiment of the present application.

[0258] Figure 6 is a schematic diagram of the structure of the training system of the cooperative semantic codec model for semantic communication. As shown in Figure 6, the training system consists of a base station 630 and N+K terminal devices. The local data sample sets of the N+K terminal devices are D1, ..., D N , D N+1 ,…,D N+K N terminal devices form a first terminal device group 610, and K terminal devices form a second terminal device group 620. The training process shown in Figure 6 is divided into several training cycles. In training cycle t, N+K terminal devices can receive the first model broadcast by base station 630.

[0259] In S61 , the first terminal device group 610 trains the first model based on local data samples to obtain N model gradients.

[0260] In S62 , the second terminal device group 620 performs semantic coding based on the local data samples to obtain K groups of semantic signals.

[0261] In S63 , N model gradients are aggregated based on over-the-air calculation to obtain an aggregated model gradient, and the aggregated model gradient is sent to the base station.

[0262] In S64, K groups of semantic signals and target data are sent to the base station respectively.

[0263] In S65 , the base station performs centralized training on the first model according to the K groups of semantic signals and target data to obtain a centralized model gradient.

[0264] In S66 , the aggregated model gradient and the centralized model gradient are mixed and aggregated to obtain a second model.

[0265] The present application also provides a model learning system, which includes a network device and multiple terminal devices, wherein any of the multiple terminal devices executes the method described above for the first terminal device to execute, and the network device executes the method described above for the network device to execute.

[0266] The method embodiment of the present application is described in detail above in conjunction with Figures 1 to 6. The device embodiment of the present application is described in detail below in conjunction with Figures 7 to 12. It should be understood that the description of the device embodiment corresponds to the description of the method embodiment. Therefore, for portions not described in detail, reference can be made to the above method embodiment.

[0267] FIG7 is a schematic block diagram of a terminal device according to an embodiment of the present application. The apparatus 700 may be a first terminal device for model training. The first terminal device may be any of the terminal devices described above. The terminal device 700 shown in FIG7 includes a first receiving unit 710 and a first processing unit 720.

[0268] The first receiving unit 710 may be configured to receive a first model and first information sent by a network device.

[0269] The first processing unit 720 can be used to select a first training mode from multiple training modes based on the first information; wherein, the multiple training modes are used together to train the first model, and the first training mode is used for the first terminal device to perform uplink transmission and / or model training related to the first model.

[0270] Optionally, the first information is used to determine whether the first training mode includes centralized training of part or all of the first model by the network device. The terminal device 700 also includes a first sending unit, which can be used to send first data / first signal for model training to the network device when the first training mode includes centralized training; and a second processing unit, which can be used to train the first model when the first training mode does not include centralized training.

[0271] Optionally, the first information includes a first threshold related to the capability of the terminal device. When the capability of the first terminal device is lower than the first threshold, the first training mode includes centralized training of part or all of the first models by the network device.

[0272] Optionally, the multiple training modes correspond to the multiple terminal device groups one-to-one, and the first information includes grouping strategies for the multiple terminal device groups, where the grouping strategies are used by the first terminal device to determine the first training mode according to the terminal device group to which it belongs.

[0273] Optionally, the first terminal device is one of a plurality of terminal devices for training the first model, and the grouping strategy is determined based on capabilities and / or channel coefficients of the plurality of terminal devices.

[0274] Optionally, the terminal device 700 also includes a second sending unit, which can be used to send capability parameters to the network device, the capability parameters including one or more of the following: the number of floating-point operations of the first terminal device; the local data sample size of the first terminal device; the maximum processor frequency of the first terminal device; the maximum uplink transmission power of the first terminal device; and the maximum energy consumption budget of the first terminal device.

[0275] Optionally, the terminal device 700 further includes a second receiving unit, which can be used to receive a first instruction sent by the network device; wherein the first instruction is used to trigger the first terminal device to send a capability parameter.

[0276] Optionally, the terminal device 700 also includes a third receiving unit, which can be used to receive second information sent by the network device; wherein the second information is used by the first terminal device to determine the resource allocation strategy for executing the first training mode, and the second information is determined based on the capability parameters of multiple terminal devices that train the first model.

[0277] Optionally, the resource allocation strategy is used to determine the uplink transmission power, energy consumption budget and / or processor frequency of the first terminal device executing the first training mode.

[0278] Optionally, the first model includes an encoder model and a decoder model, and the multiple training modes include: distributed training in which multiple terminal devices train the encoder model and the decoder model respectively; and centralized training in which a network device trains the decoder model.

[0279] Optionally, when the first training mode is distributed training, the terminal device 700 also includes a third processing unit, which can be used to train the first model based on local data samples to obtain a first local model gradient; wherein the first local model gradient is aggregated with other local model gradients based on over-the-air calculation.

[0280] Optionally, when the first training mode is centralized training, the terminal device 700 also includes a fourth processing unit, which can be used to input local data samples into the encoder model to obtain the encoded first data / first signal; a third sending unit, which can be used to send the target data and the first data / first signal to the network device, and the target data and the first data / first signal are used by the network device to train the decoder model.

[0281] Optionally, distributed training is used to determine the aggregated model gradient after aggregating multiple local model gradients in the current training cycle, and centralized training is used to determine the centralized model gradient after training the decoder model in the current training cycle. The aggregated model gradient and the centralized model gradient are jointly used to determine the second model to be trained in the next training cycle.

[0282] Optionally, the encoder model in the first model is φ t , the decoder model in the first model is θ t , the encoder model φ in the second model t+1 and decoder model θ t+1 Expressed as:

[0283] Among them, γ represents the training learning rate of the first model in the current training cycle, D P represents the total number of data samples of all terminal devices performing distributed training, D W represents the total number of data samples of all terminal devices performing centralized training, D = D P +D W , represents the encoder model gradient in the aggregated model gradient, represents the decoder model gradient in the aggregated model gradient, represents the centralized model gradient.

[0284] Optionally, the first model is a codec model for semantic communication.

[0285] Figure 8 is a schematic diagram of the structure of a control device for the terminal device shown in Figure 7. When the first model is a codec model for semantic communication, the control device 800 can be a control device for a terminal device in a collaborative semantic codec model training system for semantic communication. As shown in Figure 8, the control device 800 of the terminal device can include a category classification and resource allocation module 810, a calculation module 820 for collaborative model training, and a transmission module 830 for collaborative model training. The collaborative model can be a collaborative semantic codec model.

[0286] The category classification and resource allocation module 810 can be used to control the terminal device to use the received grouping strategy to determine the device category, and allocate the central processing unit frequency and uplink transmission power according to the central processing unit frequency and uplink transmission power allocation strategy. After completing the resource allocation, each terminal device sends a resource allocation completion instruction to the base station.

[0287] The computing module 820 for collaborative model training can be used to control the terminal devices of the first terminal device group to use the local data sample set to train the global semantic codec model at the allocated central processing unit frequency to obtain the local semantic codec model gradient; and can be used to control the terminal devices of the second terminal device group to use the local data sample set to extract the semantic signal of the local data sample using the local data sample set at the allocated central processing unit frequency.

[0288] The transmission module 830 for collaborative model training can be used to control the terminal devices of the first terminal device group to upload the normalized local semantic codec model gradient signal on the same time-frequency resources under the allocated uplink transmission power; and can be used to control the terminal devices of the second terminal device group to upload the semantic signals of local data samples to the base station on orthogonal time-frequency resources under the allocated uplink transmission power.

[0289] FIG9 is a schematic block diagram of a network device according to an embodiment of the present application. The network device 900 may be any of the network devices for model training described above. The network device 900 shown in FIG9 includes a first sending unit 910.

[0290] The first sending unit 910 can be used to send a first model and first information to a first terminal device; wherein the first information is used by the first terminal device to select a first training mode from multiple training modes, multiple training modes are used together to train the first model, and the first training mode is used by the first terminal device to perform uplink transmission and / or model training related to the first model.

[0291] Optionally, the first information is used to determine whether the first training mode includes centralized training of part or all of the first model by the network device. The network device 900 also includes a first receiving unit, which can be used to receive the first data / first signal for model training sent by the first terminal device when the first training mode includes centralized training; and a second receiving unit, which can be used to receive the result of training the first model by the first terminal device when the first training mode does not include centralized training.

[0292] Optionally, the first information includes a first threshold related to the capability of the terminal device. When the capability of the first terminal device is lower than the first threshold, the first training mode includes centralized training of part or all of the first models by the network device.

[0293] Optionally, the multiple training modes correspond to the multiple terminal device groups one-to-one, and the first information includes grouping strategies for the multiple terminal device groups, where the grouping strategies are used by the first terminal device to determine the first training mode according to the terminal device group to which it belongs.

[0294] Optionally, the first terminal device is one of multiple terminal devices for training the first model, and the grouping strategy is determined based on capabilities and / or channel coefficients of the multiple terminal devices.

[0295] Optionally, the network device 900 also includes a third receiving unit, which can be used to receive capability parameters sent by the first terminal device, the capability parameters including one or more of the following: the number of floating-point operations of the first terminal device; the local data sample size of the first terminal device; the maximum processor frequency of the first terminal device; the maximum uplink transmission power of the first terminal device; and the maximum energy consumption budget of the first terminal device.

[0296] Optionally, the network device 900 further includes a second sending unit, which can be used to send a first instruction to the first terminal device; wherein the first instruction is used to trigger the first terminal device to send a capability parameter.

[0297] Optionally, the network device 900 also includes a third sending unit, which can be used to send second information to the first terminal device; wherein the second information is used by the first terminal device to determine the resource allocation strategy for executing the first training mode, and the second information is determined based on the capability parameters of multiple terminal devices that train the first model.

[0298] Optionally, the resource allocation strategy is used to determine the uplink transmission power, energy consumption budget and / or processor frequency of the first terminal device executing the first training mode.

[0299] Optionally, the first model includes an encoder model and a decoder model, and the multiple training modes include: distributed training in which multiple terminal devices train the encoder model and the decoder model respectively; and centralized training in which a network device trains the decoder model.

[0300] Optionally, when the first training mode is distributed training, the network device 900 also includes a fourth receiving unit, which can be used to receive the first local model gradient sent by the first terminal device; wherein the first local model gradient is aggregated with other local model gradients based on over-the-air calculation.

[0301] Optionally, when the first training mode is centralized training, the network device 900 also includes a fifth receiving unit, which can be used to receive the target data and the encoded first data / first signal sent by the first terminal device; and a processing unit, which can be used to train the decoder model based on the target data and the first data / first signal.

[0302] Optionally, the network device 900 also includes a first determination unit, which can be used to determine the aggregated model gradient after aggregating multiple local model gradients in the current training cycle; a second determination unit, which can be used to determine the centralized model gradient after training the decoder model in the current training cycle; and a third determination unit, which can be used to determine the second model for the next training cycle based on the aggregated model gradient and the centralized model gradient.

[0303] Optionally, the encoder model in the first model is φ t , the decoder model in the first model is θ t , the encoder model φ in the second model t+1 and decoder model θ t+1 Expressed as:

[0304] Among them, γ represents the training learning rate of the first model in the current training cycle, D P represents the total number of data samples of all terminal devices performing distributed training, D W represents the total number of data samples of all terminal devices performing centralized training, D = D P +D W , represents the encoder model gradient in the aggregated model gradient, represents the decoder model gradient in the aggregated model gradient, represents the centralized model gradient.

[0305] Optionally, the first model is a codec model for semantic communication.

[0306] Figure 10 is a schematic diagram of the structure of a control device for the network device shown in Figure 9. When the first model is a codec model for semantic communication, control device 1000 can be a control device for a base station in a collaborative semantic codec model training system for semantic communication. As shown in Figure 10, control device 1000 can include a resource allocation strategy generation module 1010, a model gradient signal and semantic signal receiving module 1020, a global semantic decoding model centralized training module 1030, and a global semantic codec model hybrid aggregation module 1040.

[0307] The resource allocation strategy generating module 1010 may be configured to broadcast a resource data collection instruction (ie, a first instruction) to each terminal device and generate a terminal device grouping and resource allocation strategy.

[0308] The model gradient signal and semantic signal receiving module 1020 can be used to receive the federated learning aggregated semantic codec model gradient from the first terminal device group and the wireless semantic signal and target data from the second terminal device group.

[0309] The global semantic decoding model centralized training module 1030 can be used to perform centralized training of the global semantic decoding model using the wireless semantic signal and target data received from the second terminal device group.

[0310] The global semantic codec model hybrid aggregation module 1040 can be used to utilize the received federated learning aggregated semantic codec model gradient and the trained centralized semantic decoder model gradient to perform hybrid aggregation to obtain a global semantic codec model.

[0311] Optionally, the global semantic encoder model can be obtained by updating the gradient of the aggregated semantic encoder model through federated learning, and the global semantic decoder model can be obtained by updating the gradient of the aggregated semantic decoder model through mixed aggregation of the gradient of the aggregated semantic decoder model through federated learning and the gradient of the centralized semantic decoder model.

[0312] FIG11 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device is used to implement any step of the model training method described above. The following description uses a collaborative semantic codec model for semantic communication as an example. As shown in FIG11 , the electronic device structure includes a processor 1110, a memory 1120, a communication interface 1130, and a communication bus 1140.

[0313] The processor 1110 can be used to execute the program stored in the memory 1120 to implement any step in the collaborative semantic codec model training mechanism process for semantic communication provided in the above-mentioned embodiment of the present application.

[0314] In the embodiment of the present application, the processor 1110 may be a general-purpose processor, such as a CPU, a network processor (NP), etc.; it may also be other general-purpose processors or special-purpose processors, such as a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0315] The memory 1120 may be used to store programs related to the training of a collaborative semantic codec model for semantic communication.

[0316] In the embodiment of the present application, the memory 1120 may be a random access memory (RAM) or may include a non-volatile memory (NVM), such as at least one disk storage. Alternatively, the memory may be at least one storage device located remote from the aforementioned processor. This embodiment of the present application does not impose any specific limitations on this.

[0317] The communication bus 1140 can be used to implement communication between the processor 1110, the memory 1120 and the communication interface 1130.

[0318] In the embodiment of the present application, the communication bus 1140 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The communication bus may be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, only one thick line is used in the figure, but this does not mean that there is only one bus or only one type of bus. This embodiment of the present application does not impose any specific limitations on this.

[0319] The communication interface 1130 can be used for communication between the electronic device 1100 and other devices, including but not limited to: collaborative semantic codec model training maintenance personnel for semantic communication, management equipment, etc., which are not specifically limited in this embodiment of the application.

[0320] In an embodiment of the present application, the communication interface 1130 may be an interface circuit for direct digital communication between a computer system and another system. It typically includes a serial communication interface and a parallel communication interface. A serial communication interface may be, for example, an asynchronous transmission standard interface (EIA-RS-232, RS232), a universal serial bus (USB), or the like. A parallel communication interface may be, for example, a peripheral component interconnect express (PCI Express). This embodiment of the present application does not impose any specific limitations on this.

[0321] FIG12 is a schematic diagram of the structure of a communication device according to an embodiment of the present application. The dashed lines in FIG12 indicate that the unit or module is optional. Apparatus 1200 may be used to implement the method described in the above method embodiment. Apparatus 1200 may be a chip, a terminal device, or a network device.

[0322] The device 1200 may include one or more processors 1210. The processor 1210 may support the device 1200 to implement the method described in the above method embodiment. Like the processor 1110, the processor 1210 may also be a general-purpose processor or a special-purpose processor, which will not be described in detail. As described for the processor 1110. For example, the processor may be a central processing unit (CPU). Alternatively, the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0323] The apparatus 1200 may further include one or more memories 1220. The memories 1220 store programs that can be executed by the processor 1210, causing the processor 1210 to perform the methods described in the above method embodiments. The memories 1220 may be independent of the processor 1210 or integrated into the processor 1210.

[0324] The apparatus 1200 may further include a transceiver 1230. The processor 1210 may communicate with other devices or chips via the transceiver 1230. For example, the processor 1210 may transmit and receive data with other devices or chips via the transceiver 1230.

[0325] The present application also provides a computer-readable storage medium for storing a program. The computer-readable storage medium can be applied to a terminal device or network device provided in the present application, and the program enables a computer to execute the method performed by the terminal device or network device in each embodiment of the present application.

[0326] The computer-readable storage medium can be any available medium that can be read by a computer or a data storage device such as a server or data center that includes one or more available media. The available medium can be a magnetic medium, an optical medium, or a semiconductor medium. Examples of computer storage media include, but are not limited to: phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc-read only memory (CD-ROM), solid state disk (SSD), digital versatile disc (DVD) or other optical storage, magnetic cassette, tape / disk storage or other magnetic storage device or any other non-transmission medium. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0327] Computer-readable media includes both permanent and non-permanent, removable and non-removable media, and can be implemented using any method or technology for information storage. Computer-readable media can be used to store information that can be accessed by a computing device. This information can be computer-readable instructions, data structures, program modules, or other data.

[0328] The present application also provides a readable storage medium having a program or instruction stored thereon. When executed by a processor, the program or instruction can implement the various processes of the above-mentioned method embodiment or achieve the same technical effect. To avoid repetition, the program or instruction is not further described here. Optionally, when executed by the processor, the program or instruction can implement any step of the above-mentioned collaborative semantic codec model training mechanism process for semantic communication.

[0329] The embodiments of the present application also provide a computer program product. The computer program product includes a program. The computer program product can be applied to the terminal device or network device provided in the embodiments of the present application, and the program causes the computer to execute the method performed by the terminal or network device in each embodiment of the present application. Optionally, the computer program product includes instructions that, when executed on a computer, cause the computer to execute any step of the above-mentioned collaborative semantic codec model training mechanism process for semantic communication.

[0330] Each embodiment in this specification is described in a related manner. References to the same or similar parts between the various embodiments are sufficient. Each embodiment focuses on the differences from the other embodiments. In particular, since the apparatus embodiments, electronic device embodiments, computer-readable storage medium embodiments, and computer program product embodiments are generally similar to the method embodiments, their descriptions are relatively simple. For relevant parts, references to the descriptions of the method embodiments are sufficient.

[0331] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means.

[0332] The present application also provides a computer program that can be applied to a terminal device or network device provided in the present application, and enables a computer to execute the method performed by the terminal or network device in each embodiment of the present application.

[0333] In this application, the terms "system" and "network" can be used interchangeably. In addition, the terms used in this application are only used to explain the specific embodiments of this application, and are not intended to limit this application.

[0334] Relational terms such as "first," "second," and "third" in this application are merely used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations.

[0335] It should be understood that the terms "comprises," "comprising," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not preclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.

[0336] In the embodiments of this application, the term "indication" may refer to a direct indication, an indirect indication, or an indication of an association. For example, "A indicates B" may refer to a direct indication of B, e.g., B can obtain information through A; it may refer to an indirect indication of B, e.g., A indicates C, e.g., B can obtain information through C; or it may refer to an association between A and B.

[0337] In the embodiments of the present application, the term "corresponding" may indicate a direct or indirect correspondence between the two, or an association relationship between the two, or a relationship between indication and indication, configuration and configuration, etc.

[0338] In the embodiments of the present application, the "protocol" may refer to a standard protocol in the communication field, for example, it may include an LTE protocol, a NR protocol, and related protocols used in future communication systems, and this application does not limit this.

[0339] In the embodiments of the present application, determining B based on A does not mean determining B only based on A. B can also be determined based on A and / or other information.

[0340] In the embodiments of this application, the term "and / or" is simply a description of the association relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this document generally indicates that the related objects are in an "or" relationship.

[0341] In the embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0342] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0343] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0344] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0345] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above method embodiments can be implemented by means of software plus the necessary general hardware platform. Of course, hardware can also be used, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the existing technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, disk, CD), and includes a number of instructions for enabling a service classification device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0346] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any modifications, equivalent replacements, or improvements that can be easily conceived by any person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A model training method, characterized in that: include: The first terminal device receives the first model and the first information sent by the network device; The first terminal device selects a first training mode from multiple training modes according to the first information; The multiple training modes are used together to train the first model, and the first training mode is used by the first terminal device to perform uplink transmission and / or model training related to the first model.

2. The method according to claim 1, characterized in that The first information is used to determine whether the first training mode includes centralized training of part or all of the first models by the network device, and the method further includes: When the first training mode includes the centralized training, the first terminal device sends first data / first signal for the centralized training to the network device; When the first training mode does not include the centralized training, the first terminal device trains the first model.

3. The method according to claim 1 or 2, characterized in that The first information includes a first threshold related to the capability of the terminal device. When the capability of the first terminal device is lower than the first threshold, the first training mode includes centralized training of part or all of the first models by the network device.

4. The method according to any one of claims 1 to 3, characterized in that The multiple training modes correspond one-to-one to multiple terminal device groups, and the first information includes a grouping strategy for the multiple terminal device groups, where the grouping strategy is used for the first terminal device to determine the first training mode according to the terminal device group to which it belongs.

5. The method according to claim 4, characterized in that The first terminal device is one of a plurality of terminal devices for training the first model, and the grouping strategy is determined based on capabilities and / or channel coefficients of the plurality of terminal devices.

6. The method according to any one of claims 1 to 5, characterized in that The method further comprises: The first terminal device sends a capability parameter to the network device, where the capability parameter includes one or more of the following: The number of floating-point operations of the first terminal device; The local data sample volume of the first terminal device; the maximum processor frequency of the first terminal device; The maximum uplink transmission power of the first terminal device; and The maximum energy consumption budget of the first terminal device.

7. The method according to claim 6, characterized in that The method further comprises: The first terminal device receives a first instruction sent by the network device; The first instruction is used to trigger the first terminal device to send the capability parameter.

8. The method according to any one of claims 1 to 7, characterized in that The method further comprises: The first terminal device receives the second information sent by the network device; The second information is used by the first terminal device to determine a resource allocation strategy for executing the first training mode, and the second information is determined based on capability parameters of multiple terminal devices that train the first model.

9. The method according to claim 8, characterized in that The resource allocation strategy is used to determine the uplink transmission power, energy consumption budget and / or processor frequency of the first terminal device when executing the first training mode.

10. The method according to any one of claims 1 to 9, characterized in that The first model includes an encoder model and a decoder model, and the multiple training modes include: Distributed training in which a plurality of terminal devices respectively train the encoder model and the decoder model; The network device performs centralized training on the decoder model.

11. The method according to claim 10, characterized in that When the first training mode is the distributed training, the method further includes: The first terminal device trains the first model based on the local data sample to obtain a first local model gradient; The first local model gradient is aggregated with other local model gradients based on over-the-air calculation.

12. The method according to claim 10, characterized in that When the first training mode is the centralized training, the method further includes: The first terminal device inputs the local data sample into the encoder model to obtain encoded first data / first signal; The first terminal device sends the target data and the first data / first signal to the network device, and the target data and the first data / first signal are used by the network device to train the decoder model.

13. The method according to any one of claims 10 to 12, characterized in that The distributed training is used to determine the aggregated model gradient after aggregating multiple local model gradients in the current training cycle, and the centralized training is used to determine the centralized model gradient after training the decoder model in the current training cycle. The aggregated model gradient and the centralized model gradient are jointly used to determine the second model to be trained in the next training cycle.

14. The method according to claim 13, characterized in that The encoder model in the first model is φ t , the decoder model in the first model is θ t , the encoder model φ in the second model t+1 and decoder model θ t+1 Expressed as: Where γ represents the training learning rate of the first model in the current training cycle, D P represents the total number of data samples of all terminal devices that perform the distributed training, D W represents the total number of data samples of all terminal devices that perform the centralized training, D = D P +D w , represents the encoder model gradient in the aggregated model gradient, represents the decoder model gradient in the aggregated model gradient, represents the centralized model gradient.

15. The method according to any one of claims 1 to 14, characterized in that The first model is a codec model for semantic communication.

16. A model training method, characterized in that: include: The network device sends the first model and the first information to the first terminal device; Among them, the first information is used by the first terminal device to select a first training mode from multiple training modes, the multiple training modes are commonly used to train the first model, and the first training mode is used by the first terminal device to perform uplink transmission and / or model training related to the first model.

17. The method according to claim 16, characterized in that The first information is used to determine whether the first training mode includes centralized training of part or all of the first models by the network device, and the method further includes: When the first training mode includes the centralized training, the network device receives first data / first signal for model training sent by the first terminal device; When the first training mode does not include the centralized training, the network device receives a result of the first terminal device training the first model.

18. The method according to claim 16 or 17, characterized in that The first information includes a first threshold related to the capability of the terminal device. When the capability of the first terminal device is lower than the first threshold, the first training mode includes centralized training of part or all of the first models by the network device.

19. The method according to any one of claims 16 to 18, characterized in that The multiple training modes correspond one-to-one to multiple terminal device groups, and the first information includes a grouping strategy for the multiple terminal device groups, where the grouping strategy is used for the first terminal device to determine the first training mode according to the terminal device group to which it belongs.

20. The method according to claim 19, characterized in that The first terminal device is one of a plurality of terminal devices for training the first model, and the grouping strategy is determined based on capabilities and / or channel coefficients of the plurality of terminal devices.

21. The method according to any one of claims 16 to 20, characterized in that The method further comprises: The network device receives a capability parameter sent by the first terminal device, where the capability parameter includes one or more of the following: The number of floating-point operations of the first terminal device; The local data sample volume of the first terminal device; the maximum processor frequency of the first terminal device; The maximum uplink transmission power of the first terminal device; and The maximum energy consumption budget of the first terminal device.

22. The method according to claim 21, characterized in that The method further comprises: The network device sends a first instruction to the first terminal device; The first instruction is used to trigger the first terminal device to send the capability parameter.

23. The method according to any one of claims 16 to 22, characterized in that The method further comprises: The network device sends second information to the first terminal device; The second information is used by the first terminal device to determine a resource allocation strategy for executing the first training mode, and the second information is determined based on capability parameters of multiple terminal devices that train the first model.

24. The method according to claim 23, wherein The resource allocation strategy is used to determine the uplink transmission power, energy consumption budget and / or processor frequency of the first terminal device when executing the first training mode.

25. The method according to any one of claims 16 to 24, characterized in that The first model includes an encoder model and a decoder model, and the multiple training modes include: Distributed training in which a plurality of terminal devices respectively train the encoder model and the decoder model; The network device performs centralized training on the decoder model.

26. The method according to claim 25, characterized in that When the first training mode is the distributed training, the method further includes: The network device receives a first local model gradient sent by the first terminal device; The first local model gradient is aggregated with other local model gradients based on over-the-air calculation.

27. The method according to claim 25, characterized in that When the first training mode is the centralized training, the method further includes: The network device receives the target data and the encoded first data / first signal sent by the first terminal device; The network device trains the decoder model according to the target data and the first data / first signal.

28. The method according to any one of claims 25 to 27, characterized in that The method further comprises: The network device determines an aggregated model gradient after aggregating multiple local model gradients in a current training cycle; The network device determines a centralized model gradient after training the decoder model in a current training cycle; The network device determines a second model for a next training cycle based on the aggregated model gradient and the centralized model gradient.

29. The method according to claim 28, characterized in that The encoder model in the first model is φ t , the decoder model in the first model is θ t , the encoder model φ in the second model t+1 and decoder model θ t+1 Expressed as: Where γ represents the training learning rate of the first model in the current training cycle, D P represents the total number of data samples of all terminal devices that perform the distributed training, D W represents the total number of data samples of all terminal devices that perform the centralized training, D = D P +D W , represents the encoder model gradient in the aggregated model gradient, represents the decoder model gradient in the aggregated model gradient, represents the centralized model gradient.

30. The method according to any one of claims 16 to 29, wherein: The first model is a codec model for semantic communication.

31. A terminal device, characterized in that: The terminal device is a first terminal device, and the terminal device includes: A first receiving unit, configured to receive a first model and first information sent by a network device; a first processing unit, configured to select a first training mode from a plurality of training modes according to the first information; The multiple training modes are used together to train the first model, and the first training mode is used by the first terminal device to perform uplink transmission and / or model training related to the first model.

32. The terminal device according to claim 31, characterized in that The first information is used to determine whether the first training mode includes centralized training of part or all of the first models by the network device, and the terminal device further includes: A first sending unit, configured to send first data / first signal for model training to the network device when the first training mode includes the centralized training; The second processing unit is configured to train the first model when the first training mode does not include the centralized training.

33. The terminal device according to claim 31 or 32, characterized in that: The first information includes a first threshold related to the capability of the terminal device. When the capability of the first terminal device is lower than the first threshold, the first training mode includes centralized training of part or all of the first models by the network device.

34. The terminal device according to any one of claims 31 to 33, characterized in that: The multiple training modes correspond one-to-one to multiple terminal device groups, and the first information includes a grouping strategy for the multiple terminal device groups, where the grouping strategy is used for the first terminal device to determine the first training mode according to the terminal device group to which it belongs.

35. The terminal device according to claim 34, characterized in that The first terminal device is one of a plurality of terminal devices for training the first model, and the grouping strategy is determined based on capabilities and / or channel coefficients of the plurality of terminal devices.

36. The terminal device according to any one of claims 31 to 35, characterized in that: The terminal device further includes: The second sending unit is configured to send a capability parameter to the network device, where the capability parameter includes one or more of the following: The number of floating-point operations of the first terminal device; The local data sample volume of the first terminal device; the maximum processor frequency of the first terminal device; The maximum uplink transmission power of the first terminal device; and The maximum energy consumption budget of the first terminal device.

37. The terminal device according to claim 36, characterized in that The terminal device further includes: A second receiving unit, configured to receive a first instruction sent by the network device; The first instruction is used to trigger the first terminal device to send the capability parameter.

38. The terminal device according to any one of claims 31 to 37, characterized in that: The terminal device further includes: a third receiving unit, configured to receive second information sent by the network device; The second information is used by the first terminal device to determine a resource allocation strategy for executing the first training mode, and the second information is determined based on capability parameters of multiple terminal devices that train the first model.

39. The terminal device according to claim 38, characterized in that The resource allocation strategy is used to determine the uplink transmission power, energy consumption budget and / or processor frequency of the first terminal device when executing the first training mode.

40. The terminal device according to any one of claims 31 to 39, characterized in that: The first model includes an encoder model and a decoder model, and the multiple training modes include: Distributed training in which a plurality of terminal devices respectively train the encoder model and the decoder model; The network device performs centralized training on the decoder model.

41. The terminal device according to claim 40, characterized in that When the first training mode is the distributed training, the terminal device further includes: a third processing unit, configured to train the first model based on the local data sample to obtain a first local model gradient; The first local model gradient is aggregated with other local model gradients based on over-the-air calculation.

42. The terminal device according to claim 40, characterized in that When the first training mode is the centralized training, the terminal device further includes: a fourth processing unit, configured to input the local data sample into the encoder model to obtain encoded first data / first signal; The third sending unit is used to send the target data and the first data / first signal to the network device, where the target data and the first data / first signal are used by the network device to train the decoder model.

43. The terminal device according to any one of claims 40 to 42, characterized in that: The distributed training is used to determine the aggregated model gradient after aggregating multiple local model gradients in the current training cycle, and the centralized training is used to determine the centralized model gradient after training the decoder model in the current training cycle. The aggregated model gradient and the centralized model gradient are jointly used to determine the second model to be trained in the next training cycle.

44. The terminal device according to claim 43, characterized in that The encoder model in the first model is φ t , the decoder model in the first model is θ t , the encoder model φ in the second model t+1 and decoder model θ t+1 Expressed as: Where γ represents the training learning rate of the first model in the current training cycle, D P represents the total number of data samples of all terminal devices that perform the distributed training, D W represents the total number of data samples of all terminal devices that perform the centralized training, D = D P +D W , represents the encoder model gradient in the aggregated model gradient, represents the decoder model gradient in the aggregated model gradient, represents the centralized model gradient.

45. The terminal device according to any one of claims 31 to 44, characterized in that: The first model is a codec model for semantic communication.

46. ​​A network device, characterized in that include: A first sending unit, configured to send a first model and first information to a first terminal device; Among them, the first information is used by the first terminal device to select a first training mode from multiple training modes, the multiple training modes are commonly used to train the first model, and the first training mode is used by the first terminal device to perform uplink transmission and / or model training related to the first model.

47. The network device according to claim 46, wherein: The first information is used to determine whether the first training mode includes centralized training of part or all of the first models by the network device, and the network device further includes: A first receiving unit, configured to receive first data / first signal for model training sent by the first terminal device when the first training mode includes the centralized training; The second receiving unit is configured to receive a result of training the first model by the first terminal device when the first training mode does not include the centralized training.

48. The network device according to claim 46 or 47, characterized in that The first information includes a first threshold related to the capability of the terminal device. When the capability of the first terminal device is lower than the first threshold, the first training mode includes the network device training part or all of the first models.

49. The network device according to any one of claims 46 to 48, characterized in that: The multiple training modes correspond one-to-one to multiple terminal device groups, and the first information includes a grouping strategy for the multiple terminal device groups, where the grouping strategy is used for the first terminal device to determine the first training mode according to the terminal device group to which it belongs.

50. The network device according to claim 49, wherein: The first terminal device is one of a plurality of terminal devices for training the first model, and the grouping strategy is determined based on capabilities and / or channel coefficients of the plurality of terminal devices.

51. The network device according to any one of claims 46 to 50, characterized in that: The network device further includes: A third receiving unit is configured to receive a capability parameter sent by the first terminal device, where the capability parameter includes one or more of the following: The number of floating-point operations of the first terminal device; The local data sample volume of the first terminal device; the maximum processor frequency of the first terminal device; The maximum uplink transmission power of the first terminal device; and The maximum energy consumption budget of the first terminal device.

52. The network device according to claim 51, wherein: The network device further includes: A second sending unit, configured to send a first instruction to the first terminal device; The first instruction is used to trigger the first terminal device to send the capability parameter.

53. The network device according to any one of claims 46 to 52, characterized in that: The network device further includes: A third sending unit, configured to send second information to the first terminal device; The second information is used by the first terminal device to determine a resource allocation strategy for executing the first training mode, and the second information is determined based on capability parameters of multiple terminal devices that train the first model.

54. The network device according to claim 53, wherein: The resource allocation strategy is used to determine the uplink transmission power, energy consumption budget and / or processor frequency of the first terminal device when executing the first training mode.

55. The network device according to any one of claims 46 to 54, characterized in that: The first model includes an encoder model and a decoder model, and the multiple training modes include: Distributed training in which a plurality of terminal devices respectively train the encoder model and the decoder model; The network device performs centralized training on the decoder model.

56. The network device according to claim 55, characterized in that When the first training mode is the distributed training, the network device further includes: A fourth receiving unit, configured to receive the first local model gradient sent by the first terminal device; The first local model gradient is aggregated with other local model gradients based on over-the-air calculation.

57. The network device according to claim 55, characterized in that When the first training mode is the centralized training, the network device further includes: A fifth receiving unit, configured to receive the target data and the encoded first data / first signal sent by the first terminal device; A processing unit is used to train the decoder model based on the target data and the first data / first signal.

58. The network device according to any one of claims 55 to 57, characterized in that: The network device further includes: A first determining unit is used to determine an aggregated model gradient after aggregating multiple local model gradients in a current training cycle; a second determining unit, configured to determine a centralized model gradient after training the decoder model in a current training cycle; A third determining unit is configured to determine a second model for a next training cycle based on the aggregated model gradient and the centralized model gradient.

59. The network device according to claim 58, characterized in that The encoder model in the first model is φ t , the decoder model in the first model is θ t , the encoder model φ in the second model t+1 and decoder model θ t+1 Expressed as: Where γ represents the training learning rate of the first model in the current training cycle, D P represents the total number of data samples of all terminal devices that perform the distributed training, D W represents the total number of data samples of all terminal devices that perform the centralized training, D = D P +D W , represents the encoder model gradient in the aggregated model gradient, represents the decoder model gradient in the aggregated model gradient, represents the centralized model gradient.

60. The network device according to any one of claims 46 to 59, characterized in that: The first model is a codec model for semantic communication.

61. A communication device, characterized in that The system comprises a memory and a processor, wherein the memory is used to store a program, and the processor is used to call the program in the memory to execute the method according to any one of claims 1 to 30.

62. A device, characterized in that The device comprises a processor configured to call a program from a memory to execute the method according to any one of claims 1 to 30.

63. A chip, characterized in that: The device comprises a processor configured to call a program from a memory so that a device equipped with the chip executes the method according to any one of claims 1 to 30.

64. A computer-readable storage medium, characterized in that A program is stored thereon, the program causing a computer to execute the method according to any one of claims 1 to 30.

65. A computer program product, characterized in that The method comprises a program for causing a computer to execute the method according to any one of claims 1 to 30.

66. A computer program, characterized in that The computer program causes a computer to execute the method according to any one of claims 1 to 30.

Citation Information

Patent Citations

  • Wireless federal learning method for large-scale Internet of Things cooperative intelligence

    CN116306915A

  • Federal learning equipment scheduling method based on model loss tolerance

    CN117172333A

  • Machine learning model training method, terminal device and network device

    CN117546180A

  • Resource scheduling method and apparatus, and readable storage medium

    WO2021142627A1

  • Model training method and communication apparatus

    WO2024017001A1

Cited By

  • Model scheduling method and system for end-cloud collaboration

    CN122332076A

  • Apparatus and method for dual-path-based on-board signal processing for artificial semantic communication networks

    KR102998509B1