Semantic communication-based model training method, electronic equipment and storage medium
By using signal compensation and mixed signal processing in multi-user semantic communication, the balance problem between communication efficiency and data security is solved, privacy protection is improved, communication overhead is reduced, and the accuracy and security of the model are ensured.
Patent Information
- Application Number
- CN202510651425.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-10-17
AI Technical Summary
It is difficult to effectively balance communication efficiency and data security in multi-user semantic communication. The existing technology suffers from a significant drop in model accuracy when the channel and user data change, and its privacy protection capabilities are insufficient.
By determining the scaling coefficients of the server and user sides, performing signal compensation and mixed signal processing, and utilizing the physical properties of signal superposition to achieve privacy protection, and transmitting semantic information in the wireless channel, the server side generates encoder gradients for the user side to update the encoder.
While maintaining model accuracy, it significantly improves privacy protection and reduces communication overhead, achieving efficient and secure multi-user semantic communication.
Smart Images

Figure CN120805983A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of distributed learning, and in particular to a model training method based on semantic communication, an electronic device and a storage medium. BACKGROUND
[0002] In the field of multi-user semantic communication, semantic communication encodes and decodes information at the semantic level, has become an efficient and context-aware data transmission paradigm, and can provide enhanced spectral efficiency and adaptability to dynamic environments for the sixth generation (6G) network. Semantic communication is a new communication paradigm that focuses on information meaning. Its goal is not to simply ensure "lossless transmission" of data, but to efficiently complete tasks by transmitting "understandable semantic information".
[0003] In related technologies, whether transmitting raw data or transmitting inference results of raw data, when the channel changes and the user data changes, the model accuracy will decrease significantly, which seriously affects the performance of multi-user semantic communication. On the other hand, privacy issues are also very prominent. The exposed semantic features are used to reconstruct private data, which destroys the trust in decentralized training. Therefore, the multi-user semantic communication model training framework provided in related technologies has the technical problem of being difficult to effectively balance communication efficiency and data security. SUMMARY
[0004] The embodiments of the present application provide a model training method based on semantic communication, an electronic device and a storage medium to alleviate or solve the technical problem of being difficult to effectively balance communication efficiency and data security in related technologies.
[0005] In a first aspect, the embodiments of the present application provide a model training method based on semantic communication, comprising: determining a first scaling coefficient for signal compensation of a server and a second scaling coefficient for signal compensation of a plurality of same-group user terminals based on channel gains respectively corresponding to the plurality of same-group user terminals; the plurality of same-group user terminals are associated with the server and belong to a target group; respectively sending the corresponding second scaling coefficients to the plurality of same-group user terminals; receiving a mixed signal sent by the plurality of same-group user terminals through a target channel using the first scaling coefficient; the mixed signal is obtained by mixing semantic information respectively corresponding to the plurality of same-group user terminals, and the semantic information is obtained by the corresponding same-group user terminal using an encoder based on local training data and sent using the second scaling coefficient; the target channel corresponds to the target group; processing the mixed signal to obtain encoder gradients respectively corresponding to the plurality of same-group user terminals; The plurality of encoder gradients are respectively sent to corresponding same-group user terminals; and the plurality of encoder gradients are used for the plurality of same-group user terminals to respectively update corresponding encoders.
[0006] In a second aspect, an electronic device is provided, which includes a memory, a processor, and a computer program stored in the memory. The processor implements the method of any of the embodiments of the present application when executing the computer program.
[0007] In a third aspect, a computer-readable storage medium is provided, which stores a computer program. The computer program is executed by a processor to implement the method of any of the embodiments of the present application.
[0008] Based on the model training method for semantic communication based on the above first aspect, the present application has at least the following beneficial effects or advantages: based on the dynamic channel state of the same group of users, the first scaling coefficient of the server and the second scaling coefficient of the user terminal are calculated, the precise compensation of channel attenuation and noise interference is realized, and the stable signal transmission quality is maintained under the power constraint. Through the natural mixed transmission of the semantic information of the in-group users in the wireless channel, the physical characteristics of signal superposition are utilized to realize privacy protection, so that it is difficult for an attacker to reverse the original data of a single user from the mixed signal. The server generates the encoder gradient by processing the mixed signal, which not only retains the key information of the semantic features, but also avoids the security risks brought by directly transmitting the original data or the complete model. The present application provides an efficient and secure multi-user semantic communication model training framework, which aims to solve the problems of large communication overhead and poor privacy protection capability in the multi-user semantic communication model training framework. While maintaining the similar accuracy as the multi-user semantic communication model training framework based on federated learning, the technical effect of significantly improving the privacy protection level and reducing the communication overhead of training is realized.
[0009] The above description is only a summary of the technical solutions of the present application. In order to more clearly understand the technical means of the present application, the description can be implemented in accordance with the content of the specification, and in order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the following specific embodiments of the present application are described. BRIEF DESCRIPTION OF DRAWINGS
[0010] In the drawings, the same reference numbers in the several drawings represent the same or similar elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings only depict some embodiments according to the present application and should not be considered as limiting the scope of the present application.
[0011] Figure 1 A first flowchart of the model training method for semantic communication based on the embodiments of the present application is shown; Figure 2A second flowchart of the model training method based on semantic communication is shown. Figure 3 A third flowchart of the model training method based on semantic communication is shown. Figure 4 A flow interaction diagram of the model training method based on semantic communication is shown. Figure 5 A first effect schematic diagram of the model training method based on semantic communication is shown. Figure 6 A second effect schematic diagram of the model training method based on semantic communication is shown. Figure 7 A third effect schematic diagram of the model training method based on semantic communication is shown. Figure 8 A first schematic diagram of the model training apparatus based on semantic communication is shown. Figure 9 A second schematic diagram of the model training apparatus based on semantic communication is shown. Figure 10 A third schematic diagram of the model training apparatus based on semantic communication is shown. Figure 11 A block diagram of the electronic device provided by the embodiment of the present application is shown. DETAILED DESCRIPTION
[0012] In the following, only certain exemplary embodiments are simply described. As those skilled in the art can recognize, the described embodiments can be modified in various different ways without departing from the concept or scope of the present application. Therefore, the drawings and the description are considered to be exemplary in nature, rather than limiting.
[0013] In order to facilitate understanding of the technical solutions of the embodiments of the present application, the related technologies of the embodiments of the present application are described below. The following related technologies can be combined with the technical solutions of the embodiments of the present application in any way as optional solutions, which all belong to the protection scope of the embodiments of the present application.
[0014] In the following, the following terms will be used: Semantic communication is a new communication paradigm with information meaning as the core. Its goal is no longer to simply ensure the "lossless transmission" of data, but to efficiently complete the task by delivering "understandable semantic information". Semantic communication extracts the semantic information contained in the original information, and then transmits it to the receiver through the wireless channel. The receiver restores the information sent by the sender.
[0015] Model inversion attack, inversion attack is a common attack method in the field of AI, which aims to restore the original data through the semantic information output by the front-end model. If the attack is successful, the user's privacy data will be leaked, which will pose a serious threat to the security of the entire system.
[0016] It should be noted that the above application scenarios or application examples provided in the embodiments of the present application are for the purpose of understanding, and the application of the technical solutions by the embodiments of the present application is not specifically limited. In addition, the user information (including but not limited to user equipment information, user personal information, etc., training information, channel information corresponding to multiple user terminals respectively) and data (including but not limited to data for analysis, stored data, displayed data, etc., data used by user terminals and servers respectively) involved in the present application are all information and data authorized by users or fully authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose authorization or refusal.
[0017] Currently, there are mainly two forms of semantic communication: one is to transmit raw data (common semantic communication paradigm), and the other is to transmit the inference result of raw data (task-oriented semantic communication). The model training of semantic communication follows the traditional model training of deep neural network, which generally distributes the trained semantic decoder and encoder to the receiving parties after local training. Under the conditions of channel change and user data change, the model accuracy will decrease significantly.
[0018] In a related technology, the model training of semantic communication is completed by differential privacy. The differential privacy technology adds a noise to the semantic information to resist the attack of inversion based on the multi-user semantic communication model of federated learning. To some extent, it can reduce the success rate of the attack of inversion, but it will affect the communication efficiency and accuracy. Another related technology provides an unlearning scheme, which is a semantic communication model training method for adapting to user data changes. The core idea is to minimize the mutual information between the learned semantic representation and the deleted samples, and customize the joint learning method for semantic encoder and decoder. In a multi-user scenario, the model may need to be frequently updated, which will increase the communication overhead. Especially in the case of frequent changes in large-scale user data, the communication cost of model updating will increase significantly. In another related technology, an information bottleneck and instance-based adversarial learning (IBAL) privacy protection method is used to resist the threat of inversion attack. Based on the information bottleneck theory, task-related features are extracted from the input, and IBAL is used to protect the privacy of users. The IBAL method needs a large amount of calculation and communication resources in the process of extracting features and adversarial training. Especially in a multi-user scenario, this overhead will further increase.
[0019] The technical solutions of the present application and how the technical solutions of the present application solve the foregoing technical problems will be described in detail below with specific embodiments. Several specific embodiments listed can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0020] Figure 1 A flowchart of a model training method based on semantic communication according to an embodiment of the present application is shown, as shown in Figure 1 The method can include steps S101 to S105.
[0021] Step S101: determining a first scaling coefficient for signal compensation of the server and a second scaling coefficient for signal compensation of the plurality of same-group user terminals based on the channel gains corresponding to the plurality of same-group user terminals, respectively; the plurality of same-group user terminals are associated with the server and belong to a target group; Step S102: sending the corresponding second scaling coefficient to the plurality of same-group user terminals, respectively; Step S103: receiving a mixed signal sent by the plurality of same-group user terminals through a target channel by using a first scaling coefficient; the mixed signal is obtained by mixing semantic information corresponding to the plurality of same-group user terminals, and the semantic information is obtained by the corresponding same-group user terminal based on local training data by using an encoder and sent by using a second scaling coefficient, and the target channel corresponds to a target group; Step S104: processing based on the mixed signal to obtain an encoder gradient corresponding to each of the plurality of same-group user terminals; Step S105: sending the plurality of encoder gradients to the corresponding same-group user terminals respectively; the plurality of encoder gradients are used for the plurality of same-group user terminals to update the corresponding encoders respectively.
[0022] In the embodiments provided in the present application, the execution subject is a server, and the server can be disposed in a base station (Base Station) or a cloud server for processing.
[0023] In the embodiments provided in the present application, in a multi-user semantic communication system, the server calculates two scaling coefficients according to the channel gain of each of the plurality of same-group user terminals associated with the server and belonging to a target group: the first scaling coefficient is used for compensating the signal received by the server itself, and the second scaling coefficient corresponds to each same-group user terminal and is used for compensating the signal sent by the same-group user terminal. Based on the difference of the channel conditions of different same-group user terminals, a communication foundation is laid for reducing the loss of subsequent signal transmission. The plurality of same-group user terminals are user terminals belonging to the same group, i.e., target user terminals, and the user terminals in the same group share the same wireless channel, i.e., a target channel.
[0024] The server sends the second scaling coefficient of each same-group user terminal to the same-group user terminal respectively, and each same-group user terminal can compensate the signal sending process of itself according to the second scaling coefficient of itself after receiving the second scaling coefficient. For each same-group user terminal, the semantic information is obtained by the same-group user terminal by using the local training data through the encoder, and is sent out through the target channel after being adjusted by the second scaling coefficient. When the semantic information of each same-group user terminal is transmitted and converged in the target channel, a mixed signal is formed. The mixed signal carries the semantic information of the plurality of same-group user terminals, and is also transmitted by signal compensation, which not only ensures the effective transmission of information, but also enhances the privacy protection to a certain extent.
[0025] The server receives the mixed signal by using the first scaling coefficient. The first scaling coefficient enables the server to more accurately receive and analyze the signal content. The server can obtain an encoder gradient corresponding to each same-group user terminal based on the mixed signal. The encoder gradient is crucial information in the model training process, and reflects the update direction and degree of the encoder of each user terminal in the current training stage.
[0026] The server sends the encoder gradient back to the corresponding user terminal in the same group, and the user terminal in the same group updates the local encoder model according to the encoder gradient. In this way, the encoder of each user terminal can be optimized according to the latest training information, thereby continuously improving the performance and accuracy of the model.
[0027] It should be noted that in the model training based on semantic communication, the encoder, the decoder and the model interact with each other. The encoder is set in the user terminal, and its main function is to convert the original data (such as images, texts, voices, etc.) of the user terminal into semantic information. The semantic information is the data representation after encoding, which can capture the core semantic features of the original data. The encoder is locally trained in the user terminal to adapt to the data distribution and task requirements of the user terminal. The decoder is set in the server, and the decoder restores the semantic information after encoding to the semantic content of the original data, so that the server can understand the information sent by the user terminal. The semantic information is reconstructed into a format similar to the original data for further processing and analysis. The model in the semantic communication includes two parts of the encoder and the decoder, and the distributed training is performed by using multiple user terminals. The encoder is locally trained in the user terminal to adapt to the data distribution of the user terminal, and the decoder is globally optimized in the server to adapt to the semantic information of multiple user terminals. Through the separation of the encoder and the decoder, the model can effectively protect the privacy of the user data. The user terminal only sends semantic information, rather than the original data, thereby reducing the risk of privacy leakage.
[0028] According to the embodiments provided in the present application, before performing step S101: determining the first scaling coefficient of the server and the second scaling coefficient corresponding to the plurality of user terminals in the same group based on the channel gain corresponding to the plurality of user terminals in the same group, the method further includes a preprocessing stage, and the specific steps are as follows: Group all user terminals according to the label distribution and channel gain corresponding to all user terminals, and determine the grouping result; the grouping result indicates that the user terminals in the same group belong to a target group.
[0029] In the embodiments of the present application, before performing the interaction between the user terminal and the server, all user terminals need to be grouped. The user terminals in the same group use the same wireless channel to communicate with the server. The label distribution of the user terminal shows the proportion of different categories in the local data of each user terminal, and the channel gain represents the channel quality, which is used to reflect that these user terminals can use the same channel for signal mixing, reduce the problem of signal flooding or noise amplification caused by channel difference, and improve the accuracy of semantic recovery.
[0030] Exemplarily, all the above service ends can include multiple different groups, the service end can use multiple different wireless channels, and the service end assigns an appropriate wireless channel to each group according to the grouping result.
[0031] According to the embodiments provided in the present application, the grouping of all user ends according to the label distribution and the channel gain corresponding to each user end respectively, and the determination of the grouping result can include the following specific steps: determining label difference information between any two user ends included in all user ends based on the label distribution corresponding to each user end respectively; determining channel compatibility information between any two user ends based on the channel gain corresponding to each user end respectively; grouping all user ends in a manner that the label difference information meets a predetermined difference condition and the channel compatibility information meets a predetermined compatibility condition, to obtain the grouping result.
[0032] In the embodiments provided in the present application, the service end calculates the difference information between all user ends based on the label distribution, and evaluates the channel compatibility information between any two user ends, considering the similarity or compatibility of the channel gain. After obtaining the label difference information and the channel compatibility information, the service end groups the user ends according to the preset difference condition and the compatibility condition. Specifically, only when the label difference of two user ends meets the predetermined difference condition and their channel compatibility also meets the predetermined compatibility condition, the two user ends are divided into the same group. The user ends with large label difference are divided into the same group to ensure that it is difficult to reverse the original data of a single user after semantic mixing in the group, and the user ends with small channel difference are divided into the same group to avoid signal mixing distortion caused by channel difference. By considering the label distribution and the channel gain in two dimensions for grouping, it can be ensured that the user ends in the same group not only have difference in data features, but also have good compatibility in communication conditions.
[0033] Exemplarily, there can be multiple user end grouping methods, such as a greedy clustering algorithm, a dynamic grouping algorithm based on reinforcement learning, a divergence and channel gain algorithm, and a Max-Clique algorithm based on a graph.
[0034] The greedy clustering algorithm randomly selects a seed user end, and divides it into the same group with all user ends meeting the double conditions, preferentially adds user ends with the largest label difference and the best channel compatibility, and repeats the above process for the remaining user ends until there is no user end meeting the conditions.
[0035] The dynamic grouping algorithm based on reinforcement learning encodes the user label distribution and channel state into a vector as the state definition, selects the user pairing or grouping strategy through the action space, and designs the label heterogeneity gain (label entropy) and communication efficiency (channel utilization) as the reward function to obtain the grouping result.
[0036] The divergence and channel gain algorithm can be exemplified as follows. Assuming that four user terminals u1, u2, u3, and u4 correspond to the following label parts: u1: cat (80%), dog (20%) u2: dog (90%), bird (10%) u3: cat (70%), bird (30%) u4: cat (60%), dog (40%) The Jensen-Shannon divergence (JS divergence) can be used to measure the difference between two label distributions to calculate the label distribution difference between different user terminals.
[0037]
[0038]
[0039]
[0040]
[0041]
[0042]
[0043]
[0044]
[0045]
[0046]
[0047]
[0048] The channel compatibility information between any two user terminals is calculated. The channel compatibility information can be measured by the ratio of channel gains. The closer the value is to 1, the better the channel compatibility. The channel gains corresponding to u1, u2, u3, and u4 are 1, 1.1, 0.9, and 1.2, respectively. The channel compatibility is calculated as follows:
[0049]
[0050]
[0051]
[0052]
[0053]
[0054] Setting predetermined difference condition , predetermined compatibility condition When the label difference of two user terminals is less than and the channel compatibility is within , the two user terminals are divided into the same group.
[0055] The graph-based maximum clique algorithm takes each user terminal as a vertex, and if two user terminals satisfy the above two conditions at the same time, an edge is established. The Bron-Kerbosch algorithm (used to enumerate all maximum cliques in an undirected graph, that is, a complete subgraph that cannot be expanded) is used to find the maximum complete subgraph (that is, a subgraph in which all vertices are connected to each other) in the graph, each maximum clique is a group, and the grouped users are removed, and the search is repeated until all users are assigned. The server represents the label difference and channel compatibility information between the user terminals as the edge weight of a graph, where each user terminal is a node, and the edge weight reflects the heterogeneity and compatibility between the two user terminals. By finding the maximum clique in the graph, that is, finding a group of nodes such that the edge weight between these nodes satisfies the predetermined condition, the server can effectively divide the user terminals into multiple groups.
[0056] According to the embodiments provided in the present application, in step S101: based on the channel gains corresponding to the plurality of user terminals in the same group, determining the first scaling coefficient corresponding to the server and the second scaling coefficient corresponding to the plurality of user terminals in the same group, including the following specific steps: Determine the first candidate coefficient corresponding to the server as the first scaling coefficient and the second candidate coefficient corresponding to the plurality of user terminals in the same group as the corresponding second scaling coefficient in a manner that minimizes the deviation between the predetermined ideal signal and the expected signal; Wherein, the expected signal is determined based on the predetermined noise, the second candidate coefficient, and the channel gains corresponding to the plurality of user terminals in the same group and the first candidate coefficient, and the first candidate coefficient corresponding to the plurality of user terminals in the same group is subject to the power constraint of the corresponding user terminal in the same group.
[0057] In the embodiments provided in the present application, the ideal signal can be represented as a signal mixed by the same group of user terminals according to a predetermined mixing ratio. The expected signal is obtained by modeling, taking into account the channel gain corresponding to each of the plurality of same group user terminals and the pre-calculated noise. The first candidate coefficient and the second candidate coefficient corresponding to each of the plurality of same group user terminals are constructed so as to minimize the deviation between the ideal signal and the constructed expected signal, and the first scaling coefficient and the second scaling coefficient are obtained as the solution.
[0058] Exemplarily, the semantic information extracted by the user terminals in one group is mixed through a wireless channel. It is assumed that K is the identification of the target group, k is the identification of the same group user terminal in the target group, and the semantic information of the same group user terminal k is modulated into a preprocessed signal , and the ideal signal obtained according to a predetermined mixing ratio in the target channel is defined as:
[0059] wherein, is the predetermined mixing ratio.
[0060] Since the transmission of semantic information is interfered by path loss and noise in wireless communication, the sending end (i.e., the user terminal) and the receiving end (i.e., the server) can only dynamically adjust the signal by transmitting end scaling b (i.e., the second candidate coefficient) and receiving end scaling a (i.e., the first candidate coefficient) to complete accurate transmission. The user terminal uses power limit P, that is, the second scaling coefficient must be less than the square root of P. For the expected signal, the signal actually received by the server is predicted under different candidate scaling coefficients:
[0061] wherein, is the channel gain of the same group user terminal k, and is related to the path loss, is the second candidate coefficient of the same group user terminal k, a is the first candidate coefficient of the server base station, and n is the noise power.
[0062] The ideal signal and the expected signal have an error (mainly from the noise n). The difference between the over-the-air mixing calculation result (the expected signal) and the ideal mixing result (the ideal signal) is measured by using the MSE expectation, MSE represents the mean squared error (Mean Squared Error), and MSE is an index representing the difference between the ideal signal and the expected signal, which can be defined as:
[0063]
[0064] pre-processed signal , denotes the signal amplitude restriction, and the uniform value is 0, and the variance is 1, and the variance (power) of the noise n is denoted as , E denotes the expected value. The MSE equation can be obtained by substituting:
[0065] Considering the maximum power constraint condition of each user terminal transmission, the second candidate coefficient of the same group user terminal k is related to its maximum power, specifically , and the expression of the MSE optimization is:
[0066]
[0067] In the optimization problem, the goal is to minimize the deviation, that is, the value of the first candidate coefficient a under the condition of minimizing , to obtain the first scaling coefficient, and the value of the second candidate coefficient under the condition of the same, to obtain the second scaling coefficient. The channel state is updated before each round of training, and and are recalculated to adapt to the time-varying environment.
[0068] According to the embodiments provided in the present application, in step S104: based on the mixed signal, the encoder gradient corresponding to each user terminal in the same group is obtained, including the following specific steps: using the decoder corresponding to the target group to decode the mixed signal to obtain the recovered information; based on the recovered information and the pre-stored real information to calculate the loss and determine the group gradient of the multiple user terminals in the same group; deriving the group gradient to obtain the encoder gradient corresponding to each user terminal in the same group.
[0069] In the embodiments provided in the present application, the server is provided with a corresponding decoder according to different groups, and the decoder matched with the target group is used to decode the received mixed signal, which is formed by the same group of user terminals through wireless channel in the air. The decoder generates high-quality recovered information by analyzing the channel characteristics and semantic features The server compares the recovered information with the pre-stored real information to calculate the group-level loss function. Based on this, it derives a set of gradients that reflects the common semantic error for the entire group of users. Using the chain rule, the server decomposes the encoder gradients specific to each user in the group from the set of gradients. These encoder gradients are then securely transmitted back to the corresponding device for local encoder updates on each client in the group.
[0070] For example, upon receiving the actual mixed signal After the predicted signal with the first and second scaling factors is substituted, the server (i.e., the BS) performs forward propagation using the decoder corresponding to the target group to recover the information, i.e.:
[0071] in It represents the decoder of the target group on the server side in the tth round of training.
[0072] The server performs backpropagation based on the recovered signal. The base station calculates the loss function and its gradient with respect to the decoder parameters and updates the decoder using Adam. Adam (Adaptive Moment Estimation) is a gradient descent algorithm that combines momentum and an adaptive learning rate. In round t of communication, the decoder for the target group is updated as follows:
[0073] Where η is the learning rate, is the decoder gradient of target group K during round t of communication (i.e., the gradient of the entire group).
[0074] By chain-deriving the entire set of gradients, we can determine the encoder gradients corresponding to multiple user terminals in the same group included in the target group. The server sends the corresponding encoder gradient To the kth user in the same group.
[0075] User k performs back propagation, and each user uses the received encoder gradient Update its local semantic encoder:
[0076] Where η represents the learning rate, is the encoder gradient of the kth client in the tth round of communication.
[0077] In the model training based on semantic communication, a periodic processing method is adopted, that is, after the encoder of the user end is updated multiple times, the encoder parameters sent by all the user ends are aggregated to obtain the global parameters. According to the embodiment provided by the present application, the method also includes: In response to receiving the encoder parameters respectively sent by all user terminals associated with the service end, the encoder parameters respectively sent by all user terminals are aggregated to obtain a global parameter; The global parameter is respectively sent to all user terminals; the global parameter is used by all user terminals to update the corresponding encoder.
[0078] In the embodiments provided in the present application, there are multiple associated user terminals in the service end, and these user terminals can belong to different groups. After each user terminal completes training and multiple updates of the encoder locally, the updated encoder parameters are sent to the service end at a predetermined training round, which can reflect the semantic features and model state of the user terminal. After receiving the encoder parameters sent by all user terminals, the service end fuses the encoder parameters of all user terminals into a global parameter according to a certain aggregation strategy. The global parameter integrates the optimization results of all user terminals and can better reflect the global semantic features. The service end sends the aggregated global parameter back to all user terminals respectively. After receiving the global parameter, each user terminal updates the local encoder using the global parameter, so as to better adapt to the global semantic features and model state, thereby improving the performance and generalization ability of the model. Multiple local updates of the user terminal enable the encoder to better adapt to changes in local data, thereby better reflecting the data features in the dynamic environment during global aggregation.
[0079] Through this periodic processing method, the user terminal sends the encoder parameters only after multiple local updates, reducing the communication frequency with the service end and thereby reducing the communication overhead. The parameters transmitted in each communication are the results of multiple optimizations, which are more stable and accurate, further improving the communication efficiency. The user terminal only sends the encoder parameters instead of the original data, reducing the risk of privacy data exposure.
[0080] Exemplarily, the server can use various methods to perform aggregation, such as FedAvg (federated averaging) technology, GeFL (Generative Model-based Federated Learning), and dynamic weighted aggregation. The GeFL method is that the server collects and aggregates the encoder parameters to form a generative model with global knowledge, which is shared as a global parameter to all clients. The dynamic weighted aggregation method is to perform weighted aggregation on the encoder parameters of different user terminals according to different weight types (such as data quantity weighting, data quality weighting, and time weighting). The FedAvg method is that the service end receives the encoder parameters of all user terminals, performs weighted averaging on all encoder parameters according to a predetermined weight setting to obtain a global parameter. The following will illustrate the FedAvg method.
[0081] The server (denoted as BS) performs model aggregation to obtain a global semantic encoder θ and a global semantic decoder Φ. All user terminals upload their respective local semantic encoders to the BS, where t denotes the current training round, c denotes the identity of the different user terminals among all user terminals, C denotes the number of all user terminals, and the BS aggregates the model parameters according to the FedAvg technique as follows:
[0082] where I is the local training communication round, N is the total number of data sets, the number of data sets of the user terminal c, denotes a set of positive integers, denotes a predetermined periodic interval. is a conditional expression, when the current training round reaches , the server uses the same method to aggregate the semantic decoder, and the BS multicasts the aggregated semantic encoder to the user terminals for the next round of model training.
[0083] Exemplarily, after reaching the predetermined training round, the server will also perform model aggregation on the decoders corresponding to different groups respectively to obtain global decoder parameters. The global decoder parameters are used as the initial decoder corresponding to each group in the next round of training.
[0084] Figure 2 A second flowchart of the model training method based on semantic communication according to an embodiment of the present application is shown in FIG. 2B. As shown in FIG. 2B, the method can include steps S201 to S204. Figure 2
[0085] Step S201: receiving a second scaling factor for signal compensation of a target user terminal; the target user terminal is any one of a plurality of user terminals in the same group, and the plurality of user terminals in the same group are associated with the server and belong to a target group; Step S202: using the encoder of the target user terminal to process based on the local training data to obtain semantic information; Step S203: using the second scaling factor to send the semantic information through a target channel; the target channel corresponds to the target group, and the target channel is used to mix and transmit the semantic information corresponding to the plurality of user terminals in the same group to obtain mixed information; the mixed information is used for the server to receive and process using a first scaling factor to obtain the encoder gradient corresponding to the plurality of user terminals in the same group; the first scaling factor and the second scaling factor are determined by the server based on the channel gain corresponding to the plurality of user terminals in the same group; Step S204: in response to receiving the encoder gradient corresponding to the target user terminal, updating the encoder using the encoder gradient.
[0086] Exemplarily, the execution subject is a target user terminal, which can be a mobile terminal device such as a mobile phone, or a computer terminal, etc.
[0087] In the embodiments provided in the present application, the target user terminal (as one of a plurality of user terminals in the same group) receives a second scaling coefficient from the server for compensating the signal of the target user terminal. The plurality of user terminals in the same group are associated with the server and belong to the same target group and share the same target channel. The server determines the first scaling coefficient and the second scaling coefficient according to the channel gain of each user terminal to optimize the signal transmission process.
[0088] The target user terminal uses the local training data to process through the encoder thereof to extract semantic information. The target user terminal uses the received second scaling coefficient to send the semantic information through the target channel. The target channel is responsible for mixing and transmitting the semantic information of all user terminals in the same group to form mixed information. The server uses the first scaling coefficient to receive and process the mixed information to obtain the encoder gradient corresponding to each user terminal in the same group.
[0089] When the target user terminal receives the encoder gradient corresponding thereto, the gradient information is used to update the local encoder. In this way, the target user terminal can continuously optimize the encoder locally to better adapt to the requirements of global model training.
[0090] According to the embodiments provided in the present application, in order to perform periodic aggregation, the method further includes the following steps: When the target user terminal reaches a predetermined training round, the encoder parameter is sent to the server; In response to receiving the global parameter, the encoder is updated using the global parameter; the global parameter is obtained by aggregating the encoder parameters respectively sent by the server based on all associated user terminals.
[0091] In the embodiments provided in the present application, when the target user terminal reaches a predetermined training round, the encoder parameter obtained by local training is sent to the server. After receiving the encoder parameters respectively corresponding to all associated user terminals, the server performs aggregation processing of the encoder parameters using a predetermined aggregation algorithm to obtain a global parameter. The target user terminal receives the global parameter sent back by the server and updates the local encoder using the global parameter. This updating process enables the target user terminal to absorb the optimization results from other user terminals, thereby improving the performance of the target user terminal in the global model. In this way, the target user terminal can not only use local data for training, but also learn more rich semantic information from global data, thereby improving the generalization ability and accuracy of the model.
[0092] Figure 3 A third flowchart of a model training method based on semantic communication of an embodiment of the application is shown, as Figure 3 The method can include steps S301 to S309, as shown.
[0093] Step S301: The server determines a first scaling coefficient for signal compensation of the server and a second scaling coefficient for signal compensation of a plurality of user terminals in the same group based on the channel gains corresponding to the plurality of user terminals in the same group, respectively; the plurality of user terminals in the same group are associated with the server and belong to a target group; Step S302: The server sends the corresponding second scaling coefficients to the plurality of user terminals in the same group, respectively; Step S303: A target user terminal in the plurality of user terminals in the same group receives the corresponding second scaling coefficient, which is used for signal compensation of the target user terminal; Step S304: The target user terminal processes the local training data based on an encoder to obtain semantic information; Step S305: The target user terminal sends the semantic information through a target channel using the second scaling coefficient; Step S306: The server receives mixed signals sent by the plurality of user terminals in the same group through the target channel using the first scaling coefficient; Step S307: The server processes the mixed signals to obtain a plurality of encoder gradients corresponding to the plurality of user terminals in the same group, respectively; Step S308: The server sends the plurality of encoder gradients to the corresponding user terminals in the same group, respectively; Step S309: The target user terminal updates the encoder using the corresponding encoder gradient in response to receiving the corresponding encoder gradient.
[0094] The execution subject of the above embodiment is the server and the target user terminal.
[0095] According to the above embodiment, the application further provides an optional implementation, Figure 4 A flowchart of a model training method based on semantic communication of an embodiment of the application is shown, as Figure 4 As shown, an efficient and secure multi-user semantic communication model training framework is designed based on over-the-air mixing, aiming to solve the problems of large communication overhead and poor privacy protection capability in the multi-user semantic communication model training framework. The semantic information to be transmitted is mixed through over-the-air computing, and then transmitted to the base station end for model training. While maintaining a similar accuracy as the multi-user semantic communication model training framework based on federated learning, the privacy protection level is significantly improved and the communication overhead of training is reduced.
[0096] The multi-user semantic communication model training framework (SIMix) proposed by the embodiments of the present application is as shown in the following figure Figure 4 The framework follows the basic steps of semantic information extraction, semantic transmission, forward propagation, server-side back propagation, user-side back propagation, and model aggregation. Specifically, the process is as follows. The following training process description is mainly for the user side of a group. Training of all group users is completed according to the same method, which is regarded as completing the training of one communication round.
[0097] The preprocessing stage performs user grouping, and the mobile users are divided into Y groups. The user terminals in the same group can mix their semantic information through wireless transmission. The goal of the user grouping result is to improve security and accuracy: (1) Label diversity. Prioritize users with different label distributions to reduce the success rate of inversion attacks. (2) Channel compatibility. User terminals are paired to meet power constraints and minimize computational distortion to improve accuracy. Based on the above indicators, the graph-based maximum clique algorithm is preferably used for user grouping.
[0098] Step 1 performs user terminal semantic information extraction. Each user terminal k loads its data, performs forward propagation, and extracts semantic information from the original data x .
[0099]
[0100] wherein is the semantic encoder of user terminal k in the tth round of training.
[0101] Step 2: The user terminal sends the semantic information through the target channel according to the second scaling coefficient informed by the server. After the air mixing of the semantic information corresponding to multiple user terminals in the same group, the server obtains the mixed information according to the first scaling coefficient .
[0102] Step 3: The server uses the decoder corresponding to the target group to perform forward propagation to recover the information, i.e.
[0103] wherein denotes the decoder of the target group on the server in the tth round of training.
[0104] Step 4: The server performs back propagation based on the recovered signal. The server has a corresponding decoder for the target group. In the tth round of communication, the decoder of the target group is updated as follows:
[0105] wherein η is the learning rate, is the decoder gradient of the target group K during the t-th round of communication (i.e., the whole group gradient). By taking the chain derivative of the whole group gradient, the encoder gradient corresponding to each user terminal in the target group can be determined . The corresponding encoder gradient is sent by the server to the k-th user terminal in the same group.
[0106] Step 5: User terminal k performs back propagation. Each user terminal updates its local semantic encoder using the received encoder gradient :
[0107] where η represents the learning rate, is the encoder gradient of the k-th user terminal in the t-th round of communication.
[0108] Step 6: In the model training based on semantic communication, a periodic processing method is adopted, and the global parameters are obtained by aggregating the encoder parameters sent by all user terminals.
[0109] According to the above processing method, the over-the-air mixing (OAM) based on over-the-air computing can improve the privacy protection capability in the training process with little accuracy reduction, and can reduce the communication overhead in the training process. The convergence speed of the framework (SIMix) provided by the above embodiment, the multi-user semantic communication based on federated learning (Sem-FL), the multi-user semantic communication based on differential privacy (SCDP), and the SIMix without the optimization module (SIMix-W-O) is shown below. The transformer is used as the encoder and decoder of the semantic communication in the experiment, the cifar10 (a standard data set) is used as the data set, the number of users in a group in SIMix and SIMix-W-O is 2, and the PSNR is used to measure the effect of semantic transmission by default. PSNR (Peak Signal-to-Noise Ratio) is an index for measuring the quality of images or signals, commonly used to evaluate the quality loss in the process of image or video compression, reconstruction or transmission. It is calculated by comparing the difference between the original signal (or image) and the reconstructed signal (or image), and is usually expressed in decibels (dB).
[0110] Figure 5 Fig. 1 shows a first effect diagram of the model training method based on semantic communication according to an embodiment of the present application, as Figure 5As shown, the accuracy of the semantic transmission of the SIMix scheme provided by the optional embodiment of the present application is close to that of Sem-FL (only reduced by 1.83 dB), but SIMix-W-O and SCDP are 10.47 dB and 7.68 dB away from Sem-FL, respectively. In terms of communication efficiency, the communication overhead when reaching the same accuracy (25 dB) during training is calculated, and the results show that SIMix is improved by 25% compared with Sem-FL and by 46.43% compared with SCDP (the communication overhead of SMix is 6750 MB, the communication overhead of Sem-FL is 9000 MB, and the communication overhead of SCDP is 14400 MB). In terms of communication overhead, the SIMix framework provided by the optional embodiment of the present application has obvious advantages.
[0111] Figure 6 A second effect diagram of the semantic communication-based model training method of the embodiment of the present application is shown in FIG. 6. Figure 6 As shown, the actual semantic transmission effects of different multi-user semantic communication model training frameworks after convergence are shown. It can be seen that the semantic transmission effects of SIMix and Sem-FL are almost the same, but the effects of the remaining two training frameworks are obviously worse. This also confirms the experimental results in Figure 5 . Figure 6 The red box in .
[0112] Figure 7 A third effect diagram of the semantic communication-based model training method of the embodiment of the present application is shown in FIG. 7. Figure 7 As shown, the privacy protection capabilities of different multi-user semantic communication model training frameworks are shown. It is assumed that an attacker first intercepts the semantic information transmitted by a user, and then performs a communication inversion attack to try to recover the original semantic information that the user wants to transmit. The similarity between the original information and the information recovered by the attacker is measured by PSNR. The Sem-FL framework does not contain any security mechanism, and the attacker can easily recover the original information from the semantic information, resulting in serious privacy leakage (the PSNR of the recovered picture is as high as 38.42 dB). Compared with Sem-FL, the other three frameworks implement different levels of privacy protection measures to protect user data (SIMix is improved by 20.86 dB, SIMix-W-O is improved by 13.39 dB, and SCDP is improved by 18.01 dB), and it can also be seen intuitively that SIMix has the largest difference from the original image compared with SIMix-W-O and SCDP. In terms of privacy protection capability, the SIMix framework provided by the optional embodiment of the present application has obvious advantages.
[0113] Compared with the multi-user semantic communication model training basic framework, the optional embodiment adds a semantic information mixing and user pairing. The semantic information mixing mixes the semantic information based on air computing, which reduces the leakage of semantic information containing user privacy information and saves communication overhead. The user pairing serves the semantic information mixing module and aims to improve the performance of semantic information mixing, improve the accuracy and security.
[0114] Corresponding to the application scenarios and methods of the method provided by the embodiments of the present application, the embodiments of the present application also provide a model training device based on semantic communication, Figure 8 A first schematic diagram of the model training device based on semantic communication of the embodiments of the present application is shown, as Figure 8 shown, comprising: A first scaling strategy determination module 801 is configured to determine a first scaling coefficient for signal compensation of a server and a second scaling coefficient for signal compensation of a plurality of same-group user terminals based on channel gains respectively corresponding to the plurality of same-group user terminals; the plurality of same-group user terminals are user terminals associated with the server and belonging to a target group; A first sending module 802 is configured to send the corresponding second scaling coefficients to the plurality of same-group user terminals respectively; A first receiving module 803 is configured to receive mixed signals sent by the plurality of same-group user terminals through a target channel by using the first scaling coefficient; the mixed signals are obtained by mixing semantic information respectively corresponding to the plurality of same-group user terminals, and the semantic information is obtained by using an encoder based on local training data by the corresponding same-group user terminal and sent by using the second scaling coefficient; the target channel corresponds to the target group; A first forward propagation module 804 is configured to process based on the mixed signals to obtain encoder gradients respectively corresponding to the plurality of same-group user terminals; A second sending module 805 is configured to send the plurality of encoder gradients to the corresponding same-group user terminals respectively; the plurality of encoder gradients are used for the plurality of same-group user terminals to update the corresponding encoders.
[0115] The embodiments of the present application also provide another model training device based on semantic communication, Figure 9 A second schematic diagram of the model training device based on semantic communication of the embodiments of the present application is shown, as Figure 9 shown, comprising: A second receiving module 901 is configured to receive a second scaling coefficient for signal compensation of a target user terminal; the target user terminal is any one of a plurality of same-group user terminals, and the plurality of same-group user terminals are user terminals associated with a server and belonging to a target group; A second forward propagation module 902 is configured to use an encoder of the target user terminal to process based on local training data to obtain semantic information; The third sending module 903 is configured to send the semantic information through a target channel by using a second scaling coefficient; the target channel corresponds to a target group; the target channel is used to mix and transmit semantic information corresponding to a plurality of user terminals in the same group to obtain mixed information; the mixed information is used for the server to receive and process by using a first scaling coefficient to obtain a plurality of encoder gradients corresponding to the plurality of user terminals in the same group; the first scaling coefficient and the second scaling coefficient are determined by the server based on channel gains corresponding to the plurality of user terminals in the same group. The first back propagation module 904 is configured to update the encoder by using the encoder gradient in response to receiving the encoder gradient corresponding to the target user terminal.
[0116] The embodiment of the present application further provides another model training device based on semantic communication, Figure 10 A third schematic diagram of the model training device based on semantic communication is shown in the embodiment of the present application, as shown in Figure 10 The device comprises: The second scaling strategy determination module 1001 is configured to determine, by the server, a first scaling coefficient for signal compensation of the server and a second scaling coefficient for signal compensation of a plurality of user terminals in the same group based on channel gains corresponding to the plurality of user terminals in the same group; the plurality of user terminals in the same group are user terminals associated with the server and belonging to a target group; The fourth sending module 1002 is configured to send the corresponding second scaling coefficient to the plurality of user terminals in the same group by the server respectively; The third receiving module 1003 is configured to receive the corresponding second scaling coefficient by a target user terminal in the plurality of user terminals in the same group, and the second scaling coefficient is used for signal compensation of the target user terminal; The semantic extraction module 1004 is configured to process, by the target user terminal, local training data based on an encoder to obtain semantic information; The fifth sending module 1005 is configured to send the semantic information through a target channel by the target user terminal by using a second scaling coefficient; The fourth receiving module 1006 is configured to receive, by the server, mixed signals sent by the plurality of user terminals in the same group through the target channel by using a first scaling coefficient; The third forward propagation module 1007 is configured to process, by the server, the mixed signals to obtain a plurality of encoder gradients corresponding to the plurality of user terminals in the same group; The sixth sending module 1008 is configured to send the plurality of encoder gradients to the corresponding user terminals in the same group by the server respectively; The second back propagation module 1009 is configured to update the encoder by the target user terminal in response to receiving the corresponding encoder gradient.
[0117] The functions of each module in each device of the embodiments of the present application can be referred to the corresponding description in the above method, and has the corresponding beneficial effects, which will not be repeated here.
[0118] Figure 11 is a block diagram of an electronic device used to implement the embodiments of the present application. As shown in the figure, the electronic device includes a memory 1101 and a processor 1102, and the memory 1101 stores a computer program that can run on the processor 1102. The processor 1102 implements the method in the above embodiments when executing the computer program. The number of the memory 1101 and the processor 1102 can be one or more. In a specific implementation, the electronic device can also include a communication interface 1103 for communicating with external devices and transmitting data. Figure 11
[0119] In a specific implementation, if the memory 1101, the processor 1102 and the communication interface 1103 are independently implemented, the memory 1101, the processor 1102 and the communication interface 1103 can be connected to each other through a bus and complete the communication between them. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, Figure 11 only one thick line is used in the figure, but it does not mean that there is only one bus or only one type of bus.
[0120] Optionally, in a specific implementation, if the memory 1101, the processor 1102 and the communication interface 1103 are integrated on a chip, the memory 1101, the processor 1102 and the communication interface 1103 can complete the communication between them through an internal interface.
[0121] The embodiments of the present application provide a computer readable storage medium, which stores a computer program, and the program is executed by a processor to implement the method provided in the embodiments of the present application.
[0122] The embodiments of the present application provide a computer program product, which includes a computer program, and the program is executed by a processor to implement the method provided in the embodiments of the present application.
[0123] The embodiments of the present application also provide a chip, which includes a processor, is used to call and run instructions stored in a memory, so that a communication device installed with the chip executes the method provided in the embodiments of the present application.
[0124] The embodiment of the application further provides a chip, comprising an input interface, an output interface, a processor and a memory, the input interface, the output interface, the processor and the memory are connected through internal connection channels, the processor is used for executing the code in the memory, and when the code is executed, the processor is used for executing the method provided by the embodiment of the application.
[0125] It should be understood that the processor described above can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc. It should be noted that the processor can be a processor supporting an advanced RISC machine (ARM) architecture.
[0126] Further, the memory can include a read-only memory and a random access memory, optionally. The memory can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memory. The non-volatile memory can include a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically EPROM (EEPROM), or a flash memory, among others. The volatile memory can include a random access memory (RAM), which is used as an external cache. By way of example, and not limitation, many forms of RAM are available. The RAM can include a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate SDRAM (DDR SDRAM), an enhanced SDRAM (ESDRAM), a Sync link DRAM (SLDRAM), and a direct Rambus RAM (DR RAM), among others.
[0127] In the above-described embodiments, all or part of the embodiments can be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions according to the present disclosure are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable apparatus. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another computer-readable storage medium.
[0128] In the description of the application, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the application. Moreover, the specific features, structures, materials or characteristics described can be combined in any suitable manner in one or more embodiments or examples. In addition, different embodiments or examples described in the specification and characteristics of different embodiments or examples can be combined and combined by those skilled in the art without contradiction.
[0129] In addition, the terms "first", "second" are only for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include at least one of the features. In the description of the application, the meaning of "multiple" is two or more, unless otherwise specifically limited.
[0130] Any process or method described in the flowchart or otherwise described herein can be understood as representing a module, a segment or a portion of code including one or more executable instructions for implementing specific logical functions or processes. And the scope of the preferred embodiments of the application includes additional implementations, in which the functions can be performed in the order shown or discussed, including in a substantially simultaneous manner or in reverse order according to the functions involved.
[0131] The logic and / or steps described in the flowchart or otherwise described herein, for example, can be considered as a list of executable instructions for implementing logical functions, which can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus, such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute the instructions, or in conjunction with these instruction execution systems, devices or apparatus.
[0132] It should be understood that parts of the application can be implemented in hardware, software, firmware or a combination thereof. In the above embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. All or part of the steps of the above-described embodiment method can be instructed by the relevant hardware through a program, which can be stored in a computer-readable storage medium, and the program includes one or a combination of the steps of the method embodiment when executed.
[0133] In addition, each of the functional units in the various embodiments of the present application can be integrated in one processing module, or each of the units can exist separately, or two or more units can be integrated in one module. The above-mentioned integrated module can be implemented in the form of hardware or in the form of a software function module. When the above-mentioned integrated module is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer readable storage medium. The storage medium can be a read-only memory, a magnetic disk, an optical disk or the like.
[0134] The above is only an exemplary embodiment of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of various changes or replacements within the technical scope disclosed in the present application, and these should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A model training method based on semantic communication, characterized in that: include: Determining, based on the channel gains corresponding to the plurality of user terminals in the same group, a first scaling factor for performing signal compensation on the server, and a second scaling factor for performing signal compensation on the plurality of user terminals in the same group; The multiple user terminals in the same group are user terminals associated with the server and belonging to the target group; sending corresponding second scaling coefficients to the plurality of user terminals in the same group respectively; receiving, using the first scaling factor, mixed signals transmitted by the plurality of user terminals in the same group via a target channel; the mixed signal being obtained by mixing semantic information corresponding to the plurality of user terminals in the same group, the semantic information being obtained by the corresponding user terminals in the same group using an encoder based on local training data and transmitted using the second scaling factor; the target channel corresponding to the target group; Processing the mixed signal to obtain encoder gradients corresponding to the plurality of user terminals in the same group; A plurality of encoder gradients are respectively sent to corresponding user terminals in the same group; the plurality of encoder gradients are used by the plurality of user terminals in the same group to update corresponding encoders respectively.
2. The method according to claim 1, characterized in that The method further comprises: In response to receiving encoder parameters respectively sent by all the user terminals associated with the server, aggregating the encoder parameters respectively sent by all the user terminals to obtain global parameters; The global parameters are sent to all the user terminals respectively; the global parameters are used by all the user terminals to update corresponding encoders.
3. The method according to claim 2, characterized in that Before determining the first scaling factor of the server and the second scaling factors corresponding to the plurality of user terminals in the same group based on the channel gains corresponding to the plurality of user terminals in the same group, the method further includes: All the user terminals are grouped according to the label distributions and channel gains respectively corresponding to all the user terminals, and a grouping result is determined; the grouping result indicates that the user terminals in the same group belong to the target group.
4. The method according to claim 3, characterized in that Grouping all the user terminals according to the label distributions and channel gains corresponding to all the user terminals, and determining a grouping result, including: Determining label difference information between any two user terminals included in all the user terminals based on the label distributions corresponding to all the user terminals; determining channel compatibility information between any two user terminals based on the channel gains corresponding to all the user terminals; All the user terminals are grouped in a manner that the label difference information satisfies a predetermined difference condition and the channel compatibility information satisfies a predetermined compatibility condition to obtain the grouping result.
5. The method according to any one of claims 1 to 4, characterized in that The determining, based on the channel gains corresponding to the plurality of user terminals in the same group, a first scaling factor of the server and a second scaling factor corresponding to the plurality of user terminals in the same group, includes: Determining, in a manner that minimizes a deviation between a predetermined ideal signal and an expected signal, a first candidate coefficient corresponding to the server as the first scaling coefficient, and second candidate coefficients corresponding to the plurality of user terminals in the same group as corresponding second scaling coefficients; The expected signal is determined based on predetermined noise, the second candidate coefficient, and the channel gains and first candidate coefficients corresponding to the multiple user terminals in the same group, and the first candidate coefficients corresponding to the multiple user terminals in the same group are respectively subject to power constraints of the corresponding user terminals in the same group.
6. The method according to any one of claims 1 to 4, characterized in that The processing based on the mixed signal to obtain encoder gradients corresponding to the plurality of user terminals in the same group includes: Decoding the mixed signal using a decoder corresponding to the target group to obtain restored information; Perform loss calculation based on the recovery information and pre-stored real information to determine the entire group gradient of the plurality of user terminals in the same group; The entire set of gradients is derived to obtain encoder gradients corresponding to the multiple user terminals in the same group.
7. A method for training a semantic communication model, characterized in that: include: receiving a second scaling factor for signal compensation for a target user terminal; The target user terminal is any one of a plurality of user terminals in the same group, wherein the plurality of user terminals in the same group are user terminals associated with the server and belonging to the target group; Using the encoder of the target user end to process based on local training data to obtain semantic information; transmitting the semantic information through a target channel using the second scaling factor; The target channel corresponds to the target group, and the target channel is used for mixed transmission of semantic information corresponding to the multiple user terminals in the same group to obtain mixed information; the mixed information is used by the server to receive and process using a first scaling factor to obtain encoder gradients corresponding to the multiple user terminals in the same group; the first scaling factor and the second scaling factor are determined by the server based on the channel gains corresponding to the multiple user terminals in the same group; In response to receiving the encoder gradient corresponding to the target user end, the encoder is updated using the encoder gradient.
8. The method according to claim 7, characterized in that The method further comprises: When the target user terminal reaches a predetermined training round, sending encoder parameters to the server terminal; In response to receiving the global parameters, the encoder is updated using the global parameters; the global parameters are obtained by aggregating encoder parameters respectively sent by all associated user terminals by the server.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory, wherein the processor implements the method according to any one of claims 1 to 8 when executing the computer program. 10 . A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the method according to claim 1 is implemented.