Semantic communication joint training method and device

The semantic communication training method, which uses cross-platform interactive chains for data interaction, solves the problem of limited transmission performance caused by the lack of wireless communication links in traditional joint training, and achieves better semantic communication transmission effects and anti-interference capabilities.

CN121907402APending Publication Date: 2026-04-21PENG CHENG LAB
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

In existing technologies, traditional joint training schemes do not actually introduce wireless communication links, resulting in limited transmission performance.

Method used

Data interaction is achieved through a cross-platform interaction chain between the training platform and the wireless communication link platform. By using a specific model to jointly train the source channel, semantic communication training is completed, which can resist the effects of time-varying fading, interference, noise and other factors in the wireless environment.

Benefits of technology

It achieves better semantic communication transmission performance, improving transmission effect and anti-interference capability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121907402A_ABST
    Figure CN121907402A_ABST
Patent Text Reader

Abstract

The invention discloses a semantic communication joint training method and device, and relates to the technical field of semantic communication, and the method comprises the steps: a training platform carries out semantic coding based on a multi-modal information source, obtains an original coding result, and transmits the original coding result to a wireless communication link platform through a cross-platform interaction chain; the wireless communication link platform wirelessly transmits the original coding result to obtain a wireless transmission coding result, and transmits the wireless transmission coding result to a training platform through a cross-platform interaction chain; the training platform performs semantic decoding based on a wireless transmission coding result to obtain a multi-modal information sink; and the training platform carries out training on the semantic communication training target based on the multi-modal information sink and the multi-modal information source to obtain an optimized semantic communication training target. Through the above mode, joint training is carried out on the information source channel through the specific model, semantic communication training is completed, influences of time-varying fading, interference, noise and the like of a wireless environment can be resisted, and better transmission performance of semantic communication is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of semantic communication technology, and in particular to semantic communication joint training methods and apparatus. Background Technology

[0002] In semantic communication systems, to ensure better transmission effects and performance, and to combat the effects of time-varying fading, interference, and noise in the wireless environment, and to achieve better transmission performance, it is possible to co-model the semantic features of the source and the transmission characteristics of the channel. However, when co-modeling the semantic features of the source and the transmission characteristics of the channel, it is necessary to establish a transmission channel between the source and the wireless communication link. The current conventional approach is to jointly train the source and the channel, but this approach does not actually introduce a real wireless communication link, and the transmission performance is still limited.

[0003] The above content is only used to help understand the technical solution of the present invention and does not represent an admission that the above content is prior art. Summary of the Invention

[0004] The main purpose of this application is to provide a semantic communication joint training method and apparatus, which aims to solve the technical problem that traditional joint training schemes in the prior art do not actually introduce wireless communication links, resulting in limited transmission performance.

[0005] To achieve the above objectives, this application provides a semantic communication joint training device, which includes a training platform and a wireless communication link platform. The training platform and the wireless communication link platform interact with each other through a cross-platform interaction chain. The wireless communication link platform is connected to a wireless channel. The training platform includes a semantic model module, and the semantic model module includes a semantic communication training target. The training platform is used to perform semantic coding based on multimodal information sources to obtain the original coding result, and transmit the original coding result to the wireless communication link platform through the cross-platform interaction chain; The wireless communication link platform is used to wirelessly transmit the original encoding result in the wireless channel to obtain the wireless transmission encoding result, and transmit the wireless transmission encoding result to the training platform through the cross-platform interaction chain; The training platform is also used to perform semantic decoding based on the wireless transmission coding results to obtain multimodal sinks; The training platform is also used to train the semantic communication training target based on the multimodal sink and the multimodal source to obtain an optimized semantic communication training target.

[0006] In one embodiment, the training platform further includes a performance evaluation module and a gradient calculation module. The performance evaluation module is connected to the semantic model module and the gradient calculation module, respectively, and the gradient calculation module is connected to the semantic model module. The performance evaluation module is used to evaluate the semantic transmission performance based on the multimodal sink and the multimodal source, obtain the performance evaluation result, and feed the performance evaluation result back to the gradient calculation module. The gradient calculation module is used to calculate the corresponding gradient based on the performance evaluation result and feed the gradient back to the semantic model module. The semantic model module is used to optimize the parameters of the semantic communication training objective based on the gradient, so as to obtain an optimized semantic communication training objective.

[0007] In one embodiment, the training platform further includes a sending interface and a receiving interface, and the semantic model module is connected to the cross-platform interaction chain through the sending interface and the receiving interface; The semantic model module is also used to perform semantic encoding based on the multimodal information source to obtain the original encoding result, and to feed the original encoding result back to the sending interface; The sending interface is used to transmit the original encoded result to the wireless communication link platform through the cross-platform interaction chain; The receiving interface is used to receive the wireless transmission coding result transmitted by the wireless communication link platform through the cross-platform interaction chain, and to feed the wireless transmission coding result back to the semantic model module. The semantic model module is also used to perform semantic decoding based on the wireless transmission encoding result to obtain the multimodal sink.

[0008] In one embodiment, the semantic communication training objective includes a semantic encoding model and a semantic decoding model, wherein the semantic encoding model is connected to the sending interface and the semantic decoding model is connected to the receiving interface; The semantic coding model is used to perform semantic coding based on the multimodal information source to obtain the original coding result, and to feed the original coding result back to the sending interface; The semantic decoding model is used to obtain the wireless transmission coding result fed back by the receiving interface, and to perform semantic decoding based on the wireless transmission coding result to obtain the multimodal sink.

[0009] In one embodiment, the semantic communication training objective further includes a semantic knowledge base, which is connected to both the semantic encoding model and the semantic decoding model. The semantic coding model is also used to encode the multimodal information source based on the knowledge data of the semantic knowledge base when the semantic knowledge base is in the open state, so as to obtain the original coding result; The semantic coding model is also used to perform semantic coding on the multimodal information source when the semantic knowledge base is in a closed state, so as to obtain the original coding result; The semantic decoding model is also used to perform semantic decoding on the wireless transmission encoding result based on the knowledge data of the semantic knowledge base when the semantic knowledge base is in the open state, so as to obtain the multimodal sink; The semantic decoding model is also used to perform semantic decoding on the wireless transmission encoding result to obtain the multimodal sink when the semantic knowledge base is in a closed state.

[0010] In one embodiment, the wireless communication link platform includes a wireless communication link transmitting module and a wireless communication link receiving module. The wireless communication link transmitting module is connected to the transmitting interface, and the wireless communication link receiving module is connected to the receiving interface. The wireless communication link transmitting module and the wireless communication link receiving module communicate through the wireless channel.

[0011] Furthermore, to achieve the above objectives, this application also proposes a semantic communication joint training method, applied to the semantic communication joint training system described above, wherein the semantic communication joint training method includes: The training platform performs semantic encoding based on multimodal information sources to obtain the original encoding result, and transmits the original encoding result to the wireless communication link platform through a cross-platform interaction chain; The wireless communication link platform wirelessly transmits the original encoding result in a wireless channel to obtain a wireless transmission encoding result, and then transmits the wireless transmission encoding result to the training platform through the cross-platform interaction chain. The training platform performs semantic decoding based on the wireless transmission encoding results to obtain multimodal information sinks; The training platform trains the semantic communication training objective based on the multimodal sink and the multimodal source to obtain an optimized semantic communication training objective.

[0012] In one embodiment, the semantic communication training objective includes a semantic encoding model, a semantic decoding model, and a semantic knowledge base. The training platform trains the semantic communication training objective based on the multimodal sink and the multimodal source to obtain the optimized semantic communication training objective. The steps include: The training platform evaluates the semantic transmission performance based on the multimodal sink and the multimodal source, and obtains the performance evaluation results. The training platform calculates the corresponding gradient based on the performance evaluation results; When the semantic knowledge base is enabled, the training platform optimizes the parameters of the semantic encoding model and the semantic decoding model based on the gradient, and updates the knowledge data of the semantic knowledge base to obtain an optimized semantic communication training objective. When the semantic knowledge base is closed, the training platform optimizes the parameters of the semantic encoding model and the semantic decoding model based on the gradient to obtain an optimized semantic communication training objective.

[0013] In one embodiment, the training platform performs semantic encoding based on multimodal information sources to obtain the original encoding result, including the following steps: The training platform adjusts the on / off state of the semantic knowledge base based on semantic training requirements, wherein the on / off state is either an on state or a off state. When the semantic knowledge base is enabled, the training platform encodes multimodal information sources based on the knowledge data of the semantic knowledge base to obtain the original encoding result. The training platform performs semantic encoding on multimodal information sources when the semantic knowledge base is closed, and obtains the original encoding result.

[0014] In one embodiment, the step of the training platform performing semantic decoding based on the wireless transmission coding result to obtain the multimodal sink includes: When the semantic knowledge base is enabled, the training platform performs semantic decoding on the wireless transmission encoding result based on the knowledge data of the semantic knowledge base to obtain a multimodal sink. When the semantic knowledge base is closed, the training platform performs semantic decoding on the wireless transmission encoding results to obtain a multimodal sink.

[0015] In addition, to achieve the above objectives, the present invention also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the semantic communication joint training method described above.

[0016] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the semantic communication joint training method described above.

[0017] This application provides a joint training method for semantic communication. The training platform performs semantic encoding based on multimodal sources to obtain the original encoding result, which is then transmitted to the wireless communication link platform via a cross-platform interaction chain. The wireless communication link platform wirelessly transmits the original encoding result to obtain the wireless transmission encoding result, which is then transmitted to the training platform via the cross-platform interaction chain. The training platform performs semantic decoding based on the wireless transmission encoding result to obtain the multimodal destination. The training platform then trains the semantic communication training objective based on the multimodal destination and multimodal sources to obtain an optimized semantic communication training objective. The training platform and wireless communication link of this application can interact via the cross-platform interaction chain, and jointly train the source and channel using a specific model to complete semantic communication training. This method can resist the effects of time-varying fading, interference, and noise in the wireless environment, achieving better transmission performance for semantic communication. It solves the technical problem of limited transmission performance in traditional joint training schemes due to the lack of actual wireless communication links. Attached Figure Description

[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a schematic diagram of the structure of Embodiment 1 of the semantic communication joint training device of this application; Figure 2 A simplified architectural diagram of the semantic communication joint training device provided in Embodiment 1 of this application; Figure 3 This is a schematic diagram of the structure of Embodiment 2 of the semantic communication joint training device of this application; Figure 4 This is a schematic diagram showing the semantic model module of the semantic communication joint training device provided in Embodiment 2 of this application; Figure 5 This is a schematic diagram of the internal modules of the wireless communication link platform of the semantic communication joint training device provided in Embodiment 2 of this application; Figure 6 This is a schematic diagram of a joint training scenario for the semantic communication joint training device provided in Embodiment 2 of this application; Figure 7 This is a flowchart illustrating an embodiment of the semantic communication joint training method of this application; Figure 8A schematic diagram of the overall architecture of the training platform for the semantic communication joint training method provided in Embodiment 1 of this application; Figure 9 This is a schematic diagram of the semantic model module of the semantic communication joint training method provided in Embodiment 1 of this application.

[0021] Explanation of icon numbers: 10. Training Platform; 20. Wireless Communication Link Platform; 30. Cross-Platform Interaction Link; 40. Wireless Channel; 101. Semantic Model Module; 102. Performance Evaluation Module; 103. Gradient Calculation Module; 104. Transmitting Interface; 105. Receiving Interface; 201. Wireless Communication Link Transmitting Module; 202. Wireless Communication Link Receiving Module; 1011. Semantic Encoding Model; 1012. Semantic Decoding Model; 1013. Semantic Knowledge Base.

[0022] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0023] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0024] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0025] The main solution of this application embodiment is as follows: the training platform performs semantic encoding based on multimodal information sources to obtain the original encoding result, and transmits the original encoding result to the wireless communication link platform through a cross-platform interaction chain; the wireless communication link platform wirelessly transmits the original encoding result to obtain the wireless transmission encoding result, and transmits the wireless transmission encoding result to the training platform through a cross-platform interaction chain; the training platform performs semantic decoding based on the wireless transmission encoding result to obtain the multimodal information sink; the training platform trains the semantic communication training target based on the multimodal information sink and the multimodal information source to obtain the optimized semantic communication training target.

[0026] When co-modeling the semantic features of the information source and the transmission characteristics of the channel, it is necessary to establish a transmission channel between the information source and the wireless communication link. The current conventional approach is to jointly train the information source and the channel, but this approach does not actually introduce a real wireless communication link, and the transmission performance is still limited.

[0027] This application provides a solution in which the training platform and the wireless communication link can interact with each other through a cross-platform interactive chain. By jointly training the source channel through a specific model, semantic communication training can be completed. This can resist the effects of time-varying fading, interference, noise and other factors in the wireless environment, and achieve better transmission performance of semantic communication. This solves the technical problem that the traditional joint training scheme does not actually introduce a wireless communication link, which limits the transmission performance.

[0028] This application provides a semantic communication joint training device, referring to... Figure 1 , Figure 1 This is a schematic diagram of the structure of the first embodiment of the semantic communication joint training device of this application; In this embodiment, the semantic communication joint training device includes a training platform 10 and a wireless communication link platform 20. The training platform 10 and the wireless communication link platform 20 interact with each other through a cross-platform interaction chain 30. The wireless communication link platform 20 is connected to a wireless channel 40. The training platform 10 includes a semantic model module 101, which includes a semantic communication training target.

[0029] Training platform 10 is used to perform semantic encoding based on multimodal information sources to obtain the original encoding result, and transmit the original encoding result to the wireless communication link platform through cross-platform interaction chain 30; wireless communication link platform 20 is used to wirelessly transmit the original encoding result to obtain the wireless transmission encoding result, and transmit the wireless transmission encoding result to the training platform through cross-platform interaction chain 30; training platform 10 is also used to perform semantic decoding based on the wireless transmission encoding result to obtain the multimodal information sink; training platform 10 is also used to train the semantic communication training target based on the multimodal information sink and multimodal information sources to obtain the optimized semantic communication training target.

[0030] It should be noted that the multimodal information source can be source data of modalities such as images, videos, text, and speech, and can be determined according to the actual semantic transmission / semantic training requirements. This embodiment does not impose specific limitations on this. The training platform 10 will perform semantic encoding on the input multimodal information source, and the result obtained is the original encoding result.

[0031] Additionally, it should be noted that the wireless communication link platform 20 has both wireless communication link transmission and reception functions. Between transmission and reception, the signal passes through wireless channel 40. If lossless transmission is desired, the signal through wireless channel 40 can be ignored. However, if increased bandwidth, noise, or latency are required, wireless channel 40 can be a 38901 channel, Rayleigh, Rician, or similar channel. The wireless transmission encoding result is the encoded result after wireless transmission.

[0032] Understandably, since the training platform 10 and the wireless communication link platform 20 engage in bidirectional data interaction through the cross-platform interaction chain 30, the raw encoding results generated by the training platform 10 can be transmitted to the wireless communication link platform 20 using the cross-platform interaction chain 30. Correspondingly, the wireless transmission encoding results obtained by the wireless communication link platform 20 can be transmitted to the training platform 10 using the cross-platform interaction chain 30.

[0033] It should be understood that the training platform 10 performs semantic decoding on the wireless transmission encoding results to obtain a multimodal sink, whose modality corresponds to the input multimodal source. By comparing the multimodal source and the multimodal sink, the performance of semantic transmission can be evaluated for the current semantic transmission requirements, thereby enabling the training of the semantic communication training objective and obtaining an optimized semantic communication training objective.

[0034] It should be noted that the semantic communication training objective is the object to be trained, such as a semantic model or a semantic knowledge base. The semantic model includes a semantic encoding model and a semantic decoding model. Optimizing the semantic communication training objective means obtaining a more optimized model or knowledge base after training.

[0035] Currently, the most widely used deep learning frameworks include TensorFlow and PyTorch. In this setup, data transfer and interaction are required between two platforms or environments. On the same hardware platform, a common tensor data representation format and interface can be used, allowing different frameworks to directly access tensor data in the same memory block without data copying, thus reducing memory overhead and data transfer time. However, on different hardware platforms, or across platforms, a new transmission scheme needs to be designed to complete the loop of data transmission, reception, processing, and retransmission.

[0036] Understandably, reference Figure 2 The training platform and the wireless communication link platform are different hardware platforms. Both require independent interfaces for reading and writing to ensure correct data transmission. They form a semantic communication joint training device through a cross-platform interaction chain to complete system training. The entire training process is as follows: the wireless communication link platform is the main server, and the training platform acts as the client server. The training platform uploads data to the main server at specific times for wireless communication link transmission, and downloads the data transmitted from the main server via the wireless communication link to its local machine for further gradient calculation and performance evaluation.

[0037] It should be understood that the training platform includes two interfaces: Upload and Download, while the wireless communication link platform includes two interfaces: Wireless_Input and Wireless_Output. The training platform uses the Upload interface to transmit source data in various modalities, such as images, videos, text, and speech, to the Wireless_Input interface of the wireless communication link simulation platform. This interface quantizes the source data (through bit conversion and other operations) to adapt the data stream to physical layer transmission. The data is then processed by the physical layer and sent to the loopback / air interface / channel simulator, transferring the results to the Wireless_Output interface of the wireless communication link simulation platform. The training platform uses the Download interface to detect the data in the Wireless_Output interface and performs a preliminary comparison. If the sizes are the same, the generated results are downloaded locally for further gradient calculations, performance evaluation, and to improve the semantic encoding / decoding model and / or semantic knowledge base. Taking the training platform's training of image data a.png for channel quality ranging from -20dB to 10dB as an example, the parameters -20dB and a.npy are first sent to the wireless communication link platform. After encoding, transmitting, and decoding, the wireless communication link platform generates a file containing channel quality information and a name matching the source, such as a_-20dB.npy. The training platform receives this file and performs gradient calculation and performance evaluation. To proceed with training on the next channel quality setting, the parameters are changed to "-15dB, -10dB, -5dB, 0dB, 5dB, 10dB…", with different configurations depending on actual needs. After transmitting and receiving, the wireless communication link platform generates data such as "a_-20dB.npy, a_-15dB.npy, a_-10dB.npy, a_0dB.npy, a_5dB.npy, a_10dB.npy", which are then compared with the source data a.npy for gradient calculation and performance evaluation. System-level parameters such as channel type, system bandwidth, and transmission antenna can also be modified.

[0038] Understandably, to avoid situations where data uploaded to the training platform arrives before being retrieved by the wireless link platform, filenames are used to distinguish each data transmission. For example, for training data on -20dB to 10dB channel quality, with 100 sets trained each time, the data could be named file_snr_num_size.npy (file_-20_1_20M.npy, file_-20_2_20M.npy, etc.). When processing the data, the wireless communication link platform reads the data according to the SNR list and num, while simultaneously verifying the file size. If the file size does not meet the requirements, it indicates a transmission failure, and the data is retrained. To prevent repeated failures of the same data transmission, a maximum number of retransmissions is set. If the number of retransmissions exceeds this limit, the current data transmission ends, and the next transmission begins. If multiple data transmissions fail, a status indicator is provided for easy monitoring of the training progress.

[0039] This embodiment provides a semantic communication joint training device, which includes a training platform and a wireless communication link platform. The training platform and the wireless communication link platform interact with each other through a cross-platform interaction chain. The wireless communication link platform is connected to a wireless channel. The training platform includes a semantic model module, which includes a semantic communication training objective. The training platform performs semantic encoding based on multimodal sources to obtain the original encoding result, and transmits the original encoding result to the wireless communication link platform through the cross-platform interaction chain. The wireless communication link platform wirelessly transmits the original encoding result to obtain the wireless transmission encoding result, and transmits the wireless transmission encoding result to the training platform through the cross-platform interaction chain. The training platform performs semantic decoding based on the wireless transmission encoding result to obtain the multimodal destination. The training platform trains the semantic communication training objective based on the multimodal destination and multimodal sources to obtain an optimized semantic communication training objective. In this embodiment, the training platform and the wireless communication link can interact with each other through a cross-platform interaction chain, and jointly train the source and channel using a specific model to complete semantic communication training. This can resist the effects of time-varying fading, interference, noise, etc. in the wireless environment, and achieve better transmission performance for semantic communication.

[0040] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in Embodiment 1 above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 3 The training platform 10 also includes a performance evaluation module 102 and a gradient calculation module 103. The performance evaluation module 102 is connected to the semantic model module 101 and the gradient calculation module 103 respectively, and the gradient calculation module 103 is connected to the semantic model module 101.

[0041] The performance evaluation module 102 is used to evaluate the semantic transmission performance based on the multimodal sink and multimodal source, obtain the performance evaluation result, and feed the performance evaluation result back to the gradient calculation module 103; the gradient calculation module 103 is used to calculate the corresponding gradient based on the performance evaluation result and feed the gradient back to the semantic model module 101; the semantic model module is used to optimize the parameters of the semantic communication training objective based on the gradient to obtain the optimized semantic communication training objective.

[0042] It should be noted that the performance evaluation results are the results obtained by evaluating the semantic transmission performance.

[0043] Understandably, the semantic model module 101 trains a semantic communication training objective for multiple modal sources such as images, videos, text, and speech, tailored to semantic transmission requirements. The semantic model module 101 uses multimodal sources as its primary input and multimodal destinations as its primary output. The performance evaluation module 102 evaluates the performance of semantic transmission for specific semantic transmission requirements by comparing multimodal destinations and sources, and feeds the evaluation results back to the gradient calculation module 103. The gradient calculation module 103 calculates the corresponding gradients based on the semantic transmission performance evaluation results and feeds them into the semantic model module 101 to improve the semantic communication training objective (semantic encoding / decoding model and / or semantic knowledge base).

[0044] In one feasible implementation, the training platform 10 further includes a sending interface 104 and a receiving interface 105, through which the semantic model module 101 is connected to the cross-platform interaction chain 30.

[0045] The semantic model module 101 is also used to perform semantic encoding based on the multimodal information source to obtain the original encoding result, and feed the original encoding result back to the sending interface 104; the sending interface 104 is used to transmit the original encoding result to the wireless communication link platform 20 through the cross-platform interaction chain 30; the receiving interface 105 is used to receive the wireless transmission encoding result transmitted by the wireless communication link platform 20 through the cross-platform interaction chain 30, and feed the wireless transmission encoding result back to the semantic model module 101; the semantic model module 101 is also used to perform semantic decoding based on the wireless transmission encoding result to obtain the multimodal information sink.

[0046] It is understood that the semantic model module 101 is connected to the cross-platform interaction chain 30 through the sending interface 104 and the receiving interface 405 to realize real-time interaction with the wireless communication link platform 20, thereby optimizing the semantic communication training objectives (semantic encoding and decoding model and / or semantic knowledge base) for specific transmission environments.

[0047] In one feasible implementation, refer to Figure 4The semantic communication training objectives include a semantic encoding model 1011 and a semantic decoding model 1012. The semantic encoding model 1011 is connected to the sending interface 104, and the semantic decoding model 1012 is connected to the receiving interface 105. The semantic communication training objectives also include a semantic knowledge base 1013, which is connected to both the semantic encoding model 1011 and the semantic decoding model 1012.

[0048] Semantic coding model 1011 is used to perform semantic coding based on multimodal information sources to obtain the original coding result and feed the original coding result back to the transmitting interface 104; semantic decoding model 1012 is used to obtain the wireless transmission coding result fed back by the receiving interface 105, perform semantic decoding based on the wireless transmission coding result, and obtain the multimodal information sink.

[0049] The semantic coding model 1011 is also used to encode the multimodal information source based on the knowledge data of the semantic knowledge base 1013 when the semantic knowledge base 1013 is in the open state, to obtain the original coding result; the semantic coding model 1011 is also used to perform semantic coding on the multimodal information source when the semantic knowledge base 1013 is in the closed state, to obtain the original coding result; the semantic decoding model 1012 is also used to perform semantic coding on the wireless transmission coding result based on the knowledge data of the semantic knowledge base 1013 when the semantic knowledge base 1013 is in the open state, to obtain the multimodal information destination; the semantic decoding model 1012 is also used to perform semantic decoding on the wireless transmission coding result when the semantic knowledge base 1013 is in the closed state, to obtain the multimodal information destination.

[0050] It should be noted that the semantic encoding model (semantic encoder) 1011 and the semantic decoding model (semantic decoder) 1012 are composed of neural networks and are connected to the semantic knowledge base 1013 through bidirectional interfaces. The semantic knowledge base can be composed entirely of neural networks, partially of neural networks, or defined by other mathematical methods. It is worth noting that for certain specific training needs (such as training semantic communication models not supported by the knowledge base), the semantic knowledge base 1013 can be turned off. Therefore, the on / off state of the semantic knowledge base 1013 can be set according to semantic training needs, and the on / off state can be either an on or off state.

[0051] Understandably, the training platform 10 alternately executes forward and backward operations based on the direction of data flow. For the forward operation, if semantic knowledge 1013 is in a closed state, the semantic coding model 1011 uses the multimodal information source as the main input; if semantic knowledge 1013 is in a closed state, the semantic coding model 1011 uses the knowledge data of the semantic knowledge base and the multimodal information source as input. The semantic coding model 1011 performs semantic coding and feeds the coding result into the transmitting interface 104, so that the coding result is input to the wireless communication link platform, thereby allowing the original coding result to undergo wireless transmission to obtain the wireless transmission coding result. The wireless transmission coding result is fed into the semantic decoding model 1012 by the receiving interface 105. At this time, if semantic knowledge 1013 is in a closed state, the wireless transmission coding result is used as the input of the semantic decoding model 1012; if semantic knowledge base 1013 is in a closed state, the knowledge data of the semantic knowledge base and the wireless transmission coding result are used as the input of the semantic decoding model 1012. Semantic decoding model 1012 performs semantic decoding and outputs a multimodal destination.

[0052] When the semantic knowledge base 1013 is in the open state, the parameters of the semantic encoding model 1011 and the semantic decoding model 1012 are optimized based on the gradient, and the knowledge data of the semantic knowledge base 1013 is updated to obtain the optimized semantic communication training objective; when the semantic knowledge base 1013 is in the closed state, the parameters of the semantic encoding model and the semantic decoding model are optimized based on the gradient to obtain the optimized semantic communication training objective.

[0053] It should be understood that, for backward operations, the semantic model module 101 receives the gradient from the gradient calculation module 103 and then performs backpropagation. Specifically, if the semantic knowledge base 1013 is composed of a neural network and participates in the forward operation, the gradient propagation path is "semantic decoding model, semantic knowledge base, semantic encoding model"; otherwise, the gradient propagation path is "semantic decoding model, semantic encoding model". After the gradient propagation is complete, all neural networks update their own weights using the gradient descent algorithm.

[0054] Furthermore, in one feasible implementation, the wireless communication link platform 20 includes a wireless communication link transmitting module 201 and a wireless communication link receiving module 202. The wireless communication link transmitting module 201 is connected to the transmitting interface 104, and the wireless communication link receiving module 202 is connected to the receiving interface 105. The wireless communication link transmitting module 201 and the wireless communication link receiving module 202 communicate via a wireless channel 40.

[0055] It should be noted that the wireless communication link platform 20 is compatible with traditional communication and semantic communication links, encompassing the complete physical layer process, primarily including physical layer transmission and reception, namely the wireless communication link transmission module 201 and the wireless communication link reception module 202. Internal modules can be found in [reference needed]. Figure 5 For example, code block segmentation unit, check unit, channel coding unit, rate matching unit, scrambling unit, adjustment unit, layer mapping unit, precoding unit, and mapping unit can be selected according to actual needs, supporting the opening and closing of channel coding and decoding to adapt to the transmission characteristics of semantic communication and to achieve the best link performance.

[0056] It is understandable that the wireless communication link platform 20 can be a simulated link platform or a real hardware platform. (Reference) Figure 6 The three scenarios shown are for joint training. Figure 6 There are two different types of links: a purely simulated wireless communication link and a wireless communication hardware link. The hardware link has two modes: one is to connect directly through the air interface, and the other is to connect to the channel simulator. Together with the training platform, these form three different training scenarios.

[0057] In Scenario 1, the training platform and the wireless communication simulation link reside on separate servers. The training platform transmits data (images, videos, etc.) via Ethernet to another server, the wireless communication simulation link platform. The receiving interface converts the training platform's data into a bitstream. This bitstream undergoes wireless communication processing through different channels with varying quality before being returned to the training platform to evaluate the semantic transmission performance. The evaluation result is fed back to the gradient calculation module, which calculates the corresponding gradient based on the evaluation result and feeds it into the semantic model module to improve the semantic encoding / decoding model or semantic knowledge base. This scenario addresses the issue of different hardware and environmental requirements for the training platform and the wireless communication link software platform within the same server.

[0058] In Scenario 2, the wireless communication simulation link is replaced by physical hardware. This physical hardware includes both the transmitting and receiving components of the hardware link. Training data is transmitted to the transmitting end of the wireless communication link via the local network. The transmitting end processes the data through physical layer encoding and transmits it to the receiving end of the wireless communication link via an over-the-air interface or feeder. The receiving end processes the data through the complete physical layer and then transmits it back to the training platform. The training platform improves the encoding / decoding model or semantic knowledge base based on the feedback results. Because the physical hardware is a real physical layer transceiver link, the real-time performance of the training can be better guaranteed, solving the latency problem caused by network transmission in some cases, as seen in Scenario 1.

[0059] Scenario 3, based on Scenario 2, adds a channel simulator. Since the channel simulator can accurately configure key channel parameters, including propagation characteristics, interference conditions, dynamic characteristics, etc., and can save test configurations to ensure the stability of repeated tests, training can be carried out in a controllable and reproducible environment, which can meet the targeted training of source knowledge base, channel knowledge base and other content.

[0060] This embodiment provides a semantic communication joint training device. The training platform further includes a performance evaluation module and a gradient calculation module. The performance evaluation module is connected to both the semantic model module and the gradient calculation module, and the gradient calculation module is connected to the semantic model module. The performance evaluation module evaluates the semantic transmission performance based on multimodal sinks and multimodal sources, obtains the performance evaluation result, and feeds the performance evaluation result back to the gradient calculation module. The gradient calculation module calculates the corresponding gradient based on the performance evaluation result and feeds the gradient back to the semantic model module. The semantic model module optimizes the parameters of the semantic communication training objective based on the gradient to obtain an optimized semantic communication training objective. In this embodiment, the training platform and the wireless communication link can interact with each other through a cross-platform interactive chain. By jointly training the source and channel using a specific model, semantic communication training can be completed. This can resist the effects of time-varying fading, interference, and noise in the wireless environment, achieving better transmission performance for semantic communication.

[0061] This application provides a semantic communication joint training method, referring to... Figure 7 , Figure 7 This is a flowchart illustrating the first embodiment of the semantic communication joint training method of this application.

[0062] In this embodiment, the semantic communication joint training method includes steps S10 to S40: Step S10: The training platform performs semantic encoding based on multimodal information sources to obtain the original encoding result, and transmits the original encoding result to the wireless communication link platform through a cross-platform interaction chain; It should be noted that this embodiment applies to a semantic communication joint training device, which includes a training platform and a wireless communication link platform. The training platform and the wireless communication link platform interact with each other through a cross-platform interaction chain. The wireless communication link platform is connected to a wireless channel. The training platform includes a semantic model module, which includes a semantic communication training objective. For the specific structure, please refer to [reference needed]. Figures 1 to 6 This will not be elaborated upon here.

[0063] Additionally, it should be noted that the multimodal information source can be source data of modalities such as images, videos, text, and speech, and can be determined according to the actual semantic transmission / semantic training requirements. This embodiment does not impose specific limitations on this. The training platform will perform semantic encoding on the input multimodal information source, and the result obtained is the original encoding result.

[0064] In one feasible implementation, the step of the training platform performing semantic encoding on multimodal sources to obtain the original encoding result includes: the training platform adjusting the on / off state of the semantic knowledge base based on semantic training requirements, wherein the on / off state is either an on state or a off state; when the semantic knowledge base is in the on state, the training platform encodes the multimodal sources based on the knowledge data of the semantic knowledge base to obtain the original encoding result; when the semantic knowledge base is in the off state, the training platform performs semantic encoding on the multimodal sources to obtain the original encoding result.

[0065] It should be noted that the reference Figure 8 The semantic communication training objectives include a semantic encoding model, a semantic decoding model, and a semantic knowledge base. The semantic encoding model (semantic encoder) and semantic decoding model (semantic decoder) are composed of neural networks and are connected to the semantic knowledge base through bidirectional interfaces. The semantic knowledge base can be entirely composed of neural networks, partially composed of neural networks, or defined using other mathematical methods. It is worth noting that for certain specific training needs (such as training semantic communication models not supported by a knowledge base), the semantic knowledge base can be turned off. Therefore, the on / off state of the semantic knowledge base can be set according to the semantic training requirements; the on / off state can be either an on or off state.

[0066] Understandably, the training platform alternates between forward and backward operations based on the data flow direction. For the forward operation, if semantic knowledge is disabled, the semantic coding model uses multimodal sources as its primary input; if semantic knowledge is enabled, the model uses both the semantic knowledge base and multimodal sources as input. The semantic coding model performs semantic coding and feeds the coding result into the transmission interface, allowing the encoded result to be input to the wireless communication link platform. This enables the original coding result to undergo wireless transmission, resulting in the wireless transmission coding result.

[0067] Step S20: The wireless communication link platform wirelessly transmits the original encoding result in the wireless channel to obtain the wireless transmission encoding result, and transmits the wireless transmission encoding result to the training platform through the cross-platform interaction chain; It should be noted that the wireless communication link platform has both wireless communication link transmission and reception functions. Between transmission and reception, a wireless channel is used. If the wireless channel is used for lossless transmission, this can be ignored. However, if it requires increased bandwidth, noise, latency, etc., the wireless channel can be a 38901 channel, Rayleigh, Rician, or similar channel. The wireless transmission encoding result is the encoded result after wireless transmission.

[0068] Understandably, the wireless communication link platform is compatible with traditional communication and semantic communication links, encompassing the complete physical layer process, mainly including physical layer transmission and reception. Internal modules may include code block segmentation units, verification units, channel coding units, rate matching units, scrambling units, adjustment units, layer mapping units, precoding units, and mapping units, which can be selected according to actual needs. It supports the opening and closing of channel encoding and decoding to adapt to the transmission characteristics of semantic communication and achieve optimal link performance.

[0069] It should be understood that, since the training platform and the wireless communication link platform exchange data bidirectionally through a cross-platform interaction chain, the raw encoding results generated by the training platform can be transmitted to the wireless communication link platform using the cross-platform interaction chain. Correspondingly, the wireless transmission encoding results obtained by the wireless communication link platform can be transmitted to the training platform using the cross-platform interaction chain.

[0070] Step S30: The training platform performs semantic decoding based on the wireless transmission coding results to obtain the multimodal sink. Understandably, the training platform performs semantic decoding on the wireless transmission encoding results to obtain a multimodal sink, whose modality corresponds to the input multimodal source. By comparing the multimodal source and the multimodal sink, the performance of semantic transmission can be evaluated for the current semantic transmission requirements, thereby enabling the training of the semantic communication training objective and obtaining an optimized semantic communication training objective.

[0071] In one feasible implementation, step S30 may include: when the semantic knowledge base is in an open state, the training platform performs semantic decoding on the wireless transmission encoding result based on the knowledge data of the semantic knowledge base to obtain a multimodal destination; when the semantic knowledge base is in a closed state, the training platform performs semantic decoding on the wireless transmission encoding result to obtain a multimodal destination.

[0072] Understandably, the wireless transmission encoding result will be fed into the semantic decoding model via the receiving interface. At this point, if the semantic knowledge base is disabled, the wireless transmission encoding result serves as the input to the semantic decoding model; if the semantic knowledge base is enabled, the knowledge data from the semantic knowledge base and the wireless transmission encoding result will be used as the input to the semantic decoding model. The semantic decoding model performs semantic decoding and outputs a multimodal destination.

[0073] In step S40, the training platform trains the semantic communication training target based on the multimodal sink and the multimodal source to obtain an optimized semantic communication training target.

[0074] It should be noted that the semantic communication training objective is the object to be trained, such as a semantic model or a semantic knowledge base. The semantic model includes a semantic encoding model and a semantic decoding model. Optimizing the semantic communication training objective means obtaining a more optimized model or knowledge base after training.

[0075] In one feasible implementation, step S40 may include: the training platform evaluating the semantic transmission performance based on the multimodal sink and the multimodal source to obtain a performance evaluation result; the training platform calculating the corresponding gradient based on the performance evaluation result; when the semantic knowledge base is in an open state, the training platform optimizing the parameters of the semantic encoding model and the semantic decoding model based on the gradient, and updating the knowledge data of the semantic knowledge base to obtain an optimized semantic communication training objective; when the semantic knowledge base is in a closed state, the training platform optimizing the parameters of the semantic encoding model and the semantic decoding model based on the gradient to obtain an optimized semantic communication training objective.

[0076] It should be noted that the performance evaluation results are the results obtained by evaluating the semantic transmission performance.

[0077] Additionally, it should be noted that the reference Figure 9 The semantic model module trains semantic communication training objectives for multiple modal sources, including images, videos, text, and speech, to meet semantic transmission requirements. The semantic model module uses multimodal sources as its primary input and multimodal destinations as its primary output. The performance evaluation module compares the performance of multimodal sources and destinations to assess semantic transmission performance for specific requirements and feeds the evaluation results back to the gradient calculation module. The gradient calculation module calculates the corresponding gradients based on the semantic transmission performance evaluation results and feeds them into the semantic model module to improve the semantic communication training objectives (semantic encoding / decoding model and / or semantic knowledge base).

[0078] It should be understood that for backward operations, the semantic model module receives the gradient from the gradient calculation module and then performs backpropagation. Specifically, if the semantic knowledge base is composed of neural networks and participates in the forward operations, the gradient propagation path is "semantic decoder, semantic knowledge base, semantic encoder"; otherwise, the gradient propagation path is "semantic decoder, semantic encoder". After the gradient propagation is complete, all neural networks update their own weights using the gradient descent algorithm.

[0079] For example, assume the training platform includes two interfaces: Upload and Download, and the wireless communication link platform includes two interfaces: Wireless_Input and Wireless_Output. The training platform uses the Upload interface to transmit source data in various modalities, such as images, videos, text, and speech, to the Wireless_Input interface of the wireless communication link simulation platform. This interface quantizes the source data (through bit conversion and other operations) to adapt the data stream to physical layer transmission. Then, the data is processed by the physical layer and sent to the loopback / air interface / channel simulator, transferring the results to the Wireless_Output interface of the wireless communication link simulation platform. The training platform uses the Download interface to detect the data in the Wireless_Output interface and performs a preliminary comparison. If the sizes are the same, the generated results are downloaded locally for further gradient calculations, performance evaluation, and to improve the semantic encoding / decoding model and / or semantic knowledge base. Taking the training platform's training of image data a.png for channel quality ranging from -20dB to 10dB as an example, the parameters -20dB and a.npy are first sent to the wireless communication link platform. After encoding, transmitting, and decoding, the wireless communication link platform generates a file containing channel quality information and a name matching the source, such as a_-20dB.npy. The training platform receives this file and performs gradient calculation and performance evaluation. To proceed with training on the next channel quality setting, the parameters are changed to "-15dB, -10dB, -5dB, 0dB, 5dB, 10dB…", with different configurations depending on actual needs. After transmitting and receiving, the wireless communication link platform generates data such as "a_-20dB.npy, a_-15dB.npy, a_-10dB.npy, a_0dB.npy, a_5dB.npy, a_10dB.npy", which are then compared with the source data a.npy for gradient calculation and performance evaluation. System-level parameters such as channel type, system bandwidth, and transmission antenna can also be modified.

[0080] Understandably, to avoid situations where data uploaded to the training platform arrives before being retrieved by the wireless link platform, filenames are used to distinguish each data transmission. For example, for training data on -20dB to 10dB channel quality, with 100 sets trained each time, the data could be named file_snr_num_size.npy (file_-20_1_20M.npy, file_-20_2_20M.npy, etc.). When processing the data, the wireless communication link platform reads the data according to the SNR list and num, while simultaneously verifying the file size. If the file size does not meet the requirements, it indicates a transmission failure, and the data is retrained. To prevent repeated failures of the same data transmission, a maximum number of retransmissions is set. If the number of retransmissions exceeds this limit, the current data transmission ends, and the next transmission begins. If multiple data transmissions fail, a status indicator is provided for easy monitoring of the training progress.

[0081] This embodiment provides a joint training method for semantic communication. The training platform performs semantic encoding based on multimodal sources to obtain the original encoding result, which is then transmitted to the wireless communication link platform via a cross-platform interaction chain. The wireless communication link platform wirelessly transmits the original encoding result to obtain the wireless transmission encoding result, which is then transmitted to the training platform via the cross-platform interaction chain. The training platform performs semantic decoding based on the wireless transmission encoding result to obtain the multimodal destination. The training platform then trains the semantic communication training objective based on the multimodal destination and multimodal sources to obtain an optimized semantic communication training objective. In this embodiment, the training platform and the wireless communication link can interact via the cross-platform interaction chain, and jointly train the source and channel using a specific model to complete semantic communication training. This method can resist the effects of time-varying fading, interference, and noise in the wireless environment, achieving superior transmission performance for semantic communication.

[0082] The semantic communication joint training method provided in this application, applied to the semantic communication joint training device in the above embodiments, can solve the technical problem that traditional joint training schemes do not actually introduce wireless communication links, resulting in limited transmission performance. Compared with the prior art, the beneficial effects of the semantic communication joint training method provided in this application are the same as those of the semantic communication joint training device provided in the above embodiments, and other technical features in the semantic communication joint training method are the same as those disclosed in the device of the above embodiments, and will not be repeated here.

[0083] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0084] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0085] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, which are used to execute the semantic communication joint training method in the above embodiments.

[0086] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor devices, apparatuses, or components, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution apparatus, device, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0087] The aforementioned computer-readable storage medium may be included in the semantic communication joint training device; or it may exist independently and not be assembled into the semantic communication joint training device.

[0088] The aforementioned computer-readable storage medium carries one or more programs. When these programs are executed by the semantic communication joint training device, the semantic communication joint training device performs the following: a training platform performs semantic encoding based on multimodal information sources to obtain the original encoding result, and transmits the original encoding result to the wireless communication link platform through a cross-platform interaction chain; the wireless communication link platform wirelessly transmits the original encoding result to obtain the wireless transmission encoding result, and transmits the wireless transmission encoding result to the training platform through a cross-platform interaction chain; the training platform performs semantic decoding based on the wireless transmission encoding result to obtain the multimodal information sink; and the training platform trains the semantic communication training objective based on the multimodal information sink and the multimodal information source to obtain an optimized semantic communication training objective.

[0089] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0090] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using dedicated hardware-based apparatus to perform the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0091] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0092] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described semantic communication joint training method. This solves the technical problem that traditional joint training schemes, which do not actually introduce a wireless communication link, have limited transmission performance. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the semantic communication joint training method provided in the above embodiments, and will not be repeated here.

[0093] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the semantic communication joint training method described above.

[0094] The computer program product provided in this application can solve the technical problem that traditional joint training schemes lack actual wireless communication links, resulting in limited transmission performance. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the semantic communication joint training method provided in the above embodiments, and will not be repeated here.

[0095] The above are only some embodiments of this application and do not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A semantic communication joint training device, characterized in that, The semantic communication joint training device includes a training platform and a wireless communication link platform. The training platform and the wireless communication link platform interact with each other through a cross-platform interaction chain. The wireless communication link platform is connected to a wireless channel. The training platform includes a semantic model module, and the semantic model module includes a semantic communication training target. The training platform is used to perform semantic coding based on multimodal information sources to obtain the original coding result, and transmit the original coding result to the wireless communication link platform through the cross-platform interaction chain; The wireless communication link platform is used to wirelessly transmit the original encoding result in the wireless channel to obtain the wireless transmission encoding result, and transmit the wireless transmission encoding result to the training platform through the cross-platform interaction chain; The training platform is also used to perform semantic decoding based on the wireless transmission coding results to obtain multimodal sinks; The training platform is also used to train the semantic communication training target based on the multimodal sink and the multimodal source to obtain an optimized semantic communication training target.

2. The apparatus as claimed in claim 1, characterized in that, The training platform also includes a performance evaluation module and a gradient calculation module. The performance evaluation module is connected to the semantic model module and the gradient calculation module, respectively, and the gradient calculation module is connected to the semantic model module. The performance evaluation module is used to evaluate the semantic transmission performance based on the multimodal sink and the multimodal source, obtain the performance evaluation result, and feed the performance evaluation result back to the gradient calculation module. The gradient calculation module is used to calculate the corresponding gradient based on the performance evaluation result and feed the gradient back to the semantic model module. The semantic model module is used to optimize the parameters of the semantic communication training objective based on the gradient, so as to obtain an optimized semantic communication training objective.

3. The apparatus as described in claim 2, characterized in that, The training platform also includes a sending interface and a receiving interface, and the semantic model module is connected to the cross-platform interaction chain through the sending interface and the receiving interface. The semantic model module is also used to perform semantic encoding based on the multimodal information source to obtain the original encoding result, and to feed the original encoding result back to the sending interface; The sending interface is used to transmit the original encoded result to the wireless communication link platform through the cross-platform interaction chain; The receiving interface is used to receive the wireless transmission encoding result transmitted by the wireless communication link platform through the cross-platform interaction chain, and to feed back the wireless transmission encoding result to the semantic model module. The semantic model module is also used to perform semantic decoding based on the wireless transmission encoding result to obtain the multimodal sink.

4. The apparatus as described in claim 3, characterized in that, The semantic communication training objective includes a semantic encoding model and a semantic decoding model. The semantic encoding model is connected to the sending interface, and the semantic decoding model is connected to the receiving interface. The semantic coding model is used to perform semantic coding based on the multimodal information source to obtain the original coding result, and to feed the original coding result back to the sending interface; The semantic decoding model is used to obtain the wireless transmission coding result fed back by the receiving interface, and to perform semantic decoding based on the wireless transmission coding result to obtain the multimodal sink.

5. The apparatus as described in claim 4, characterized in that, The semantic communication training objective also includes a semantic knowledge base, which is connected to the semantic encoding model and the semantic decoding model respectively. The semantic coding model is also used to encode the multimodal information source based on the knowledge data of the semantic knowledge base when the semantic knowledge base is in the open state, so as to obtain the original coding result; The semantic coding model is also used to perform semantic coding on the multimodal information source when the semantic knowledge base is in a closed state, so as to obtain the original coding result; The semantic decoding model is also used to perform semantic decoding on the wireless transmission encoding result based on the knowledge data of the semantic knowledge base when the semantic knowledge base is in the open state, so as to obtain the multimodal sink; The semantic decoding model is also used to perform semantic decoding on the wireless transmission encoding result to obtain the multimodal sink when the semantic knowledge base is in a closed state.

6. The apparatus as claimed in claim 3, characterized in that, The wireless communication link platform includes a wireless communication link transmitting module and a wireless communication link receiving module. The wireless communication link transmitting module is connected to the transmitting interface, and the wireless communication link receiving module is connected to the receiving interface. The wireless communication link transmitting module and the wireless communication link receiving module communicate through the wireless channel.

7. A semantic communication joint training method, characterized in that, The semantic communication joint training apparatus, as described in any one of claims 1 to 6, comprises the following semantic communication joint training method: The training platform performs semantic encoding based on multimodal information sources to obtain the original encoding result, and transmits the original encoding result to the wireless communication link platform through a cross-platform interaction chain; The wireless communication link platform wirelessly transmits the original encoding result in a wireless channel to obtain a wireless transmission encoding result, and then transmits the wireless transmission encoding result to the training platform through the cross-platform interaction chain. The training platform performs semantic decoding based on the wireless transmission encoding results to obtain multimodal information sinks; The training platform trains the semantic communication training objective based on the multimodal sink and the multimodal source to obtain an optimized semantic communication training objective.

8. The method as described in claim 7, characterized in that, The semantic communication training objective includes a semantic encoding model, a semantic decoding model, and a semantic knowledge base. The training platform trains the semantic communication training objective based on the multimodal sink and the multimodal source to obtain the optimized semantic communication training objective. The steps include: The training platform evaluates the semantic transmission performance based on the multimodal sink and the multimodal source, and obtains the performance evaluation results. The training platform calculates the corresponding gradient based on the performance evaluation results; When the semantic knowledge base is enabled, the training platform optimizes the parameters of the semantic encoding model and the semantic decoding model based on the gradient, and updates the knowledge data of the semantic knowledge base to obtain an optimized semantic communication training objective. When the semantic knowledge base is closed, the training platform optimizes the parameters of the semantic encoding model and the semantic decoding model based on the gradient to obtain an optimized semantic communication training objective.

9. The method as described in claim 7, characterized in that, The training platform performs semantic coding based on multimodal information sources, and the steps to obtain the original coding results include: The training platform adjusts the on / off state of the semantic knowledge base based on semantic training requirements, wherein the on / off state is either an on state or a off state. When the semantic knowledge base is enabled, the training platform encodes the multimodal information source based on the knowledge data of the semantic knowledge base to obtain the original encoding result. The training platform performs semantic encoding on multimodal information sources when the semantic knowledge base is closed, and obtains the original encoding result.

10. The method as described in claim 9, characterized in that, The training platform performs semantic decoding based on the wireless transmission coding results to obtain the multimodal information sink, including the following steps: When the semantic knowledge base is enabled, the training platform performs semantic encoding on the wireless transmission encoding result based on the knowledge data of the semantic knowledge base to obtain a multimodal sink. When the semantic knowledge base is closed, the training platform performs semantic decoding on the wireless transmission encoding results to obtain a multimodal sink.