Multi-modal semantic communication method and device oriented to multiple tasks
By employing multimodal semantic coding, channel mapping, and synchronization modules, the problems of model universality and signal transmission distortion in wireless communication are solved, enabling efficient and reliable semantic communication for multiple tasks and adapting to real hardware environments.
Patent Information
- Application Number
- CN202511662317.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-13
- Publication Date
- 2026-01-30
Smart Images

Figure CN121442011A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of wireless communication and artificial intelligence technology, and particularly relates to a method and apparatus for multi-modal semantic communication oriented towards multiple tasks. Background Technology
[0002] Currently, the integration of wireless communication and artificial intelligence is transitioning from a trend to a reality. Semantic communication, which transmits information semantically rather than bit-level error-free data, can achieve more efficient communication in noisy, bandwidth-limited channels. However, current research and experiments have the following limitations: **Single-Task Modeling:** Most current semantic communication models are designed and trained for specific tasks, such as image classification. Models that perform well in image transmission may not be directly applicable to speech tasks, limiting their versatility. **Lack of Real-World Training Data:** Current model training data is primarily generated by simulation tools, lacking the interference and noise of real-world communication environments. This can lead to poor performance in actual communication scenarios, compromising their effectiveness. **Signal Transmission Distortion:** Signals may experience amplitude or phase distortion during channel transmission, affecting data transmission accuracy, causing the receiver to receive incorrect information, and reducing communication reliability.
[0003] A similar prior art: CN119478386A (“Multimodal Generative Semantic Communication Method and Apparatus”) discloses a multimodal semantic communication scheme for feedback of perceptual data. It mainly improves the transmission efficiency of perceptual data in emergency scenarios with limited communication bandwidth by performing semantic segmentation, compression coding and generating reconstructed images based on diffusion models for visible light and infrared images.
[0004] However, this technology still has the following technical problems:
[0005] First, it focuses on semantic generation and reconstruction between image modalities (visible light and infrared), and does not fully cover the joint transmission of multimodal inputs such as images, text, and voice, nor does it explicitly support the function of completing multiple tasks (such as classification, generation, and sentiment analysis) simultaneously within a unified framework. Second, its semantic transmission is still based on the process of "semantic segmentation → compression coding → generation and reconstruction". The implementation details of the communication link layer, such as signal power normalization, real-time hardware frame structure design, pilot insertion, and multi-task decoding in the physical channel, are not fully covered, resulting in gaps in the robustness and efficiency of multimodal semantic joint transmission, multi-task adaptation, and real-time transmission in actual wireless hardware channels. Summary of the Invention
[0006] To address the problems existing in the prior art, this invention provides a method and apparatus for multi-task-oriented multimodal semantic communication.
[0007] This invention is implemented as follows: A method for multi-task-oriented multimodal semantic communication includes:
[0008] Step 1 involves feature extraction and compression mapping of input data from different modalities using a multimodal semantic encoding module, which includes three sub-modules: the image encoding sub-module, based on a visual Transformer structure, divides the input image into several image blocks and encodes them into semantic vectors; the text encoding sub-module calls a pre-trained BERT model to extract text semantic features; and the speech encoding sub-module uses a self-attention acoustic Transformer to model time-frequency features.
[0009] Step 2: The semantic features are mapped into complex symbol signals through the channel mapping module, the power is normalized, and channel noise is loaded.
[0010] Step 3: Perform actual signal transmission in the hardware test environment through the wireless channel module, call the radio frequency data transmission function to send IQ signals (in-phase and quadrature signals) to the software radio device, and use the data reception function to receive the returned data;
[0011] Step 4: The channel estimation and frame synchronization module inserts pilot signals before each frame to provide channel information; the receiver uses the pilot ratio to perform channel response estimation and amplitude calibration.
[0012] Step 5: Using a unified decoder structure, the multi-task semantic decoding module takes the feature vector from the channel output as input and executes different tasks through task embedding control, including image reconstruction and classification; text understanding and generation; and speech sentiment analysis.
[0013] Step 6: The training and inference engine module executes the training loop, loss function calculation, gradient normalization, and learning rate scheduling, supporting distributed parallel training and inference.
[0014] Furthermore, the feature extraction:
[0015] First, semantic features are extracted. A deep neural network encoder is used to extract features from the input modal data (images, text, etc.). The extracted semantic features are then compressed into a fixed-dimensional channel symbol matrix via linear mapping. The signal power is adjusted using batch normalization and power normalization modules, and the feature signals are input as modulation signals to a real-time radio frequency hardware channel for signal transmission and recovery at the physical layer. The receiving end performs inverse normalization and deep decoding on the returned semantic features and outputs the task results. The core design principle is to map the input data to a high-dimensional semantic feature space and then transmit it directly through the physical channel, rather than using bit-level data.
[0016] Furthermore, the signal transmission
[0017] 1) Frame structure design: Each transmission frame consists of a pilot segment, a data segment, and a padding segment, i.e., [Pilot|Data|Padding]; the pilot symbol length is a fixed value and is used for channel estimation and equalization at the receiver, the data segment is used to carry semantic features, and the padding segment is used for sampling alignment and timing synchronization;
[0018] 2) Pilot generation and insertion: At the signal transmission end, the system embeds a complex pilot sequence of unit amplitude before the transmission sequence. The pilot signal is orthogonally modulated after passing through the Hilbert transform to achieve spectral separation from the main data stream.
[0019] 3) Receiver channel estimation and compensation: The receiver estimates the complex channel response by using the pilot ratio and calculates the mean normalization factor; the estimation result is used to perform amplitude and phase compensation on the main data signal to achieve channel equalization and self-synchronization.
[0020] Another object of the present invention is to provide an apparatus for multi-task-oriented multimodal semantic communication, comprising:
[0021] The multimodal semantic coding module is used for feature extraction and compression mapping of input data from different modalities. It consists of three sub-modules: the image coding sub-module is based on the visual Transformer structure, which divides the input image into several image blocks and encodes them into semantic vectors; the text coding sub-module calls the pre-trained BERT model to extract text semantic features; and the speech coding sub-module uses a self-attention acoustic Transformer to model time-frequency features.
[0022] The channel mapping module is used to map semantic features into complex symbol signals, perform power normalization, and load channel noise.
[0023] The wireless channel module is used to perform actual signal transmission in the hardware test environment. It calls the radio frequency data transmission function to send IQ signals to the software wireless device and uses the data reception function to receive the returned data.
[0024] The channel estimation and frame synchronization module inserts pilot signals before each frame to provide channel information; the receiver uses the pilot ratio to estimate the channel response and perform amplitude calibration.
[0025] The multi-task semantic decoding module adopts a unified decoder structure. It takes feature vectors from the channel output as input and executes different tasks through task embedding control, including image reconstruction and classification; text understanding and generation; and speech sentiment analysis.
[0026] The training and inference engine module performs training loops, loss function calculations, gradient normalization, and learning rate scheduling, and supports distributed parallel training and inference.
[0027] Another object of the present invention is to provide a computer device including a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the steps of the method for multi-tasking multimodal semantic communication.
[0028] Another object of the present invention is to provide a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to perform the steps of the method for multi-task-oriented multimodal semantic communication.
[0029] Another object of the present invention is to provide an information data processing terminal for implementing the device for multi-task-oriented multimodal semantic communication.
[0030] Based on the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solution to be protected by this invention are as follows:
[0031] This invention designs a deep learning-based multimodal semantic communication system capable of transmitting, classifying, and reconstructing multiple tasks including image, vision, and speech. It primarily addresses the following technical challenges: directly transmitting semantic features rather than raw data over a wireless channel; achieving a unified semantic communication architecture for multimodal and multi-task communication through deep learning; implementing semantic communication in a real radio frequency hardware environment; and optimizing the transmission frame structure to achieve real-time channel estimation and automatic synchronization.
[0032] 1. Semantic-level transmission replaces bit-level transmission
[0033] This technology constructs a semantic communication framework. Its core lies in bypassing the transmission of the original data stream and utilizing deep neural networks to extract and transmit "semantic twins" (i.e., highly abstract semantic features). The receiving end reconstructs the information accordingly, achieving a paradigm shift from "bit-fidelity" to "semantic fidelity." Through end-to-end joint optimization and channel adaptation, this framework significantly reduces bandwidth consumption and greatly improves robustness against interference in adverse channel conditions.
[0034] 2. Multimodal, multi-task unified architecture
[0035] This technology constructs a unified multimodal semantic communication architecture by sharing a neural network encoder. It deeply integrates heterogeneous data such as images, text, and speech into a unified semantic space for transmission and task collaboration, breaking down semantic barriers between modalities. This architecture ensures that various core semantic tasks achieve significantly enhanced anti-interference capabilities and execution reliability under complex channel conditions.
[0036] 3. Supports communication in real hardware environments
[0037] The transmission and reception of radio frequency data are completed on a real hardware platform, realizing the physical combination of the algorithm and the radio frequency link. The implementation of deep semantic communication in a real wireless link demonstrates that the invention has engineering deployability and scalability.
[0038] 4. Frame structure with pilot and padding fields
[0039] By introducing pilot signals and padding fields, this technology constructs an intelligent physical layer signal structure with self-calibration and self-synchronization capabilities. This design enables real-time and accurate estimation of the wireless channel, providing a highly reliable and adaptive signal foundation for the stable operation of deep semantic communication models in real, time-varying physical environments.
[0040] The expected benefits and commercial value of the technical solution of this invention after transformation are as follows:
[0041] The "Multi-task-oriented Multimodal Semantic Communication and Intelligent Processing System" proposed in this invention advances deep semantic communication from theoretical simulation to real hardware implementation, and innovatively designs a unified architecture and robust transmission scheme. Its technological breakthrough will be transformed into multi-level, high-potential commercial value.
[0042] 1. Significantly reduce costs and increase efficiency, resolving the contradiction between scarce spectrum resources and explosive data growth.
[0043] 1) Bandwidth costs are drastically reduced: Semantic-level transmission can significantly reduce bandwidth usage compared to traditional bit transmission, allowing operators to carry more users and data within a fixed spectrum, directly reducing capital expenditure (CAPEX) and operating expenditure (OPEX) for network expansion.
[0044] 2) Improved energy efficiency: The significant reduction in the amount of data transmitted directly reduces transmission power and computing energy consumption, which is in line with the urgent need for green communication under the "dual carbon" target, bringing huge green competitiveness to equipment manufacturers and operators.
[0045] 2. Build a reliable machine interaction network to empower key IoT and industrial internet applications.
[0046] 1) Industrial Automation: In noisy factory environments, it provides high noise immunity and low latency instruction transmission for tasks such as robot vision guidance and collaborative control, ensuring the stability and safety of the production line.
[0047] 2) Autonomous driving and vehicle-to-everything (V2X) communication: In scenarios with high-speed movement and rapid signal fading, ensure reliable interaction of key semantic information (such as obstacle location and intent) between vehicles and infrastructure (V2I) and between vehicles (V2V) to improve the safety of autonomous driving.
[0048] 3) Telemedicine: Provides ultra-reliable and lossless transmission of critical data (such as medical images and vital sign semantics) for applications such as remote surgery and intensive care, reducing the risk of bit errors.
[0049] 3. Build a universal communication AI foundation.
[0050] A unified architecture reduces development complexity. The unified architecture for multimodal and multitasking eliminates the need for developers to design separate communication systems for different modalities and data, such as images, speech, and text. A single system can support multiple intelligent tasks, significantly shortening the development cycle and reducing the cost of complex AI applications.
[0051] 4. It has engineering deployability, which accelerates the maturity and industrialization of technology.
[0052] The successful implementation on a real RF hardware platform is a key advantage that distinguishes this technology from many semantic communication solutions that remain at the research paper stage. It demonstrates the maturity and feasibility of the technology to investors and customers, shortening the transformation path from laboratory to product.
[0053] In summary, this invention is not only an integration of a series of advanced technologies, but also a solution capable of directly creating economic value. Its commercial value lies in direct cost reduction and efficiency improvement, reliable assurance for critical applications, scalability as a general-purpose platform, and the market advantage brought by its proven engineering feasibility. This technology is expected to become a core component of the information infrastructure of the future intelligent society, with extremely broad market prospects. Attached Figure Description
[0054] Figure 1 This is a flowchart of a method for multi-task-oriented multimodal semantic communication provided in an embodiment of the present invention.
[0055] Figure 2 This is a structural block diagram of a device for multi-task-oriented multimodal semantic communication provided in an embodiment of the present invention.
[0056] Figure 3 This is a flowchart of a multi-task-oriented semantic communication system provided in an embodiment of the present invention.
[0057] Figure 4 This is a schematic diagram of a hardware-in-the-loop communication structure for multi-tasking provided in an embodiment of the present invention.
[0058] Figure 5 These are physical images of the related products provided in the embodiments of the present invention.
[0059] Figure 6 This is a channel stability feature map provided in an embodiment of the present invention.
[0060] Figure 7 These are the images sent and received according to embodiments of the present invention.
[0061] Figure 8 This is a display of training and testing parameters provided in the embodiments of the present invention.
[0062] Figure 9 This is the text reconstruction result provided by the embodiments of the present invention.
[0063] Figure 10 This is the text classification result provided in the embodiments of the present invention.
[0064] Figure 11 This is a display of training and testing parameters provided in the embodiments of the present invention.
[0065] Figure 12 This is a test result diagram provided in an embodiment of the present invention. Detailed Implementation
[0066] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0067] To address the limitations of traditional communication systems in multimodal data fusion and semantic-level transmission, this multi-task-oriented multimodal semantic communication method proposes a cross-modal information transmission mechanism based on end-to-end semantic layer learning. Existing systems generally rely on independent compression coding and source-channel separation designs, resulting in low bandwidth utilization and poor semantic consistency when dealing with heterogeneous data such as images, text, and speech. Especially in multi-task scenarios, traditional architectures cannot simultaneously support joint optimization of classification, generation, and recognition tasks at the physical layer, leading to interference and performance degradation between tasks. This proposed method constructs a unified multimodal semantic coding and decoding framework, achieving feature sharing and semantic collaboration among multiple tasks, thereby improving overall transmission efficiency and perception quality in constrained channel environments.
[0068] Semantic feature extraction and compression mapping. The system performs unified representation learning on input data from different modalities through a multimodal semantic encoding module. The visual branch uses a Transformer-based self-attention mechanism to achieve global dependency modeling, encoding image patch sequences into context-dependent semantic vectors. The text branch extracts contextual semantic embeddings based on a pre-trained language model, and the speech branch captures time-frequency dynamic features through an acoustic Transformer. The semantic vectors output from each modality are linearly transformed and normalized, then mapped to a fixed-dimensional channel symbol matrix, providing compatible input for subsequent physical layer transmission.
[0069] Channel mapping and physical transmission of semantic signals. At the transmitter, the system performs power normalization and complex symbol modulation on the semantic features, converting them into IQ signals that can be carried by the radio frequency link. To achieve end-to-end training, channel noise is explicitly modeled during the mapping phase, using Rayleigh or Gaussian noise models to simulate fading and interference characteristics in real wireless environments. This process breaks through the hierarchical mode of traditional communication in bit-level coding and modulation, enabling direct and continuous mapping and transmission of semantic features in the channel space.
[0070] The system employs a pilot-assisted frame structure, with each frame containing a pilot segment, a data segment, and a padding segment to simultaneously ensure synchronization and equalization performance. The pilot signal, after undergoing Hilbert transform to generate orthogonal components, is embedded into the main signal stream to achieve spectral isolation. The receiver estimates the complex channel response based on the pilot ratio and performs amplitude and phase compensation to ensure the recoverability of semantic features after passing through the physical channel. This design enables semantic communication to possess self-synchronization and channel awareness capabilities, improving robustness in practical hardware environments.
[0071] Multi-task decoding of semantic features. The receiver adopts a unified decoder architecture, inputting the feature vector from the channel output and controlling the decoding path of different tasks through task embedding vectors. This module can perform tasks such as image reconstruction, text generation, and speech sentiment analysis within the same semantic space, thereby achieving knowledge sharing and semantic transfer between tasks. Through this cross-task semantic alignment mechanism, the system significantly improves the performance of multimodal fusion and task collaboration.
[0072] Channel estimation and adaptive correction. The channel estimation module calculates the instantaneous channel response using the amplitude-phase ratio of the pilot symbols and performs amplitude calibration and noise suppression using a normalization factor. The corrected semantic features are mapped back to the higher-level task semantic space via a deep decoder, thereby achieving end-to-end semantic recovery and task output. This mechanism ensures the model's adaptability under different channel conditions and significantly reduces semantic information distortion.
[0073] The system optimizes the entire process through a training and inference engine module. This module performs backpropagation and gradient normalization of the multi-task joint loss, and combined with a dynamic learning rate scheduling strategy, supports distributed parallel training and asynchronous inference. Through online training in a real RF hardware closed loop, the model gradually learns the implicit mapping relationship between the physical channel and semantic features, thereby achieving integrated collaborative optimization of communication and intelligent tasks. Overall, this method establishes a cross-modal joint transmission mechanism at the semantic layer, providing a new architecture and application direction for intelligent wireless communication.
[0074] like Figure 1 As shown in the figure, a method for multi-task-oriented multimodal semantic communication provided by an embodiment of the present invention includes the following steps:
[0075] S101 performs feature extraction and compression mapping on input data of different modalities through a multimodal semantic coding module, which includes three sub-modules: the image coding sub-module is based on the visual Transformer structure, which divides the input image into several image blocks and encodes them into semantic vectors; the text coding sub-module calls the pre-trained BERT model to extract text semantic features; and the speech coding sub-module uses a self-attention acoustic Transformer to model time-frequency features.
[0076] S102, the semantic features are mapped into complex symbol signals through the channel mapping module, the power is normalized and channel noise is loaded;
[0077] S103 performs actual signal transmission in the hardware test environment through the wireless channel module, calls the radio frequency data transmission function to send IQ signals to the software wireless device, and uses the data reception function to receive the returned data.
[0078] S104 inserts pilot signals before each frame through the channel estimation and frame synchronization module to provide channel information; the receiver uses the pilot ratio to perform channel response estimation and amplitude calibration.
[0079] S105 employs a unified decoder structure through a multi-task semantic decoding module. It takes feature vectors from the channel output as input and executes different tasks through task embedding control, including image reconstruction and classification; text understanding and generation; and speech sentiment analysis.
[0080] S106 performs training loops, loss function calculations, gradient normalization, and learning rate scheduling through the training and inference engine modules, supporting distributed parallel training and inference.
[0081] In this embodiment of the invention, the signal data processing follows an end-to-end semantic communication link from multimodal semantic coding to channel transmission and then to task decoding, specifically including the entire process of feature extraction, complex mapping, channel transmission, channel estimation and semantic recovery.
[0082] First, in the semantic feature extraction stage, the system receives multi-source input data, including image frames, text sentences, and speech signals. Image input is divided into several fixed-size image patches. Each image patch is linearly projected to form a sequence of feature vectors, which are then input into a visual Transformer structure based on a self-attention mechanism for global dependency modeling, thereby generating image semantic vectors. Text data is processed through word embedding and a pre-trained language model (BERT) to extract context-dependent deep semantic representations. Speech signals undergo Mel-spectrum analysis to obtain time-frequency feature maps, which are then input into an acoustic Transformer. A multi-head self-attention structure is used to extract semantic dependency features in the time and frequency domains. The high-dimensional semantic vectors output from these three modalities are uniformly mapped to the semantic feature space through a fusion layer and then a fixed-dimensional feature matrix is generated through a linear compression layer.
[0083] Secondly, in the channel mapping stage, the system transforms the fused semantic feature matrix into a complex signal representation. Specifically, each element in the feature matrix is mapped to a complex symbol, with its real part corresponding to the semantic amplitude and its imaginary part corresponding to the semantic phase. After batch normalization and power normalization, a power-controlled complex signal sequence is obtained. To simulate the real transmission process, channel noise is added to the system, and the noise follows a complex Gaussian distribution with zero mean and variance σ squared. The signal sequence is then converted from digital to analog and modulated by radio frequency to form an IQ signal, which is sent to the wireless channel module for physical transmission at a sampling rate Fs.
[0084] During the channel transmission and estimation phase, each frame structure includes a pilot segment, a data segment, and a padding segment. The pilot segment is used for channel estimation and equalization. The pilot signal uses a unit amplitude complex sequence, which is transformed by Hilbert to obtain orthogonal components and then superimposed on the main signal to achieve spectrum separation and synchronization marking. The receiver obtains the returned IQ signal through a data reception function and recovers it into a complex signal sequence through analog-to-digital conversion. Subsequently, the system uses the complex ratio of the pilot symbols and the received symbols to calculate the instantaneous channel response, obtaining estimates of amplitude and phase. The amplitude normalization factor is determined by the ratio of the mean amplitudes of the transmitted and received pilots, and the channel phase offset is calculated from the argument of the complex ratio. The compensated signal is processed by an equalizer to correct channel distortion and timing drift.
[0085] In the semantic decoding stage, the corrected feature vectors are input to the multi-task semantic decoding module. The decoder adopts a unified structure, using task embedding vectors to control different task paths. For image tasks, the decoder performs deconvolution and inverse Transformer transformation to restore the image; for text tasks, the decoder generates text sequences through attention masks; for speech tasks, the decoder extracts acoustic feature distributions and maps them to sentiment tags.
[0086] Finally, during the training and inference phases, the training engine weighted and fused the semantic reconstruction error, classification error, and sentiment recognition error to calculate the joint loss function. The encoder, decoder, and channel mapping parameters were then updated via backpropagation. Gradient normalization and adaptive learning rate scheduling ensured convergence stability. The entire signal processing process achieved a differentiable mapping from semantic layer features to physical layer signals, ensuring consistency and robustness of multimodal semantic information in real-world channel environments, thereby completing end-to-end communication for multi-task joint optimization.
[0087] Feature extraction provided by embodiments of the present invention:
[0088] First, semantic features are extracted. A deep neural network encoder is used to extract features from the input modal data (images, text, etc.). The extracted semantic features are then compressed into a fixed-dimensional channel symbol matrix via linear mapping. The signal power is adjusted using batch normalization and power normalization modules, and the feature signals are input as modulation signals to a real-time radio frequency hardware channel for signal transmission and recovery at the physical layer. The receiving end performs inverse normalization and deep decoding on the returned semantic features and outputs the task results. The core design principle is to map the input data to a high-dimensional semantic feature space and then transmit it directly through the physical channel, rather than using bit-level data.
[0089] Signal transmission provided in the embodiments of the present invention
[0090] 1) Frame structure design: Each transmission frame consists of a pilot segment, a data segment, and a padding segment, i.e., [Pilot|Data|Padding]; the pilot symbol length is a fixed value and is used for channel estimation and equalization at the receiver, the data segment is used to carry semantic features, and the padding segment is used for sampling alignment and timing synchronization;
[0091] 2) Pilot generation and insertion: At the signal transmission end, the system embeds a complex pilot sequence of unit amplitude before the transmission sequence. The pilot signal is orthogonally modulated after Hilbert transform to achieve spectral separation from the main data stream.
[0092] 3) Receiver channel estimation and compensation: The receiver estimates the complex channel response by using the pilot ratio and calculates the mean normalization factor; the estimation result is used to perform amplitude and phase compensation on the main data signal to achieve channel equalization and self-synchronization.
[0093] like Figure 2 As shown, an embodiment of the present invention provides an apparatus for multi-task-oriented multimodal semantic communication, comprising:
[0094] The multimodal semantic coding module is used for feature extraction and compression mapping of input data from different modalities. It consists of three sub-modules: the image coding sub-module is based on the visual Transformer structure, which divides the input image into several image blocks and encodes them into semantic vectors; the text coding sub-module calls the pre-trained BERT model to extract text semantic features; and the speech coding sub-module uses a self-attention acoustic Transformer to model time-frequency features.
[0095] The channel mapping module is used to map semantic features into complex symbol signals, perform power normalization, and load channel noise.
[0096] The wireless channel module is used to perform actual signal transmission in the hardware test environment. It calls the radio frequency data transmission function to send IQ signals to the software wireless device and uses the data reception function to receive the returned data.
[0097] The channel estimation and frame synchronization module inserts pilot signals before each frame to provide channel information; the receiver uses the pilot ratio to estimate the channel response and perform amplitude calibration.
[0098] The multi-task semantic decoding module adopts a unified decoder structure. It takes feature vectors from the channel output as input and executes different tasks through task embedding control, including image reconstruction and classification; text understanding and generation; and speech sentiment analysis.
[0099] The training and inference engine module performs training loops, loss function calculations, gradient normalization, and learning rate scheduling, and supports distributed parallel training and inference.
[0100] Another object of the present invention is to provide a computer device including a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the steps of the method for multi-tasking multimodal semantic communication.
[0101] Another object of the present invention is to provide a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to perform the steps of the method for multi-task-oriented multimodal semantic communication.
[0102] Another object of the present invention is to provide an information data processing terminal for implementing the device for multi-task-oriented multimodal semantic communication.
[0103] Specific implementation of the present invention:
[0104] This invention proposes a multimodal semantic communication system and formulates an engineering-feasible technical solution, specifically including the following:
[0105] 1. Design of communication methods for semantic feature transmission
[0106] First, semantic feature extraction is performed. Deep neural network encoders are used to extract features from the input modal data (images, text, etc.). The extracted semantic features are then compressed into a fixed-dimensional channel symbol matrix via linear mapping. The signal power is adjusted through batch normalization and power normalization modules, and the feature signal is input as a modulation signal to a real-time RF hardware channel for signal transmission and recovery at the physical layer. At the receiving end, the returned semantic features are inversely normalized and deeply decoded to output the task results. The core design principle lies in mapping the input data to a high-dimensional semantic feature space and then transmitting it directly through the physical channel, rather than using bit-level data.
[0107] 2. Multimodal, multi-task unified architecture design
[0108] Design a unified semantic communication system architecture for multimodal and multitasking systems. The system includes the following modules:
[0109] 1) Multimodal semantic coding module: It adopts a unified deep feature extraction structure, assigns corresponding encoders for different modal inputs, and maps them uniformly to the channel symbol space at the coding output.
[0110] 2) Multi-task control module: The system has a built-in task identifier embedding vector, which automatically selects the corresponding embedding based on the task type during the training and inference phases. Different tasks share the decoder and prediction head, achieving unified communication task joint optimization.
[0111] 3) Joint channel optimization and multimodal fusion: After all modal features are transmitted uniformly through the channel, they are spliced together at the decoding end to achieve cross-modal information complementarity. Furthermore, the system simultaneously optimizes semantic representation, channel robustness, and task performance through backpropagation.
[0112] 3. Training and testing on a real RF hardware platform
[0113] This system enables the coordinated operation of deep learning models and physical channels, improving the control interface of the Software-Defined Radio (SDR). Within the hardware closed-loop channel module, the control interface, based on UDP (User Datagram Protocol) communication, transmits the encoded IQ signal to the SDR via Ethernet, and then transmits the wireless signal through the RF transmit port. Simultaneously, the RF data receive port can receive the complex signal returned via the physical channel. The system performs channel estimation based on the pilot signal, calculates the time delay, and reconstructs the signal response matrix, achieving an end-to-end hardware feedback closed loop.
[0114] 4. Design of the complete transmission frame structure
[0115] A complete frame structure design method including pilot and padding fields is proposed at the channel transport layer to solve the problems of channel estimation and frame synchronization. The specific content is as follows:
[0116] 1) Frame Structure Design: Each transmission frame consists of a pilot segment, a data segment, and a padding segment, i.e., [Pilot|Data|Padding]. The pilot symbol length is a fixed value and is used for channel estimation and equalization at the receiver. The data segment is used to carry semantic features, and the padding segment is used for sampling alignment and timing synchronization.
[0117] 2) Pilot generation and insertion: At the signal transmitting end, the system embeds a complex pilot sequence of unit amplitude before the transmission sequence. The pilot signal is orthogonally modulated after Hilbert transform to achieve spectral separation from the main data stream.
[0118] 3) Receiver-end channel estimation and compensation: The receiver estimates the complex channel response using the pilot ratio and calculates the mean normalization factor. This estimation result is then used to perform amplitude and phase compensation on the main data signal, achieving channel equalization and self-synchronization.
[0119] The system structure diagram of the present invention is as follows: Figure 2 As shown, the functional modules of the semantic communication system mainly include: a multimodal semantic coding module, a channel mapping module, a wireless channel module, a channel estimation and frame synchronization module, a multi-task semantic decoding module, and a training and inference engine module. The connection relationships and functional descriptions of these parts are as follows:
[0120] The multimodal semantic coding module is used for feature extraction and compression mapping of input data from different modalities, and consists of three sub-modules. The image coding sub-module is based on the visual Transformer structure, which divides the input image into several image patches and encodes them into semantic vectors; the text coding sub-module calls the pre-trained BERT model to extract text semantic features; and the speech coding sub-module uses a self-attention acoustic Transformer to model time-frequency features.
[0121] The channel mapping module is used to map semantic features into complex symbol signals, perform power normalization, and load channel noise.
[0122] The wireless channel module is used to perform actual signal transmission in the hardware test environment. It calls the radio frequency data transmission function to send IQ signals to the software radio device and uses the data reception function to receive the returned data.
[0123] The channel estimation and frame synchronization module inserts pilot signals before each frame to provide channel information. The receiver uses the pilot ratio to estimate the channel response and perform amplitude calibration.
[0124] The multi-task semantic decoding module adopts a unified decoder structure, takes feature vectors from the channel output as input, and executes different tasks through task embedding control, including image reconstruction and classification; text understanding and generation; and speech sentiment analysis.
[0125] The training and inference engine module performs training loops, loss function calculations, gradient normalization, and learning rate scheduling, and supports distributed parallel training and inference.
[0126] The communication process of this invention is as follows: Figure 3 As shown, user input includes images, text, and voice data, which are preprocessed and then fed into the corresponding encoding modules. A deep neural network then extracts features from each modality, compresses them into channel symbol dimensions through linear mapping, and performs power normalization to ensure the transmitted signal power meets channel requirements. After feature extraction, the system automatically creates a complete transmission frame containing pilot sequences, data signals, and padding for subsequent signal transmission. The system performs Hilbert transform modulation or quadrature modulation to transform the real signal into a complex signal, then sends the IQ signal to the software-defined radio device via the UDP interface, shifting the baseband signal to a suitable frequency band for wireless transmission, enabling wireless transmission in a real-world environment. At the receiving end, the returned IQ signal is acquired, quadrature demodulation is performed to recover the complex signal, the current channel state is estimated using the pilot ratio, and the data signal undergoes corresponding amplitude and phase compensation to effectively overcome channel distortion. The compensated received signal is then restored to deep features and input to the Transformer decoder for multi-task output, such as image reconstruction and text generation. The system uses a backpropagation algorithm to jointly optimize the entire encoder-channel-decoder link, achieving end-to-end parameter updates and enabling the system to automatically learn the optimal semantic communication strategy under wireless channel constraints.
[0127] like Figure 4 As shown, in a real communication test environment, the system constructs a complete hardware-in-the-loop experimental platform, which consists of three core components: a master control computing node (PC) responsible for running the deep learning model and data processing program; a software-defined radio device responsible for signal modulation, transmission, and reception functions; and an Ethernet interface that uses the UDP protocol to achieve high-speed data transmission and reception between the master control node and the radio frequency device. The experimental process is as follows: First, multimodal data such as images and voice are input into the deep learning model. The model processes the received data to generate a transmission signal, packages the IQ data through the UDP communication interface, and sends it to the radio frequency device for signal transmission and reception. After parsing the received signal, the system calculates the delay compensation (delay=58) and saves the channel response. Finally, during the training or testing phase, the model can load the real channel matrix by calling a function, thereby achieving end-to-end hardware closed-loop testing that is completely equivalent to the simulated channel.
[0128] This invention is mainly applied in the field of education and teaching, specifically supporting comprehensive applications such as course experiments and course design, intelligent communication integrated design, and graduation projects.
[0129] The related product of this invention is the ES2201 Artificial Intelligence Communication Technology Experimental Development Platform, such as... Figure 5 As shown, the platform is based on the Ascend 310 AI processor, features rich peripheral interfaces, and is based on the high-performance heterogeneous computing architecture CANN, supporting collaborative processing of multiple computing nodes and significantly improving computing efficiency. The platform integrates the industry-leading AI framework MindSpore, adapts to various computing scenarios, and, together with comprehensive software development tools and API interfaces (Application Programming Interfaces), builds an experimental and application development environment for artificial intelligence and intelligent communication technologies.
[0130] This platform utilizes artificial intelligence methods to design and train intelligent communication models, endowing them with intelligent processing capabilities for wireless signals. Its functions encompass analog / digital modulation identification, channel coding rate identification, channel encoding and decoding, channel state indication, channel estimation, intelligent receivers, and semantic communication. Its core objective is to optimize the physical layer algorithms and implementation methods of mobile communication air interfaces through AI technology, driving a shift from traditional module-by-module optimization to end-to-end transceiver joint optimization.
[0131] During the training phase, the software-defined radio platform receives air interface wireless signals and converts them into baseband data. This data is then sent to a server or PC as training and testing datasets for training, testing, and optimizing the intelligent communication AI algorithm. After training is complete, the model is adapted and deployed to the artificial intelligence communication technology platform using the appropriate toolchain.
[0132] During the inference phase, the software-defined radio platform continues to convert the received wireless signals into baseband data, which is then input into the artificial intelligence communication technology platform. The platform runs the deployed intelligent communication model, performs real-time inference on the baseband data, and outputs the results, thereby realizing the expected functions of the intelligent communication device or system.
[0133] In practical applications, this invention can be used for model inference, model training, and testing. During model inference, the stability of the real channel is tested and visualized. Figure 6 As shown in the diagram. The blue curve represents the random signal, the red curve represents the modulated signal, the green curve represents the received signal, and the black curve represents the demodulated signal.
[0134] In image tasks, by executing an image reconstruction inference procedure, the model can reconstruct a new image with the same features as the original image based on the sent original image, such as... Figure 7 The system shown can reconstruct the original image; the left side is the transmitted image, and the right side is the received image.
[0135] During model training and testing, the program will print out the number of model parameters, loaded pre-trained weights, data preprocessing methods, optimizer parameters, etc. During testing, it will print out the testing progress, loss, and custom criteria, such as... Figure 8 As shown.
[0136] The text task includes text reconstruction and text classification. The inference program runs and prints the reconstruction and classification results for the text samples. The text reconstruction results are as follows: Figure 9 As shown, the reconstruction result is preds, and the label is targets. It can be seen that the text reconstruction result is consistent with the original label, demonstrating good ability to reconstruct text data through semantic communication; the text classification result is as follows. Figure 10 As shown, the classification result is predicted, and the true category is true_label. The classification category is consistent with the true category result, and the system can correctly perform the text classification task.
[0137] During model training and testing, the program will print out information such as the number of model parameters, the pre-trained weights loaded, the data preprocessing method, and the optimizer parameters. Figure 11 As shown. During testing, the test progress, losses, and custom criteria will be printed out, such as... Figure 12 As shown.
[0138] It should be noted that embodiments of the present invention can be implemented in hardware, software, or a combination of both. The hardware portion can be implemented using dedicated logic; the software portion can be stored in memory and executed by a suitable instruction execution system, such as a microprocessor or dedicated-design hardware. Those skilled in the art will understand that the above-described devices and methods can be implemented using computer-executable instructions and / or included in processor control code, for example, such code provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuitry such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field-programmable gate arrays, programmable logic devices, etc., or by software executed by various types of processors, or by a combination of the above-described hardware circuitry and software, such as firmware.
[0139] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions, and improvements made by those skilled in the art within the scope of the technology disclosed in the present invention, and within the spirit and principles of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A multi-task oriented multi-modal semantic communication method, characterized in that, The method comprises: Through a multi-modal semantic encoding module, semantic feature extraction and compression mapping are performed on image, text and voice inputs; The image encoding adopts a visual Transformer structure based on a self-attention mechanism, and after the input image is divided into a plurality of image blocks, it is encoded into a semantic vector; The text encoding is based on a pre-trained language model to extract context semantic embedding; The voice encoding utilizes an acoustic Transformer to capture time-frequency semantic features; After the multi-modal semantic features are linearly mapped and power normalized, they are input into a channel mapping module, and actual physical transmission is performed in the channel; The receiving end performs reverse normalization and unified decoding on the transmitted semantic features, realizing multi-task output of image reconstruction, text generation and voice emotion recognition.
2. The method of claim 1, wherein, The linear mapping compresses the multi-modal features into a fixed-dimension channel symbol matrix, and each element of the matrix is a complex number, whose real part and imaginary part correspond to feature amplitude and phase information respectively.
3. The method of claim 1, wherein, The channel mapping module performs complex modulation on the semantic features, realizes robust transmission of end-to-end semantic features through random channel noise modeling, and the noise obeys a Gaussian distribution with a mean of 0 and a variance of σ square.
4. The method of claim 1, wherein, The unified decoder structure controls the output paths of different tasks based on task embedding, controls the activation of decoding parameters through a task identification vector, and realizes semantic recovery of multi-task cooperation.
5. A multi-task oriented multi-modal semantic communication device, characterized by, It comprises: A multi-modal semantic encoding module for extracting and compressing the semantic features of different modal inputs; A channel mapping module for converting semantic features into complex symbol signals and performing power normalization; A wireless channel module for transmitting and receiving radio frequency signals in a hardware channel; A channel estimation and frame synchronization module for estimating channel response and amplitude calibration using pilot signals; A multi-task semantic decoding module for multi-task decoding of semantic vectors output by the channel at the receiving end; A training and inference engine module for performing training cycles, gradient normalization and learning rate scheduling to realize end-to-end joint optimization.
6. The apparatus of claim 5, wherein, The channel estimation and frame synchronization module inserts a pilot signal before each frame of data, and the pilot signal is modulated orthogonally to realize spectral separation, which is used for complex ratio estimation of the channel response at the receiving end.
7. The apparatus of claim 5, wherein, The training and inference engine module calculates a multi-task joint loss function, which is composed of the weighted sum of semantic reconstruction loss, classification accuracy loss and voice emotion recognition loss.
8. A channel estimation and self-synchronization method for multi-modal semantic communication, characterized in that, It comprises: A transmission frame structure composed of a pilot segment, a data segment and a padding segment is constructed; A unit amplitude pilot sequence is inserted at the head of each frame, and the orthogonal components are obtained by Hilbert transform and are complexly transmitted with the main signal; The receiving end estimates the channel complex response according to the pilot ratio, and uses the estimation result to compensate the amplitude and phase of the main data signal, realizing channel equalization and timing self-synchronization.
9. The method of claim 8, wherein, The estimation of the channel complex response is realized by calculating the complex ratio of the received pilot and the transmitted pilot, and the amplitude calibration coefficient after compensation is the ratio of the pilot amplitude mean and the received amplitude mean.
10. The method of claim 8, wherein, The self-synchronization process realizes real-time correction of timing drift by detecting the pilot phase difference between consecutive frames after compensation is completed.
Citation Information
Patent Citations
Multi-modal generative semantic communication method and device for sensing data return
CN119478386A
Cited By
Semantic communication method and system, electronic equipment and storage medium
CN122120840A
Semantic communication method and system, electronic device, storage medium
CN122120840B