Semantic communication method and apparatus, electronic device, and storage medium

CN122824352APending Publication Date: 2026-09-25TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610910709.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-23
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0004]本申请实施例的目的是提供一种语义通信方法、装置、电子设备和存储介质,能够解决传输鲁棒性差、场景适配能力严重不足的问题

Benefits of technology

[0012]在本申请实施例中,通过各模态对应的编码器对多模态数据中相应模态数据进行特征提取,得到各模态对应的初始特征,通过语义解耦网络将各模态对应的初始特征解耦为多类语义分量,多类语义分量逼近两两相互正交,根据无线通信信道的信道参数和目标业务需求,将多类语义分量中至少一类语义分量通过无线通信信道发送至接收端,通过将多模态数据的语义特征解耦为逼近两两相互正交的多类语义分量,解决了现有技术中单一整体语义表征中存在的语义特征高度纠缠的问题,不会因为信道噪声、衰落等损伤而导致全域扩散,可以提高低信噪比、带宽受限场景下的传输性能,提高传输的鲁棒性,而且可以基于信道参数和目标业务需求来选择要传输的语义分量,可以自适应调整传输策略,提高了场景适配能力,可以促进多模态语义通信的规模化落地应用。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122824352A_ABST
    Figure CN122824352A_ABST
Patent Text Reader

Abstract

The application discloses a semantic communication method and device, electronic equipment and storage medium, and belongs to the technical field of communication. The method comprises the following steps: performing feature extraction on corresponding modal data in multi-modal data through an encoder corresponding to each mode to obtain initial features corresponding to each mode; decoupling the initial features corresponding to each mode into multi-class semantic components through a semantic decoupling network, wherein the multi-class semantic components are mutually orthogonal; and transmitting at least one semantic component in the multi-class semantic components to a receiving end through a wireless communication channel according to channel parameters of the wireless communication channel and target service requirements. The application can improve the transmission performance in a low signal-to-noise ratio and bandwidth limited scene, improve the robustness of transmission, select the semantic component to be transmitted based on the channel parameters and the target service requirements, adaptively adjust the transmission strategy, improve the scene adaptation capability, and promote the large-scale landing application of multi-modal semantic communication.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of communication technology, specifically relating to a semantic communication method, device, electronic device, and storage medium. Background Technology

[0002] With the rapid iteration and development of 6G mobile communication technology, emerging intelligent services such as autonomous driving, remote smart healthcare, immersive AR / VR interaction, smart cities, and integrated air-space-ground communication are being rapidly implemented. These service scenarios continuously generate massive amounts of heterogeneous multimodal data, including text, images, audio, video, and sensor data. Traditional wireless communication systems use a "bit-equal transmission" mechanism, which does not distinguish between the semantic value of data and its relevance to the task. This results in significant problems such as large data transmission redundancy, low bandwidth utilization, and high transmission latency, making it difficult to adapt to the high-concurrency, low-latency, and high-reliability requirements of 6G communication. Semantic communication, as a core technology to replace traditional bit communication, extracts key semantic features of the task through neural networks, eliminating the original redundant bit transmission and effectively improving communication compression efficiency.

[0003] Existing multimodal semantic communication solutions mostly employ simple feature fusion methods to integrate heterogeneous multimodal data into a single overall semantic representation, which has significant technical drawbacks. These drawbacks are mainly manifested in the following ways: semantic features are highly entangled, and impairments such as channel noise and fading spread throughout the entire semantic representation, polluting the overall semantic representation. This leads to a sharp deterioration in transmission performance under low signal-to-noise ratio and bandwidth-constrained scenarios, resulting in extremely poor robustness. Furthermore, the lack of a dynamic channel adaptation mechanism means that transmission strategies cannot be adaptively adjusted based on real-time channel conditions and service requirements, resulting in severely insufficient scenario adaptability and greatly limiting the large-scale application of multimodal semantic communication technology. Summary of the Invention

[0004] The purpose of this application is to provide a semantic communication method, apparatus, electronic device, and storage medium that can solve the problems of poor transmission robustness and severe lack of scene adaptability.

[0005] To solve the above-mentioned technical problems, this application is implemented as follows: In a first aspect, embodiments of this application provide a semantic communication method applied at a sending end, the method comprising: By using the encoder corresponding to each modality, feature extraction is performed on the corresponding modal data in the multimodal data to obtain the initial features corresponding to each modality; The initial features corresponding to each modality are decoupled into multiple semantic components by a semantic decoupling network, and the multiple semantic components are approximately mutually orthogonal to each other. Based on the channel parameters of the wireless communication channel and the target service requirements, at least one of the multiple semantic components is transmitted to the receiving end through the wireless communication channel.

[0006] Secondly, embodiments of this application provide a semantic communication method applied at a receiving end, the method comprising: The receiving end sends at least one type of semantic component through a wireless communication channel. The at least one type of semantic component is at least one of the multiple semantic components corresponding to the multimodal data. The multiple semantic components are approximately pairwise orthogonal to each other. Perform the operation corresponding to the target business requirement on the at least one type of semantic component.

[0007] Thirdly, embodiments of this application provide a semantic communication device applied at a sending end, comprising: The feature extraction module is used to extract features from the corresponding modal data in the multimodal data through the encoder corresponding to each modality, so as to obtain the initial features corresponding to each modality; The semantic decoupling module is used to decouple the initial features corresponding to each modality into multiple semantic components through a semantic decoupling network. The multiple semantic components are approximately pairwise orthogonal to each other. The semantic transmission module is used to transmit at least one type of semantic component among the multiple types of semantic components to the receiving end through the wireless communication channel according to the channel parameters of the wireless communication channel and the target service requirements.

[0008] Fourthly, embodiments of this application provide a semantic communication device applied at a receiving end, comprising: The receiving module is used to receive at least one type of semantic component sent by the transmitting end through the wireless communication channel. The at least one type of semantic component is at least one type of semantic component among the multiple types of semantic components corresponding to the multimodal data. The multiple types of semantic components are approximately pairwise orthogonal to each other. The operation execution module is used to perform operations corresponding to the target business requirements on the at least one type of semantic components.

[0009] Fifthly, embodiments of this application provide an electronic device including a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the method described in the first aspect.

[0010] In a sixth aspect, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.

[0011] In a seventh aspect, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.

[0012] In this embodiment, the encoder corresponding to each modality extracts features from the corresponding modal data in the multimodal data to obtain the initial features corresponding to each modality. The initial features corresponding to each modality are decoupled into multiple semantic components through a semantic decoupling network. The multiple semantic components are approximately pairwise orthogonal. According to the channel parameters of the wireless communication channel and the target service requirements, at least one of the multiple semantic components is sent to the receiving end through the wireless communication channel. By decoupling the semantic features of the multimodal data into multiple semantic components that are approximately pairwise orthogonal, the problem of highly entangled semantic features in the single overall semantic representation in the prior art is solved. It will not cause global diffusion due to channel noise, fading and other damage. It can improve the transmission performance in low signal-to-noise ratio and bandwidth-limited scenarios, improve the robustness of transmission, and can select the semantic components to be transmitted based on the channel parameters and target service requirements. It can adaptively adjust the transmission strategy, improve the scenario adaptability, and promote the large-scale application of multimodal semantic communication. Attached Figure Description

[0013] Figure 1 This is a flowchart of a semantic communication method provided in an embodiment of this application; Figure 2 This is a schematic diagram of the structure of a semantic communication system that implements the semantic communication method in the embodiments of this application; Figure 3 This is a schematic diagram of the semantic decoupling network in the embodiments of this application; Figure 4a and Figure 4b This is a schematic diagram of channel-aware adaptive semantic transmission mode selection in an embodiment of this application; Figure 5 This is a flowchart of a semantic communication method provided in an embodiment of this application; Figure 6 This is a schematic diagram of the structure of a semantic communication device provided in an embodiment of this application; Figure 7 This is a schematic diagram of the structure of a semantic communication device provided in an embodiment of this application; Figure 8 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application. Detailed Implementation

[0014] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0015] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0016] The semantic communication method, apparatus, electronic device, and storage medium provided in this application will be described in detail below with reference to the accompanying drawings and through specific embodiments and application scenarios.

[0017] Figure 1 This is a flowchart of a semantic communication method provided in an embodiment of this application. This semantic communication method is applied at the sending end, such as... Figure 1 As shown, the method may include steps 110 to 130.

[0018] Step 110: Extract features from the corresponding modal data in the multimodal data using the encoders corresponding to each modality to obtain the initial features corresponding to each modality.

[0019] Multimodal data can include at least two types of heterogeneous modal data, such as text data, audio data, and visual data.

[0020] The original input multimodal data is acquired, and the corresponding modal data in the multimodal data is extracted by the encoder corresponding to each modality to obtain the initial features corresponding to each modality.

[0021] Step 120: Decouple the initial features corresponding to each modality into multiple semantic components through a semantic decoupling network. The multiple semantic components are approximately pairwise orthogonal to each other.

[0022] Based on the principles of Partial Information Decomposition (PID) and Minimum Necessary Information (MNI), and combined with information bottleneck theory, a semantic decoupling network is used to decouple the initial features of each modality through information theory. This decomposes the multimodal global semantics into multiple classes of approximately orthogonal semantic components, resulting in a multi-class semantic component. These multi-class semantic components can be categorized into four types: cross-modal shared semantic components, modality-specific semantic components, multimodal collaborative semantic components, and modality redundant semantic components. Specifically, cross-modal shared semantic components carry unified core task information across all modalities; modality-specific semantic components carry effective task information specific to a single modality; multimodal collaborative semantic components carry high-order interaction information generated by multimodal fusion; and modality redundant semantic components carry detailed information only used for original modality reconstruction and irrelevant to downstream tasks.

[0023] Using partial information decomposition to guide task-related information decomposition:

[0024] Where X1 and X2 represent the original modal data (taking two types of modal data as an example), that is, the initial features of each modality, Y represents the task information, C represents the shared task information across modalities, U1 and U2 represent the modality-specific task information, and S represents the collaborative task information.

[0025] Modal redundancy semantic component ZR is a reconstruction-oriented residual semantics used to preserve details that are weakly related to downstream tasks but useful for recovering the original modality.

[0026] MNI employs the minimum necessary information in multimodal representation learning and information bottleneck contexts, meaning "a semantic representation that is sufficient for the task and as minimal as possible." The goal of cross-modal semantic component sharing (ZC) is:

[0027] in, This indicates cross-modal sharing of semantic components. These represent the network parameters of the shared semantic branches. This represents the overall network parameters (including the semantically decoupled network, the encoders for each modality, and the decoder at the receiver). The objective is to preserve cross-modal consistency information useful for task Y, while suppressing task-independent dependencies that still exist given Y.

[0028] Step 130: Based on the channel parameters of the wireless communication channel and the target service requirements, at least one semantic component among the multiple semantic components is transmitted to the receiving end through the wireless communication channel.

[0029] Channel parameters are key physical or logical indicators describing the transmission characteristics of wireless communication channels, used to evaluate channel quality, plan system performance, and achieve link adaptation. Target service requirements refer to the type of service the receiver needs to perform using this multimodal data, such as task inference or modality reconstruction.

[0030] Based on the channel parameters of the wireless communication channel and the target service requirements, a target semantic transmission mode is determined. At least one semantic component corresponding to the target semantic transmission mode is selected from the multiple semantic components obtained through decoupling. This at least one semantic component is encoded into a channel symbol, and the encoded channel symbol is transmitted to the receiving end via the wireless communication channel to complete the wireless transmission. This allows the receiving end to execute service operations corresponding to the target service requirements based on this at least one semantic component. The target semantic transmission mode is one of several pre-set semantic transmission modes.

[0031] The semantic communication method provided in this application extracts features from the corresponding modal data in multimodal data through encoders corresponding to each modality, obtaining initial features corresponding to each modality. A semantic decoupling network decouples these initial features into multiple semantic components, which are approximately pairwise orthogonal. Based on the channel parameters of the wireless communication channel and the target service requirements, at least one semantic component is transmitted to the receiving end through the wireless communication channel. By decoupling the semantic features of multimodal data into multiple semantic components that are approximately pairwise orthogonal, this method solves the problem of highly entangled semantic features in the single overall semantic representation in existing technologies. It avoids global diffusion due to channel noise, fading, and other impairments, improving transmission performance in low signal-to-noise ratio and bandwidth-constrained scenarios, enhancing transmission robustness. Furthermore, it allows selection of the semantic components to be transmitted based on channel parameters and target service requirements, enabling adaptive adjustment of transmission strategies and improving scenario adaptability. This promotes the large-scale application of multimodal semantic communication.

[0032] In some embodiments of this application, the multiple semantic components include cross-modal shared semantic components, modality-specific semantic components, multimodal collaborative semantic components, and modality-redundant semantic components; the semantic decoupling network includes shared semantic branches, unique semantic branches, collaborative semantic branches, and reconstructed residual semantic branches; The step of decoupling the initial features corresponding to each modality into multiple semantic components through a semantic decoupling network includes: extracting the cross-modal shared semantic components from the initial features corresponding to each modality through the shared semantic branch; extracting the modality-specific semantic components corresponding to each modality from the initial features corresponding to each modality through the unique semantic branch under the condition of the cross-modal shared semantic components; fusing the initial features corresponding to each modality through the collaborative semantic branch to obtain the multimodal collaborative semantic components; and extracting modality reconstruction information from the initial features corresponding to each modality through the reconstruction residual semantic branch to obtain the modality redundant semantic components.

[0033] Figure 2 This is a schematic diagram of the structure of a semantic communication system that implements the semantic communication method in an embodiment of this application. For example... Figure 2 As shown, each mode undergoes feature extraction via its corresponding encoder to obtain initial features for each mode, which can be represented as follows:

[0034] in, This represents the initial feature corresponding to the m-th mode. This represents the data of the m-th mode. This represents the encoder for the m-th mode.

[0035] Shared semantic branches include variational autoencoders based on the vMF distribution. For example... Figure 2 As shown, the initial features of each modality are processed by a shared semantic branch and a variational autoencoder based on the vMF distribution to extract the shared semantic component Z. Cm The shared semantic components obtained from all modalities constitute the cross-modal shared semantic component Z. C The shared semantic branch provides cross-modal consistency during training through the mutual information lower bound loss function (InfoNCE) and suppresses task-independent conditional dependencies through the mutual information lower bound loss function (NCE-CLUB).

[0036] Modality-specific semantic components are task-specific information that is exclusive to a single modality and independent of shared semantics. The unique semantic branch, through conditional encoding, extracts the modality-specific semantic components from the initial features corresponding to each modality under the constraint of shared semantic components across modalities. This can be represented as follows:

[0037] in, This represents the mode-specific semantic component of the m-th mode. This represents the initial feature corresponding to the m-th mode. This indicates that semantic components are shared across modalities.

[0038] The advantage of using cross-modal shared semantic components as constraints to extract modality-specific semantic components is that, given the cross-modal shared semantic components, the unique semantic branch only learns the incremental information of that modality for task Y, and avoids copying shared semantics through constraints.

[0039] The variational autoencoder based on the vMF distribution predicts the mean direction and concentration parameters by an MLP (Multilayer Perceptron) and samples shared and unique semantics from the vMF distribution based on the mean direction and concentration parameters.

[0040] like Figure 2 As shown, the collaborative semantic branch can include a fusion layer, which can be a Transformer network. The multimodal collaborative semantic components are high-order interactive semantic information generated after multimodal data fusion. The collaborative semantic branch uses a Transformer network to jointly encode the initial features corresponding to each modality to obtain the multimodal collaborative semantic components. This process can be represented as follows:

[0041] in, Represents multimodal collaborative semantic components. This represents the initial characteristics of each mode (taking two modes as an example). Network parameters representing collaborative semantic branches.

[0042] The reconstruction residual semantic branch may include a variational autoencoder. The modal redundancy semantic component is detailed information that is irrelevant to downstream tasks and is only used for repairing and reconstructing the original modal data. Modal reconstruction information is extracted from the initial features corresponding to each modality by the variational autoencoder to obtain the modal redundancy semantic component.

[0043] Figure 3 This is a schematic diagram of the semantic decoupling network structure in an embodiment of this application, such as... Figure 3 As shown, modal data (text modality is used as an example in the figure) X t The initial feature H is obtained by encoding through the backbone encoder (i.e., the encoder corresponding to the mode). t The shared semantic layer (i.e., the shared semantic branch) extracts the shared semantic components corresponding to the modality from the initial features, while the unique semantic layer (i.e., the unique semantic branch) extracts the unique semantic components corresponding to the modality from the initial features. Figure 3 The collaborative semantic branch and the reconstruction residual semantic branch are not shown in the figure.

[0044] By extracting the corresponding semantic components from each semantic branch, semantic decoupling of multimodal data is achieved, making each semantic component approximately orthogonal to each other, thus avoiding entanglement between different semantic components during transmission.

[0045] In some embodiments of this application, the channel parameters include signal-to-noise ratio and available bandwidth; The step of transmitting at least one semantic component from the multiple semantic components to the receiving end through the wireless communication channel according to the channel parameters of the wireless communication channel and the target service requirements includes: When the target service requirement is a task inference service, and the signal-to-noise ratio is less than the signal-to-noise ratio threshold and / or the available bandwidth is less than the bandwidth threshold, the cross-modal shared semantic component among the multiple semantic components is transmitted to the receiving end through the wireless communication channel; or When the target service requirement is a task inference service, and the signal-to-noise ratio is greater than or equal to the signal-to-noise ratio threshold and the available bandwidth is greater than or equal to the bandwidth threshold, the cross-modal shared semantic component, the modality-specific semantic component, and the multimodal collaborative semantic component among the multi-type semantic components are transmitted to the receiving end through the wireless communication channel.

[0046] Channel parameters and target service requirements for each of the various semantic transmission modes can be pre-configured. For example... Figure 2 As shown, channel parameters can be acquired for channel-aware transmission. Multiple semantic transmission modes can include shared semantic transmission mode, full-task semantic transmission mode, and modality reconstruction transmission mode. Shared semantic transmission mode is... Figure 2 The transmission of the minimum necessary semantics is shown to transmit cross-modal shared semantic components; the full-task semantic transmission mode is... Figure 2 The transmission enhancement semantics shown are used to transmit cross-modal shared semantic components C, modality-specific semantic components U, and multimodal cooperative semantic components S; the modality reconstruction transmission mode is... Figure 2 The reconstructed information shown is used to transmit modal redundancy semantic components.

[0047] When the target service requirement is task inference, and the signal-to-noise ratio (SNR) in the channel parameters is less than the SNR threshold and / or the available bandwidth is less than the bandwidth threshold, a shared semantic transmission mode is adopted. This mode transmits the cross-modal shared semantic components among multiple semantic components to the receiver through the wireless communication channel, which can ensure the robustness of basic task inference under poor channel conditions and limited bandwidth.

[0048] When the target service requirement is task-based inference, if the signal-to-noise ratio (SNR) and available bandwidth are both greater than or equal to the SNR threshold and the bandwidth threshold, it indicates good channel quality and sufficient bandwidth resources. In this case, a full-task semantic transmission mode can be adopted, simultaneously transmitting cross-modal shared semantic components, modality-specific semantic components, and multi-modal collaborative semantic components to maximize the inference accuracy of downstream tasks. Alternatively, when the target service requirement is also task-based inference, if the SNR and available bandwidth are both greater than or equal to the SNR threshold and the bandwidth threshold, the task utility gain can be estimated. If the estimated task utility gain exceeds the transmission cost, cross-modal shared semantic components, modality-specific semantic components, and multi-modal collaborative semantic components can be transmitted simultaneously. If the estimated task utility gain does not exceed the transmission cost, a shared semantic transmission mode can be used, transmitting only the cross-modal shared semantic components.

[0049] In some other embodiments of this application, the step of transmitting at least one type of semantic component among the multiple types of semantic components to the receiving end through the wireless communication channel according to the channel parameters of the wireless communication channel and the target service requirements includes: when the target service requirement is a mode reconstruction service and the occupied bandwidth of the modal redundancy semantic component is less than the available bandwidth in the channel parameters, transmitting the modal redundancy semantic component among the multiple types of semantic components to the receiving end through the wireless communication channel.

[0050] When the target service requirement is modal reconstruction, a modal reconstruction transmission mode can be selected. In this case, the available bandwidth in the channel parameters needs to be greater than or equal to the bandwidth occupied by the modal redundancy semantic component to be transmitted. Then, the modal redundancy semantic component can be sent to the receiving end through the wireless communication channel to achieve accurate reconstruction of the multimodal raw data.

[0051] When the target service requirement is modal reconstruction, if the channel quality is poor (e.g., the signal-to-noise ratio is less than the signal-to-noise ratio threshold) or the available bandwidth is less than the bandwidth occupied by the modal redundancy semantic components, you can choose to reduce the reconstruction quality (e.g., reduce the amount of data in the modal redundancy semantic components), enhance channel protection, or delay transmission.

[0052] When comparing the bandwidth occupied by the data to be transmitted with the available bandwidth, the bandwidth compression ratio can also be used to make a comprehensive judgment. For example, the available bandwidth compression ratio of the channel can be determined, and the bandwidth compression ratio occupied by the data to be transmitted can be compared with the available bandwidth compression ratio. Figure 4a and Figure 4b This is a schematic diagram of channel-aware adaptive semantic transmission mode selection in an embodiment of this application. Figure 4a The channel environment represents additive white Gaussian noise. Figure 4b This indicates the channel environment in which Rayleigh fading occurs. Figure 4a and Figure 4bIn the diagram, yellow indicates shared semantic transport mode, and blue indicates full-task semantic transport mode, such as... Figure 4a and Figure 4b As shown, when the target service requirement is task-oriented inference, if the signal-to-noise ratio (SNR) in the channel parameters is less than the SNR threshold and / or the available bandwidth is less than the bandwidth threshold (or the available bandwidth compression rate is less than the bandwidth compression rate threshold), a shared semantic transmission mode is used to transmit cross-modal shared semantic components. If the SNR in the channel parameters is greater than or equal to the SNR threshold and the available bandwidth is greater than or equal to the bandwidth threshold (or the available bandwidth compression rate is greater than or equal to the bandwidth compression rate threshold), a full-task semantic transmission mode is used to transmit cross-modal shared semantic components, modality-specific semantic components, and multimodal collaborative semantic components. Mode 1 represents the shared semantic transmission mode, and Mode 2 represents the full-task semantic transmission mode.

[0053] In some embodiments of this application, the training process of the semantic decoupling network includes: The shared semantic branch is trained based on the training data. During the training process, the network parameters of the shared semantic branch are updated based on the first mutual information lower bound loss function and the mutual information upper bound loss function. The network parameters of the shared semantic branch are frozen, and the unique semantic branch and the collaborative semantic branch are trained based on the training data. During the training process, the network parameters of the unique semantic branch and the collaborative semantic branch are updated based on the second mutual information lower bound loss function, the supervised contrast loss function, and the orthogonal regularization loss function. Based on the training data, the shared semantic branch, the unique semantic branch, the collaborative semantic branch, and the reconstructed residual semantic branch are trained.

[0054] The semantic decoupling network adopts a three-stage orthogonal regularization training mechanism, which performs hierarchical training on the encoding networks (each branch) of the four semantic components respectively. The loss function of each semantic component is constructed using the lower bound of InfoNCE mutual information and the upper bound of CLUB mutual information to constrain each semantic component to have no information overlap and no redundant interference.

[0055] In the first stage, the shared semantic branch is trained based on the training data. During training, the network parameters of the shared semantic branch are updated using the first mutual information lower bound (InfoNCE) loss function and the mutual information upper bound (NCE-CLUB) loss function. The InfoNCE loss maximizes the mutual information consistency of cross-modal shared semantics, while the NCE-CLUB loss constrains the task-conditional mutual information. A task loss function can also be used to optimize the task relevance of the shared semantics. Simultaneously, a warm-up strategy is employed to dynamically adjust the regularization hyperparameters (network parameters of the shared semantic branch). During training, shared semantic pairs within the same sample (data pairs composed of shared semantic components from different modalities in the same multimodal data) are used as positive samples, while unpaired shared semantics within a batch (shared semantic components from different modalities in different multimodal data, such as shared semantic components of text data in the first multimodal data and shared semantic components of audio data in the second multimodal data) are used as negative samples. These are used to form an optimizable lower bound for cross-modal shared semantic consistency. Here, InfoNCE can be used for self-supervised cross-modal pairing learning.

[0056] The shared semantic loss function in the first stage includes a lower bound mutual information loss function, an upper bound mutual information loss function, and a task loss function. This shared semantic loss function can be expressed as follows:

[0057] Wherein, represents the shared semantic loss function. Denotes the loss function of the lower bound of the first mutual information. These represent the shared semantic components of the two modalities (using two modalities as an example here). The loss function represents the upper bound of mutual information. Let Y represent the task loss function, and let Y represent the task information. This represents the regularization hyperparameter. This represents the task loss hyperparameter.

[0058] When dynamically adjusting the regularization hyperparameter using a warm-up strategy, initially set the regularization hyperparameter β to 0 or a small value during the early stages of training, training only the InfoNCE mutual information lower bound and task prediction term. After Tw epochs, gradually increase it to the maximum value βmax.

[0059] This avoids the loss of effective task information in the shared semantic component ZC due to excessively strong task-independent dependency compression when the early network (critic) has not converged. This indicates the total number of training rounds.

[0060] In the second stage, the network parameters of the shared semantic branch are frozen, and the unique semantic branch and the collaborative semantic branch are trained based on the training data. During the training process, the task effectiveness of the unique semantic component is maximized based on the second mutual information lower bound (conditional InfoNCE) loss function, the higher-order interaction capability of the collaborative semantic component is optimized by the supervised contrastive loss function, and at the same time, the orthogonal regularization loss function is applied to eliminate information redundancy between different semantic components.

[0061] The loss function in the second stage includes a unique semantic loss function and a collaborative semantic loss function. The unique semantic loss function includes a second mutual information lower bound (conditional InfoNCE) loss function, which is used to maximize the task effectiveness of the unique semantic components. The unique semantic loss function is expressed as follows:

[0062] in, This represents a unique semantic loss function. This represents the second mutual information lower bound loss function, i.e., the InfoNCE loss under the condition of shared semantic components. This indicates a unique semantic component. and shared semantic components Apply orthogonal regularization constraints. This represents the adjustment factor.

[0063] The collaborative semantic loss function includes a supervised contrastive loss function, used to optimize the higher-order interaction capabilities of collaborative semantic components. The collaborative semantic loss function is expressed as follows:

[0064] in, Represents the collaborative semantic loss function. This represents the supervised contrastive loss function. Represents the cooperative semantic components Other semantic components Orthogonal regularization constraints are applied to (both unique and shared semantic components) to constrain collaborative semantic components. It cannot degenerate into a copy of a unique semantic component or a shared semantic component. , Indicates the adjustment factor. Indicates the cooperative gain. , Indicates that no cooperative semantic components are used. The task prediction cross-entropy loss function at that time. This indicates a prediction result without synergy. Indicates the addition of collaborative semantic components The subsequent task prediction cross-entropy loss function, This indicates a prediction with synergy. It is typically derived from shared semantic components and unique semantic components, and is used to measure whether it relies solely on cross-modal shared semantic components. Modality-specific semantic components At that time, the model corresponds to the true label. The prediction error. Typically derived from the complete task semantics, it is used to measure the model's performance on the ground truth labels after simultaneously using cross-modal shared semantic components, modality-specific semantic components, and multimodal collaborative semantic components. The prediction error.

[0065] In the third stage, all network parameters are unfrozen, and the reconstruction residual semantic branch is trained. In this stage, all network parameters are updated, namely the network parameters of the shared semantic branch, the unique semantic branch, the collaborative semantic branch, and the reconstruction residual semantic branch. At the same time, the parameters of the encoder of each modality and the decoding network of the receiver are also updated. In this stage, the modality reconstruction accuracy is optimized by using the L2 reconstruction loss function, the cosine similarity loss function, and the KL divergence loss function.

[0066] The third stage involves fine-tuning the end-to-end service conditions, which allows for joint optimization of all trainable network parameters, including the encoders of each modality, the branches of the semantic decoupling network, the channel encoder, and the channel decoder. The loss function in the third stage includes the reconstruction residual loss function, and the reconstruction parameter loss function comprises the L2 reconstruction loss function, the cosine similarity loss function, and the KL divergence loss function, which can be expressed as follows:

[0067] in, This represents the reconstruction residual loss function. This represents the L2 reconstruction loss function and the cosine similarity loss function, used to constrain reconstruction accuracy. This represents the KL divergence loss function, used to constrain the distribution of residual latent variables. , Indicates hyperparameters, This indicates that the task prediction information in the modal redundancy semantic component ZR is suppressed by the upper bound of CLUB. Used to suppress the predictability of modal redundancy semantic components (ZR) for task labels.

[0068] The above orthogonal regularization constraint is calculated as follows: After centering the batch features, the correlation between the feature matrices of different semantic components is calculated using the Frobenius norm. The correlation value is used as the regularization loss to constrain the subspaces of each semantic component to be mutually orthogonal, thus preventing semantic information overlap. The feature representations Za and Zb (initial features of two different modalities) of each multimodal data within the batch are then centered.

[0069] in, Let represent the decentralized representation of the initial features of mode a in the i-th multimodal data. Let B represent the initial features of mode a in the i-th multimodal data, and let B represent the number of samples in the batch. This represents the decentralized representation of the initial features of mode b in the i-th multimodal data set. This represents the initial feature of mode b in the i-th multimodal data.

[0070] Then, the intra-batch cross-covariance penalty for the two types of semantic features is calculated, i.e., the orthogonal regularization constraint:

[0071] The smaller the value of this formula, the weaker the linear correlation between the two semantic subspaces, thereby reducing the information overlap between the shared semantic component ZC, the unique semantic component ZU, and the collaborative semantic component ZS.

[0072] By employing a three-stage hierarchical training approach and orthogonal regularization constraints, the four semantic components are ensured to be free of redundancy and highly discriminative, significantly improving the model's semantic learning efficiency and completely resolving the training defects of semantic overlap and chaotic representation in traditional models.

[0073] In the third stage of training, the receiver decoder also needs to be trained. This can be done by simulating Gaussian white noise, Rayleigh fading, and 3GPP TDL-C / CDL-D multipath fading channel transmission environments to achieve noise and fading superposition transmission of channel symbols. This can adapt to various wireless channels such as AWGN (Additive White Gaussian Noise), Rayleigh fading, and 3GPP multipath fading, and is compatible with multiple modal data such as text, image, and audio. It has strong versatility and engineering feasibility. Compared with traditional solutions, it significantly improves task inference accuracy under the same bandwidth constraints and greatly reduces bandwidth consumption under the same channel conditions, which has extremely high economic value and industry promotion value.

[0074] Figure 5 This is a flowchart of a semantic communication method provided in an embodiment of this application. This semantic communication method is applied at the receiving end, such as... Figure 5As shown, the method may include steps 510 to 520.

[0075] Step 510: The receiving end sends at least one type of semantic component through the wireless communication channel. The at least one type of semantic component is at least one of the multiple semantic components corresponding to the multimodal data. The multiple semantic components are approximately pairwise orthogonal to each other.

[0076] The receiver first performs equalization, denoising, or channel decoding on the received channel symbols based on the estimated channel state to obtain a semantic representation. Then, it reads the semantic transmission mode identifier in the semantic frame header to determine at least one type of semantic component corresponding to the current payload, that is, to determine which specific semantic component it is.

[0077] The multiple semantic components are obtained by semantic decoupling of multimodal data. For specific decoupling methods, please refer to the above embodiments, which will not be repeated here.

[0078] Optionally, the multiple semantic components include cross-modal shared semantic components, modality-specific semantic components, multimodal collaborative semantic components, and modality redundant semantic components. Specifically, cross-modal shared semantic components carry unified core task information across all modalities; modality-specific semantic components carry effective task information specific to a single modality; multimodal collaborative semantic components carry high-order interaction information generated by multimodal fusion; and modality redundant semantic components carry detailed information used only for original modality reconstruction and unrelated to downstream tasks.

[0079] Step 520: Perform the operation corresponding to the target business requirement on the at least one type of semantic component.

[0080] Perform the operation corresponding to the target business requirement on at least one type of semantic component, that is, send it into the corresponding task inference head or reconstruction decoder according to the semantic transmission mode representation to perform the corresponding operation.

[0081] In some embodiments of this application, performing the operation corresponding to the target business requirement on the at least one type of semantic component includes one of the following: If the at least one type of semantic component includes the cross-modal shared semantic component, then the first task reasoning operation is performed on the cross-modal shared semantic component; If the at least one type of semantic component includes the cross-modal shared semantic component, the modality-specific semantic component, and the multimodal collaborative semantic component, then a second task reasoning operation is performed on the cross-modal shared semantic component, the modality-specific semantic component, and the multimodal collaborative semantic component; If the at least one type of semantic component includes the modal redundancy semantic component, then a modal recovery operation is performed on the modal redundancy semantic component to obtain the recovered multimodal data.

[0082] When at least one type of semantic component is a cross-modal shared semantic component, the cross-modal shared semantic component can be used to perform the first task reasoning operation, which is the basic task reasoning operation, to obtain the first task result.

[0083] If at least one type of semantic component includes cross-modal shared semantic components, modality-specific semantic components, and multimodal collaborative semantic components, the second task reasoning operation can be performed using cross-modal shared semantic components, modality-specific semantic components, and multimodal collaborative semantic components. The second task reasoning operation is an enhanced task reasoning operation, and the second task result is obtained.

[0084] When at least one type of semantic component is a modal redundancy semantic component, modal recovery operation can be performed using the modal redundancy semantic component to obtain the recovered multimodal data.

[0085] The semantic communication method provided in this application embodiment receives at least one type of semantic component sent by the transmitting end through a wireless communication channel. The at least one type of semantic component is at least one of the multiple semantic components corresponding to multimodal data. The multiple semantic components are approximately pairwise orthogonal. The method performs operations corresponding to the target service requirements on the at least one type of semantic component. By decoupling the semantic features of multimodal data into multiple semantic components that are approximately pairwise orthogonal, the method solves the problem of highly entangled semantic features in the single overall semantic representation in the prior art. It will not cause global diffusion due to channel noise, fading, or other damage. It can improve the transmission performance in low signal-to-noise ratio and bandwidth-limited scenarios, improve the robustness of transmission, and select the semantic components to be transmitted based on channel parameters and target service requirements. It can adaptively adjust the transmission strategy, improve the scenario adaptability, and promote the large-scale application of multimodal semantic communication.

[0086] The semantic communication method provided in this application embodiment can be implemented through a multimodal semantic communication implementation system, which includes a transmitter module, a wireless channel module, and a receiver module.

[0087] The transmitter module includes a multimodal feature coding unit, an information-theoretic semantic decoupling unit, an adaptive transmission control unit, and a channel coding unit. The multimodal feature coding unit is used to extract features from various heterogeneous modal data. The information-theoretic semantic decoupling unit is used to decompose and encode four types of orthogonal semantic components. The adaptive transmission control unit is used to collect channel parameters, match and switch semantic transmission modes. The channel coding unit is used to convert semantic representations (at least one type of semantic component) into transmissible channel symbols.

[0088] The wireless channel module is used to simulate Gaussian white noise, Rayleigh fading, and 3GPP TDL-C / CDL-D multipath fading channel transmission environments during training, and to complete the transmission of channel symbols with noise and fading superposition.

[0089] The receiver module includes a channel decoding unit, a semantic parsing unit, a task reasoning unit, and a modality reconstruction unit. The channel decoding unit is used to restore the semantic representation of the transmission. The semantic parsing unit is used to distinguish and parse four types of semantic components. The task reasoning unit is used to complete intelligent decision reasoning based on task-type semantics. The modality reconstruction unit is used to reconstruct the original multimodal data based on redundant semantic components.

[0090] The information-theoretic semantic decoupling unit includes a shared semantic extraction submodule, a unique semantic extraction submodule, a collaborative semantic extraction submodule, and a redundant semantic extraction submodule. These four submodules operate in parallel, and their output semantic components are mutually orthogonal. Each submodule corresponds to an independent loss optimization function and training logic. For details, please refer to the above embodiment; further elaboration is omitted here.

[0091] The adaptive transmission control unit has a built-in channel parameter detection submodule, service requirement identification submodule, and mode switching decision submodule, which can sense the dynamic changes of the channel and the service type (target service requirement) in real time, and realize millisecond-level adaptive switching of transmission mode.

[0092] This application's embodiments can be widely applied to multiple core fields such as 6G full-domain intelligent communication, vehicle-to-everything (V2X) intelligent sensing communication, industrial internet equipment data transmission, remote smart healthcare communication, immersive multimedia interaction, integrated air-space-ground communication, emergency rescue communication, and intelligent security monitoring. In V2X scenarios, it can achieve highly reliable, low-bandwidth transmission of multimodal data from vehicle vision, radar sensing, and voice interaction, adapting to complex vehicle communication environments with high-speed movement and rapid channel attenuation, ensuring the accuracy and real-time performance of autonomous driving decisions. In remote healthcare scenarios, it can achieve layered transmission of multimodal information such as medical images, patient vital signs data, and diagnostic voice, balancing the accuracy of diagnostic task inference with the need for complete reconstruction of medical data. In industrial internet scenarios, it can adapt to multi-source heterogeneous data from industrial cameras, sensors, and control commands, achieving priority transmission of key semantics in industrial environments with strong electromagnetic interference and limited bandwidth. It can also adapt to extreme channel scenarios such as satellite communication, maritime communication, and emergency offline communication, meeting the needs of intelligent semantic communication across multiple scenarios, multiple services, high reliability, and high efficiency.

[0093] Compared to existing technologies, this application possesses several significant technical advantages and beneficial effects. First, based on information theory PID decomposition and the MNI principle, this application achieves four-layer orthogonal decoupling of multimodal semantics, completely overcoming the entanglement defects of traditional integrated semantic fusion. Each semantic component is transmitted independently without interference, and channel impairments only affect local semantic components without global spread, significantly improving transmission robustness in complex fading channels and low signal-to-noise ratio extreme scenarios, effectively solving the problem of poor channel adaptability in traditional solutions. Second, this application achieves refined hierarchical filtering of semantic information, filtering task-irrelevant redundant semantics and prioritizing the transmission of high-value core semantics, greatly improving bandwidth resource utilization and adapting to the stringent bandwidth-constrained, low-latency transmission scenarios of 6G. Third, through three-stage hierarchical training and orthogonal regularization constraints, this application ensures that the four types of semantic components are free of redundancy and have high discriminative power, significantly improving the model's semantic learning efficiency and completely solving the training defects of semantic overlap and chaotic representation in traditional models. Fourth, the three-mode adaptive transmission strategy designed in this application can dynamically adapt to channel conditions and service requirements, while simultaneously being compatible with both intelligent task inference and original modal reconstruction services, breaking the limitation of traditional semantic communication services being singular. Fifth, this application is compatible with various wireless channels such as AWGN, Rayleigh fading, and 3GPP multipath fading, and is compatible with multiple modal data such as text, image, and audio. It has extremely strong versatility and engineering feasibility. Compared with traditional solutions, it significantly improves task inference accuracy under the same bandwidth constraints and greatly reduces bandwidth usage under the same channel conditions, possessing extremely high economic and industry promotion value.

[0094] It should be noted that the semantic communication method provided in this application embodiment can be executed by a semantic communication device, or a control module within the semantic communication device for executing the loading semantic communication method. This application embodiment uses the execution of the loading semantic communication method by a semantic communication device as an example to illustrate the semantic communication method provided in this application embodiment.

[0095] Figure 6 This is a schematic diagram of a semantic communication device provided in an embodiment of this application. This semantic communication device is applied at the sending end, such as... Figure 6 As shown, the device includes: The feature extraction module 610 is used to extract features from the corresponding modal data in the multimodal data through the encoder corresponding to each modality, so as to obtain the initial features corresponding to each modality. The semantic decoupling module 620 is used to decouple the initial features corresponding to each modality into multiple semantic components through a semantic decoupling network. The multiple semantic components are approximately pairwise orthogonal to each other. The semantic transmission module 630 is used to transmit at least one type of semantic component among the multiple types of semantic components to the receiving end through the wireless communication channel according to the channel parameters of the wireless communication channel and the target service requirements.

[0096] Optionally, the multiple semantic components include cross-modal shared semantic components, modality-specific semantic components, multimodal collaborative semantic components, and modality-redundant semantic components; the semantic decoupling network includes shared semantic branches, unique semantic branches, collaborative semantic branches, and reconstructed residual semantic branches; The semantic decoupling module is specifically used for: The cross-modal shared semantic components are extracted from the initial features corresponding to each modality through the shared semantic branches; Using the unique semantic branches, each modality-specific semantic component is extracted from the initial features corresponding to each modality under the condition of cross-modal shared semantic components; The initial features corresponding to each modality are fused through the collaborative semantic branch to obtain the multimodal collaborative semantic components; The modality reconstruction information is extracted from the initial features corresponding to each modality through the reconstruction residual semantic branch, and the modality redundancy semantic component is obtained.

[0097] Optionally, the channel parameters include signal-to-noise ratio and available bandwidth; The semantic transmission module is used for: When the target service requirement is a task inference service, and the signal-to-noise ratio is less than the signal-to-noise ratio threshold and / or the available bandwidth is less than the bandwidth threshold, the cross-modal shared semantic component among the multiple semantic components is transmitted to the receiving end through the wireless communication channel; or When the target service requirement is a task inference service, and the signal-to-noise ratio is greater than or equal to the signal-to-noise ratio threshold and the available bandwidth is greater than or equal to the bandwidth threshold, the cross-modal shared semantic component, the modality-specific semantic component, and the multimodal collaborative semantic component among the multi-type semantic components are transmitted to the receiving end through the wireless communication channel.

[0098] Optionally, the semantic transmission module is used for: When the target service requirement is modal reconstruction service and the bandwidth occupied by the modal redundancy semantic component is less than or equal to the available bandwidth in the channel parameters, the modal redundancy semantic component among the multiple semantic components is transmitted to the receiving end through the wireless communication channel.

[0099] Optionally, the training process of the semantic decoupling network includes: The shared semantic branch is trained based on the training data. During the training process, the network parameters of the shared semantic branch are updated based on the first mutual information lower bound loss function and the mutual information upper bound loss function. The network parameters of the shared semantic branch are frozen, and the unique semantic branch and the collaborative semantic branch are trained based on the training data. During the training process, the network parameters of the unique semantic branch and the collaborative semantic branch are updated based on the second mutual information lower bound loss function, the supervised contrast loss function, and the orthogonal regularization loss function. Based on the training data, the shared semantic branch, the unique semantic branch, the collaborative semantic branch, and the reconstructed residual semantic branch are trained.

[0100] The semantic communication device provided in this application embodiment can achieve... Figure 1 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.

[0101] The semantic communication device provided in this application extracts features from the corresponding modal data in multimodal data through encoders corresponding to each modality to obtain initial features corresponding to each modality. A semantic decoupling network decouples the initial features corresponding to each modality into multiple semantic components, which are approximately pairwise orthogonal. Based on the channel parameters of the wireless communication channel and the target service requirements, at least one semantic component is transmitted to the receiving end through the wireless communication channel. By decoupling the semantic features of multimodal data into multiple semantic components that are approximately pairwise orthogonal, the device solves the problem of highly entangled semantic features in the single overall semantic representation in the prior art. It avoids global diffusion due to channel noise, fading, and other impairments, improving transmission performance in low signal-to-noise ratio and bandwidth-constrained scenarios, enhancing transmission robustness. Furthermore, it allows selection of the semantic components to be transmitted based on channel parameters and target service requirements, enabling adaptive adjustment of transmission strategies and improving scenario adaptability. This promotes the large-scale application of multimodal semantic communication.

[0102] Figure 7 This is a schematic diagram of a semantic communication device provided in an embodiment of this application. This semantic communication device is applied at the receiving end, such as... Figure 7 As shown, the device includes: The receiving module 710 is used to receive at least one type of semantic component sent by the transmitting end through the wireless communication channel. The at least one type of semantic component is at least one type of semantic component among the multiple types of semantic components corresponding to multimodal data. The multiple types of semantic components are approximately pairwise orthogonal to each other. The operation execution module 720 is used to perform operations corresponding to the target business requirements on the at least one type of semantic components.

[0103] Optionally, the multiple semantic components include cross-modal shared semantic components, modality-specific semantic components, multimodal collaborative semantic components, and modality-redundant semantic components.

[0104] Optionally, the operation execution module is used to perform one of the following: If the at least one type of semantic component includes the cross-modal shared semantic component, then the first task reasoning operation is performed on the cross-modal shared semantic component; If the at least one type of semantic component includes the cross-modal shared semantic component, the modality-specific semantic component, and the multimodal collaborative semantic component, then a second task reasoning operation is performed on the cross-modal shared semantic component, the modality-specific semantic component, and the multimodal collaborative semantic component; If the at least one type of semantic component includes the modal redundancy semantic component, then a modal recovery operation is performed on the modal redundancy semantic component to obtain the recovered multimodal data.

[0105] The semantic communication device provided in this application embodiment can achieve... Figure 5 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.

[0106] The semantic communication method provided in this application embodiment receives at least one type of semantic component sent by the transmitting end through a wireless communication channel. The at least one type of semantic component is at least one of the multiple semantic components corresponding to multimodal data. The multiple semantic components are approximately pairwise orthogonal. The method performs operations corresponding to the target service requirements on the at least one type of semantic component. By decoupling the semantic features of multimodal data into multiple semantic components that are approximately pairwise orthogonal, the method solves the problem of highly entangled semantic features in the single overall semantic representation in the prior art. It will not cause global diffusion due to channel noise, fading, or other damage. It can improve the transmission performance in low signal-to-noise ratio and bandwidth-limited scenarios, improve the robustness of transmission, and select the semantic components to be transmitted based on channel parameters and target service requirements. It can adaptively adjust the transmission strategy, improve the scenario adaptability, and promote the large-scale application of multimodal semantic communication.

[0107] The semantic communication device in this application embodiment can be a device, or a component, integrated circuit, or chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. For example, mobile electronic devices can be mobile phones, tablets, laptops, PDAs, in-vehicle electronic devices, wearable devices, ultra-mobile personal computers (UMPCs), netbooks, or personal digital assistants (PDAs), etc., while non-mobile electronic devices can be servers, network-attached storage (NAS), personal computers (PCs), televisions (TVs), ATMs, or self-service machines, etc. This application embodiment does not impose specific limitations.

[0108] The semantic communication device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.

[0109] Optionally, embodiments of this application also provide an electronic device, including a processor, a memory, and a program or instructions stored in the memory and executable on the processor. When the program or instructions are executed by the processor, they implement the various processes of the above-described semantic communication method embodiments and achieve the same technical effects. To avoid repetition, they will not be described again here.

[0110] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0111] Figure 8 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.

[0112] The electronic device 800 includes, but is not limited to, components such as: a radio frequency unit 801, a network module 802, an audio output unit 803, an input unit 804, a sensor 805, a display unit 806, a user input unit 807, an interface unit 808, a memory 809, and a processor 810. The input unit 804 may include an image processor 8041 and a microphone 8042; the display unit 806 may include a display panel 8061; and the user input unit 807 may include a touch panel 8071 and other input devices (such as a keyboard) 8072.

[0113] Those skilled in the art will understand that the electronic device 800 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 810 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 8 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0114] The memory 809 stores a program or instruction that can run on the processor. When the program or instruction is executed by the processor 810, it implements the various processes of the above-described semantic communication method embodiment and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0115] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described semantic communication analysis method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0116] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0117] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described semantic communication method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0118] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0119] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0120] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0121] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A semantic communication method, characterized in that, Applied to the sending end, the method includes: By using the encoder corresponding to each modality, feature extraction is performed on the corresponding modal data in the multimodal data to obtain the initial features corresponding to each modality; The initial features corresponding to each modality are decoupled into multiple semantic components by a semantic decoupling network, and the multiple semantic components are approximately mutually orthogonal to each other. Based on the channel parameters of the wireless communication channel and the target service requirements, at least one of the multiple semantic components is transmitted to the receiving end through the wireless communication channel.

2. The method according to claim 1, characterized in that, The multiple semantic components include cross-modal shared semantic components, modality-specific semantic components, multimodal collaborative semantic components, and modality-redundant semantic components; the semantic decoupling network includes shared semantic branches, unique semantic branches, collaborative semantic branches, and reconstructed residual semantic branches; The process of decoupling the initial features corresponding to each modality into multiple semantic components through a semantic decoupling network includes: The cross-modal shared semantic components are extracted from the initial features corresponding to each modality through the shared semantic branches; Using the unique semantic branches, each modality-specific semantic component is extracted from the initial features corresponding to each modality under the condition of cross-modal shared semantic components; The initial features corresponding to each modality are fused through the collaborative semantic branch to obtain the multimodal collaborative semantic components; Modality reconstruction information is extracted from the initial features corresponding to each modality through the reconstruction residual semantic branch, and the modality redundancy semantic component is obtained.

3. The method according to claim 2, characterized in that, The channel parameters include signal-to-noise ratio and available bandwidth; The step of transmitting at least one semantic component from the multiple semantic components to the receiving end through the wireless communication channel according to the channel parameters of the wireless communication channel and the target service requirements includes: When the target service requirement is a task inference service, and the signal-to-noise ratio is less than the signal-to-noise ratio threshold and / or the available bandwidth is less than the bandwidth threshold, the cross-modal shared semantic component among the multiple semantic components is transmitted to the receiving end through the wireless communication channel; or When the target service requirement is a task inference service, and the signal-to-noise ratio is greater than or equal to the signal-to-noise ratio threshold and the available bandwidth is greater than or equal to the bandwidth threshold, the cross-modal shared semantic component, the modality-specific semantic component, and the multimodal collaborative semantic component among the multi-type semantic components are transmitted to the receiving end through the wireless communication channel.

4. The method according to claim 2, characterized in that, The step of transmitting at least one semantic component from the multiple semantic components to the receiving end through the wireless communication channel according to the channel parameters of the wireless communication channel and the target service requirements includes: When the target service requirement is modal reconstruction service and the bandwidth occupied by the modal redundancy semantic component is less than or equal to the available bandwidth in the channel parameters, the modal redundancy semantic component among the multiple semantic components is transmitted to the receiving end through the wireless communication channel.

5. The method according to claim 2, characterized in that, The training process of the semantic decoupling network includes: The shared semantic branch is trained based on the training data. During the training process, the network parameters of the shared semantic branch are updated based on the first mutual information lower bound loss function and the mutual information upper bound loss function. The network parameters of the shared semantic branch are frozen, and the unique semantic branch and the collaborative semantic branch are trained based on the training data. During the training process, the network parameters of the unique semantic branch and the collaborative semantic branch are updated based on the second mutual information lower bound loss function, the supervised contrast loss function, and the orthogonal regularization loss function. Based on the training data, the shared semantic branch, the unique semantic branch, the collaborative semantic branch, and the reconstructed residual semantic branch are trained.

6. A semantic communication method, characterized in that, Applied to the receiving end, the method includes: The receiving end sends at least one type of semantic component through a wireless communication channel. The at least one type of semantic component is at least one of the multiple semantic components corresponding to the multimodal data. The multiple semantic components are approximately pairwise orthogonal to each other. Perform the operation corresponding to the target business requirement on the at least one type of semantic component.

7. The method according to claim 6, characterized in that, The various semantic components include cross-modal shared semantic components, modality-specific semantic components, multimodal collaborative semantic components, and modality redundant semantic components.

8. The method according to claim 7, characterized in that, The operation corresponding to the target business requirement performed on the at least one type of semantic component includes one of the following: If the at least one type of semantic component includes the cross-modal shared semantic component, then the first task reasoning operation is performed on the cross-modal shared semantic component; If the at least one type of semantic component includes the cross-modal shared semantic component, the modality-specific semantic component, and the multimodal collaborative semantic component, then a second task reasoning operation is performed on the cross-modal shared semantic component, the modality-specific semantic component, and the multimodal collaborative semantic component; If the at least one type of semantic component includes the modal redundancy semantic component, then a modal recovery operation is performed on the modal redundancy semantic component to obtain the recovered multimodal data.

9. A semantic communication device, characterized in that, Applied to the sending end, including: The feature extraction module is used to extract features from the corresponding modal data in the multimodal data through the encoder corresponding to each modality, so as to obtain the initial features corresponding to each modality; The semantic decoupling module is used to decouple the initial features corresponding to each modality into multiple semantic components through a semantic decoupling network. The multiple semantic components are approximately pairwise orthogonal to each other. The semantic transmission module is used to transmit at least one type of semantic component among the multiple types of semantic components to the receiving end through the wireless communication channel according to the channel parameters of the wireless communication channel and the target service requirements.

10. A semantic communication device, characterized in that, Applied to the receiving end, including: The receiving module is used to receive at least one type of semantic component sent by the transmitting end through the wireless communication channel. The at least one type of semantic component is at least one type of semantic component among the multiple types of semantic components corresponding to the multimodal data. The multiple types of semantic components are approximately pairwise orthogonal to each other. The operation execution module is used to perform operations corresponding to the target business requirements on the at least one type of semantic components.

11. An electronic device, characterized in that, It includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the semantic communication method as described in any one of claims 1-5 or 6-8.

12. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the semantic communication method as described in any one of claims 1-5 or 6-8.