A token-based semantic communication method, electronic device and storage medium

By designing a token-based semantic communication method, leveraging the capabilities of point Transformers and differentiable modulators, the efficiency and reliability issues in point cloud transmission are resolved, achieving efficient and reliable point cloud transmission and overcoming the compatibility and performance bottlenecks of traditional methods.

CN120729800BActive Publication Date: 2025-11-04TSINGHUA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511141485.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-11-04
Estimated Expiration
2045-08-15

AI Technical Summary

Technical Problem

Existing technologies suffer from insufficient efficiency and reliability in point cloud transmission, especially in the transmission of large amounts of point cloud data. Existing communication networks are unable to meet the requirements, and traditional point cloud compression and deep learning-based JSCC methods are incompatible with digital communication systems.

Method used

A token-based semantic communication method is adopted, and a joint semantic-channel coding and modulation scheme is designed. By utilizing the capabilities of point Transformers, point tokens are mapped to a finite set of digital constellation points through master and slave encoders. Differentiable modulators are used to generate modulation tokens with strong semantic representation capabilities and robustness, thereby achieving efficient and reliable point cloud transmission.

Benefits of technology

It overcomes the cliff effect and compatibility issues, achieving efficient and reliable point cloud transmission, and improving the performance and robustness of point cloud reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120729800B_ABST
    Figure CN120729800B_ABST
Patent Text Reader

Abstract

The application provides a token-based semantic communication method, an electronic device and a storage medium, and relates to the field of semantic communication. The method comprises the following steps: a sending end in a semantic communication system processes an original point cloud to obtain point tokens; the point tokens are input into a main encoder to obtain main point cloud semantic features; the point tokens are input into an auxiliary encoder to obtain auxiliary point cloud semantic features; the main point cloud semantic features and the auxiliary point cloud semantic features are input into a differentiable modulator to obtain modulation tokens; a signal power normalization process is performed on the modulation tokens by a power normalizer to obtain to-be-sent modulation tokens; the to-be-sent modulation tokens are sent to a receiving end through a wireless channel; the receiving end demodulates the received modulation tokens to obtain received point cloud semantic features, decodes the received point cloud semantic features to obtain down-sampled point tokens, and up-samples the down-sampled point tokens to obtain reconstructed point clouds, so that token-based semantic communication is realized to complete efficient and reliable point cloud transmission.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of semantic communication, and in particular to a token-based semantic communication method, an electronic device and a storage medium. BACKGROUND

[0002] In the era of Artificial Intelligence (AI), the Transformer is the main backbone model choice in various tasks of various modalities. After achieving success in the field of Natural Language Processing (NLP) at the beginning, the Transformer, with its excellent ability to capture complex relationships between sequences, has been further applied to other fields such as audio, vision, 3D data and multimodal fields, and has shown superior performance in various tasks. In addition, the parallel computing efficiency and unprecedented scalability of the Transformer have given rise to the emergence of Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs). They have profoundly shaped the development paradigm of AI and constructed a generation-based information processing approach. Token, as a unified representation form of the input and output of the Transformer, is gradually becoming a new basic information unit. Data from different modalities are converted into tokens for processing. Token / s and Token / J are considered as key indicators of reasoning speed and energy efficiency. In order to fully tap the potential of deep learning models, a lot of research needs to be done on the processing method of Token. As important as processing Token is, effectively and reliably transmitting Token is also a necessary condition for establishing an integrated communication AI.

[0003] Point cloud is an important representation form of three-dimensional world, which represents a three-dimensional object through a set of three-dimensional (x, y, z) coordinates. The three-dimensional coordinates are referred to as point cloud geometry (Geometry) information. In addition to the geometry information, optional attribute (Attribute) information such as color, normal vector and reflectivity can also be attached to the coordinates. With the advancement of acquisition and display hardware and related algorithms, point cloud is increasingly used in immersive media, autonomous driving and robotics. However, the huge amount of data required for point cloud transmission poses a huge challenge to existing communication networks. Through calculation, it can be obtained that a dense point cloud containing 1 million points captured by a camera array at a speed of 30 frames per second, if its coordinates are represented by 32-bit floating-point numbers and the color channel is represented by 8-bit precision, the transmission bandwidth requirement is 3.6 Gbps.

[0004] Therefore, how to complete efficient and reliable point cloud transmission based on token semantic communication is crucial. SUMMARY

[0005] In order to overcome the above problems or at least partially solve the above problems, the present application provides a token-based semantic communication method, an electronic device and a storage medium.

[0006] The first aspect of the present application provides a token-based semantic communication method, the method comprising:

[0007] The sending end in the semantic communication system processes the original point cloud to obtain point tokens;

[0008] The point tokens are input into a main encoder in a main branch of the sending end to obtain main point cloud semantic features;

[0009] The point tokens are input into an auxiliary encoder in an auxiliary branch of the sending end to obtain auxiliary point cloud semantic features; the main encoder and the auxiliary encoder are used for extracting semantic features at different levels for the point tokens;

[0010] The main point cloud semantic features and the auxiliary point cloud semantic features are input into a differentiable modulator of the sending end to obtain modulation tokens;

[0011] The modulation tokens are subjected to signal power normalization processing by a power normalizer of the sending end to obtain to-be-sent modulation tokens;

[0012] The sending end sends the to-be-sent modulation tokens to a receiving end in the semantic communication system through a wireless channel;

[0013] The receiving end demodulates the received modulation tokens to obtain received point cloud semantic features, decodes the received point cloud semantic features to obtain down-sampled point tokens, and up-samples the down-sampled point tokens to obtain reconstructed point cloud.

[0014] The second aspect of the present application provides an electronic device, which comprises a memory, a processor and a computer program stored on the memory and executable on the processor, wherein the computer program is executed by the processor to implement the token-based semantic communication method according to the first aspect of the present application.

[0015] The third aspect of the present application provides a computer-readable storage medium, which stores a computer program, wherein the computer program is executed by a processor to implement the token-based semantic communication method according to the first aspect of the present application.

[0016] In the token-based semantic communication method provided by the application, in order to obtain a token (Token) representation form rich in semantic information and robust, instead of directly using the bit form of token (Token) index, a joint semantic-channel and modulation (JSCCM) scheme is designed in the semantic communication system of the application, mainly including: a Token encoder (i.e. primary and secondary encoders) and a differentiable modulator at the sending end, which are used for the primary and secondary encoders at the sending end, map the obtained point Token to a limited set of digital constellation points, to generate modulation Token (constellation points), and transmit the semantic Token with strong semantic representation ability and robustness, and perform point cloud reconstruction based on the received modulation Token at the receiving end. In this way, the primary and secondary encoders of the joint semantic-channel coding JSCC fully utilize the capabilities of point Transformer, and through the differentiable modulator, the advantages of the probability sampling method are combined, so that the output of the primary and secondary encoders is interpretable, thereby overcoming the cliff effect faced by current Token (Token) communication transmission index, traditional point cloud compression, separate coding, and the incompatibility between the deep learning-based JSCC method and the current digital communication system, and realizing token (Token)-based semantic communication to complete efficient and reliable point cloud transmission. BRIEF DESCRIPTION OF DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the application, the following will briefly introduce the drawings needed to be used in the description of the embodiments of the application. Obviously, the drawings in the following description are only some embodiments of the application, and for those skilled in the art, other drawings can also be obtained without creative labor.

[0018] Figure 1 is a step flow chart of a token-based semantic communication method according to an embodiment of the application;

[0019] Figure 2 is a schematic diagram of the model architecture of the encoder and the point transformer according to an embodiment of the application;

[0020] Figure 3 is a structural schematic diagram of a channel adapter according to an embodiment of the application;

[0021] Figure 4 is a structural schematic diagram of a rate allocator according to an embodiment of the application;

[0022] Figure 5 is a modulation process schematic diagram of a differentiable modulator according to an embodiment of the application;

[0023] Figure 6 is a schematic diagram of an overall framework of a semantic communication system according to an embodiment of the present application;

[0024] Figure 7 is a schematic diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0025] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative effort are within the scope of the present application.

[0026] Currently, the related point cloud compression technologies V-PCC, G-PCC and the point cloud compression method based on the deep learning model need to cascade channel coding. This separate coding method has a serious cliff effect, and the performance decreases greatly or even cannot be decoded at a low signal-to-noise ratio.

[0027] And the related technology using Token Communication directly represents the token (Token) output by the Tokenizer (tokenizer) as a bit, and generates an image at the receiving end. This is a simplified design of the token (Token) encoder and token (Token) decoder. In addition, this method also needs channel coding to protect the bit stream, which is a separate coding design. In the presence of bit errors, the error of some bits may seriously affect the token (Token) index, and thus affect the reconstruction performance at the receiving end.

[0028] In addition, the current point cloud semantic communication technology focuses on the design of the joint source-channel coding (Joint Source-Channel Coding, JSCC) structure. However, it still faces challenges to apply these JSCC designs to existing digital communication systems, because the current semantic communication systems are implemented with deep learning-based JSCC, and their encoder outputs are floating-point numbers. In these works, the floating-point numbers are paired two by two to form a constellation point, one of which represents the I-channel component and the other represents the Q-channel component. This method is difficult to implement with existing hardware and is incompatible with current communication protocols. To solve this challenge, it is necessary to map the JSCC output to a finite set of channel symbols, so that the semantic communication system can be compatible with existing digital communication systems. Existing researches have explored methods to quantize the JSCC output to bits, but after introducing digital modulation, these methods have encountered difficulties due to their non-differentiable nature.

[0029] Based on this, in order to at least partially solve one or more of the above-mentioned problems and other potential problems, this embodiment of the invention proposes a token-based semantic communication method to construct an effective token communication framework for point cloud geometric transmission, and to apply this framework to point cloud sources to achieve effective and robust point cloud transmission: In order to obtain a semantically rich and robust token representation, instead of directly using the bit form of token index, this embodiment designs a joint semantic-channel coding and modulation scheme in the token communication framework for the token encoder (i.e., master and auxiliary encoder) at the transmitting end to map point tokens to a finite set of digital constellation points to generate modulation tokens (constellation points), and transmit the semantically rich and robust modulation tokens (modulation tokens). The master and slave encoders of this Joint Semantic-Channel Coding (JSCC) fully utilize the capabilities of the point Transformer. The differentiable modulator combines the advantages of the probabilistic sampling method, making the outputs of the master and slave encoders interpretable. This overcomes the shortcomings of current token communication transmission indexing, traditional point cloud compression, the cliff effect faced by split coding, and the incompatibility between deep learning-based JSCC methods and current digital communication systems. It realizes token-based semantic communication to achieve efficient and reliable point cloud transmission.

[0030] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating the steps of a token-based semantic communication method according to an embodiment of the present invention. Figure 1 As shown, the token-based semantic communication method provided in this embodiment includes at least the following steps:

[0031] Step S11: The sending end in the semantic communication system processes the original point cloud to obtain point tokens.

[0032] In this embodiment, the semantic communication system includes a sender and a receiver. For the original point cloud, the sender can process the original point cloud to obtain a point token corresponding to the original point cloud. Alternatively, the sender includes a point tokenizer to convert the original point cloud into point tokens represented by vectors.

[0033] Step S12: Input the point token into the main encoder in the main branch of the sending end to obtain the main point cloud semantic features.

[0034] In this embodiment, the sending end further includes two branches: a main branch and an auxiliary branch. The main branch at least includes a main encoder. The sending end can input the point token into the main encoder in the main branch for semantic encoding to obtain a main point cloud semantic feature.

[0035] Step S13: inputting the point token into an auxiliary encoder in the auxiliary branch of the sending end to obtain an auxiliary point cloud semantic feature.

[0036] In this embodiment, the auxiliary branch at least includes an auxiliary encoder. The sending end can also input the point token into the auxiliary encoder in the auxiliary branch for semantic encoding to obtain an auxiliary point cloud semantic feature. The point token is input into the parallel main encoder and auxiliary encoder to obtain a point cloud semantic feature with richer semantic information, wherein the main encoder and the auxiliary encoder are used to extract semantic features at different levels (such as extracting geometric features at different levels) for the point token.

[0037] Step S14: inputting the main point cloud semantic feature and the auxiliary point cloud semantic feature into a differentiable modulator of the sending end to obtain a modulation token.

[0038] In this embodiment, the sending end further includes a differentiable modulator, which is used to generate a constellation point according to the semantic feature. After obtaining the main point cloud semantic feature and the auxiliary point cloud semantic feature, the main point cloud semantic feature and the auxiliary point cloud semantic feature can be input into the differentiable modulator of the sending end together to obtain a modulation token (modulation Token). In this embodiment, the modulation token (modulation Token) is the constellation point.

[0039] In this embodiment, in order to generate a token (Token) representation with rich semantic information and robustness, a joint semantic-channel coding and modulation (JSCCM) scheme is designed for the token (Token) encoder in the sending end: the JSCCM is composed of two parallel JSCC encoders and a differentiable modulator. It can be understood that the sending end at least includes a token (Token) encoder for further refining the point token to convert it into a modulation form suitable for transmission. The token (Token) encoder includes two parallel deep learning-based JSCC encoders, i.e., a parallel main encoder (main JSCC encoder) and an auxiliary encoder (auxiliary JSCC encoder), which are used to obtain the probability of the constellation point position, and then guide the differentiable modulator to generate the modulation Token (constellation point). After being processed by the main encoder, the auxiliary encoder and the differentiable modulator, the point token (Point Token) to be transmitted is converted into a modulation token (modulation Token) with fewer symbols, rich semantics and suitable for transmission.

[0040] Step S15: signal power normalization processing is performed on the modulated token by a power normalizer of the sending end, to obtain a to-be-sent modulated token.

[0041] In this embodiment, the sending end further comprises a power normalizer. After obtaining the modulated token output by the modulator, the sending end can perform signal power normalization processing on the modulated token by the power normalizer, to obtain a to-be-sent modulated token.

[0042] Step S16: the sending end sends the to-be-sent modulated token to a receiving end in the semantic communication system through a wireless channel.

[0043] In this embodiment, after obtaining the to-be-sent modulated token, the sending end can send the to-be-sent modulated token to the receiving end in the semantic communication system through a wireless channel.

[0044] Step S17: the receiving end demodulates the received modulated token, to obtain a received point cloud semantic feature, decodes the received point cloud semantic feature, to obtain a down-sampled point token, and up-samples the down-sampled point token, to obtain a reconstructed point cloud.

[0045] In this embodiment, the receiving end can receive the to-be-sent modulated token based on the wireless channel, to obtain a received modulated token. After obtaining the received modulated token, the receiving end can demodulate the received modulated token, to obtain a received point cloud semantic feature, then decode the received point cloud semantic feature, such as inputting the received point cloud semantic feature into a decoder (such as a JSCC decoder) in the receiving end for semantic decoding, to obtain a down-sampled point token, and after obtaining the down-sampled point token, up-sample the down-sampled point token, to obtain a reconstructed point cloud.

[0046] In an optional embodiment, the receiving end comprises a point token inverse generator (Point De-tokenizer). The receiving end can input the down-sampled point token into the point token inverse generator for a final reconstruction task, i.e., up-sampling, to obtain a reconstructed point cloud output by the point token inverse generator.

[0047] In the present embodiment, in order to obtain a token representation form rich in semantic information and robust, instead of directly using the bit form of token index, a joint semantic-channel coding and modulation scheme is designed in the semantic communication system of the present application, involving the primary and secondary encoders and the differentiable modulator of the sending end, to map the obtained point token to a limited set of digital constellation points for the primary and secondary encoders of the sending end to generate modulation tokens, and transmit the modulation tokens which are rich in semantic representation ability and robust, and reconstruct the point cloud based on the received modulation tokens at the receiving end. In this way, the present application fully utilizes the ability of point Transformer through the primary and secondary encoders of joint semantic-channel coding, and combines the advantages of the probability sampling method through the differentiable modulator, so that the output of the primary and secondary encoders is interpretable, thereby overcoming the defects of the current token communication transmission index, traditional point cloud compression, separate coding, facing cliff effect, deep learning-based JSCC method and current digital communication system difficult to be compatible, and realizing token-based semantic communication to complete efficient and reliable point cloud transmission.

[0048] In an optional embodiment, only considering the geometric information of the original point cloud, the original point cloud can be represented as where N is the number of points in the original point cloud, represents the three-dimensional coordinates of each point. The original point cloud is first processed by the point token generator of the sending end, which reduces the number of points in the original point cloud and enhances the feature dimension, and the output of the point token generator is the point token where represents the number of point tokens, is the embedding of Token. is the feature dimension of the point token. Subsequently, is sent into two parallel deep learning-based JSCC encoders, namely the primary and secondary JSCC encoders, to obtain point cloud semantic features richer in semantic information.

[0049] In combination with the above embodiments, in an implementation, the application further provides a token-based semantic communication method. In this embodiment, the main encoder and the auxiliary encoder are both Point Transformer-based JSCC encoders, both of which take the point tokens obtained from the Set Abstraction of the point token generator as input, and further represent the relationship between the point tokens through a point Transformer to enhance the representation ability of the tokens. Among them, the point Transformer no longer uses dot product to calculate attention score, but calculates the difference between query vector and key vector, and then performs nonlinear transformation through multi-layer perception (MLP) to obtain attention score, which is a kind of vector attention mechanism. In addition, since the point cloud itself contains position information, the position encoding of the point Transformer is obtained by calculating the difference of the point cloud coordinates and then performing nonlinear transformation through MLP.

[0050] The main encoder at least includes: a first point Transformer block, a Set Abstraction block, and a second point Transformer block connected in sequence; and the auxiliary encoder at least includes: a first Linear layer and a third point Transformer block connected in sequence. Among them, the first point Transformer block, the second point Transformer block and the third point Transformer block are all point Transformer blocks (Point Transformer Block) with the same structure and keeping the input and output dimensions consistent.

[0051] The Set Abstraction block in the main encoder is used to aggregate the features of the point tokens; the first point Transformer block in the main encoder and the third point Transformer block in the auxiliary encoder are respectively used to capture the relationship between different local regions of the point tokens; and the second point Transformer block in the main encoder is used to further capture the relationship between different local regions of the point tokens.

[0052] In the embodiment, the main encoder includes an additional point set abstraction block compared to the auxiliary encoder, and the point set abstraction block of the main encoder is used to aggregate the features of the point tokens. The first point Transformer block in the main encoder and the third point Transformer block in the auxiliary encoder are respectively used to capture the relationship between different local regions of the point tokens; the main encoder further includes an additional second point Transformer block compared to the auxiliary encoder, and the second point Transformer block is used to further capture the relationship between different local regions of the point tokens.

[0053] In addition, the main encoder and the auxiliary encoder respectively further include adjacent flatten layers and linear layers, which are used to distribute the semantic features of the point tokens on different constellation points.

[0054] In the embodiment, in addition to the first point Transformer block, the point set abstraction block and the second point Transformer block connected in sequence, the main encoder further includes adjacent flatten layers and linear layers. In addition to the first linear layer and the third point Transformer block connected in sequence, the auxiliary encoder further includes adjacent flatten layers and linear layers. The adjacent flatten layers and linear layers in the two encoders (the main encoder and the auxiliary encoder) are used to distribute as much as possible the semantic information of the point cloud (such as the point token) on different constellation points.

[0055] In addition, in an embodiment, the two parallel main encoders and auxiliary encoders can further include a multi-layer perception, which is used to adjust the feature dimension so that the output can be compatible with the subsequent modulation module.

[0056] In an embodiment, as shown in FIG. 6, Figure 2 Figure 2 is a schematic diagram of the model architecture of the encoder and the point transformer shown in an embodiment of the present application. Wherein, Figure 2 ​Figure (a) shows a schematic diagram of two parallel encoders, including a master encoder (i.e., the master JSCC encoder) and an auxiliary encoder (i.e., the auxiliary JSCC encoder). The master encoder includes: a first point transformer block (the aforementioned first point transformer block), a set abstraction block, a second point transformer block (the aforementioned second point transformer block), a multilayer perceptron (MLP), a flattened layer, a linear layer, and a ReLU layer, connected in sequence. The auxiliary encoder includes: a first linear layer, a third point transformer block (the aforementioned third point transformer block), a multilayer perceptron (MLP), a flattened layer, a linear layer, and a ReLU layer, connected in sequence.

[0057] like Figure 2 As shown in (a), the final outputs of the two parallel encoders are modulated into logits representing the probability of each constellation point location. and .

[0058] in, The parameters of the main encoder, The parameters of the main encoder; For tokens, Main encoder, The semantic features of the main point cloud output by the main encoder; For auxiliary encoder, The semantic features of the auxiliary point cloud output by the auxiliary encoder. Figure 2 In (a) of the middle, and These represent the number of modulation tokens in each main and auxiliary branch, respectively. M represents the number of constellation points in Quadrature Amplitude Modulation (QAM), which determines the number of available choices in the token codebook. for The coordinates, i.e., the position in space; and similar, This indicates that the sampled data was obtained through the Set Abstraction block in the main encoder. The coordinates of the center point.

[0059] Figure 2(c) in the diagram is a schematic diagram of the Point Transformer Block. In order to control the feature dimension and alleviate the gradient vanishing problem, a linear layer is added before and after the Point Transformer Layer in the Point Transformer Block, and the input and output are connected through residual connections to form a complete Point Transformer Block.

[0060] Figure 2 Figure (b) shows a schematic diagram of the Point Transformer Layer, which includes three linear layers, two multilayer perceptrons, and one aggregation layer. The linear layers... Linear layer and linear layer These represent the linear layers that project the point token onto the query vector, key vector, and value vector, respectively; the multilayer perceptron. and multilayer perceptron Both are multilayer perceptrons (MLPs) containing two linear layers and a ReLU activation function.

[0061] For example, the point converter layer in the first point converter block of the main encoder (i.e., the first point converter layer in the main encoder) can be represented as:

[0062] ;

[0063] in, For tokens; and Through in Top Obtained by performing a k-nearest neighbors (kNN) index; where, for The coordinates, i.e., the position in space, where j is the index number. , , yes The index number of the inner center point; The feature dimension of the dot token. , and These represent linear layers that project point tokens onto the query vector, key vector, and value vector, respectively. They are and The coordinates of the point, and It is the point token with index numbers i and j, and X is the original point cloud. and are multi-layer perceptrons comprising two linear layers and a ReLU activation function; is a Hadamard product (element-wise multiplication); softmax is a softmax function, i.e. Figure 2 is a flexible normalization exponential function in (b); is a positional encoding of a Point Transformer (i.e., Point Transformer); is an output of the first Point Transformer layer in the main encoder.

[0064] In combination with any of the above embodiments, the application further provides a token-based semantic communication method, in which, in addition to the above steps, steps S21 to S23 can be included, and the above step S14 can specifically include the following step S24:

[0065] Step S21: determining the signal-to-noise ratio of the wireless channel as the channel condition.

[0066] In this embodiment, two independent channel adapters can be embedded in the main and auxiliary branches of the sending end to realize the corresponding channel adaptation adaptive capability. Specifically, the signal-to-noise ratio of the wireless channel can be determined, and the signal-to-noise ratio of the wireless channel is input into the two independent channel adapters as the channel condition, so as to adjust the output of the main and auxiliary encoders according to the channel condition.

[0067] In an optional embodiment, the signal-to-noise ratio SNR of the wireless channel can be determined by the following formula:

[0068] ;

[0069] wherein, is the noise power of the noise of the wireless channel, is the signal power of the received modulated token in the wireless channel.

[0070] In an optional embodiment, the wireless channel can be an additive white Gaussian noise (AWGN) channel or a Rayleigh fading channel. The signal-to-noise ratios of these two types of wireless channels can also be calculated according to the above formula.

[0071] Step S22: inputting the main point cloud semantic feature and the channel condition into the main channel adapter in the main branch to obtain an adjusted main point cloud semantic feature.

[0072] In this embodiment, in addition to the main encoder, the main branch also includes a main channel adapter. After the main encoder obtains the main point cloud semantic feature, the main point cloud semantic feature and the channel condition can be input into the main channel adapter. The main channel adapter adjusts the main point cloud semantic feature according to the channel condition, so that it can better adapt to the transmission of the wireless channel, and then obtains the adjusted main point cloud semantic feature output by the main channel adapter.

[0073] Step S23: inputting the auxiliary point cloud semantic feature and the channel condition into an auxiliary channel adapter in the auxiliary branch to obtain an adjusted auxiliary point cloud semantic feature.

[0074] In this embodiment, in addition to the auxiliary encoder, the auxiliary branch also includes an auxiliary channel adapter. After the auxiliary encoder obtains the auxiliary point cloud semantic feature, the auxiliary point cloud semantic feature and the channel condition can be input into the auxiliary channel adapter. The auxiliary channel adapter adjusts the auxiliary point cloud semantic feature according to the channel condition, so that it can better adapt to the transmission of the wireless channel, and then obtains the adjusted auxiliary point cloud semantic feature output by the auxiliary channel adapter. The main channel adapter and the auxiliary channel adapter are two channel adapters with the same structure but independent.

[0075] Step S24: inputting the adjusted main point cloud semantic feature and the adjusted auxiliary point cloud semantic feature into a differentiable modulator of the sending end to obtain the modulation token.

[0076] In this embodiment, after obtaining the adjusted main point cloud semantic feature and the adjusted auxiliary point cloud semantic feature, the sending end can input the adjusted main point cloud semantic feature and the adjusted auxiliary point cloud semantic feature into a differentiable modulator. The differentiable modulator modulates the adjusted main point cloud semantic feature and the adjusted auxiliary point cloud semantic feature to obtain the modulation token.

[0077] In this embodiment, in order to enable the proposed model to adapt to different channel environments, a channel adapter is introduced. Since the encoder output represents the probability of a constellation point in the JSCCM scheme, this embodiment concatenates the JSCC output and the channel condition to generate refined JSCC output, instead of implicitly modulating the output feature at each layer of the JSCC model. Through this simple method, excellent performance is achieved under various channel conditions using a single model.

[0078] In an embodiment, the output of the main branch may be represented as:

[0079] ;

[0080] wherein, and respectively represent trainable parameters primary encoder and a secondary channel adapter with trainable parameters . representing channel conditions, is the original point cloud. Similarly, for the secondary branch, the output of the secondary branch can be represented as:

[0081] ;

[0082] where, and represent a secondary encoder with trainable parameters and a secondary channel adapter with trainable parameters , respectively, representing channel conditions.

[0083] After that, and are concatenated to obtain Y and fed into a differentiable modulator to generate modulated tokens . C is the complex field, represents a complex space with dimension , is the total number of modulated tokens generated by the modulator.

[0084] In an embodiment, as shown in Figure 3 , the channel adapter is a primary channel adapter. Figure 3 is a structural schematic diagram of a channel adapter according to an embodiment of the present application. Figure 3 The channel adapter in may be a primary channel adapter or a secondary channel adapter. When it is a primary channel adapter, its input is the primary point cloud semantic feature and the signal-to-noise ratio (SNR), and its output is the adjusted primary point cloud semantic feature ; when it is a secondary channel adapter, its input is the secondary point cloud semantic feature and the signal-to-noise ratio (SNR), and its output is the adjusted secondary point cloud semantic feature Figure 3 . In , the channel adapter uses a connection method to fuse the semantic information of the point cloud (the primary point cloud semantic feature or the secondary point cloud semantic feature ) with the signal-to-noise ratio (SNR), and uses the signal-to-noise ratio (SNR) to refine the features containing the position information of the constellation points (i.e., the primary point cloud semantic feature or the secondary point cloud semantic feature ). First, the primary point cloud semantic feature or the secondary point cloud semantic featureReshape, then Scaled & Repeat SNR (divide SNR by 10 to avoid numerical issues), to match the reshaped feature dimension. Concatenate the Scaled & Repeated SNR and the reshaped feature, pass through a multi-layer perceptron (MLP) including linear layer & linear rectification =, linear layer & linear rectification 2, and linear layer & linear rectification =, and finally pass through a sigmoid function to produce the final output: adjusted primary point cloud semantic feature or adjusted secondary point cloud semantic feature .

[0085] In combination with any of the above embodiments, the application also provides a token-based semantic communication method, in which, in addition to the above steps, steps S31 and S32 can be included, the above step S15 can specifically include the following step S33, and in addition to the above steps, step S34 can be included, and the "demodulation of the received modulation token by the receiving end" in the above step S17 can specifically include step S35:

[0086] Step S31: the sending end inputs the modulation token, the concatenated feature of the primary point cloud semantic feature and the secondary point cloud semantic feature into the rate allocator of the sending end, to obtain a mask.

[0087] In this embodiment, the sending end also includes a rate allocator to dynamically adjust the number of constellation points (modulation tokens) transmitted according to the semantic features of the original point cloud. Before the transmission of the modulation token, the rate allocator can generate a mask according to the outputs of the primary branch and the secondary branch to discard part of the modulation token. The rate allocator of this embodiment is between the differentiable modulator and the power normalizer.

[0088] The sending end can input the modulation token output by the differentiable modulator, and the concatenated feature of the primary point cloud semantic feature and the secondary point cloud semantic feature into the rate allocator of the sending end. The rate allocator can first obtain a mask according to the concatenated feature of the primary point cloud semantic feature and the secondary point cloud semantic feature, and the mask represents discarding part of the modulation token from the secondary branch.

[0089] Step S32: process the modulation token according to the mask to obtain the rate-allocated modulation token.

[0090] In this embodiment, after the rate allocator obtains the mask, it can process the input modulation token according to the mask, discard part of the input modulation token, and obtain the rate-allocated modulation token. In this embodiment, the modulation token is a constellation point.

[0091] In an optional embodiment, the rate allocator only adjusts the number of constellation points (modulation tokens) from the secondary branch, without affecting the number of constellation points (modulation tokens) from the primary branch.

[0092] Step S33: performing signal power normalization processing on the rate-allocated modulation tokens by a power normalizer of the sending end, to obtain the to-be-sent modulation tokens.

[0093] In this embodiment, the sending end inputs the rate-allocated modulation tokens output by the rate allocator into the power normalizer, and performs signal power normalization processing on the rate-allocated modulation tokens by the power normalizer, to obtain the to-be-sent modulation tokens.

[0094] Step S34: performing zero padding on the received modulation tokens by the receiving end, to obtain modulation tokens with fixed length.

[0095] In this embodiment, after the number of modulation tokens is adjusted by the rate allocator at the sending end, at the receiving end, the receiving end performs zero padding on the received modulation tokens obtained through the wireless channel, to obtain modulation tokens with fixed length. This is because the mask of this embodiment only discards the constellation points at the tail from the secondary branch, so zero padding can be applied at the receiving end without transmitting the mask.

[0096] Step S34: performing demodulation on the modulation tokens with fixed length by the receiving end.

[0097] In this embodiment, after obtaining the modulation tokens with fixed length, the receiving end performs demodulation on the constellation points on the I path and the constellation points on the Q path in the modulation tokens with fixed length, respectively, to obtain the received point cloud semantic feature. For example, the receiving end includes an I path demodulator and a Q path demodulator, the receiving end can perform demodulation on the constellation points on the I path in the modulation tokens with fixed length by the I path demodulator, to obtain the output of the I path demodulator, perform demodulation on the constellation points on the Q path in the modulation tokens with fixed length by the Q path demodulator, to obtain the output of the Q path demodulator, and then obtain the received point cloud semantic feature based on the output of the I path demodulator and the output of the Q path demodulator.

[0098] In an optional example, after the rate allocation operation of the rate allocator, the obtained rate-allocated modulation tokens (constellation point representation) are wherein, is the total number of modulation tokens after rate allocation, i.e., the number of modulation tokens sent in the wireless channel, is the number of modulation tokens output by the modulator, i.e., the total number of modulation tokens generated before rate allocation. The above process can be represented as: wherein, The spliced result of the main branch output and the auxiliary branch output (such as the spliced feature of the main point cloud semantic feature and the auxiliary point cloud semantic feature, or the spliced feature of the adjusted main point cloud semantic feature and the adjusted auxiliary point cloud semantic feature). is a modulator with parameters . denotes a rate allocator with parameters . It should be noted that is trainable, while denotes a pre-set standard digital constellation point as a token codebook. The output of the power normalizer: the modulated token to be transmitted is denoted as , will be transmitted through a wireless channel.

[0099] In combination with any of the above embodiments, in an implementation, the present application further provides a token-based semantic communication method. In the method, and the above step S31 can specifically include the following steps S41 to S45:

[0100] Step S41: performing dimension adjustment on the spliced feature of the main point cloud semantic feature and the auxiliary point cloud semantic feature to obtain a dimension-adjusted spliced feature.

[0101] In the present embodiment, in the rate allocator, the rate allocator can first perform dimension adjustment on the spliced feature of the main point cloud semantic feature and the auxiliary point cloud semantic feature to obtain a dimension-adjusted spliced feature, so that the max-pooling layer can run along the dimension representing the constellation point position probability. In an optional manner, under the condition that the main branch and the auxiliary branch respectively include a main channel adapter and an auxiliary channel adapter, the rate allocator can first perform dimension adjustment on the spliced feature of the adjusted main point cloud semantic feature and the adjusted auxiliary point cloud semantic feature to obtain a dimension-adjusted spliced feature.

[0102] Step S42: inputting the dimension-adjusted spliced feature into the max-pooling layer in the rate allocator to obtain a max-pooling processing result.

[0103] In the present embodiment, the rate allocator inputs the dimension-adjusted spliced feature into the max-pooling layer (MaxPooling layer) in the rate allocator to obtain a max-pooling processing result.

[0104] Step S43: inputting the max-pooling processing result into the progressive multilayer perceptron in the rate allocator until the feature dimension of the output result of the progressive multilayer perceptron reaches the number of target rate grades.

[0105] In this embodiment, after the rate allocator obtains the maximum pooling result, the maximum pooling result is input into the progressive multi-layer perceptron (progressive MIP) in the rate allocator until the feature dimension of the output result of the progressive multi-layer perceptron reaches the number of target rate bins, wherein the target rate bin is set in advance, and the target rate is the transmission rate of the modulation token that the sending end / receiving end hopes to set.

[0106] Step S44: According to the output result of the progressive multi-layer perceptron, the cutoff position of the mask is determined to obtain a one-hot vector.

[0107] In this embodiment, until the feature dimension of the output result of the progressive multi-layer perceptron in the rate allocator reaches the number of target rate bins, the output result of the progressive multi-layer perceptron in the rate allocator is obtained, and according to the output result of the progressive multi-layer perceptron, the cutoff position of the mask is determined to obtain a one-hot vector. In an optional implementation, the Gumbel-Max method can be used to determine the cutoff position of the mask according to the output result of the progressive multi-layer perceptron.

[0108] Step S45: The one-hot vector is converted into a hot vector as the mask, wherein 1 in the mask indicates that the corresponding modulation token is transmitted, and 0 in the mask indicates that the corresponding modulation token is not transmitted.

[0109] In this embodiment, the rate allocator can further convert the one-hot vector into a hot vector as a mask, and the mask includes a plurality of 1s and 0s, wherein 1 in the mask indicates that the corresponding modulation token is transmitted, 0 in the mask indicates that the corresponding modulation token is not transmitted, and the mask only discards the constellation points from the tail of the auxiliary branch, and the mask does not process the modulation tokens from the main branch.

[0110] In this embodiment, in view of the defect of the fixed rate transmission in the current point cloud semantic communication scheme, i.e., the same number of channel transmission symbols are allocated to point clouds with different semantics, a rate allocator is designed, and a mask strategy based on a cutoff position is proposed. This method only transmits the modulation tokens before the cutoff position, and in the receiving end, the blank symbols after the cutoff position are simply filled with zeros, and there is no need to transmit the mask.

[0111] In an example, after obtaining the cutoff position of the mask, the cutoff position is set to 1 and other positions are set to 0 based on the cutoff position of the mask to obtain a one-hot vector, wherein the target rate level includes multiple levels, such as 5 levels or 10 levels, and the number of positions of the one-hot vector is the number of levels of the target rate level. When converting the one-hot vector to a hot vector, the positions before and after the position of 1 in the one-hot vector are set to 1 and 0 respectively, so as to obtain the converted hot vector as the mask. For example, the mask is 11100, and the number of 1s accounts for 60% of the number of positions of the mask, which indicates that the first 60% of the constellation points (modulation tokens) in the auxiliary branch from the head to the tail need to be transmitted; the number of 0s accounts for 40% of the number of positions of the mask, which indicates that the first 40% of the constellation points (modulation tokens) in the auxiliary branch from the tail to the head need to be discarded.

[0112] In an embodiment, as shown in Figure 4 , Figure 4 is a structural schematic diagram of a rate allocator according to an embodiment of the present application. In Figure 4 , the rate allocator uses the Gumbel-Softmax technique to realize differentiable rate selection: first, the rate allocator adjusts the dimension of the spliced feature Y so that the max-pooling layer can run along the dimension representing the probability of the constellation point position. Then, the pooled result (the maximum pooling result output by the max-pooling layer) is passed through a progressive multi-layer perceptron. Each module in the progressive multi-layer perceptron reduces the output feature dimension to 1 / 4 of the input until the feature dimension of the output result of the progressive multi-layer perceptron reaches the number of the target rate level. In this embodiment, the rate is set to 5 levels. Subsequently, the rate allocator determines the cutoff position of the mask using the Gumbel-Max method (i.e., the Gumbel random variable (Gumbel R.V.) and the maximum value index function (Argmax operation) in Figure 4 ), so as to obtain a one-hot vector. The one-hot vector is further converted into a hot vector representing the mask to realize mask generation. In the mask, 1 indicates that the corresponding constellation point should be transmitted, and 0 indicates that it should not be transmitted. And the rate allocator only adjusts the number of constellation points from the auxiliary branch. In Figure 4 , and represent the modulation tokens in each main and auxiliary branch.

[0113] In combination with any of the above embodiments, in an implementation, the application further provides a token-based semantic communication method. In the method, in the case where the wireless channel is a Rayleigh fading channel, in addition to the above steps, step S51 can be further included, and the "demodulating the received modulated token" in S17 can specifically include step S52:

[0114] Step S51: The receiving end performs zero-forcing equalization processing on the received modulated token through an equalizer to obtain an equalized modulated token.

[0115] In the embodiment, in the case where the wireless channel is a Rayleigh fading channel, the receiving end will perform zero-forcing (ZF) equalization: the receiving end performs zero-forcing equalization processing on the received modulated token through an equalizer to obtain an equalized modulated token.

[0116] Step S52: The receiving end demodulates the equalized modulated token.

[0117] In the embodiment, in the case where the receiving end does not perform zero padding, the receiving end can demodulate the equalized modulated token. In the case where the receiving end performs zero padding, after obtaining the equalized modulated token, the receiving end performs zero padding on the equalized modulated token to obtain a modulated token with a fixed length, thereby demodulating the modulated token with a fixed length.

[0118] In an optional example, at the receiving end, the received constellation point (received modulated token) is represented as For an additive white Gaussian noise (AWGN) wireless channel, It can be represented as: .

[0119] Where n is a Gaussian noise with noise power , and I is the identity matrix . The signal power of the received modulated token in the wireless channel can be calculated as:

[0120] ; where is the number of modulated tokens sent in the wireless channel, denotes the expectation operation.

[0121] For a Rayleigh fading wireless channel, It can be represented as: ; where h is the channel gain between the sending end and the receiving end, and n is defined the same as in the AWGN scenario. The signal power of the received modulated token in the wireless channel should be calculated as:

[0122] ;

[0123] In the Rayleigh fading scenario, the receiver will perform zero-forcing equalization, as shown in the following formula:

[0124] ;

[0125] wherein, is the output of the equalizer, is the complex conjugate of h. When the transmitter contains a rate allocator, the receiver will perform a zero-padding operation. By padding zeros after to restore the number of symbols to , so that the modulation token with a fixed length matches the subsequent demodulation model. Subsequently, the modulation token with a fixed length is demodulated on the I and Q paths respectively, completing the conversion from the constellation point to the received point cloud semantic feature . The received point cloud semantic feature is further processed in the JSCC decoder to obtain the down-sampled point token . The above process can be represented as:

[0126] ;

[0127] wherein, represents the zero-padding process, which does not need to transmit a mask from the transmitter. represents the demodulator (including the I path demodulator and the Q path demodulator) with trainable parameters , and represents the JSCC decoder with trainable parameters . Finally, the down-sampled point token is up-sampled by the point token inverse generator to obtain the reconstructed point cloud .

[0128] In combination with any of the above embodiments, in an implementation, the application further provides a token-based semantic communication method. In the method, the step S14 can specifically include steps S61 to S65:

[0129] Step S61: Splicing the main point cloud semantic feature and the auxiliary point cloud semantic feature to obtain a spliced feature Y, wherein the i-th row of the spliced feature Y contains the position information of the i-th constellation point.

[0130] In this embodiment, the sending end can splice the main point cloud semantic feature and the auxiliary point cloud semantic feature to obtain the spliced feature Y; in the case of including a main-auxiliary channel adapter, the sending end can splice the adjusted main point cloud semantic feature and the adjusted auxiliary point cloud semantic feature to obtain the spliced feature Y. The i-th row of the spliced feature Y of this embodiment contains the position information of the i-th constellation point (modulation token), and i is an integer greater than 0.

[0131] Step S62: for the i-th row of the spliced feature Y, a soft output of the i-th constellation point position probability is obtained using the Gumbel-Softmax method.

[0132] In this embodiment, for the i-th row of the spliced feature Y, the differentiable modulator can use the Gumbel-Softmax method to obtain a soft output of the i-th constellation point position probability.

[0133] Step S63: the inner product of the soft output of the i-th constellation point position probability and the token codebook of the modulator is calculated to generate an initial constellation point position.

[0134] In this embodiment, the token codebook of the modulator is defined in advance, such as As the token codebook of the modulator, it is composed of the standard coordinates of the MQAM modulation scheme. Then the inner product of the soft output of the i-th constellation point position probability and the token codebook of the modulator is calculated to generate the initial constellation point position of the i-th constellation point.

[0135] Step S64: the distance between the initial constellation point position and the token codebook is calculated.

[0136] In this embodiment, after obtaining the initial constellation point position of the i-th constellation point, the distance between the initial constellation point position of the i-th constellation point and the token codebook needs to be calculated. Specifically, the token codebook of the modulator includes multiple token codebooks, and therefore the distance between the initial constellation point position of the i-th constellation point and each token codebook needs to be calculated.

[0137] Step S65: according to the distance and the token codebook, the Q-path coordinate and the I-path coordinate of the i-th constellation point are obtained to generate the modulation token of the i-th constellation point.

[0138] In this embodiment, the modulator can obtain the Q-path coordinate and the I-path coordinate of the i-th constellation point according to the distance between the initial constellation point position of the i-th constellation point and the token codebook, and the token codebook of the modulator, and finally generate the modulation token of the i-th constellation point.

[0139] In an optional embodiment, the The i-th row of the concatenation feature Y, denoted as Y i, contains the position information of the i-th constellation point. In addition, let , which means the first components of Y i represent the position logits of the i-th constellation point for the I-arm component , while the last components of Y i represent the position logits of the i-th constellation point for the Q-arm component . Taking the I-arm component of the i-th constellation point as an example, in the forward stage, the I-arm component soft output of the i-th constellation point position probability is obtained using the Gumbel-Softmax method , as follows:

[0140] ;

[0141] wherein is the j-th element of , and is the j-th element of . is sampled from Gumbel(0, 1), which is the Gumbel-Softmax trick that allows reparameterized sampling for discrete distributions. It converts the process of sampling from a discrete distribution into a process of finding the maximum element of a vector with added randomness (argmax). T is a temperature hyperparameter that controls the steepness of the distribution. is the k-th element of , and is sampled from Gumbel(0, 1), and exp represents the exponential operation with the natural constant e as the base number, for example, exp(2) is the 2nd power of e.

[0142] After that, based on the I-arm component soft output of the i-th constellation point position probability and the token codebook , the inner product is calculated to generate the initial constellation point position of the I-arm component of the i-th constellation point : . Then, based on the initial constellation point position of the I-arm component of the i-th constellation point and the token codebook , the distance d between the initial constellation point position of the I-arm component of the i-th constellation point and each token codebook is calculated: .

[0143] Then, based on the above distance d and the token codebook c, the quantization operation is performed as follows:

[0144] ;

[0145] in, It is the I-way coordinate of the i-th constellation point. for The j-th element in the array, T is the transpose, one-hot is the one-hot encoding operation, and argmin is the operation to get the minimum position index.

[0146] Similarly, using As input, the Q-path coordinates of the i-th constellation point are generated according to the above process. . and The pairs will form constellation points. At this point, the differentiable modulator will generate logits. Mapping to modulation Thus, the output Z of the modulator is obtained.

[0147] Since the argmin operation is not differentiable when performing quantization based on the distance d and the token codebook c, a soft quantization method is used in the reverse phase, and the soft output corresponding to the I-way coordinates of the i-th constellation point is calculated. as follows:

[0148] ;

[0149] in, for The k-th element in, and , They are respectively The j-th and k-th elements in the expression, exp represents the exponential operation with the natural constant e as the base. For example, exp(2) is e to the power of 2.

[0150] In an alternative embodiment, the computations of the forward and backward phases can be integrated as follows:

[0151] ;

[0152] in, This can represent the output of the i-th constellation point in the forward and backward phases of the I-way component. It is the I-way coordinate of the i-th constellation point. This is the soft output of the I-path coordinates of the i-th constellation point. Since the detach operation has no effect during the forward pass, therefore... In the reverse phase, the detach operation will... Separate from the computational graph. The gradient will pass through differentiable... To spread.

[0153] In one embodiment, such as Figure 5 As shown, Figure 5This is a schematic diagram illustrating the modulation process of a differentiable modulator according to an embodiment of the present invention. Figure 5 In this example, MQAM is considered as the modulation scheme, employing a differentiable modulation method combining Gumbel-Softmax and soft quantization. 64QAM is used, where M=64. The input to the differentiable modulator is the splicing features. ,in This represents the total number of modulation tokens generated before rate allocation. Unlike converting bits into constellation points, the modulator generates constellation points based on Y. After processing by the JSCC encoder and modulator, the point tokens to be transmitted are converted into modulation tokens with fewer symbols, richer semantics, and suitable for transmission.

[0154] exist Figure 5 In this context, the Gumbel-Softmax method can be used for the splicing feature Y, based on Gumbel random variables (Gumbel RV). The soft output of the constellation point position probability is obtained by using a soft normalized exponential function (i.e., softmax operation) of the temperature hyperparameter, and then the soft output of the constellation point position probability is calculated by inner product with a preset token codebook. The initial constellation point position is obtained.

[0155] Then, multiple distances d between the initial constellation point positions and the token codebook are calculated. During the forward pass, a non-differentiable argmin operation is performed based on the distances d and the token codebook (i.e., ...). Figure 5 The minimum index function in the algorithm is used to obtain the I-path and Q-path coordinates of the constellation points, and then the modulation token is obtained based on the I-path and Q-path coordinates of the constellation points. In the reverse process, a differentiable softmax operation is performed based on the distance d and the token codebook (i.e., ... Figure 5 The softmax operation is used to approximate the non-differentiable argmax operation, obtaining the soft outputs of the I-path coordinates and Q-path coordinates of the constellation points. Then, based on the soft outputs of the I-path coordinates and Q-path coordinates of the constellation points, a soft modulation token is obtained. The soft modulation token enables the gradient to be propagated during backpropagation.

[0156] In conjunction with any of the above embodiments, in one implementation, the present invention also provides a token-based semantic communication method. In addition to the steps described above, this method may further include steps S71 to S74:

[0157] Step S71: Input the original point cloud of the sample into the initial semantic communication system to obtain the point cloud after sample reconstruction.

[0158] In this embodiment, the semantic communication system is obtained based on training of an initial semantic communication system. In the training process of the initial semantic communication system, a sample original point cloud can be input into the initial semantic communication system to obtain a sample reconstructed point cloud output by the initial semantic communication system. The sample original point cloud is an original point cloud used in the training process, and the sample reconstructed point cloud is a reconstructed point cloud obtained in the training process. The structure of the initial semantic communication system can be the same as or similar to the structure of the semantic communication system in any of the foregoing embodiments. The manner in which the initial semantic communication system processes the sample original point cloud into the sample reconstructed point cloud is the same as or similar to the manner in which the semantic communication system in any of the foregoing embodiments processes the original point cloud into the reconstructed point cloud, and will not be described again.

[0159] Step S72: obtaining a distance loss based on the sample original point cloud and the sample reconstructed point cloud.

[0160] In this embodiment, the distance loss can be obtained based on the sample original point cloud X and the sample reconstructed point cloud In an optional manner, the distance loss can be obtained by using Chamfer distance (CD) , which is used to measure the reconstruction performance. The distance metric is symmetric, differentiable, and computationally efficient:

[0161] ;

[0162] wherein, , denote the number of elements in the original point cloud X and the sample reconstructed point cloud , respectively, x is an element in the original point cloud X, is an element in the sample reconstructed point cloud .

[0163] Step S73: obtaining a rate allocation loss based on the number of modulation tokens after rate allocation corresponding to the sample original point cloud and the number of modulation tokens corresponding to the sample original point cloud.

[0164] In this embodiment, the initial semantic communication system includes a rate allocator, and the rate adaptation capability is considered. Therefore, the embodiment needs to add a loss for controlling the number of modulation tokens for transmission. At this time, the number of modulation tokens after rate allocation corresponding to the sample original point cloud is , and the number of modulation tokens (i.e., the number of modulation tokens output by the modulator) corresponding to the sample original point cloud is The rate allocation loss can be obtained based on and .

[0165] In an optional embodiment, the rate allocation loss is .

[0166] Step S74: updating the model parameters of the first model deployed in the sending end of the initial semantic communication system and the model parameters of the second model deployed in the receiving end of the initial semantic communication system based on the rate allocation loss and the distance loss, to obtain the semantic communication system.

[0167] In the embodiment, the rate allocation loss and the distance loss are used to obtain a total loss L, and the model parameters of the first model deployed in the sending end of the initial semantic communication system and the model parameters of the second model deployed in the receiving end of the initial semantic communication system are updated based on the total loss until the total loss converges, to obtain the trained first model and the trained second model, and further to obtain the semantic communication system, which includes the sending end with the trained first model and the receiving end with the trained second model.

[0168] In an optional example, the total loss L is: ; wherein, is a hyperparameter for balancing the reconstruction quality and the number of transmission modulation tokens, which can be freely set according to requirements, and is not limited.

[0169] In an optional example, the first model at least includes: primary and secondary encoders, a differentiable modulator in the initial semantic communication system; and the second model at least includes: a demodulator and a decoder. In some optional examples, the first model further includes: at least one of a primary and secondary channel adapter group and a rate allocator.

[0170] In an optional embodiment, when the initial semantic communication system does not include a rate allocator, the model parameters of the first model deployed in the sending end of the initial semantic communication system and the model parameters of the second model deployed in the receiving end of the initial semantic communication system are updated based on only the distance loss, to obtain the semantic communication system.

[0171] In an optional embodiment, the model with channel adaptation capability can be trained in a signal-to-noise ratio (SNR) interval. In the training process, the SNR value of each batch in each iteration is obtained by uniform sampling in the interval. In the training, signal power normalization processing is performed for all batches, so that For a Rayleigh fading channel, the channel gain h is generated according to In all experiments, the model for the Rayleigh fading channel is obtained by fine-tuning the model trained under the same SNR setting through an AWGN channel.

[0172] In an embodiment, asFigure 6 is shown, Figure 6 is a schematic diagram of the overall framework of a semantic communication system according to an embodiment of the present application. In Figure 6 , the semantic communication system includes a transmitting end and a receiving end. The transmitting end includes a point tokenizer, a token encoder, a channel adapter, a modulator, a rate allocator and a power normalizer. The point tokenizer is configured to convert the original point cloud X into point tokens represented by vectors . The token encoder is configured to further refine the tokens and convert them into a modulated form suitable for transmission. The token encoder is composed of two parallel deep learning-based JSCC encoders (a main encoder in the main branch and a secondary encoder in the secondary branch) to obtain the probability of constellation point positions, which in turn guides the modulator to generate modulated tokens (constellation points). The channel adapter includes a main channel adapter in the main branch and a channel adapter in the secondary branch, which are configured to adjust the outputs of the main encoder and the secondary encoder based on the channel conditions of the wireless channel . The output of the main branch is , and the output of the secondary branch is . Then and are concatenated to obtain Y and sent to the modulator (differentiable) to generate modulated tokens Z. Then the modulated tokens Z and Y are input into the rate allocator to obtain the rate-allocated modulated tokens . The rate-allocated modulated tokens are then input into the power normalizer for power normalization to obtain the to-be-transmitted modulated tokens . The to-be-transmitted modulated tokens are then transmitted to the receiving end through the wireless channel.

[0173] The receiving end includes an equalizer, zero padding, a demodulator (including an I-channel demodulator and a Q-channel demodulator), a decoder and a point token inverse generator. The receiving end obtains the received modulated tokens based on the wireless channel, and then processes them through the equalizer to obtain . Then is zero-padded to obtain modulated tokens with a fixed length . The modulated tokens with a fixed length are then input into the demodulator to demodulate the constellation points on the I-channel and the Q-channel, respectively, to obtain the received point cloud semantic features . Then is processed in the decoder (e.g., a JSCC decoder) to obtain down-sampled point tokens . The final reconstruction task is performed by the point token inverse generator to obtain the reconstructed point cloud .

[0174] wherein, Figure 6 The modules in dashed boxes represent that these modules are optional: rate allocator and channel adapter are necessary when the relevant adaptive feature is needed. If Rayleigh fading channel is considered, equalizer will be used. And when adaptive transmission is done by rate allocation, the receiving end will do constellation point zero padding accordingly.

[0175] Based on this, in the semantic communication system of the embodiment, in order to obtain a semantic information-rich and robust token Token representation form, instead of directly using the bit form of the token Token index, a joint semantic-channel coding and modulation scheme JSCCM is designed for the token encoder to generate the modulated token (constellation point). The encoder of the joint semantic-channel coding fully utilizes the ability of the point Transformer. The modulator combines the advantages of the probability sampling method, in which the output of the JSCCM is interpretable, and the benefits of distance-weighted soft quantization, which can explore more available modulation positions. Based on the proposed JSCCM, a rate allocator and a channel adapter are introduced to adaptively generate the modulated token according to the semantics of the point token and the channel condition. By integrating the rate allocator and the channel adapter into the JSCCM framework, the developed system is no longer fixed rate, and can adapt to different channel conditions using a single model, achieving high-quality transmission.

[0176] And compared with the traditional G-PCC+LDPC separate coding method, the method proposed in the above embodiment has a reconstruction quality gain of more than 2dB under the same number of transmission symbols, and has a reconstruction performance gain of more than 1dB compared with the JSCC method based on deep learning.

[0177] It should be noted that the semantic communication system proposed by the present application is an end-to-end semantic communication system for point cloud geometry transmission, but in fact the semantic communication system is not only suitable for point cloud transmission, but also suitable for text, audio, image / video, three-dimensional data, multi-modal, etc.

[0178] It should be noted that for the method embodiments, in order to simply describe, they are all expressed as a series of action combinations, but those skilled in the art should know that the embodiments of the present application are not limited to the order of the described actions, because according to the embodiments of the present application, certain steps can be performed in other order or at the same time. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions involved are not necessarily necessary for the embodiments of the present application.

[0179] Based on the same inventive concept, another embodiment of the present application provides a computer readable storage medium, having stored thereon a computer program, which, when executed by a processor, implements the steps of the token-based semantic communication method according to any one of the above embodiments of the present application.

[0180] Based on the same inventive concept, another embodiment of the present application provides an electronic device, such as Figure 7 as shown in FIG. 8, Figure 7 is a schematic diagram of an electronic device according to an embodiment of the present application. The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the token-based semantic communication method according to any one of the above embodiments of the present application when executed.

[0181] Each of the embodiments in the present specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be mutually referred to.

[0182] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, device, or computer program product. Therefore, the embodiments of the present application can be in the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present application can be in the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0183] The embodiments of the present application are described with reference to flowcharts and / or block diagrams of the method, terminal device (system), and computer program product according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, so that the instructions executed by the computer or other programmable data processing terminal device produce the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in one flow or multiple flows and / or blocks Figure 1 The functions specified in one flow or multiple flows and / or blocks

[0184] These computer program instructions can also be stored in a computer readable storage medium, which can cause the computer or other programmable data processing terminal device to work in a specific manner, so that the instructions stored in the computer readable storage medium produce a product including instruction devices, which implement the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in one flow or multiple flows and / or blocksFigure 1 the function specified in the one or more blocks.

[0185] These computer program instructions can also be loaded into computer or other programmable data processing terminal devices, so that a series of operation steps are performed on the computer or other programmable terminal devices to generate a computer implemented process, so that the instructions executed on the computer or other programmable terminal devices provide a process for implementing the flow Figure 1 the flow or flows and / or blocks Figure 1 Figure 1 the function specified in the one or more blocks.

[0186] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications to these embodiments once they get the basic inventive concept. Therefore, the appended claims are intended to cover all the changes and modifications falling within the scope of the embodiments of the present application.

[0187] Finally, it should be noted that the relational terms herein, such as first and second, and the like, are used solely to distinguish one from another entity or action, without necessarily requiring or implying any such actual relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or terminal device. An element proceeded by "comprises a... " does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or terminal device that comprises the element.

[0188] The above provides a token-based semantic communication method, an electronic device and a storage medium. The principles and implementation manners of the present application are described by using specific examples. The above description of the embodiments is only used to help understand the method and core idea of the present application. For those skilled in the art, according to the idea of the present application, the specific implementation manners and application ranges can be changed. The above description of the present application should not be understood as a limitation.

Claims

1. A token-based semantic communication method, characterized by, The method comprises: The sending end in the semantic communication system processes the original point cloud to obtain point tokens; The point tokens are input into a main encoder in a main branch of the sending end to obtain main point cloud semantic features; The point tokens are input into an auxiliary encoder in an auxiliary branch of the sending end to obtain auxiliary point cloud semantic features; the main encoder and the auxiliary encoder are used for extracting semantic features at different levels for the point tokens; The main point cloud semantic features and the auxiliary point cloud semantic features are input into a differentiable modulator of the sending end to obtain modulation tokens; The modulation tokens are subjected to signal power normalization processing by a power normalizer of the sending end to obtain to-be-sent modulation tokens; The sending end sends the to-be-sent modulation tokens to a receiving end in the semantic communication system through a wireless channel; The receiving end demodulates the received modulation tokens to obtain received point cloud semantic features, decodes the received point cloud semantic features to obtain down-sampled point tokens, and up-samples the down-sampled point tokens to obtain reconstructed point cloud.

2. The token-based semantic communication method of claim 1, wherein, The main encoder at least comprises a first point Transformer block, a point set abstraction block, and a second point Transformer block connected in sequence; and the auxiliary encoder at least comprises a first linear layer and a third point Transformer block connected in sequence; The point set abstraction block in the main encoder is used for aggregating features of the point tokens; The first point Transformer block in the main encoder and the third point Transformer block in the auxiliary encoder are respectively used for capturing relationships between different local regions of the point tokens; The second point Transformer block in the main encoder is used for further capturing relationships between different local regions of the point tokens; The main encoder and the auxiliary encoder respectively further comprise adjacent flattening layers and linear layers, which are used for distributing semantic features of the point tokens on different constellation points.

3. The token-based semantic communication method of claim 1, wherein, The method further comprises: The signal-to-noise ratio of the wireless channel is determined and used as a channel condition; The main point cloud semantic features and the channel condition are input into a main channel adapter in the main branch to obtain adjusted main point cloud semantic features; The auxiliary point cloud semantic features and the channel condition are input into an auxiliary channel adapter in the auxiliary branch to obtain adjusted auxiliary point cloud semantic features; The main point cloud semantic features and the auxiliary point cloud semantic features are input into the differentiable modulator of the sending end to obtain the modulation tokens, comprising: The adjusted main point cloud semantic features and the adjusted auxiliary point cloud semantic features are input into the differentiable modulator of the sending end to obtain the modulation tokens.

4. The token-based semantic communication method of claim 1, wherein, The method further comprises: The sending end inputs the modulation tokens, splicing features of the main point cloud semantic features and the auxiliary point cloud semantic features into a rate allocator of the sending end to obtain a mask, the mask representing discarded part of the modulation tokens from the auxiliary branch; The modulation tokens are processed according to the mask to obtain rate-allocated modulation tokens; and The method further comprises: The sending end inputs the modulation tokens, splicing features of the main point cloud semantic features and the auxiliary point cloud semantic features into a rate allocator of the sending end to obtain a mask, the mask representing discarded part of the modulation tokens from the auxiliary branch; The modulation tokens are processed according to the mask to obtain rate-allocated modulation tokens; and Signal power normalization processing is performed on the modulation token by a power normalizer of the sending end to obtain a to-be-sent modulation token, including: Signal power normalization processing is performed on the modulation token by a power normalizer of the sending end to obtain a to-be-sent modulation token, including: The method further includes: The receiving end performs zero padding on the received modulation token to obtain a modulation token with a fixed length; The receiving end performs demodulation on the received modulation token, including: The receiving end performs demodulation on the modulation token with a fixed length.

5. The token-based semantic communication method of claim 4, wherein, The sending end inputs the spliced features of the main point cloud semantic features and the auxiliary point cloud semantic features into a rate allocator of the sending end to obtain a mask, including: Dimension adjustment is performed on the spliced features of the main point cloud semantic features and the auxiliary point cloud semantic features to obtain dimension-adjusted spliced features; The dimension-adjusted spliced features are input into a max pooling layer in the rate allocator to obtain a max pooling processing result; The max pooling processing result is input into a progressive multi-layer perceptron in the rate allocator until the feature dimension of the output result of the progressive multi-layer perceptron reaches the number of target rate grades; According to the output result of the progressive multi-layer perceptron, a cutoff position of the mask is determined to obtain a one-hot vector; The one-hot vector is converted into a hot vector as the mask, wherein 1 in the mask indicates that the corresponding modulation token is transmitted, and 0 in the mask indicates that the corresponding modulation token is not transmitted.

6. The token-based semantic communication method of claim 4, wherein, The method further includes: Inputting a sample original point cloud into an initial semantic communication system to obtain a sample reconstructed point cloud; Based on the sample original point cloud and the sample reconstructed point cloud, a distance loss is obtained; Based on the number of modulation tokens corresponding to the rate allocation of the sample original point cloud and the number of modulation tokens of the sample original point cloud, a rate allocation loss is obtained; Based on the rate allocation loss and the distance loss, the model parameters of a first model deployed in a sending end of the initial semantic communication system and the model parameters of a second model deployed in a receiving end of the initial semantic communication system are updated to obtain the semantic communication system; the semantic communication system includes: a sending end in which a trained first model is deployed, and a receiving end in which a trained second model is deployed.

7. The token-based semantic communication method according to any one of claims 1 to 6, characterized in that, In the case where the wireless channel is a Rayleigh fading channel, the method further includes: The receiving end performs zero-forcing equalization processing on the received modulation token by an equalizer to obtain an equalized modulation token; The receiving end performs demodulation on the received modulation token, including: The receiving end performs demodulation on the equalized modulation token.

8. The token-based semantic communication method according to any one of claims 1 to 6, characterized in that, The main point cloud semantic features and the auxiliary point cloud semantic features are input into a differentiable modulator of the sending end to obtain a modulation token, including: The main point cloud semantic features and the auxiliary point cloud semantic features are spliced to obtain spliced features Y, and the i-th row of the spliced features Y contains position information of an i-th constellation point; For the i-th row of the stitching feature Y, a soft output of the i-th constellation point position probability is obtained using a Gumbel-Softmax method; An inner product of the soft output of the i-th constellation point position probability and a token codebook of the modulator is calculated to generate an initial constellation point position; A distance between the initial constellation point position and the token codebook is calculated; According to the distance and the token codebook, Q-path coordinates and I-path coordinates of the i-th constellation point are obtained to generate a modulation token of the i-th constellation point; i is an integer greater than 0.

9. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The computer program, when executed by the processor, implements the token-based semantic communication method according to any one of claims 1 to 8.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program, when executed by the processor, implements the token-based semantic communication method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Semantic codec training method, semantic codec transmission method and semantic codec training system for point cloud transmission

    CN117135179A

  • Pre-training model determination method and device, equipment and storage medium

    CN117745944A