A semantic communication method, device and system for large-scale access

Through the combination of deep air computing and semantic communication, the problems of high computational complexity and unstable signal detection in large-scale access scenarios are solved, efficient information fusion and intelligent task processing are achieved, and the performance and reliability of the Internet of Things system are improved.

CN119324763BActive Publication Date: 2025-08-08BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411436749.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-15
Publication Date
2025-08-08
Estimated Expiration
2044-10-15

AI Technical Summary

Technical Problem

The existing semantic communication methods have high computational complexity, high energy consumption and unstable signal detection in large-scale access scenarios. The performance is degraded especially when channel conditions are poor, making it impossible to effectively handle the connection and information processing of a large number of devices in the Internet of Things.

Method used

Using a method of combining deep air computing and semantic communication, user information is obtained through the source side and semantic encoder is extracted, and the semantic information is converted into a binary code stream to be transmitted in the wireless channel. It is directly used for intelligent tasks at the receiving end. The weight parameters of the air computing encoder are optimized by using OFDM modulation and end-to-end training to achieve efficient fusion and processing of multimodal information.

Benefits of technology

It reduces the complexity of the communication system, improves the information fusion efficiency, reduces the computing pressure at the receiver, improves the accuracy of signal detection and the reliability of the system, and especially shows excellent performance under low signal-to-noise ratio conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119324763B_ABST
    Figure CN119324763B_ABST
Patent Text Reader

Abstract

The present invention discloses a semantic communication method, device, and system for large-scale access, including: obtaining user information at the source end and extracting semantic information using a semantic encoder; converting the high-dimensional features of the semantic information into a binary code stream transmitted in a physical channel; channel-encoding the converted semantic features and transmitting them to a wireless channel; aggregating local features on the same subcarrier of the semantic features; directly using the aggregated semantic features for intelligent tasks; solving the weight parameters of the over-the-air computation encoder through end-to-end joint training; and jointly optimizing and solving the semantic codec and over-the-air computation encoder. The technical solution of the present invention can effectively handle the connection and information processing of a large number of devices in the Internet of Things, achieving more efficient interconnection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of wireless communications, and in particular relates to a semantic communication method, device and system for large-scale access. Background Art

[0002] With the rapid development of IoT and AI technologies, the number of IoT devices is rapidly increasing, and IoT applications are also developing towards greater intelligence, efficiency, and personalization. For example, in urban flood warning systems, densely deployed sensors capture environmental information such as precipitation and water levels in the form of images, radar point clouds, and satellite signals. This massive amount of multimodal data is then transmitted to monitoring centers to train AI-based flood warning prediction models. With the proliferation of these AI-based, task-oriented applications, wireless communications will face enormous transmission pressure, and the computational workload of receivers will also increase significantly.

[0003] To alleviate transmission pressure, semantic communication, a novel communication paradigm, has attracted widespread attention. In semantic communication, the sender only needs to transmit semantic information relevant to the task or source, rather than all information, to the receiver. Unlike traditional bit-based communication methods, which focus on accurately transmitting symbols, semantic communication addresses the question of accurately conveying the meaning of the content while ignoring irrelevant information. As a result, semantic communication can significantly improve communication efficiency.

[0004] Semantic communication can be categorized into unimodal multi-user semantic communication and multimodal multi-user semantic communication. In the case of unimodal multi-user semantic communication, a multi-user semantic communication system for collaborative object recognition dynamically adjusts weights to fuse individual semantic features into global features. Multiple semantic features are then used for recognition, reducing system latency. In the case of multimodal multi-user semantic communication, a multi-user semantic communication approach that supports multimodal data transmission proposes a Transformer-based semantic communication framework for image retrieval, machine translation, and visual question answering tasks. Although these studies have made some progress, they share a common drawback: multi-user data fusion at the receiver follows the communication-first, computation-later paradigm. To facilitate signal fusion processing, algorithms such as zero forcing (ZF), successive interference cancellation (SIC), and parallel interference cancellation (PIC) are required to individually separate and recover the signals from each user. However, the computational complexity of these algorithms increases exponentially with the number of access nodes, placing significant processing and computational pressure on the receiver, resulting in high energy consumption and communication latency. Furthermore, signal detection accuracy cannot be guaranteed under poor channel conditions, thus impacting the reliability and performance of the communication system. Summary of the Invention

[0005] The technical problem to be solved by this invention is to provide a semantic communication method, device, and system for large-scale access. These methods not only effectively handle the connection and information processing of a large number of devices in the Internet of Things (IoT), achieving more efficient interconnection, but also process information from different users in different modalities, enabling multi-source perception and processing. Furthermore, by combining deep air computing with semantic communication and fully exploiting the orthogonality of semantic domains, they can fuse multimodal and multi-user signals during transmission. The fused information can be directly used for subsequent intelligent tasks at the receiving end, eliminating the need for multi-user detection.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] A semantic communication method for large-scale access, comprising:

[0008] Step S1: Obtain user information at the source end and extract semantic information using a semantic encoder;

[0009] Step S2: converting the high-dimensional features of the semantic information into a binary code stream transmitted in a physical channel;

[0010] Step S3: channel-encoding the converted semantic features and sending them to the wireless channel;

[0011] Step S4: Aggregate local features on the same subcarrier of the semantic features;

[0012] Step S5: directly use the aggregated semantic features for intelligent tasks; wherein the intelligent tasks include semantic segmentation, classification, and object detection;

[0013] Step S6: solving the weight parameters of the air calculation encoder through end-to-end joint training;

[0014] Step S7: jointly optimize and solve the semantic codec and the air computing encoder.

[0015] Preferably, in step S4, OFDM modulation is used to transmit M-dimensional semantic features on M subcarriers, and the signal transmitted through the channel is:

[0016]

[0017] Among them, h k is the channel coefficient of the channel between each user node and the receiver, is the identity matrix, is additive white Gaussian noise, and the channel signal-to-noise ratio is p k is the transmission power of the kth transmitter;

[0018] The fusion feature y[m] on the mth subcarrier is:

[0019]

[0020] Where N = Σ k,s n is Gaussian additive white noise.

[0021] Preferably, in step S6, the weight parameters of the aerial computation encoder are solved through end-to-end joint training based on the video data and the skeleton point data; wherein the video data is used to provide appearance information of human body movements, and the skeleton point data is used to provide spatial information of human body posture and movements.

[0022] The present invention also provides a semantic communication device for large-scale access, comprising:

[0023] The first processing module is used to obtain user information at the source end and extract semantic information using a semantic encoder;

[0024] A second processing module is used to convert the high-dimensional features of the semantic information into a binary code stream transmitted in the physical channel;

[0025] The third processing module is used to send the converted semantic features to the wireless channel after channel coding;

[0026] A fourth processing module, configured to aggregate local features on the same subcarrier of the semantic features;

[0027] A fifth processing module is used to directly apply the aggregated semantic features to intelligent tasks, wherein the intelligent tasks include semantic segmentation, classification, and object detection;

[0028] a sixth processing module, configured to solve weight parameters of an over-the-air computation encoder through end-to-end joint training;

[0029] The seventh processing module is used to jointly optimize and solve the semantic codec and the air computing encoder.

[0030] Preferably, in step S4, OFDM modulation is used to transmit M-dimensional semantic features on M subcarriers, and the signal transmitted through the channel is:

[0031]

[0032] Among them, h k is the channel coefficient of the channel between each user node and the receiver, is the identity matrix, is additive white Gaussian noise, and the channel signal-to-noise ratio is p k is the transmission power of the kth transmitter;

[0033] The fusion feature y[m] on the mth subcarrier is:

[0034]

[0035] Where N = Σ k,s n is Gaussian additive white noise.

[0036] Preferably, the sixth processing module solves the weight parameters of the aerial computation encoder through end-to-end joint training based on the video data and the skeleton point data; wherein the video data is used to provide appearance information of human body movements, and the skeleton point data is used to provide spatial information of human body posture and movements.

[0037] The present invention also provides a semantic communication system for large-scale access, comprising: a memory and a processor, wherein the memory stores a computer program run by the processor, and the computer program executes a semantic communication method for large-scale access when run by the processor.

[0038] Compared with the prior art, the present invention has the following beneficial effects:

[0039] The present invention can effectively handle the connection and information processing of a large number of devices in the Internet of Things, achieving more efficient interconnection; it can process information of different modalities from different users, realizing multi-source perception and processing; it realizes "comprehensive calculation combination", fully explores the orthogonality of semantic domains, realizes efficient information fusion, and reduces the complexity of the communication system; by designing appropriate semantic codecs, it can be flexibly applied to various intelligent tasks; under different channel conditions, its performance is better than that of traditional multi-user communication methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0041] Figure 1 It is a block diagram of the task-oriented multi-user semantic communication system based on air computing of the present invention;

[0042] Figure 2 It is a joint optimization framework of semantic codec and air computing encoder of the present invention;

[0043] Figure 3 This is a flow chart of the semantic communication method for large-scale access according to the present invention;

[0044] Figure 4 This is a graph showing the convergence trend of the reward function of the DDQN algorithm and the loss function of the semantic decoder training.

[0045] Figure 5 This is a comparison chart of the motion recognition accuracy results of the present invention and other methods under different signal-to-noise ratio conditions;

[0046] Figure 6 This is a diagram showing changes in transmission energy consumption under different access quantity scenarios of the present invention;

[0047] Figure 7 It is the result graph of the semantic segmentation task performed by the present invention;

[0048] Figure 8 This is a comparison chart of semantic segmentation mean intersection over union (mIoU) between the present invention and other methods under different signal-to-noise ratio conditions;

[0049] Figure 9 This is a comparison chart of the average accuracy of semantic segmentation between the present invention and other methods under different signal-to-noise ratio conditions. DETAILED DESCRIPTION

[0050] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0051] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0052] Example 1:

[0053] The embodiment of the present invention provides a semantic communication method for large-scale access, such as Figure 1 、 2 As shown in Figure 2, the multi-user semantic communication framework designed by the present invention consists of K single-antenna transmitters and a receiver equipped with M antennas, and all transmitters share the same channel. The local device collects data Extracting semantic features Then, the feature vector encoded by the encoder is calculated over the air through the SIMO-OFDM channel Transmitted to the receiver. Semantic features of different dimensions are transmitted using different subcarriers. For all local features of the same dimension, the same subcarrier is used for transmission. The method of the embodiment of the present invention supports multiple local devices in the same channel to simultaneously transmit feature vectors at the same frequency. By using air computing technology, the feature vectors are naturally superimposed in multiple access channels. Finally, the global semantic feature Y is obtained at the receiving end, and the global features of different intelligent tasks are directly subjected to task-oriented semantic decoding to obtain the task result z.

[0054] like Figure 3 As shown, the semantic communication method for large-scale access in an embodiment of the present invention includes:

[0055] Step S1: Build a semantic transmitter. The multi-user semantic communication framework in this embodiment of the present invention is equipped with K single-antenna transmitters. User information is collected at the source end and semantic information is extracted using a semantic encoder. The specific steps are as follows:

[0056] Step 101: The source collects user data c s Indicates the data modality type. During the communication process, a total of S types of modal data need to be transmitted. Each type of modal data can be provided by one or more users. The obtained semantic feature vector is Where M is the feature dimension. Here, a deep neural network is used as a semantic encoder to extract semantic features:

[0057]

[0058] in, is the semantic encoder, is the encoding parameter.

[0059] In the specific implementation process of the present invention, the human action recognition task is used as an example. The task includes two types of modal data: appearance and posture. For appearance data, ResNet is used as the backbone network for feature extraction. The ResNet output is globally average pooled to obtain the global features of the image. For posture data, the same ResNet model is used to extract posture features, but a front-end convolutional block is added. This module consists of a convolution with a kernel size of 3, batch normalization, and a rectified linear unit (ReLU) activation function. This module is added to adapt to the input channel dimension of the first ResNet block.

[0060] Step 102: Use an air-computing encoder to calculate local semantic features. Encode and get The expression is as follows:

[0061]

[0062] Among them, ψ k is the kth element of Ψ, Ψ=[ψ1,ψ2,...,ψ k ,...,ψ K ] T is the encoder weight parameter calculated over the air. Ψ includes the modality alignment parameter and the weights between different modalities.

[0063] Step S2: Convert the high-dimensional features of the semantic information into a binary code stream transmitted in the physical channel.

[0064] For categorical features (discrete features), you can use one of the following methods to encode:

[0065] One-Hot Encoding: Each category is assigned a binary vector whose length is equal to the number of categories and has 1 in the position corresponding to the category and 0 in the rest of the positions.

[0066] Binary Encoding: Map each category directly to a binary number.

[0067] For continuous features, one of the following methods can be used for encoding:

[0068] Linear quantization: Mapping the value of a feature to a finite range of binary numbers.

[0069] Nonlinear quantization: Use piecewise functions to map eigenvalues to binary codes.

[0070] Step S3: Wireless channel modeling. The processed semantic features are channel-coded and then sent to the wireless channel. Using OFDM modulation, the M-dimensional semantic features are transmitted on M subcarriers. The signal transmitted through the channel can be expressed as:

[0071]

[0072] Among them, h k is the channel coefficient of the channel between each user node and the receiver. For Rayleigh fading channel, Among them I Nr is the identity matrix, is additive white Gaussian noise, and the channel signal-to-noise ratio is p k is the transmission power of the kth transmitter.

[0073] Step S4: The transmitter sends the semantic information of multiple users to the wireless channel. At the receiving end, the local features on the same subcarrier are aggregated. The fused feature y[m] on the mth subcarrier can be expressed as:

[0074]

[0075] Where N = Σ k,s n is Gaussian additive white noise.

[0076] Step S5: Aggregate the local feature vectors across all dimensions of M subcarriers, and the global semantic feature obtained at the receiving end is Y = [y[1], y[2], ..., y[M]] T The fused semantic features are then directly used for intelligent tasks:

[0077] z=F se (Y;ω),

[0078] Among them, z is the execution result of the intelligent task, F se (·; ω) is a task-oriented semantic decoder, and ω is a learnable parameter.

[0079] These intelligent tasks may include semantic segmentation, classification, object detection, etc. The specific learnable parameters depend on the type of task.

[0080] Step S6: Construct an optimization problem and solve the weight parameters of the air-computing encoder through end-to-end joint training. The specific steps are as follows:

[0081] Step 601: Construct an optimization problem:

[0082] To illustrate this process more clearly, we use the action recognition task as an illustrative example. This task involves two data modalities: video and skeleton points. Video provides appearance information about human actions, while skeleton points provide spatial information about human posture and motion. The fusion of these two types of information helps improve the accuracy and robustness of action recognition.

[0083] Assume that there are k1 users providing video data and k2 users providing skeleton point data in the system, where k1 + k2 = K. Before fusion, the features are first aligned to ensure that their dimensions are consistent. Then, they are used to construct the m-th dimension fused feature y[m] as:

[0084]

[0085] Next, a softmax function with a fully connected layer is used as a semantic decoder at the receiving end to obtain the decoding result:

[0086]

[0087] Among them, ω and b are the weight and bias parameters of the fully connected layer, respectively, and M′ is the dimension of the output of the fully connected layer.

[0088] The next step is to use cross entropy loss to train the semantic decoder and construct the cross entropy loss function:

[0089]

[0090] in is the true value, and T is the number of samples.

[0091] The loss function is determined by the semantic encoder parameters α, the over-the-air encoder parameters ψ, and the semantic encoder parameters ω and b. To minimize the loss, α, ω, and b are optimized using the stochastic gradient descent (SGD) algorithm. For the aggregate weight ψ, when α, ω, and b are determined, the following optimization problem P is constructed for optimization.

[0092]

[0093] in is the semantic feature of video data, is the semantic feature of the skeleton point data. Ψ1 and Ψ2 are the sum of the weights of the video data and the skeleton point data respectively. min and Ψ max are their minimum and maximum values respectively.

[0094] Step 602: Given the in-flight computation encoder and semantic decoder, use the stochastic gradient descent algorithm to solve the semantic encoder parameter α so as to minimize the cross entropy function.

[0095] Step 603: Given the semantic encoder and semantic decoder, solve the encoder parameters in the air. Due to the non-convexity of the target loss function and the existence of the ResNet neural network, problem P cannot be solved using traditional methods. Therefore, the deep reinforcement learning algorithm DDQN is used to solve it. The key elements are defined as follows:

[0096] Agent: Set the edge server as the agent, responsible for calculating the fusion weight;

[0097] State space: Calculate the fusion weights assigned by the encoder to each user on the fly:

[0098]

[0099] Action space: The change of each user's fusion weight. Assuming that the minimum step size of the change is Δψ and the number of connected users is K=3, the action space can be expressed as:

[0100] A={a|a∈{(1,2),(1,3),(2,1),(2,3),(3,1),(3,2)}}

[0101] Among them, (i, j) represents (ψ i +Δψ,ψ j -Δψ),i,j∈{1,2,3},

[0102] Environmental feedback: The environmental feedback is set as the difference between the reward function R of the next state and the current state, where ΔR = RR′.

[0103] The negative value of the optimization objective given by P is used as the reward function R. The greater the accuracy of the task during training, the higher the value of R.

[0104] Therefore, we want to optimize in the direction of increasing R. When R increases, the feedback function ΔR becomes positive, generating rewards; when R decreases, ΔR becomes negative, resulting in penalties. Training ends when the number of training rounds reaches a predetermined maximum or the model converges.

[0105] Build the main Q network to select actions based on the maximum Q value;

[0106] a=argmaxQ(s,a;θ),

[0107] Build a target Q network to generate Q values for calculating the loss function of each action during training;

[0108] Q t (s,a;θ t )=ΔR+λmaxQ(s′,a′;θ t ),

[0109] Among them, θ t is the target Q network parameter, s, a are the current state and action, s′, a′ are the next state and action, and λ is the reward decay factor;

[0110] Calculate the loss function of the double Q network:

[0111] L q =(Q t (s,a;θ t )-Q(s,a;θ)) 2 .

[0112] Step 604: Given the parameters of the semantic encoder and the over-the-air computation encoder, the semantic decoder is trained by minimizing the loss function using the SGD algorithm.

[0113] Step S7: jointly optimize and solve the semantic codec and the air computing encoder.

[0114] First, the semantic encoder is optimized using the stochastic gradient descent algorithm while freezing ψ, w, and b. Subsequently, the air computation encoder is optimized using the DDQN algorithm while freezing the other models. Finally, we use the SGD algorithm to update the parameters w and b of the semantic decoder, keeping α and ψ fixed to minimize the loss value. The semantic decoder processes the fused features and obtains the results. Based on these results, we calculate the reward function ΔR and provide feedback on the parameter update to the air computation encoder and the semantic encoder. This iterative process continues until the loss function converges, at which point the joint optimization is completed. The convergence results are shown in Figure 2. Figure 4 shown.

[0115] The comparison results of the motion recognition accuracy of the present invention with other methods under different signal-to-noise ratio conditions are as follows: Figure 5 shown. Figure 5 In the formula, SC-DAC is the method proposed by the present invention; BPG+LDPC represents the traditional communication method based on BPG image compression and LDPC channel coding; Error-free Transmission represents error-free transmission; SC-FR represents the feature reconstruction method, in which the semantic encoder extracts the global semantic information of the data and reconstructs the features based on the received semantic information; CIF represents the channel-level information fusion method, in which multimodal data from multiple users are fused through the wireless channel during the transmission phase and weightedly aggregated according to the channel gain. Under the framework proposed by the present invention, the task recognition accuracy is much higher than that of other communication methods, especially when the signal-to-noise ratio is lower than 30dB, its advantage is more obvious. Under low signal-to-noise ratio conditions, the accuracy of action recognition is improved by about 20% compared with traditional methods and by 50% compared with other multi-user semantic communication methods. At the same time, it can be seen from the result curve that the present invention exhibits stronger stability to changes in channel conditions.

[0116] The transmission energy consumption changes under different numbers of access users in the present invention are as follows: Figure 6 As shown in Figure 3, as the number of connected users increases, transmission energy consumption increases slightly. Assuming that each user provides consistent data, an exponential increase in the number of users will result in a proportional increase in the amount of data. Notably, the increase in transmission energy consumption is negligible when the number of users increases from 10 to 20. This demonstrates the nonlinear response of the system to exponential user growth.

[0117] The results of applying this invention to semantic segmentation tasks are as follows: Figure 7 As shown in the figure, MFnet is a CNN architecture used for real-time semantic segmentation of autonomous vehicles. Compared with MFnet, our proposed method reduces the mean Intersection Over Union (MIoU) by 0.06 and improves the average class precision by 0.04. The segmentation accuracy of the "Car Stop" and "Bump" classes is significantly higher than that of MFnet.

[0118] Comparison of semantic segmentation mean intersection over union (mIoU) and class average accuracy of the proposed method with other methods under different signal-to-noise ratio conditions Figure 8 and Figure 9 The semantic segmentation mIoU and class average precision achieved by the present invention are much higher than those of other methods, and the accuracy of semantic segmentation is improved by 0.2 to 0.34.

[0119] In summary, this method can effectively handle the connection and information processing of a large number of devices in IoT scenarios, achieving multi-source perception while reducing the complexity of the communication system through "comprehensive computing integration." Furthermore, by designing appropriate semantic codecs, this method can be flexibly applied to various intelligent tasks and demonstrates superior performance to traditional multi-user communication methods under different channel conditions, thereby helping to achieve more efficient and intelligent IoT interconnection.

[0120] Example 2:

[0121] An embodiment of the present invention further provides a semantic communication device for large-scale access, comprising:

[0122] The first processing module is used to obtain user information at the source end and extract semantic information using a semantic encoder;

[0123] A second processing module is used to convert the high-dimensional features of the semantic information into a binary code stream transmitted in the physical channel;

[0124] The third processing module is used to send the converted semantic features to the wireless channel after channel coding;

[0125] A fourth processing module, configured to aggregate local features on the same subcarrier of the semantic features;

[0126] A fifth processing module is used to directly apply the aggregated semantic features to intelligent tasks, wherein the intelligent tasks include semantic segmentation, classification, and object detection;

[0127] a sixth processing module, configured to solve weight parameters of an over-the-air computation encoder through end-to-end joint training;

[0128] The seventh processing module is used to jointly optimize and solve the semantic codec and the air computing encoder.

[0129] As an implementation method of an embodiment of the present invention, in step S4, OFDM modulation is used to transmit M-dimensional semantic features on M subcarriers, and the signal transmitted through the channel is:

[0130]

[0131] Among them, h k is the channel coefficient of the channel between each user node and the receiver, is the identity matrix, is additive white Gaussian noise, and the channel signal-to-noise ratio is p k is the transmission power of the kth transmitter;

[0132] The fusion feature y[m] on the mth subcarrier is:

[0133]

[0134] Where N = ∑ k,s n is Gaussian additive white noise.

[0135] As an implementation method of an embodiment of the present invention, the sixth processing module solves the weight parameters of the aerial computing encoder through end-to-end joint training based on video data and skeleton point data; wherein, the video data is used to provide appearance information of human body movements, and the skeleton point data is used to provide spatial information of human body posture and movement.

[0136] Example 3:

[0137] An embodiment of the present invention also provides a semantic communication system for large-scale access, comprising: a memory and a processor, wherein the memory stores a computer program executed by the processor, and the computer program executes a semantic communication method for large-scale access when executed by the processor.

[0138] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.

Claims

1. A semantic communication method for large-scale access, characterized in that: include: Step S1: Obtain user information at the source end and extract semantic information using a semantic encoder; Step S2: converting the high-dimensional features of the semantic information into a binary code stream transmitted in a physical channel; Step S3: channel-encoding the converted semantic features and sending them to the wireless channel; Step S4: Aggregate local features on the same subcarrier of the semantic features; Step S5: directly use the aggregated semantic features for intelligent tasks; wherein the intelligent tasks include semantic segmentation, classification, and object detection; Step S6: solving the weight parameters of the air calculation encoder through end-to-end joint training; Step S7: jointly optimize and solve the semantic codec and the air computing encoder; Wherein, step S6 includes: Step 601: Construct an optimization problem: Assume that there are k1 users providing video data and k2 users providing skeleton point data in the system, where k1+k2=K. Before fusion, feature alignment is first performed, and then they are used to construct the m-th dimension fusion feature y[m] as follows: At the receiving end, a softmax function with a fully connected layer is used as a semantic decoder to obtain the decoding result: Among them, ω and b are the weight and bias parameters of the fully connected layer, respectively, and M′ is the dimension of the output of the fully connected layer; Use cross entropy loss to train the semantic decoder and construct the cross entropy loss function: in, is the true value, T is the number of samples; The loss function is determined by the semantic encoder parameters α, the air-computation encoder parameters ψ, and the semantic encoder parameters ω and b. α, ω, and b are optimized using the stochastic gradient descent (SGD) algorithm. For the aggregation weight ψ, when α, ω, and b are determined, the following optimization problem P is constructed for optimization: Ψ1+Ψ2=1, C2:P min ≤Ψ1,Ψ2≤Ψ max . in, is the semantic feature of video data, is the semantic feature of the skeleton point data; Ψ1 and Ψ2 are the sum of the weights of the video data and the skeleton point data respectively, min and Ψ max are their minimum and maximum values, respectively; Step 602: Given an air computation encoder and a semantic decoder, use a stochastic gradient descent algorithm to solve for semantic encoder parameters α so as to minimize the cross entropy function. Step 603: Given a semantic encoder and a semantic decoder, use the deep reinforcement learning algorithm DDQN to solve the encoder parameters in the air. The key elements are defined as follows: Agent: Set the edge server as the agent, responsible for calculating the fusion weight; State space: Calculate the fusion weights assigned by the encoder to each user on the fly: Action space: The change of each user's fusion weight. Assuming that the minimum step size of the change is Δψ and the number of connected users is K=3, the action space can be expressed as: A={a|a∈{(1,2),(1,3),(2,1),(2,3),(3,1),(3,2)}} where (i, j) represents (ψ i + Δψ, ψ j - Δψ), i, j ∈ {1, 2, 3} Environmental feedback: The environmental feedback is set to the difference between the reward function R of the next state and the current state, where ΔR = RR′; The negative value of the optimization objective given by P is used as the reward function R; When R increases, the feedback function ΔR is positive, generating rewards; when R decreases, ΔR is negative, resulting in penalties; training ends when the number of training rounds reaches a predetermined maximum value or the model converges; Build the main Q network to select actions based on the maximum Q value; a=arg max Q(s,a;θ), Build a target Q network to generate Q values for calculating the loss function of each action during training; Q t (s,a;θ t )=ΔR+λmaxQ(s′,a′;θ t ), Among them, θ t is the target Q network parameter, s, a are the current state and action, s′, a′ are the next state and action, and λ is the reward decay factor; Calculate the loss function of the double Q network: L q =(Q t (s, a;θ t )-Q(s,a;θ)) 2 , Step 604: Given the parameters of the semantic encoder and the over-the-air computation encoder, the semantic decoder is trained by minimizing the loss function using the SGD algorithm.

2. The semantic communication method for large-scale access according to claim 1, wherein: In step S4, OFDM modulation is used to transmit M-dimensional semantic features on M subcarriers, and the signal transmitted through the channel is: Among them, h k is the channel coefficient of the channel between each user node and the receiver, is the identity matrix, is additive white Gaussian noise, and the channel signal-to-noise ratio is p k is the transmission power of the kth transmitter; The fusion feature y[m] on the mth subcarrier is: Where N = ∑ k,s n is Gaussian additive white noise.

3. A semantic communication device for large-scale access that implements the semantic communication method for large-scale access according to claim 1, characterized in that: include: The first processing module is used to obtain user information at the source end and extract semantic information using a semantic encoder; A second processing module is used to convert the high-dimensional features of the semantic information into a binary code stream transmitted in the physical channel; The third processing module is used to send the converted semantic features to the wireless channel after channel coding; A fourth processing module, configured to aggregate local features on the same subcarrier of the semantic features; A fifth processing module is used to directly apply the aggregated semantic features to intelligent tasks, wherein the intelligent tasks include semantic segmentation, classification, and object detection; a sixth processing module, configured to solve weight parameters of an over-the-air computation encoder through end-to-end joint training; The seventh processing module is used to jointly optimize and solve the semantic codec and the air computing encoder.

4. The semantic communication device for large-scale access according to claim 3, characterized in that: In step S4, OFDM modulation is used to transmit M-dimensional semantic features on M subcarriers, and the signal transmitted through the channel is: Among them, h k is the channel coefficient of the channel between each user node and the receiver, is the identity matrix, is additive white Gaussian noise, and the channel signal-to-noise ratio is p k is the transmission power of the kth transmitter; The fusion feature y[m] on the mth subcarrier is: Where N = ∑ k,s n is Gaussian additive white noise.

5. A semantic communication system for large-scale access, characterized in that: include: A memory and a processor, wherein the memory stores a computer program executed by the processor, and when the computer program is executed by the processor, the computer program executes the semantic communication method for large-scale access according to any one of claims 1 to 2.

Citation Information

Patent Citations

  • Joint information source channel coding method for image semantic communication

    CN117879765A

  • Task-driven image semantic communication method and system and training method thereof

    CN118298834A