A real-time three-dimensional semantic map construction method based on semantic communication

By building a semantic communication encoding-decoding framework on edge devices, combined with cloud-edge collaborative computing and attention mechanisms, the problem of real-time semantic mapping under limited resources is solved, achieving efficient and real-time semantic map construction and localization, adapting to different channel conditions.

CN119516069BActive Publication Date: 2025-10-24NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411463043.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-20
Publication Date
2025-10-24
Estimated Expiration
2044-10-20

AI Technical Summary

Technical Problem

Achieving high-precision and real-time semantic mapping tasks on resource-constrained edge devices presents challenges related to computational resources and data transmission efficiency, issues that existing technologies have failed to effectively address.

Method used

We adopt a semantic communication-based encoder-decoder framework, combined with cloud-edge collaborative computing, to leverage the powerful computing capabilities of edge servers to assist edge devices in completing semantic mapping tasks. We also optimize data transmission through attention mechanisms and vector quantization modules, and design a communication system that adapts to different channel conditions.

Benefits of technology

High-precision and real-time semantic mapping was achieved under limited resources, with a mapping update time of less than 1 second and a positioning accuracy error of less than 0.1%. Stable data transmission and efficient communication were maintained under different channel conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119516069B_ABST
    Figure CN119516069B_ABST
Patent Text Reader

Abstract

The application discloses a kind of real-time three-dimensional semantic map construction methods based on semantic communication, belong to semantic communication technical field.The method described in the application includes according to the real-time three-dimensional semantic mapping task demand under resource limitation, design the communication system scene model in combination with cloud edge cooperation, and simplify basic communication transmission model;Then according to simplified version communication transmission model, design with edge node end as encoding end, with edge server as decoding end Coding and decoding structure semantic communication framework, while increasing quantization module on framework design to realize the maximum compression of semantic information, ensure that real-time semantic mapping task can be completed on the edge node of limited resources, finally, according to the different channel condition requirements of scene under communication scene, design attention module based on attention mechanism, and embed the module in the coding and decoding structure semantic communication framework, ensure that the method can be real-time, correct to complete semantic mapping task under various channel conditions.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of semantic communication, and particularly relates to a real-time three-dimensional semantic map construction method based on semantic communication. BACKGROUND

[0002] Cloud-edge collaborative computing refers to using cloud computing, edge computing and collaborative computing technologies to realize collaborative work and resource sharing between the cloud and edge devices. Through this technology, the powerful computing power of the cloud and the good mobility of the edge can be effectively combined to serve various application scenarios. It involves the fusion and collaboration of edge computing and cloud computing. Edge computing pushes computing and data storage capabilities to the network edge to reduce data transmission delay and network load and support real-time response. Cloud computing provides large-scale computing and storage resources to support complex data processing and analysis tasks. Cloud-edge collaboration needs to effectively combine the two to achieve dynamic allocation of resources and collaborative execution of tasks. Through effective implementation of cloud-edge collaboration, the performance and availability of the system can be improved to support various cloud-edge collaborative application scenarios such as smart Internet of Things and smart cities.

[0003] Semantic communication technology is an information exchange technology based on semantics. It uses structured and standardized methods to focus on the meaning of the information to be transmitted, so that computers can clearly transmit the subject information based on the meaning of the information transmitted, helping people to more effectively obtain, analyze, integrate and understand information, improve work efficiency and provide better service experience. Especially in data compression, semantic communication technology can compress image data more finely and efficiently. Compared with traditional compression methods based on prediction and transformation, which have distortion and information loss problems, semantic communication technology can analyze the semantics of image content to remove redundant information and use image features for compression, making the compression more efficient, less distorted and more informative. A large amount of effective information is transmitted under a small amount of data transmission.

[0004] Semantic mapping technology refers to creating a structured map for a robot that represents semantic information, to help the robot understand and process information in the environment, and support the execution of related tasks. This technology is crucial in the field of robotics, providing a good foundation for the robot's own precise positioning and high-level task requirements. In construction, robot semantic map construction relies on perception and perception fusion technology. Often through visual, laser radar, sound and other multi-sensing data. Perception fusion technology integrates the information obtained to form a comprehensive understanding of the environment. Secondly, semantic map construction requires semantic understanding and knowledge representation. That is, the perceived environmental information is converted into a form that the machine can understand, and a semantic representation related to the task is established. Robot semantic map construction often focuses on map construction accuracy and real-time map construction, and high-precision, real-time construction often relies on powerful edge computing nodes, and resource consumption and other issues have not been considered in previous research.

[0005] Applying the cloud-edge collaboration idea to resource-constrained real-time semantic mapping tasks can utilize the powerful computing power of the cloud to assist in the completion of real-time semantic mapping tasks and reduce the cost of completing semantic mapping tasks. Secondly, the fusion of semantic communication methods not only improves data transmission efficiency, but also provides support for real-time task completion. Combined with the attention mechanism, effective transmission under different channel conditions is achieved, ensuring efficient communication and correct completion of semantic mapping tasks. SUMMARY

[0006] The purpose of the present application is to provide a real-time three-dimensional semantic map construction method based on semantic communication, which mainly uses semantic communication to solve the problem of implementing resource-intensive real-time semantic mapping tasks on resource-limited edge devices.

[0007] Technical scheme: A real-time three-dimensional semantic map construction method based on semantic communication, comprising the following steps:

[0008] (1) According to the real-time three-dimensional semantic mapping task requirements under resource constraints, a communication system scenario model combining cloud-edge collaboration is designed, and a basic communication transmission model is simplified;

[0009] In the communication system model, two types of devices are considered, the transmitter represented by the edge device and the receiver represented by the edge server, and the communication tasks performed by the two types of devices include:

[0010] The edge server trains a semantic communication encoder-decoder model, and then distributes the semantic encoder model to all edge devices through wireless transmission;

[0011] The edge device fine-tunes the model according to the actual scene and extracts semantic features, and then sends them to the edge server for semantic recovery;

[0012] The edge server recovers data according to the received semantic information and completes a semantic mapping task;

[0013] (2) According to the communication transmission model, the coding-decoding structure semantic communication framework is designed with the edge node end as the encoding end and the edge server as the decoding end, and a vector quantization module is added to the framework design to realize the maximum compression of semantic information, so that the real-time semantic mapping task can be completed on the edge node with limited resources;

[0014] The encoding end includes a semantic extraction module, a channel encoder and a knowledge base, wherein the knowledge base is used for storing background knowledge shared between the transmitter and the receiver; an RGB-D sensor is selected as a data acquisition sensor for capturing RGB frames and depth frames, and the RGB image is sent to the semantic extraction module, and the depth image is directly sent to the channel encoding module;

[0015] The decoding end includes a channel decoder, a semantic recovery module, a frame alignment module and a semantic mapping module, wherein the frame alignment module matches the RGB frame and the depth frame according to a preset timestamp, and only the image frame pair that is successfully aligned will be sent to the subsequent semantic mapping module, and the semantic mapping module includes semantic positioning, semantic extraction and semantic map generation;

[0016] (3) According to the different channel condition requirements of the scene under the communication scene, an attention module is designed based on the attention mechanism and the perceived signal-to-noise ratio, and the attention module is embedded into the coding-decoding structure semantic communication framework in step (2), so that the method can complete the semantic mapping task in real time and correctly under various channel conditions.

[0017] Further, in the simplified model, the transmitter is responsible for extracting semantic features and transmitting through the channel, wherein the transmitter includes a semantic extraction and a channel encoder, and assuming that the image source data is s, the whole transmission process can be represented as:

[0018] x=P μ (E α (s))

[0019] Wherein, E α (·) represents a semantic extraction network with a training parameter of α, P μ (·) refers to a channel encoder network with a training parameter of μ, and the output x is sent to the channel for transmission later;

[0020] In order to consider the channel fluctuation and realize the channel variability, an adaptive white noise channel is selected as the system channel, and the formula is as follows:

[0021] y=x+n

[0022] Wherein, n represents channel noise, which follows a normal distribution N(0,σ2 ), y represents the channel output, which will be transmitted to the receiver in the subsequent step;

[0023] The receiver is responsible for recovering semantic information and constructing a semantic map, which includes a channel decoder, semantic recovery and semantic mapping, represented as follows:

[0024]

[0025] where Q ν (·) represents a channel decoder network trained with parameters v, R β (·) represents a semantic decoder network trained with parameters β;

[0026] The semantic recovery process is the inverse of the semantic feature extraction process, and the output of the semantic recovery process will serve as the input for the subsequent semantic mapping;

[0027] In terms of semantic mapping, the method includes generating a semantic map M in real time using semantic simultaneous localization and mapping technology; the simplified model also includes a knowledge base for sharing background knowledge between the sender and the receiver, which is represented by the structure and parameters of the neural network, and helps to understand and interpret the exchanged information.

[0028] Further, the semantic extraction module of step (2) includes a semantic encoder, an attention module and a vector quantization module; the semantic encoder uses a convolutional neural network-based method to extract semantic features, and after obtaining the semantic features; the attention module is designed based on the perceived signal-to-noise ratio, which is used to adapt to various channel conditions; the vector quantization module is used to compress semantic information, mapping continuous information into a finite-dimensional discrete vector, thereby achieving the greatest compression without destroying the semantic data.

[0029] Further, the semantic recovery module includes an attention module and a CNN-based semantic decoder, which is used to restore the semantic feature information to the original RGB frame.

[0030] The vector quantization model finds the closest discrete semantic embedding vector for continuous input data, ensuring that the sender only needs to send these semantic vectors, rather than the entire semantic features, thereby achieving reduced data transmission; the vector quantization module will search for the nearest vector in the semantic embedding space, and the mathematical expression is as follows:

[0031] n = arg n min‖f m -e n ||2,

[0032] where f m is the semantic encoding result, e nis a semantic embedding space shared by the sender and the receiver;

[0033] After vector quantization, the discrete semantic features and the depth frame are input into the channel encoder and then transmitted to the channel for transmission.

[0034] Further, in order to adapt to various channel conditions in the communication process, an attention module is designed based on an attention mechanism and a perceptual signal-to-noise ratio, the corresponding attention weights of different signal-to-noise ratio channels are automatically learned by training the model, a full connection layer is used, the signal-to-noise ratio is used as input, the sigmoid function is used in the attention module to obtain the attention weights, then the attention weights are expanded to match the shape of the input features, and then multiplied with the input features; the attention layer is represented as:

[0035] z f+1 =f FCA (z f ,snr)*z f

[0036] wherein z f and z f+1 represent input and output features respectively, f FCA(·) is composed of a connection layer and a sigmoid function.

[0037] Beneficial effects: the method constructs an encoding-decoding semantic communication framework for completing real-time semantic mapping tasks under limited resources. In addition, considering the influence of different channel conditions on communication, an attention mechanism-based module is designed to realize stable data transmission under different channel conditions. In terms of simulation experiments, based on the TUM dataset, it is verified that the system has an error of less than 0.1% in mapping and positioning accuracy compared with the ground truth, and is superior to traditional communication algorithms in real-time and channel adaptation. In addition, by realizing a prototype system, the effectiveness of the proposed framework and the designed module in actual indoor scenes is verified. The results show that the method can complete the real-time semantic mapping task of indoor objects (chairs, computers, people, etc.) under limited resources, and the mapping update time is less than 1 second. BRIEF DESCRIPTION OF DRAWINGS

[0038] Figure 1 is a system scene diagram;

[0039] Figure 2 is a simplified system framework diagram;

[0040] Figure 3 is a semantic communication mapping system encoding-decoding framework;

[0041] Figure 4 is an attention module structure design;

[0042] Figure 5 is a transmitter network structure diagram;

[0043] Figure 6 is a receiver network structure diagram;

[0044] Figure 7 is a real-time performance comparison of data transmission;

[0045] Figure 8 is a semantic mapping effect using the Tum dataset;

[0046] Figure 9 is the transmission quality of the ablation experiment with or without the attention module under different SNRs;

[0047] Figure 10 is a result map of semantic mapping of real indoor objects in the embodiment. DETAILED DESCRIPTION

[0048] The present application will be specifically described below in combination with the drawings and specific examples.

[0049] The present application provides a real-time three-dimensional semantic map construction method based on semantic communication. The cloud-edge collaborative computing idea is introduced into the real-time semantic mapping task under resource-limited conditions, and a semantic communication framework is designed in combination with the high data transmission efficiency of semantic communication to complete the real-time mapping task in resource-limited edge devices. The problem to be solved by the present application is how to complete the real-time semantic mapping task in the cloud-edge collaborative computing manner under the condition of limited resources of mobile devices and maintain high data transmission efficiency in the process, and correctly complete the mapping under various channel conditions. Based on the semantic communication technology, the image semantic features are extracted and then transmitted and restored, which are used as the data source of the semantic mapping task. As shown in Figure 1 The scene is considered as a resource-limited edge computing node and a powerful edge server. Figure 2 is a communication coding and decoding structure formed by simplifying the system model. Figure 3 is a coding and decoding semantic communication framework designed based on the simplified communication structure. Figure 4 is an attention design module realized based on the attention mechanism of signal-to-noise ratio perception, which realizes effective transmission of data under various channel conditions. Figure 5 is a transmitter network structure design diagram, which realizes semantic extraction and channel coding transmission of data. Figure 6 is a receiver network structure design diagram, which realizes semantic recovery of semantic features and semantic mapping. Figure 7 is a real-time performance comparison of data transmission, which proves that the present method is superior to the traditional algorithm in real-time performance. Figure 8 is the transmission performance under different signal-to-noise ratios, which proves that the present method is superior to the traditional algorithm in transmission performance under different channel conditions. Figure 9The mapping effect realized by using the Tum data set for the method is compared with the positioning, which shows that the method can correctly complete the mapping task and perform well in positioning. Figure 10 The semantic graph result constructed in the embodiment is as follows.

[0050] Step 1: Design a communication system scenario model

[0051] Two devices are considered in the communication system scenario model: a transmitter represented by an edge device and a receiver represented by an edge server. The former has limited computing resources but good mobility, and the latter has strong computing power and can generate a semantic map in real time. The specific communication process of the two types of devices can be divided into the following steps:

[0052] 1) The edge server trains a semantic communication encoder-decoder model, and then distributes the semantic encoder model to all edge devices through wireless transmission.

[0053] 2) The edge device fine-tunes the model according to the actual scene and extracts semantic features, and then sends the semantic features to the edge server for semantic recovery.

[0054] 3) The edge server recovers data according to the received semantic information and completes the semantic mapping task.

[0055] We first simplify the above model, as shown in Figure 1 and Figure 2 In the simplified model, the transmitter is responsible for extracting semantic features and transmitting them through the channel. This process mainly consists of two parts: semantic extraction and channel encoder. Assuming that the image source data is s, the entire transmission process can be represented as:

[0056] x=P μ (E α (s)), (1)

[0057] where E α (·) represents a semantic extraction network with training parameters α, and P μ (·) refers to a channel encoder network with training parameters μ. The output x is later sent to the channel for transmission.

[0058] To consider channel fluctuations and achieve channel variability, an adaptive white noise channel (AWGN) is chosen as the system channel, represented by equation (2):

[0059] y=x+n, (2)

[0060] where n represents channel noise, which follows a normal distribution N(0, σ 2 ), and y represents the channel output, which will be transmitted to the receiver in the subsequent steps.

[0061] The receiver is responsible for recovering semantic information and constructing semantic maps. This process mainly includes three parts: channel decoder, semantic recovery and semantic mapping. As shown in equation (3), the semantic recovery process is the inverse process of semantic feature extraction.

[0062]

[0063] where Q ν (·) represents the channel decoder network with training parameters v, R β (·) represents the semantic decoder network with training parameters β. The output of semantic recovery will be the input of the subsequent semantic mapping. In terms of semantic mapping, the semantic simultaneous localization and mapping (SLAM) technique is adopted to generate semantic maps M in real time.

[0064] In addition, the knowledge base is a basic component of the system model. The background knowledge shared between the sender and the receiver helps to understand and interpret the exchanged information. In this method, the knowledge base is represented by the structure and parameters of the neural network.

[0065] Step 2: Designing the codec structure semantic communication framework

[0066] 2.1 System framework structure

[0067] The entire system is composed of a coding-decoding structure. At the encoding end, the RGB-D sensor is chosen as the data acquisition sensor, which can capture both RGB frames and depth frames simultaneously. Compared with the RGB sensor, the depth information helps to achieve more effective localization. The RGB image will be sent to the semantic extraction module, while the depth image, due to the lack of semantic information, will be directly sent to the channel encoding module. The semantic extraction module is composed of three parts: semantic encoder, attention module and vector quantization module. The semantic encoder adopts a convolutional neural network (CNN) based method to extract semantic features. After obtaining the semantic features, an attention module is added to adapt to various channel conditions. This attention module is designed according to the perceived signal-to-noise ratio.

[0068] In order to further compress the semantic information, a vector quantization module is specially designed in the codec structure semantic communication framework. The vector quantization module can map continuous information to a finite-dimensional discrete vector, thereby achieving the maximum compression without destroying the semantic data. The quantized and compressed data will be concatenated with the depth frame, and then sent to the channel by the channel encoder.

[0069] At the receiving end, channel decoding is performed first after receiving the signal. It helps to restore the signal to the previous semantic information. The recovery process of semantic information is similar to the inverse process of the previous transmission process. Specifically, it includes two parts: attention module and CNN-based semantic decoder. The ultimate goal is to restore the semantic feature information to the original RGB frame. Compared with the complex processing process of RGB image, depth image can be directly transmitted and is faster than RGB image. Therefore, when restoring the RGB frame, they will be misaligned due to the time difference.

[0070] In order to solve the misalignment problem, a frame alignment module is specially designed to match the RGB frame and the depth frame according to the preset timestamp. Only the image frame pair that is successfully aligned will be sent to the subsequent semantic mapping. For semantic mapping, high map accuracy and good real-time performance are our ultimate goal. This module is divided into three parts: positioning, semantic extraction and semantic map generation. In order to achieve good real-time performance, the three modules are deployed in three threads and executed in parallel.

[0071] In this encoding and decoding architecture, the encoding end only needs to extract semantic information, and the requirement for computing power is extremely low. At the same time, the amount of quantized semantic information is very small, which can ensure good real-time performance and complete real-time semantic mapping tasks at the decoding end.

[0072] 2.2 Sender design

[0073] The role of the sender is to extract semantic information from the source data, mainly including semantic encoder, attention module and vector quantization. As shown in Figure 5 , they are almost realized by neural networks: the semantic encoder module is based on CNN, the attention module is based on the perceived signal-to-noise ratio, and the vector quantization module is based on the generation of quantization matrix in the embedded semantic space.

[0074] The semantic encoder extracts semantic features from the RGB frame S∈R H×W×3 through two down-sampling modules, where H, W and 3 are the image width, image height and channel number respectively. Specifically, the semantic encoder module includes two convolutional down-sampling modules and two residual down-sampling modules. Each convolutional down-sampling module includes a convolutional layer and a batch normalization layer. Between the two modules, there is an activation function ReLU to improve the non-linear ability. After the convolutional down-sampling operation, there will be two residual down-sampling modules, which include residual layers based on ResNet and batch normalization layers. In this way, the semantic features are extracted.

[0075] After semantic encoding, the semantic features are sent to the attention module to add attention weights. The feature tensor output by the attention module enters the vector quantization part to compress the data volume. The principle of vector quantization is that the continuous input data will find the closest discrete semantic embedding vector, ensuring that the transmitter only needs to send these semantic vectors instead of the entire semantic features, thereby greatly reducing the data transmission volume. Specifically, the vector quantization module will search for the nearest vector in the semantic embedding space, as shown in equation (4):

[0076] n = arg n min||f m -e n ||2, (4)

[0077] where f m is the semantic encoding result, e n is the semantic embedding space shared by the sender and the receiver. After vector quantization, the discrete semantic features and the depth frame are input into the channel encoder and then sent to the channel for transmission.

[0078] 3.2 Receiver design

[0079] The role of the receiver is to recover the extracted semantic information and use the recovery result for semantic mapping. Its network, as shown in Figure 6 , can be divided into three parts: attention module, semantic decoder, and semantic mapping. After passing through the attention module, the semantic decoder restores the semantic features to the original image. In network design, the components of the decoder network are similar to those of the encoder network, but the functions are opposite. Figure 6 The upsampling part shown in is connected by a deconvolution layer, a residual block, and a batch normalization layer in turn. Like the semantic encoder, the role of ReLU is to increase nonlinearity. After deconvolution and upsampling, the image will pass through tanh to change the result to [0, 1], thereby re-storing. In addition to the semantic recovery part, the receiver also includes a semantic mapping part. To ensure the real-time operation of the system, we use three threads in parallel. For the positioning thread, this method is designed according to ORB-SLAM3, which removes the mapping part and only retains self-positioning and loop detection. In the semantic extraction module, semantic segmentation is selected to obtain high-precision semantic information. In addition, considering the real-time requirement of the system, it is necessary to complete the semantic segmentation task as soon as possible. Therefore, SCTNet is selected as the semantic segmentation network, which effectively separates training and inference, ensuring real-time performance. After obtaining the semantic information, this semantic mapping module will form a semantic point cloud according to the depth information and the semantics. To better fuse semantic information and point clouds, Bayesian fusion is selected to achieve high-precision semantic point cloud fusion. Finally, the semantic map is generated, which is in the form of an octree map to save storage space and improve mapping efficiency.

[0080] Step three: attention module framework integration

[0081] In neural networks, attention is a mechanism that reallocates resources that are originally allocated evenly according to the importance of the attention object. First consider the difference between channel attention and spatial attention. Channel attention aims to calculate (C x 1 x 1) channel weights, while spatial attention calculates spatial weights of size (1 x H x W), where C represents the number of channels, and H and W represent the height and width of the features, respectively. Channel attention focuses on the input of different channel information, while spatial attention mainly focuses on the input of different position information. However, channel attention and spatial attention do not focus on resource allocation under different channel conditions. In this regard, we propose an attention module based on signal-to-noise ratio perception, which automatically learns the corresponding attention weights of different signal-to-noise ratio channels through model training. As shown in Figure 4 , using a fully connected layer, taking the signal-to-noise ratio (SNR) as input, using a sigmoid function to obtain the attention weight in the attention module. Then extend the attention weight to match the shape of the input feature, and then multiply it with the input feature.

[0082] The attention layer can be represented as:

[0083] z f+1 =f FCA (z f ,snr)*z f (5)

[0084] where z f and z f+1 represent the input and output features, respectively. f FCA(·) is composed of a connection layer and a sigmoid function.

[0085] This paper describes a real-time 3D semantic map construction method based on semantic communication. 3D semantic maps are playing an increasingly important role in high-precision robot positioning and scene understanding in recent years. However, real-time semantic map construction requires mobile edge devices with extremely high computing power, which is expensive and limits their widespread application. To address this limitation, inspired by cloud-edge collaborative computing and the high transmission efficiency of semantic communication, this paper proposes a method for implementing real-time semantic mapping tasks on resource-limited mobile devices. This paper designs an encoding-decoding semantic communication framework to accomplish real-time semantic mapping tasks under resource-limited conditions. Furthermore, considering the impact of varying channel conditions on communication, a module based on an attention mechanism is designed to ensure stable data transmission under various channel conditions. Simulation experiments, using the TUM dataset, demonstrate that this system achieves mapping and localization accuracy with an error of less than 0.1% compared to ground truth, outperforming traditional communication algorithms in terms of real-time performance and channel adaptability. Furthermore, a prototype system is implemented to verify the effectiveness of the proposed framework and designed modules in real-world indoor scenarios. The results show that this method can complete the real-time semantic mapping task of indoor objects (chairs, computers, people, etc.) under limited resources, and the mapping update time is less than 1 second.

Claims

1. A real-time three-dimensional semantic map construction method based on semantic communication, characterized in that, It comprises the following steps: (1) According to the real-time three-dimensional semantic mapping task requirements under resource constraints, a communication system scene model combining cloud-edge collaboration is designed, and a basic communication transmission model is simplified; In the communication system model, two types of devices are considered, the transmitter represented by the edge device and the receiver represented by the edge server, and the communication tasks performed by the two types of devices include: The edge server trains a semantic communication encoder-decoder model, and then distributes the semantic encoder model to all edge devices through wireless transmission; The edge device fine-tunes the model according to the actual scene and extracts semantic features, and then sends them to the edge server for semantic recovery; The edge server recovers data according to the received semantic information and completes the semantic mapping task; (2) According to the communication transmission model, the encoding and decoding structure semantic communication framework is designed with the edge node as the encoding end and the edge server as the decoding end, and a vector quantization module is added to the framework design to realize the maximum compression of semantic information, ensuring that the real-time semantic mapping task can be completed on the resource-limited edge node. The encoding end includes a semantic extraction module, a channel encoder, and a knowledge base for storing background knowledge shared between the transmitter and the receiver; an RGB-D sensor is selected as the data acquisition sensor to capture RGB frames and depth frames, and the RGB image is sent to the semantic extraction module, while the depth image is directly sent to the channel encoding module. The decoding end includes a channel decoder, a semantic recovery module, a frame alignment module, and a semantic mapping module. The frame alignment module matches the RGB frame and the depth frame according to the preset timestamp, and only the successfully aligned image frame pair will be sent to the subsequent semantic mapping module. The semantic mapping module includes semantic positioning, semantic extraction, and semantic map generation. (3) According to the different channel condition requirements of the communication scene, an attention module is designed based on the attention mechanism and the perceived signal-to-noise ratio, and the attention module is embedded in the encoding and decoding structure semantic communication framework described in step (2), ensuring that the method can complete the semantic mapping task in real time and correctly under various channel conditions. In the simplified model, the transmitter is responsible for extracting semantic features and transmitting them through a channel, which includes a semantic extractor and a channel encoder, assuming the image source data is The entire transmission process is represented as: , wherein, a semantic extraction network with training parameters , refers to a channel encoder network with training parameters , is later sent into a channel for transmission; In order to consider the channel fluctuation and realize the channel variability, an adaptive white noise channel is selected as the system channel, and the formula is as follows: , wherein, denotes the channel noise, which follows a normal distribution , denotes the channel output, which will be transmitted to the receiver in a subsequent step; The receiver is responsible for recovering semantic information and constructing a semantic map, which includes channel decoding, semantic recovery, and semantic mapping, represented as follows: , wherein represents a channel decoder network with training parameters represents a semantic decoder network with training parameters ;​ The semantic recovery process is the inverse of the semantic feature extraction, and the output of the semantic recovery will be input to the subsequent semantic mapping. In semantic mapping, the method includes adopting a semantic simultaneous localization and mapping technology to generate a semantic map in real time The simplified model also includes a knowledge base for sharing background knowledge between the sender and the receiver, which is represented by the structure and parameters of the neural network, and helps to understand and interpret the exchanged information. 2.The real-time three-dimensional semantic map construction method based on semantic communication according to claim 1, wherein, The semantic extraction module described in step (2) includes a semantic encoder, an attention module, and a vector quantization module. The semantic encoder uses a method based on convolutional neural network to extract semantic features. After obtaining the semantic features, the attention module is designed based on the perceived signal-to-noise ratio to adapt to various channel conditions. The vector quantization module is used to compress semantic information, mapping continuous information to a finite-dimensional discrete vector, thereby realizing the maximum compression without destroying the semantic data. 3.The real-time three-dimensional semantic map construction method based on semantic communication according to claim 1, wherein, The designed semantic recovery module includes an attention module and a CNN-based semantic decoder, which is used to restore the semantic feature information to the original RGB frame. 4.The real-time three-dimensional semantic map construction method based on semantic communication according to claim 1 or 2, characterized in that, The vector quantization model finds the closest discrete semantic embedding vector for continuous input data, ensuring that the sender only needs to send these semantic vectors instead of the entire semantic features, thereby achieving reduced data transmission; the vector quantization module will search for the nearest vector in the semantic embedding space, and the mathematical expression is as follows: wherein, is a semantic encoding result, is a semantic embedding space shared by the sender and the receiver; After vector quantization, the discrete semantic features and the deep frame are input into the channel encoder and then transmitted to the channel for transmission. 5.The real-time three-dimensional semantic map construction method based on semantic communication according to claim 2, wherein, In order to adapt to various channel conditions in the communication process, an attention module is designed based on attention mechanism and perceived signal-to-noise ratio, the corresponding attention weights of different signal-to-noise ratio channels are automatically learned by training the model, the signal-to-noise ratio is used as the input of the full connection layer, the sigmoid function is used in the attention module to obtain the attention weights, then the attention weights are expanded to match the shape of the input features, and then multiplied with the input features; the attention layer is represented as: , where, and denote input and output features, respectively, is composed of a connection layer and a sigmoid function.

Citation Information

Patent Citations

  • Multi-unmanned aerial vehicle cooperative mapping and sensing method and system based on semantic consistency

    CN117152249A

  • Channel adaptive semantic communication method and system for multi-vision task

    CN118691937A