Three-dimensional point cloud sampling-semantic joint optimization method and device for holographic communication and medium

By using semantic feature extraction and downsampling methods in a holographic communication system, the problem of high end-to-end latency caused by large point cloud data volume was solved, achieving efficient point cloud data transmission and processing and improving communication efficiency.

CN120635903BActive Publication Date: 2025-11-25TSINGHUA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511141478.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-11-25
Estimated Expiration
2045-08-15

AI Technical Summary

Technical Problem

In holographic communication, the large amount of point cloud data leads to high end-to-end processing latency. Furthermore, the disorder and permutation invariance of point clouds limit network sampling, making it impossible to meet the requirements of high bandwidth, low latency, and high reliability.

Method used

The semantic features of point clouds are extracted by a semantic communication system, the keyness of local semantic features and semantic information is determined, downsampling is performed, and the downsampled point cloud is restored at the receiving end to perform the target task, thus realizing joint optimization of semantic sampling.

Benefits of technology

It reduces the amount of point cloud data, lowers the burden of network processing and real-time transmission, improves end-to-end communication efficiency, and ensures that content with high semantic importance or rich information is accurately extracted and encoded.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635903B_ABST
    Figure CN120635903B_ABST
Patent Text Reader

Abstract

The application provides a three-dimensional point cloud sampling-semantic joint optimization method and device for holographic communication and a medium, and relates to the technical field of holographic communication. The method comprises the following steps: a sending end in a semantic communication system extracts semantic features of an original point cloud; for each point in the original point cloud, based on the correlation between the semantic features of a plurality of adjacent points of the point and the semantic features of the point, the local semantic features of the point are determined; the semantic information key degree of each point is determined based on the local semantic features of the point; the points in the original point cloud are sorted in descending order of the semantic information key degree, the first M points are subjected to semantic sampling, and down-sampling semantic features are obtained; the down-sampling semantic features are sent to a receiving end in the semantic communication system; the receiving end recovers a down-sampling point cloud based on the received down-sampling semantic features, and performs a target task based on the down-sampling point cloud, so that the sampling and semantic joint coding and decoding are coordinated, and the end-to-end communication efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of holographic communication technology, and in particular to a three-dimensional point cloud sampling-semantic joint optimization method, device and medium for holographic communication. Background Technology

[0002] Holographic communication can provide an immersive interactive experience, but its high-performance and differentiated requirements pose a significant challenge to existing communication networks. Improving the compression and transmission capabilities of source representations, as well as resource utilization, are crucial for holographic communication. Among various 3D data representation methods, point clouds offer six degrees of freedom of view, possessing a powerful ability to express 3D spatial geometric details, preserving rich 3D details, and enabling a true holographic visual experience. Compared to technologies such as meshes and light fields, point clouds require less data, have strong scalability, are convenient for algorithm preprocessing and post-processing, and have relatively mature acquisition and reconstruction devices. Therefore, point clouds are expected to become one of the mainstream 3D data representation methods for future holographic services.

[0003] Point clouds, as a classic method for representing 3D data, enable applications such as real-time high-definition 3D interaction, holographic video conferencing, remote projection of 3D digital humans, 3D environment modeling, and digital twins. In immersive communication scenarios, point cloud data is massive; for example, an indoor point cloud model at a university contained over 40,000,000 points, reaching a size of 3.2GB, leading to high end-to-end processing latency. Therefore, while ensuring the transmission of most content, it is necessary to reduce the data volume through sampling to alleviate the burden on network processing and real-time transmission.

[0004] The disordered and permutation-invariant properties of point clouds also impose constraints on various point cloud networks, including sampling. Immersive communication presents new requirements for high bandwidth, low latency, and high reliability. These new requirements, as well as how to coordinate sampling and semantic joint encoding and decoding to improve end-to-end communication efficiency, are also urgent issues to be addressed in the research of point cloud sampling networks. Summary of the Invention

[0005] To address the aforementioned technical problems, this invention provides a method, device, and medium for joint optimization of three-dimensional point cloud sampling and semantics for holographic communication, aiming to overcome or at least partially solve the aforementioned problems.

[0006] The first aspect of this invention provides a three-dimensional point cloud sampling-semantic joint optimization method for holographic communication, the method comprising:

[0007] In a semantic communication system, the sending end extracts semantic features from the original point cloud;

[0008] For each point in the original point cloud, the transmitting end determines the local semantic features of the point based on the correlation between the semantic features of multiple neighboring points of the point and the semantic features of the point itself.

[0009] The sending end determines the semantic information criticality of each point based on the local semantic features of each point in the original point cloud;

[0010] The points in the original point cloud are sorted in descending order of semantic information criticality, and the first M points are semantically sampled to obtain downsampled semantic features.

[0011] The sending end sends the downsampled semantic features to the receiving end in the semantic communication system;

[0012] The receiving end recovers the downsampled point cloud based on the received downsampled semantic features;

[0013] The receiving end performs a target task based on the downsampled point cloud, and the target task includes at least one of the following: classification task and semantic segmentation task.

[0014] A second aspect of the present invention provides an electronic device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the three-dimensional point cloud sampling-semantic joint optimization method for holographic communication as described in the first aspect of the present invention.

[0015] A third aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the three-dimensional point cloud sampling-semantic joint optimization method for holographic communication as described in the first aspect of the present invention.

[0016] In the proposed 3D point cloud sampling-semantic joint optimization method for holographic communication, semantic features are extracted from the original point cloud by the transmitting end in the semantic communication system. Then, for each point in the original point cloud, the local semantic features of the point are determined based on the correlation between the semantic features of multiple neighboring points and the semantic features of the point itself. Based on the local semantic features, the semantic information relevance of the point is determined. The points in the original point cloud are then sorted globally according to the semantic information criticality from largest to smallest to obtain the top M points. Semantic sampling is performed on the M points to obtain downsampled semantic features. The transmitting end then sends the downsampled semantic features to the receiving end in the semantic communication system. The receiving end then recovers the downsampled point cloud based on the received downsampled semantic features and performs the target task based on the downsampled point cloud. Thus, this invention combines sampling and semantic coding to achieve joint optimization of semantic sampling: by combining global and local semantic-driven point cloud downsampling with semantic encoding and decoding, it is beneficial to ensure that the "main" content with high semantic importance or rich semantic information is accurately extracted and encoded in subsequent processing, while other redundant parts are reduced in attention or even ignored. Based on the semantic content of the object itself, it can strategically downsample the original point cloud. The point cloud downsampling integrates semantic-driven methods, which can adaptively adjust to different point clouds, thereby reducing the source magnitude and further reducing the transmission overhead of 3D point cloud holographic communication, reducing the burden of network processing and real-time transmission, and improving end-to-end communication efficiency. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart illustrating the steps of a three-dimensional point cloud sampling-semantic joint optimization method for holographic communication according to an embodiment of the present invention;

[0019] Figure 2 This is an example diagram illustrating a comparison of the standard deviations of edge points and non-edge points in a point cloud, as shown in an embodiment of the present invention.

[0020] Figure 3 This is a schematic diagram illustrating the processing of the attention layer and semantic sampling layer in a three-dimensional point cloud sampling-semantic joint encoder according to an embodiment of the present invention;

[0021] Figure 4This is a schematic diagram illustrating the structure of a 3D point cloud sampling-semantic joint encoder and decoder in a 3D point cloud sampling-semantic joint model according to an embodiment of the present invention;

[0022] Figure 5 This is a schematic diagram of the architecture of a semantic-oriented point cloud sampling network according to an embodiment of the present invention;

[0023] Figure 6 This is a schematic diagram of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0024] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0025] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating the steps of a three-dimensional point cloud sampling-semantic joint optimization method for holographic communication, as shown in an embodiment of the present invention. Figure 1 As shown, the 3D point cloud sampling-semantic joint optimization method for holographic communication provided in this embodiment includes at least the following steps:

[0026] Step S11: The sending end in the semantic communication system extracts the semantic features of the original point cloud.

[0027] In this embodiment, the semantic communication system includes a transmitter and a receiver, which can be deployed on different devices for holographic communication of 3D point clouds. After acquiring the original point cloud, the transmitter in the semantic communication system can send the point cloud to the receiver, which then completes the downstream target task based on the received point cloud. In this embodiment, after acquiring the original point cloud, the transmitter in the semantic communication system can perform semantic encoding on the original point cloud to extract semantic features, thus obtaining the semantic features of the original point cloud, i.e., the semantic features of each point in the original point cloud.

[0028] Step S12: For each point in the original point cloud, the sending end determines the local semantic features of the point based on the correlation between the semantic features of multiple neighboring points and the semantic features of the point.

[0029] In this embodiment, the transmitting end can determine multiple neighboring points for each point in the original point cloud, and based on the semantic features of each of the multiple neighboring points and the semantic features of the point itself, obtain the correlation between the semantic features of each of the multiple neighboring points and the semantic features of the point. Then, based on the correlation between the semantic features of the multiple neighboring points and the semantic features of the point itself, the local semantic features of the point are determined. In this way, the transmitting end can obtain the local semantic features of each point in the original point cloud.

[0030] Step S13: The sending end determines the semantic information criticality of each point based on the local semantic features of each point in the original point cloud.

[0031] In this embodiment, the sending end can determine the semantic information criticality of each point based on the local semantic features of each point in the original point cloud, thereby obtaining the semantic information criticality of each point in the original point cloud. The semantic information criticality in this embodiment represents the semantic importance of points in the point cloud. The greater the semantic information criticality of a point, the richer the semantic information and the greater its importance.

[0032] Step S14: Sort the points in the original point cloud according to the semantic information keyness from largest to smallest, and perform semantic sampling on the first M points to obtain downsampled semantic features; M is an integer greater than 1.

[0033] In this embodiment, the sending end can globally sort the points in the original point cloud according to their semantic information criticality, from highest to lowest, to determine the top M points. These top M points are the sampling points, representing the main content in the original point cloud that has high semantic importance and rich semantic information. In this embodiment, M can be freely set according to project needs or experience; the specific value of M is not limited in this embodiment.

[0034] After obtaining the first M points in the original point cloud, semantic sampling is performed on the first M points, that is, semantic encoding is performed on the first M points to obtain downsampled semantic features, thereby realizing semantic downsampling of the original point cloud.

[0035] Step S15: The sending end sends the downsampled semantic features to the receiving end in the semantic communication system.

[0036] In this embodiment, after obtaining the downsampled semantic features corresponding to the original point cloud, the transmitting end can send the downsampled semantic features to the receiving end in the semantic communication system via a wireless channel.

[0037] Step S16: The receiving end recovers the downsampled point cloud based on the received downsampled semantic features.

[0038] In this embodiment, the receiver in the semantic communication system can receive downsampled semantic features sent by the transmitter via a wireless channel, thus obtaining the received downsampled semantic features. The receiver can then perform semantic decoding based on the received downsampled semantic features to obtain the downsampled point cloud, i.e., reconstruct the downsampled point cloud. However, due to noise and other factors in the wireless channel, transmission may suffer some loss; therefore, the downsampled semantic features sent by the transmitter and the downsampled semantic features received by the receiver in this embodiment are different.

[0039] Step S17: The receiving end performs a target task based on the downsampled point cloud. The target task includes at least one of the following: classification task and semantic segmentation task.

[0040] In this embodiment, the receiving end can perform downstream target tasks based on the recovered downsampled point cloud. The target tasks include at least one of the following: classification task and semantic segmentation task.

[0041] In this embodiment, for 3D point clouds used in holographic communication, sampling and semantic coding are jointly designed to achieve joint optimization of semantic sampling: by combining global and local semantic-driven point cloud downsampling and combining semantic encoding and decoding, it is beneficial to ensure that the "main" content with high semantic importance or rich semantic information is accurately extracted and encoded in subsequent processing, while other redundant parts are reduced in attention or even ignored. Based on the semantic content of the object itself, the original point cloud can be strategically downsampled. The point cloud downsampling integrates semantic-driven methods, which can adaptively adjust to different point clouds, thereby reducing the source magnitude and further reducing the transmission overhead of 3D point cloud holographic communication, reducing the burden of network processing and real-time transmission, and improving end-to-end communication efficiency.

[0042] In conjunction with the above embodiments, in one implementation, the present invention also provides a three-dimensional point cloud sampling-semantic joint optimization method for holographic communication. This method is implemented using a pre-trained three-dimensional point cloud sampling-semantic joint model. In this method, the transmitting end deploys a three-dimensional point cloud sampling-semantic joint encoder within the three-dimensional point cloud sampling-semantic joint model. The three-dimensional point cloud sampling-semantic joint encoder includes an attention layer and a semantic sampling layer. The downsampled semantic features are obtained according to the following steps:

[0043] Step S21: For each point in the original point cloud, the sending end inputs the semantic features of multiple neighboring points of the point and the semantic features of the point into the attention layer to obtain the local semantic features of the point and the V value of the point.

[0044] In this embodiment, for each point in the original point cloud, the transmitting end inputs the semantic features of multiple neighboring points of the point and the semantic features of the point itself into the attention layer of the 3D point cloud sampling-semantic joint encoder. Through the attention layer, the correlation between the semantic features of multiple neighboring points of the point and the semantic features of the point is processed to obtain the local semantic features of the point output by the attention layer and the V value corresponding to the point, thereby obtaining the local semantic features of each point in the original point cloud and the V value of each point.

[0045] Step S22: The sending end inputs the local semantic features of each point in the original point cloud and the V value of each point into the semantic sampling layer to obtain the normalized correlation map of each point in the original point cloud. Based on the normalized correlation map of each point in the original point cloud, the semantic information criticality of each point in the original point cloud is determined. Based on the semantic information criticality, the normalized correlation map and the V value of each point in the original point cloud, the downsampled semantic features are obtained.

[0046] In this embodiment, after obtaining the local semantic features and V-values ​​of each point in the original point cloud, the transmitting end inputs the local semantic features and V-values ​​of each point in the original point cloud into the semantic sampling layer of the 3D point cloud sampling-semantic joint encoder. The semantic sampling layer can first obtain the normalized correlation map of each point in the original point cloud based on the local semantic features of each point; then, based on the normalized correlation map of each point in the original point cloud, obtain the semantic information keyness of each point in the original point cloud, and thus obtain the semantic information keyness of each point in the original point cloud; then, based on the semantic information keyness of each point in the original point cloud, the normalized correlation map of each point, and the V-value of each point, the semantic sampling layer obtains the downsampled semantic features corresponding to the original point cloud.

[0047] In conjunction with any of the above embodiments, the present invention also provides a three-dimensional point cloud sampling-semantic joint optimization method for holographic communication. In this method, the step S21 above, "the transmitting end inputs the semantic features of multiple neighboring points of each point in the original point cloud and the semantic features of the point into the attention layer to obtain the local semantic features of the point," specifically includes the following steps S31 to S33:

[0048] Step S31: For each point in the original point cloud, the sending end inputs the differences between the semantic features of the point and the semantic features of the point from the semantic features of the multiple neighboring points of the point into the K-linear layer in the attention layer to obtain the K value of the point.

[0049] In this embodiment, the sending end can determine the differences between the semantic features of multiple neighboring points and the semantic features of the point for each point in the original point cloud, and input the differences between the semantic features of multiple neighboring points and the semantic features of the point into the K-linear layer in the attention layer to obtain the K value of the point, and then obtain the K value of each point in the original point cloud.

[0050] Step S32: For each point in the original point cloud, the sending end inputs the semantic features of that point into the Q-linear layer in the attention layer to obtain the Q value of that point.

[0051] In this embodiment, for each point in the original point cloud, the sending end inputs the semantic features of that point into the Q-linear layer in the attention layer to obtain the Q value of that point, and then obtains the Q value of each point in the original point cloud.

[0052] Step S33: For each point in the original point cloud, the sending end determines the correlation metric between the point and multiple neighboring points based on the point's Q value and K value, and uses it as the local semantic feature of the point.

[0053] In this embodiment, after obtaining the K value and Q value of each point in the original point cloud, the sending end can determine the correlation measure between the point and multiple neighboring points based on the Q value and K value of the point, and use the correlation measure between the point and multiple neighboring points as the local semantic feature of the point.

[0054] In an alternative implementation, the correlation metric (i.e., the local semantic features of the i-th point) between the i-th point and multiple neighboring points in the original point cloud can be determined in the following way:

[0055] ;

[0056] in, and These represent the Q-linear layer (applied to query input) and the K-linear layer (applied to key input) in the attention layer, respectively, with T being the transpose; the semantic features of the i-th point in the original point cloud are... As input to the Q-linear layer, the semantic features of the j neighboring points of the i-th point are... Semantic features of the i-th point respectively The difference is used as the input to the K-linear layer.

[0057] Furthermore, in an optional implementation, a normalized correlation map of each point in the original point cloud is obtained based on the local semantic features of each point, which can be obtained by the following formula:

[0058] ;

[0059] in, This represents the normalized correlation graph of the i-th point in the original point cloud; Softmax represents the Softmax function. This represents the local semantic features of the i-th point in the original point cloud. Let be a local point cloud consisting of j neighboring points corresponding to the i-th point; j represents the j neighboring points of the i-th point; d represents the dimension of the vector. It is used as the denominator for normalization.

[0060] In conjunction with any of the above embodiments, in one implementation, the present invention also provides a three-dimensional point cloud sampling-semantic joint optimization method for holographic communication. In this method, step S21, "the transmitting end inputs the semantic features of multiple neighboring points of each point in the original point cloud, along with the semantic features of the point itself, into the attention layer to obtain the V value of that point," specifically includes step S41. Furthermore, step S22, "based on the semantic information keyness, normalized correlation graph, and V value of each point in the original point cloud, the obtained downsampled semantic features," specifically includes steps S42 to S45.

[0061] Step S41: For each point in the original point cloud, the sending end inputs the differences between the semantic features of the point and the semantic features of the point from the semantic features of the multiple neighboring points of the point into the V linear layer in the attention layer to obtain the V value of the point.

[0062] In this embodiment, for each point in the original point cloud, the sending end determines the differences between the semantic features of the multiple neighboring points of the point and the semantic features of the point, and inputs the differences between the semantic features of the multiple neighboring points of the point and the semantic features of the point into the V linear layer in the attention layer to obtain the V value of the point, and then obtains the V value of each point in the original point cloud.

[0063] Step S42: The sending end sorts the points in the original point cloud according to the semantic information keyness from large to small, and determines the index of the first M points.

[0064] In this embodiment, the transmitting end uses the semantic sampling layer in the 3D point cloud sampling-semantic joint encoder to sort the points in the original point cloud globally based on the semantic information keyness of each point in the original point cloud, in descending order of semantic information keyness, to determine the first M points and the index of the first M points.

[0065] Step S43: The sending end selects the V values ​​of the first M points from the V values ​​of each point in the original point cloud based on the index of the first M points.

[0066] In this embodiment, the transmitting end uses the semantic sampling layer in the 3D point cloud sampling-semantic joint encoder to filter out the V values ​​corresponding to the first M points from the V values ​​of each point in the original point cloud based on the index of the first M points.

[0067] Step S44: Based on the index of the first M points, the sending end selects the normalized correlation graph of the first M points from the normalized correlation graph of each point in the original point cloud.

[0068] In this embodiment, the transmitting end uses the semantic sampling layer in the 3D point cloud sampling-semantic joint encoder to filter out the normalized correlation map corresponding to the first M points from the normalized correlation map of each point in the original point cloud based on the index of the first M points.

[0069] Step S45: The transmitting end obtains the downsampling features based on the V values ​​of the first M points and the normalized correlation graph of the first M points.

[0070] In this embodiment, the transmitting end processes the V values ​​of the first M points and the normalized correlation graph of the first M points through the semantic sampling layer in the 3D point cloud sampling-semantic joint encoder to obtain the downsampling features corresponding to the original point cloud.

[0071] In conjunction with any of the above embodiments, in one implementation, the present invention also provides a three-dimensional point cloud sampling-semantic joint optimization method for holographic communication. In this method, the step S22 above, "determining the semantic information criticality of each point in the original point cloud based on the normalized correlation graph of each point in the original point cloud," specifically includes step S51:

[0072] Step S51: For each point in the original point cloud, the sending end calculates the standard deviation based on the T elements corresponding to that point, which is used as the semantic information criticality of that point.

[0073] In this embodiment, the number of neighboring points for each point in the original point cloud is T, where T is an integer greater than 0. In an optional implementation, for each point in the original point cloud, the kNN method (K-nearest neighbor method) can be used to find T neighboring points of that point, thus obtaining the T neighboring points of that point. The normalized correlation graph of each point in the original point cloud is a vector representation including T elements, that is, the normalized correlation graph (vector representation) corresponding to that point includes T elements, and the t-th element represents the correlation between the semantic features of the t-th neighboring point and the semantic features of that point, where t∈[1,T].

[0074] In this embodiment, the transmitting end uses the semantic sampling layer in the 3D point cloud sampling-semantic joint encoder to calculate the standard deviation of each point in the original point cloud based on the T elements corresponding to that point, and uses the standard deviation of that point as the semantic information criticality of that point.

[0075] In this embodiment, considering the significant differences between the point cloud at the edge and its surrounding point cloud in the original point cloud, the standard deviation of its normalized correlation map is larger, thus possessing richer semantic information and greater importance (e.g., Figure 2 As shown, Figure 2 This is an example diagram illustrating a comparison of the standard deviations of edge points and non-edge points in a point cloud, as shown in an embodiment of the present invention. Figure 2 In the diagram, point C is an edge point, and point D is a non-edge point. It is evident that edge points differ more significantly from non-edge points and have a larger standard deviation. Therefore, the standard deviation corresponding to this point is used as its semantic information criticality. Based on the semantic information criticality in descending order, the points in the original point cloud are sorted to determine the top M points. This maximizes the preservation of edge points in the original point cloud while removing non-edge points, ensuring that semantically important or semantically rich "main" content is semantically sampled, thus achieving semantic downsampling.

[0076] In one embodiment, such as Figure 3 As shown, Figure 3 This is a schematic diagram illustrating the processing of the attention layer and semantic sampling layer in a three-dimensional point cloud sampling-semantic joint encoder according to an embodiment of the present invention. Figure 3 In the input original point cloud, for the semantic features of each point (dimension: N x d, where N represents the total number of points in the original point cloud), the attention layer uses kNN (K-nearest neighbor method) to find the k nearest neighbors of each point (dimension: N x k x d). The semantic features of each point in the original point cloud (dimension: N x 1 x d) are processed by a Q-linear layer to obtain the Q-value of each point (dimension: N x 1 x d). The differences between the semantic features of the k nearest neighbors of each point in the original point cloud (dimension: N x k x d) and the semantic features of the point itself (dimension: N x d, i.e., N x 1 x d) are input into the K-linear layer of the attention layer to obtain the K-value of each point (dimension: N x k x d). Finally, the differences between the semantic features of the k nearest neighbors of each point in the original point cloud (dimension: N x k x d) and the semantic features of the point itself (dimension: N x d, i.e., N x 1 x d) are input into the V-linear layer of the attention layer to obtain the V-value of each point (dimension: N x k x d). (xkxd). Then, the attention layer determines the local semantic features of each point in the original point cloud based on the Q-value and K-value of that point.

[0077] The local semantic features and V values ​​of each point in the original point cloud are input into the semantic sampling layer. The semantic sampling layer processes the local semantic features of each point through a softmax operation to obtain a normalized correlation map (dimension N x 1 x K) for each point. Then, the standard deviation is calculated based on the normalized correlation map of each point. The semantic information key score of each point is obtained, resulting in a total of N semantic information key scores. Then, the points in the original point cloud are sorted in descending order of the N semantic information key scores to determine the top M points. Based on the index of the top M points, the normalized correlation maps (with dimensions of M x 1 x K) of the top M points are selected from the normalized correlation maps (with dimensions of M x 1 x K) of the points in the original point cloud. Also, based on the index of the top M points, the V values ​​(with dimensions of M x 1 x d) of the top M points are selected from the V values ​​(with dimensions of N x k x d) of the points in the original point cloud. Finally, the semantic sampling layer outputs the downsampling features corresponding to the original point cloud based on the V values ​​and the normalized correlation maps of the top M points.

[0078] In conjunction with any of the above embodiments, in one implementation, the present invention also provides a three-dimensional point cloud sampling-semantic joint optimization method for holographic communication. In this method, the method is implemented using a pre-trained three-dimensional point cloud sampling-semantic joint model. The transmitting end is deployed with a three-dimensional point cloud sampling-semantic joint encoder in the three-dimensional point cloud sampling-semantic joint model, and the receiving end is deployed with a task network corresponding to the target task and a decoder in the three-dimensional point cloud sampling-semantic joint model. The training process of the three-dimensional point cloud sampling-semantic joint model includes a first stage and a second stage.

[0079] The training in the first stage yields an intermediate 3D point cloud sampling-semantic joint encoder, an intermediate decoder, and a trained task network.

[0080] In this embodiment, by training the first stage of the 3D point cloud sampling-semantic joint model, we can obtain the intermediate 3D point cloud sampling-semantic joint encoder, intermediate decoder, and the trained task network.

[0081] And the second stage may include steps S61 to S65:

[0082] Step S61: The sending end processes the original point cloud of the sample through the intermediate 3D point cloud sampling-semantic joint encoder to obtain the sample downsampled semantic features.

[0083] In this embodiment, the original point cloud of the sample is the original point cloud used for training the 3D point cloud sampling-semantic joint model, and the sample downsampled semantic features are the downsampled semantic features obtained during the training process of the 3D point cloud sampling-semantic joint model. The sending end can process the original point cloud of the sample through an intermediate 3D point cloud sampling-semantic joint encoder to obtain the sample downsampled semantic features.

[0084] In one optional implementation, the transmitting end can extract semantic features of the original point cloud of the sample using an intermediate 3D point cloud sampling-semantic joint encoder. For each sample point in the original point cloud, based on the correlation between the semantic features of multiple neighboring sample points and the semantic features of the sample point itself, the sample local semantic features of that sample point are determined. Based on the sample local semantic features of each sample point in the original point cloud, the semantic information criticality of that sample point is determined. The sample points in the original point cloud are sorted according to the semantic information criticality from largest to smallest, and semantic sampling is performed on the first M sample points to obtain the sample downsampled semantic features. The method of obtaining the sample downsampled semantic features is similar to the method of obtaining the downsampled semantic features in steps S11 to S14, and can be implemented with reference to steps S11 to S14.

[0085] Step S62: The sending end sends the sample downsampling semantic features to the receiving end.

[0086] In this embodiment, after obtaining the sample downsampling semantic features, the sending end can send the sample downsampling semantic features to the receiving end via a wireless channel.

[0087] Step S63: The receiving end recovers the sample downsampled point cloud through the intermediate decoder based on the received sample downsampled semantic features.

[0088] In this embodiment, the receiving end can receive the sample downsampling semantic features sent by the sending end through a wireless channel, thus obtaining the received sample downsampling semantic features. Based on the received sample downsampling semantic features, the receiving end can perform semantic decoding through an intermediate decoder deployed in the receiving end to obtain the recovered sample downsampling point cloud, i.e., recover the sample downsampling point cloud.

[0089] Step S64: The receiving end executes the target task through the trained task network based on the recovered sample downsampled point cloud to obtain the first task execution result.

[0090] In this embodiment, the receiving end can execute the target task based on the recovered sample downsampled point cloud, through the trained task network deployed in the receiving end, and obtain the first task execution result output by the trained task network.

[0091] Step S65: Based at least on the execution result of the first task and the labels carried by the original point cloud of the sample, fine-tune the parameters of the intermediate 3D point cloud sampling-semantic joint encoder and the intermediate decoder to obtain the 3D point cloud sampling-semantic joint model.

[0092] In this embodiment, the intermediate 3D point cloud sampling-semantic joint encoder and intermediate decoder can be fine-tuned based at least on the execution result of the first task and the labels carried by the original point cloud samples, so as to finally obtain the trained 3D point cloud sampling-semantic joint encoder and trained decoder, and thus obtain a 3D point cloud sampling-semantic joint model composed of the trained 3D point cloud sampling-semantic joint encoder and trained decoder.

[0093] Specifically, when the target task is a classification task, the execution result of the first task is the result of the first classification task, and the label carried by the original point cloud of the sample is the category label of the original point cloud of the sample; when the target task is a semantic segmentation task, the execution result of the first task is the result of the first semantic segmentation task, and the label carried by the original point cloud of the sample is the semantic label of the original point cloud of the sample.

[0094] In conjunction with any of the above embodiments, in one implementation, the present invention also provides a three-dimensional point cloud sampling-semantic joint optimization method for holographic communication. In this method, the second stage may include steps S71 to S74:

[0095] Step S71: The sending end extracts the semantic features of the original point cloud of the sample through the three-dimensional point cloud sampling-semantic joint encoder to be trained; the sending end sends the semantic features of the original point cloud of the sample to the receiving end.

[0096] In this embodiment, the transmitting end can use the three-dimensional point cloud sampling-semantic joint encoder to be trained to perform semantic encoding on the original point cloud of the sample, extract the semantic features of the original point cloud of the sample, and send the semantic features of the original point cloud of the sample to the receiving end through a wireless channel.

[0097] Step S72: The receiving end recovers the original point cloud of the sample based on the semantic features of the received sample original point cloud through the decoder to be trained.

[0098] In this embodiment, the receiving end can receive the semantic features of the original point cloud of the sample sent by the transmitting end through a wireless channel, and obtain the semantic features of the received original point cloud of the sample. Based on the semantic features of the received original point cloud of the sample, the receiving end can perform semantic decoding through the decoder to be trained in the receiving end to obtain the recovered original point cloud of the sample, that is, to recover the original point cloud of the sample.

[0099] Step S73: The receiving end executes the target task through the task network to be trained based on the recovered original point cloud of the sample to obtain the second task execution result.

[0100] In this embodiment, the receiving end can execute the target task based on the recovered original point cloud of the sample through the task network to be trained deployed in the receiving end, and obtain the second task execution result output by the task network to be trained.

[0101] Step S74: Based on the execution result of the second task and the labels carried by the original point cloud of the sample, update the parameters of the 3D point cloud sampling-semantic joint encoder to be trained, the decoder to be trained, and the task network to be trained to obtain the intermediate 3D point cloud sampling-semantic joint encoder, the intermediate decoder, and the trained task network.

[0102] In this embodiment, the parameters of the 3D point cloud sampling-semantic joint encoder, the decoder, and the task network to be trained can be updated based on the execution result of the second task and the labels carried by the original point cloud samples, so as to obtain the intermediate 3D point cloud sampling-semantic joint encoder, the intermediate decoder, and the trained task network.

[0103] Specifically, when the target task is a classification task, the result of the second task is the result of the second classification task, and the label carried by the original point cloud of the sample is the category label of the original point cloud of the sample; when the target task is a semantic segmentation task, the result of the second task is the result of the second semantic segmentation task, and the label carried by the original point cloud of the sample is the semantic label of the original point cloud of the sample.

[0104] In one embodiment, such as Figure 4 As shown, Figure 4 This is a schematic diagram illustrating the structure of a 3D point cloud sampling-semantic joint encoder and decoder in a 3D point cloud sampling-semantic joint model according to an embodiment of the present invention. Figure 4 A 3D point cloud sampling-semantic joint encoder includes at least a feature embedding layer, two attention layers, and two semantic sampling layers; the decoder in the 3D point cloud sampling-semantic joint model includes at least an attention layer, a multilayer perceptron, a max pooling layer, and an activation function layer (such as Softmax). At the 3D point cloud sampling-semantic joint encoder, the original point cloud is first... Embedded into the feature space (i.e., the input feature embedding layer), its semantic information is extracted. The input to the global attention layer (i.e., the attention layer) then guides the subsequent semantic sampling layer to perform semantic sampling, resulting in downsampled features. Finally, at the decoder, the sampled encoded information is restored to a downsampled point cloud. It also performs downstream tasks such as classification.

[0105] In conjunction with any of the above embodiments, in one implementation, the present invention also provides a three-dimensional point cloud sampling-semantic joint optimization method for holographic communication. In this method, step S65 may specifically include steps S81 to S83:

[0106] Step S81: The receiving end determines the sampling loss based on the difference between the recovered sample downsampled point cloud and the original sample point cloud.

[0107] In this embodiment, the receiving end can determine the difference between the restored downsampled point cloud and the original point cloud based on the restored sample downsampled point cloud and the original point cloud, and determine the sampling loss based on the difference between the restored downsampled point cloud and the original point cloud. .

[0108] Step S82: The receiving end determines the task loss based on the difference between the first task execution result and the second task execution result, and the difference between the first task execution result and the label carried by the original point cloud of the sample.

[0109] In this embodiment, the receiving end can further determine the difference between the first task execution result and the second task execution result based on the first task execution result and the second task execution result, and determine the difference between the first task execution result and the labels carried by the original point cloud of the sample based on the first task execution result and the labels carried by the original point cloud of the sample; then, based on the difference between the first task execution result and the second task execution result and the difference between the first task execution result and the labels carried by the original point cloud of the sample, the task loss is determined. .

[0110] Step S83: Based on the sampling loss and the task loss, fine-tune the parameters of the intermediate 3D point cloud sampling-semantic joint encoder and the intermediate decoder to obtain the 3D point cloud sampling-semantic joint model.

[0111] In this embodiment, the total loss can be obtained based on the sampling loss and task loss. The parameters of the intermediate 3D point cloud sampling-semantic joint encoder and intermediate decoder can be fine-tuned based on the total loss until the total loss converges. The parameters of the intermediate 3D point cloud sampling-semantic joint encoder and intermediate decoder are fixed when the total loss converges, and the trained 3D point cloud sampling-semantic joint encoder and trained decoder are obtained, thus obtaining the 3D point cloud sampling-semantic joint model.

[0112] In one embodiment, such as Figure 5 As shown, Figure 5 This is a schematic diagram illustrating the architecture of a semantically oriented point cloud sampling network according to an embodiment of the present invention. Figure 5In this context, the semantic-oriented point cloud sampling network includes a 3D point cloud sampling-semantic joint encoder, a decoder, and a task network. The 3D point cloud sampling-semantic joint encoder and decoder constitute the 3D point cloud sampling-semantic joint model. Training of the 3D point cloud sampling-semantic joint model is divided into a first-stage training and a second-stage training.

[0113] First, for the first stage of training, the input sample is the original point cloud. The training network performs end-to-end semantic encoding and decoding using a 3D point cloud sampling-semantic joint encoder and a decoder. Then, it performs target tasks such as classification and semantic segmentation on the task network to be trained, resulting in the second task execution result. Based on the second task execution result and the labels carried by the original point cloud samples, the parameters of the training network, the decoder, and the task network are updated to obtain an intermediate 3D point cloud sampling-semantic joint encoder and an intermediate decoder. The parameters of the task network are then frozen to obtain the trained task network.

[0114] For the second phase of training, the original point cloud of the samples is used during training. The input intermediate 3D point cloud sampling-semantic co-encoder performs downsampling. Then, through the intermediate 3D point cloud sampling-semantic co-encoder, intermediate decoder, and the trained task network, after semantic encoding and decoding, the task is executed again to obtain the reconstructed downsampled point cloud. The first task execution result output by the trained task network.

[0115] In the second stage of training, the sampling loss is determined using the recovered downsampled point cloud of the sample output from the intermediate decoder and the original point cloud of the sample. The task loss is determined by the difference between the results of the first and second tasks, and by the difference between the results of the first task and the labels carried by the original point cloud of the sample. The total loss function is: Where α is the weight, which can be set based on experience; , These are the parameters to be fine-tuned for the intermediate 3D point cloud sampling-semantic joint encoder and the intermediate decoder, respectively. These are the parameters of the trained task network. Then, based on the total loss function, the parameters of the intermediate 3D point cloud sampling-semantic joint encoder and intermediate decoder are fine-tuned to obtain the 3D point cloud sampling-semantic joint model.

[0116] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.

[0117] Based on the same inventive concept, another embodiment of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in the three-dimensional point cloud sampling-semantic joint optimization method for holographic communication as described in any of the above embodiments of the present invention.

[0118] Based on the same inventive concept, another embodiment of the present invention provides an electronic device, such as... Figure 6 As shown, Figure 6 This is a schematic diagram of an electronic device according to an embodiment of the present invention. The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When executed by the processor, the program implements the steps of the three-dimensional point cloud sampling-semantic joint optimization method for holographic communication described in any of the above embodiments of the present invention.

[0119] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0120] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, apparatus, or computer program products. Therefore, embodiments of the present invention can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of the present invention can take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0121] Embodiments of the present invention are described with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0122] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0123] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0124] Although preferred embodiments of the present invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present invention.

[0125] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.

[0126] The foregoing has provided a detailed description of a three-dimensional point cloud sampling-semantic joint optimization method, device, and medium for holographic communication provided by the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A three-dimensional point cloud sampling-semantic joint optimization method for holographic communication, characterized in that, The method includes: In a semantic communication system, the sending end extracts semantic features from the original point cloud; For each point in the original point cloud, the transmitting end determines the local semantic features of the point based on the correlation between the semantic features of multiple neighboring points of the point and the semantic features of the point itself. The sending end determines the semantic information criticality of each point based on the local semantic features of each point in the original point cloud; The points in the original point cloud are sorted in descending order of semantic information criticality, and the first M points are semantically sampled to obtain downsampled semantic features; M is an integer greater than 1. The sending end sends the downsampled semantic features to the receiving end in the semantic communication system; The receiving end recovers the downsampled point cloud based on the received downsampled semantic features; The receiving end performs a target task based on the downsampled point cloud, and the target task includes at least one of the following: classification task and semantic segmentation task.

2. The three-dimensional point cloud sampling-semantic joint optimization method for holographic communication according to claim 1, characterized in that, The method is implemented using a pre-trained 3D point cloud sampling-semantic joint model. The transmitting end is deployed with a 3D point cloud sampling-semantic joint encoder within the 3D point cloud sampling-semantic joint model. The 3D point cloud sampling-semantic joint encoder includes an attention layer and a semantic sampling layer. The downsampled semantic features are obtained according to the following steps: For each point in the original point cloud, the transmitting end inputs the semantic features of multiple neighboring points of that point and the semantic features of that point into the attention layer to obtain the local semantic features of that point and the V value of that point. The transmitting end inputs the local semantic features of each point in the original point cloud and the V value of each point into the semantic sampling layer to obtain the normalized correlation map of each point in the original point cloud. Based on the normalized correlation map of each point in the original point cloud, the semantic information criticality of each point in the original point cloud is determined. Based on the semantic information criticality, the normalized correlation map and the V value of each point in the original point cloud, the downsampled semantic features are obtained.

3. The three-dimensional point cloud sampling-semantic joint optimization method for holographic communication according to claim 2, characterized in that, For each point in the original point cloud, the transmitting end inputs the semantic features of multiple neighboring points of that point, along with the semantic features of that point itself, into the attention layer to obtain the local semantic features of that point, including: For each point in the original point cloud, the transmitting end inputs the differences between the semantic features of the point and the semantic features of the point itself from the semantic features of the multiple neighboring points of that point into the K-linear layer in the attention layer to obtain the K value of that point. For each point in the original point cloud, the transmitting end inputs the semantic features of that point into the Q-linear layer in the attention layer to obtain the Q value of that point. For each point in the original point cloud, the transmitting end determines the correlation metric between the point and multiple neighboring points based on the point's Q value and K value, and uses it as the local semantic feature of the point.

4. The three-dimensional point cloud sampling-semantic joint optimization method for holographic communication according to claim 2, characterized in that, For each point in the original point cloud, the transmitting end inputs the semantic features of multiple neighboring points of that point along with the semantic features of that point into the attention layer to obtain the V value of that point, including: For each point in the original point cloud, the transmitting end inputs the differences between the semantic features of the point and the semantic features of the point's multiple neighboring points into the V linear layer of the attention layer to obtain the V value of the point. Based on the semantic information keyness, normalized correlation graph, and V value of each point in the original point cloud, the downsampled semantic features are obtained, including: The sending end sorts the points in the original point cloud in descending order of semantic information keyness to determine the indices of the first M points; The sending end selects the V values ​​of the first M points from the V values ​​of each point in the original point cloud based on the index of the first M points. Based on the index of the first M points, the sending end filters out the normalized correlation graph of the first M points from the normalized correlation graph of each point in the original point cloud; The transmitting end obtains the downsampling semantic features based on the V values ​​of the first M points and the normalized correlation graph of the first M points.

5. The three-dimensional point cloud sampling-semantic joint optimization method for holographic communication according to claim 2, characterized in that, The number of neighboring points for each point in the original point cloud is T. The normalized correlation graph of the point is a vector representation including T elements, where the t-th element represents the correlation between the semantic features of the t-th neighboring point and the semantic features of the point. Based on the normalized correlation graph of each point in the original point cloud, the semantic information criticality of each point in the original point cloud is determined, including: For each point in the original point cloud, the sending end calculates the standard deviation based on the T elements corresponding to that point, which serves as the semantic information criticality of that point; where T is an integer greater than 0, and t∈[1,T].

6. The three-dimensional point cloud sampling-semantic joint optimization method for holographic communication according to claim 1, characterized in that, The method is implemented through a pre-trained 3D point cloud sampling-semantic joint model. The transmitting end is equipped with a 3D point cloud sampling-semantic joint encoder in the 3D point cloud sampling-semantic joint model, and the receiving end is equipped with a task network corresponding to the target task and a decoder in the 3D point cloud sampling-semantic joint model. The training process of the 3D point cloud sampling-semantic joint model includes a first stage and a second stage. Through the first stage of training, the intermediate 3D point cloud sampling-semantic joint encoder, intermediate decoder, and the trained task network are obtained. The second stage is: The transmitting end processes the original point cloud of the sample through the intermediate three-dimensional point cloud sampling-semantic joint encoder to obtain the sample downsampled semantic features; The sending end sends the sample downsampled semantic features to the receiving end; The receiving end recovers the sample downsampled point cloud through the intermediate decoder based on the received sample downsampled semantic features; The receiving end executes the target task based on the recovered sample downsampled point cloud through the trained task network to obtain the first task execution result; Based at least on the execution result of the first task and the labels carried by the original point cloud of the sample, the parameters of the intermediate 3D point cloud sampling-semantic joint encoder and the intermediate decoder are fine-tuned to obtain the 3D point cloud sampling-semantic joint model.

7. The three-dimensional point cloud sampling-semantic joint optimization method for holographic communication according to claim 6, characterized in that, The first stage is: The transmitting end extracts the semantic features of the original point cloud of the sample through the three-dimensional point cloud sampling-semantic joint encoder to be trained; the transmitting end sends the semantic features of the original point cloud of the sample to the receiving end. The receiving end recovers the original point cloud of the sample based on the semantic features of the received sample original point cloud through the decoder to be trained. The receiving end executes the target task based on the recovered original point cloud of the sample through the task network to be trained, and obtains the second task execution result; Based on the execution result of the second task and the labels carried by the original point cloud of the sample, the parameters of the 3D point cloud sampling-semantic joint encoder to be trained, the decoder to be trained, and the task network to be trained are updated to obtain the intermediate 3D point cloud sampling-semantic joint encoder, the intermediate decoder, and the trained task network.

8. The three-dimensional point cloud sampling-semantic joint optimization method for holographic communication according to claim 7, characterized in that, Based at least on the execution result of the first task and the labels carried by the original point cloud samples, the parameters of the intermediate 3D point cloud sampling-semantic joint encoder and the intermediate decoder are fine-tuned to obtain the 3D point cloud sampling-semantic joint model, including: The receiving end determines the sampling loss based on the difference between the recovered sample downsampled point cloud and the original sample point cloud; The receiving end determines the task loss based on the difference between the first task execution result and the second task execution result, and the difference between the first task execution result and the label carried by the original point cloud of the sample. Based on the sampling loss and the task loss, the parameters of the intermediate 3D point cloud sampling-semantic joint encoder and the intermediate decoder are fine-tuned to obtain the 3D point cloud sampling-semantic joint model.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the computer program is executed by the processor, it implements the three-dimensional point cloud sampling-semantic joint optimization method for holographic communication as described in any one of claims 1 to 8.

10. A computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is executed by the processor, it implements the three-dimensional point cloud sampling-semantic joint optimization method for holographic communication as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Semantic segmentation method for point cloud in outdoor large scene

    CN112560865A

  • Semantic codec training method, semantic codec transmission method and semantic codec training system for point cloud transmission

    CN117135179A