Holographic communication-oriented three-dimensional point cloud sampling-semantic joint optimization method, equipment and medium
Through the semantic joint optimization method of holographic communication, the semantic features of the point cloud are extracted and downsampled, which solves the problems of large amount and disorder of point cloud data in holographic communication and improves the communication efficiency and the accuracy of semantic content.
Patent Information
- Application Number
- CN202511141478.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-08-15
AI Technical Summary
The large amount of point cloud data in holographic communication leads to high end-to-end processing latency, and the disorder and permutation invariance of point clouds limit network sampling, making it impossible to meet the requirements of high bandwidth, low latency and high reliability.
The semantic features of the point cloud are extracted through the semantic communication system, the criticality is determined based on the local semantic features, downsampling is performed, and the point cloud is restored at the receiving end to achieve joint optimization of semantic sampling.
It reduces the amount of point cloud data, reduces the burden on network processing and real-time transmission, improves end-to-end communication efficiency, and ensures the accurate extraction and encoding of semantically important content.
Smart Images

Figure CN120635903A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of holographic communication technology, and in particular to a three-dimensional point cloud sampling-semantic joint optimization method, device and medium for holographic communication. Background Art
[0002] Holographic communication can provide an immersive interactive experience, but its differentiated high-performance requirements pose a severe challenge to existing communication networks. Improving the source representation compression and transmission capabilities, as well as resource utilization, are key to holographic communication. Among various 3D data representation methods, point clouds offer six degrees of freedom, powerful representations of 3D spatial geometric detail, and retain rich 3D detail, enabling a truly holographic visual experience. Compared to other technologies like mesh and light field, point clouds require less data, are more scalable, offer convenient algorithm pre- and post-processing, and possess relatively mature acquisition and reconstruction components. Therefore, point clouds are expected to become one of the mainstream 3D data representations for future holographic services.
[0003] Point cloud, a classic method for representing 3D data, enables applications such as real-time high-definition 3D interaction, holographic video conferencing, remote projection of 3D digital humans, 3D environment modeling, and digital twins. In immersive communication scenarios, point cloud data volumes are large. For example, an indoor point cloud model at a university contains over 40,000,000 points, totaling 3.2GB. This results in high end-to-end processing latency. Therefore, while ensuring the transmission of the majority of content, it is necessary to reduce data sampling to alleviate the burden on network processing and real-time transmission.
[0004] The disorder and permutation invariance of point clouds also place limitations on various point cloud networks, including sampling. Immersive communication demands high bandwidth, low latency, and high reliability. These new requirements, along with the coordinated sampling and semantic joint encoding and decoding to improve end-to-end communication efficiency, are pressing challenges in researching point cloud sampling networks. Summary of the Invention
[0005] Based on the above technical problems, the present invention provides a three-dimensional point cloud sampling-semantic joint optimization method, device and medium for holographic communication, aiming to overcome the above problems or at least partially solve the above problems.
[0006] A first aspect of the present invention provides a three-dimensional point cloud sampling-semantics joint optimization method for holographic communication, the method comprising: The sender in the semantic communication system extracts the semantic features of the original point cloud; The transmitting end determines, for each point in the original point cloud, a local semantic feature of the point based on correlations between the semantic features of a plurality of adjacent points of the point and the semantic feature of the point; The sending end determines the semantic information criticality of each point in the original point cloud based on the local semantic features of the point; Sort the points in the original point cloud in descending order of semantic information criticality, perform semantic sampling on the first M points, and obtain downsampled semantic features; The sending end sends the downsampled semantic features to the receiving end in the semantic communication system; The receiving end restores the downsampled point cloud based on the received downsampled semantic features; The receiving end performs a target task based on the downsampled point cloud, where the target task includes at least any one of the following: a classification task and a semantic segmentation task.
[0007] The second aspect of the present invention provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the three-dimensional point cloud sampling-semantic joint optimization method for holographic communication as described in the first aspect of the present invention.
[0008] The third aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the three-dimensional point cloud sampling-semantic joint optimization method for holographic communication as described in the first aspect of the present invention.
[0009] In the three-dimensional point cloud sampling-semantic joint optimization method for holographic communication proposed in the present invention, the original point cloud is extracted to obtain semantic features by the sending end in the semantic communication system, and then for each point in the original point cloud, the local semantic features of the point are determined based on the correlation between the semantic features of multiple adjacent points of the point and the semantic features of the point, and the semantic information relevance of the point is determined based on the local semantic features, so as to globally sort the points in the original point cloud in order of semantic information criticality from large to small to obtain the top M points, and semantically sample the M points to obtain downsampled semantic features; the sending end then sends the downsampled semantic features to the receiving end in the semantic communication system, and the receiving end restores the downsampled point cloud based on the received downsampled semantic features, and then performs the target task based on the downsampled point cloud. In this way, the present invention jointly designs sampling and semantic coding to achieve joint optimization of semantic sampling: by globally combining local semantically driven point cloud downsampling and combining semantic encoding and decoding, it is beneficial to ensure that the "main body" content with high semantic importance or rich semantic information is accurately extracted and encoded in the subsequent processing, and other redundant parts are reduced in attention or even ignored. Based on the semantic content of the object itself, the original point cloud can be strategically downsampled. Point cloud downsampling integrates semantic-driven methods and can adaptively adjust different point clouds, thereby reducing the signal source level and further reducing the transmission overhead of three-dimensional point cloud holographic communication, reducing the burden of network processing and real-time transmission, and improving end-to-end communication efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0011] Figure 1 This is a flowchart of a three-dimensional point cloud sampling-semantic joint optimization method for holographic communication according to an embodiment of the present invention; Figure 2 This is an example diagram showing a comparison of standard deviations of edge points and non-edge points in a point cloud according to an embodiment of the present invention; Figure 3 1 is a schematic diagram illustrating the processing of an attention layer and a semantic sampling layer in a three-dimensional point cloud sampling-semantic joint encoder according to an embodiment of the present invention; Figure 4 3D point cloud sampling-semantics joint encoder and decoder in a 3D point cloud sampling-semantics joint model according to an embodiment of the present invention; Figure 51 is a schematic diagram of the architecture of a semantic-oriented point cloud sampling network according to an embodiment of the present invention; Figure 6 FIG. 1 is a schematic diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0012] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0013] Please refer to Figure 1 , Figure 1 This is a flowchart of a three-dimensional point cloud sampling-semantic joint optimization method for holographic communication according to an embodiment of the present invention. Figure 1 As shown, the three-dimensional point cloud sampling-semantic joint optimization method for holographic communication provided in this embodiment includes at least the following steps: Step S11: The sending end in the semantic communication system extracts semantic features of the original point cloud.
[0014] In this embodiment, the semantic communication system includes a transmitter and a receiver, which can be deployed on different devices for holographic communication of three-dimensional point clouds. After the original point cloud is collected, the point cloud can be sent to the receiver of the semantic communication system through the transmitter in the semantic communication system, and the receiver then completes the downstream target task based on the received point cloud. In this embodiment, after obtaining the original point cloud, the transmitter in the semantic communication system can perform semantic encoding on the original point cloud to extract semantic features, and extract the semantic features of the original point cloud, that is, the semantic features of each point in the original point cloud.
[0015] Step S12: The sending end determines, for each point in the original point cloud, a local semantic feature of the point based on the correlation between the semantic features of multiple adjacent points of the point and the semantic feature of the point.
[0016] In this embodiment, the transmitting end can determine multiple adjacent points of each point in the original point cloud, and based on the semantic features of each of the multiple adjacent points of the point and the semantic features of the point, respectively obtain the correlation between the semantic features of each adjacent point of the point and the semantic features of the point, and then determine the local semantic features of the point based on the correlation between the semantic features of the multiple adjacent points of the point and the semantic features of the point. In this way, the transmitting end can obtain the local semantic features of each point in the original point cloud.
[0017] Step S13: The sending end determines the semantic information criticality of each point in the original point cloud based on the local semantic features of the point.
[0018] In this embodiment, the transmitter can determine the semantic information criticality of each point in the original point cloud based on the local semantic features of the point, thereby obtaining the semantic information criticality of each point in the original point cloud. The semantic information criticality of this embodiment represents the semantic importance of a point in the point cloud. The greater the semantic information criticality of a point, the richer the semantic information and the more important the point.
[0019] Step S14: Sort the points in the original point cloud in descending order of semantic information criticality, perform semantic sampling on the first M points, and obtain downsampled semantic features; M is an integer greater than 1.
[0020] In this embodiment, the transmitting end can globally sort the points in the original point cloud in descending order of semantic information criticality based on the semantic information criticality of each point in the original point cloud, and determine the top M points. These top M points are the points to be sampled. These top M points represent the main content of the original point cloud with high semantic importance and rich semantic information. In this embodiment, M can be freely set based on project requirements or experience, and this embodiment does not limit the specific value of M.
[0021] After obtaining the first M points in the original point cloud, semantic sampling is performed on the first M points, that is, semantic encoding is performed on the first M points to obtain downsampled semantic features, thereby realizing semantic downsampling of the original point cloud.
[0022] Step S15: The sending end sends the downsampled semantic features to the receiving end in the semantic communication system.
[0023] In this embodiment, after obtaining the downsampled semantic features corresponding to the original point cloud, the transmitting end may send the downsampled semantic features to the receiving end in the semantic communication system through a wireless channel.
[0024] Step S16: The receiving end restores the downsampled point cloud based on the received downsampled semantic features.
[0025] In this embodiment, a receiving end in a semantic communication system can receive downsampled semantic features sent by a transmitting end via a wireless channel to obtain received downsampled semantic features. The receiving end can perform semantic decoding based on the received downsampled semantic features to obtain a downsampled point cloud, i.e., restore the downsampled point cloud. However, due to noise in the wireless channel, etc., certain transmission losses may occur. Therefore, in this embodiment, the downsampled semantic features sent by the transmitting end differ from the downsampled semantic features received by the receiving end.
[0026] Step S17: The receiving end performs a target task based on the downsampled point cloud, where the target task includes at least one of the following: a classification task and a semantic segmentation task.
[0027] In this embodiment, the receiving end may perform a downstream target task based on the restored downsampled point cloud, where the target task includes at least one of the following: a classification task and a semantic segmentation task.
[0028] In this embodiment, for three-dimensional point clouds for holographic communication, sampling and semantic coding are jointly designed to achieve joint optimization of semantic sampling: by globally combining local semantically driven point cloud downsampling and combining semantic encoding and decoding, it is beneficial to ensure that the "main body" content with high semantic importance or rich semantic information is accurately extracted and encoded in the subsequent processing, and other redundant parts are reduced in attention or even ignored. Based on the semantic content of the object itself, the original point cloud can be strategically downsampled. Point cloud downsampling integrates semantic-driven methods and can adaptively adjust different point clouds, thereby reducing the signal source level and further reducing the transmission overhead of three-dimensional point cloud holographic communication, reducing the burden of network processing and real-time transmission, and improving end-to-end communication efficiency.
[0029] In conjunction with the above embodiments, in one implementation, the present invention further provides a 3D point cloud sampling-semantics joint optimization method for holographic communication. This method is implemented using a pre-trained 3D point cloud sampling-semantics joint model. In this method, the transmitting end deploys a 3D point cloud sampling-semantics joint encoder from the 3D point cloud sampling-semantics joint model. The 3D point cloud sampling-semantics joint encoder includes an attention layer and a semantic sampling layer. The downsampled semantic features are obtained according to the following steps: Step S21: For each point in the original point cloud, the sending end inputs the semantic features of multiple adjacent points of the point and the semantic features of the point into the attention layer respectively to obtain the local semantic features of the point and the V value of the point.
[0030] In this embodiment, for each point in the original point cloud, the sending end inputs the semantic features of multiple adjacent points of the point and the semantic features of the point into the attention layer in the three-dimensional point cloud sampling-semantic joint encoder. Through the attention layer, the semantic features of multiple adjacent points of the point are processed based on the correlation between the semantic features of the point and the semantic features of the point, and the local semantic features of the point and the V value corresponding to the point are obtained by the attention layer output, thereby obtaining the local semantic features of each point in the original point cloud and the V value of each point.
[0031] Step S22: The sending end inputs the local semantic features of each point in the original point cloud and the V value of each point into the semantic sampling layer to obtain the normalized correlation graph of each point in the original point cloud, and determines the semantic information criticality of each point in the original point cloud based on the normalized correlation graph of each point in the original point cloud, and obtains the downsampled semantic features based on the semantic information criticality, normalized correlation graph and V value of each point in the original point cloud.
[0032] In this embodiment, after obtaining the local semantic features and V values of each point in the original point cloud, the transmitter inputs the local semantic features and V values of each point in the original point cloud into the semantic sampling layer of the 3D point cloud sampling-semantic joint encoder. The semantic sampling layer can first obtain a normalized correlation map for each point in the original point cloud based on the local semantic features of each point in the original point cloud; then, based on the normalized correlation map of each point in the original point cloud, obtain the semantic information criticality of each point in the original point cloud, and further obtain the semantic information criticality of each point in the original point cloud; then, based on the semantic information criticality, normalized correlation map of each point, and V value of each point in the original point cloud, the semantic sampling layer obtains the downsampled semantic features corresponding to the original point cloud.
[0033] In combination with any of the above embodiments, the present invention further provides a three-dimensional point cloud sampling-semantic joint optimization method for holographic communication. In this method, the above step S21 of "the transmitting end, for each point in the original point cloud, inputs the semantic features of multiple adjacent points of the point into the attention layer respectively with the semantic features of the point to obtain the local semantic features of the point" may specifically include the following steps S31 to S33: Step S31: For each point in the original point cloud, the sending end inputs the differences between the semantic features of multiple adjacent points of the point and the semantic features of the point into the K linear layers in the attention layer to obtain the K value of the point.
[0034] In this embodiment, the sending end can determine the differences between the semantic features of multiple adjacent points of each point in the original point cloud and the semantic features of the point, and input the differences between the semantic features of multiple adjacent points of the point and the semantic features of the point into the K linear layer in the attention layer to obtain the K value of the point, and then obtain the K value of each point in the original point cloud.
[0035] Step S32: The sending end inputs the semantic features of each point in the original point cloud into the Q linear layer in the attention layer to obtain the Q value of the point.
[0036] In this embodiment, the sending end inputs the semantic features of each point in the original point cloud into the Q linear layer in the attention layer to obtain the Q value of the point, and then obtains the Q value of each point in the original point cloud.
[0037] Step S33: The sending end determines, for each point in the original point cloud, a correlation measure between the point and multiple adjacent points based on the Q value and K value of the point, and uses the correlation measure as the local semantic feature of the point.
[0038] In this embodiment, after obtaining the K value and Q value of each point in the original point cloud, the sending end can determine the correlation measure between each point in the original point cloud and multiple adjacent points based on the Q value and K value of the point, and use the correlation measure between the point and multiple adjacent points as the local semantic feature of the point.
[0039] In an optional embodiment, the correlation measure between the i-th point and multiple adjacent points in the original point cloud (i.e., the local semantic feature of the i-th point) can be determined by: ; in, and Represents the Q linear layer in the attention layer (applied to the query input) and the K linear layer in the attention layer (applied to the key input), T is the transpose; the semantic features of the i-th point in the original point cloud are transformed into As the input of the Q linear layer, the semantic features of the j adjacent points of the i-th point are Respectively with the semantic features of the i-th point The difference is used as the input of the K linear layer.
[0040] In addition, in an optional embodiment, based on the local semantic features of each point in the original point cloud, a normalized correlation map of each point in the original point cloud is obtained, which can be obtained by the following formula: ; in, represents the normalized correlation graph of the i-th point in the original point cloud; Softmax represents the Softmax function, Represents the local semantic features of the i-th point in the original point cloud, is a local point cloud consisting of j adjacent points corresponding to the i-th point; j represents the j adjacent points of the i-th point; d represents the dimension of the vector, Used as the denominator for normalization.
[0041] In combination with any of the above embodiments, in one embodiment, the present invention further provides a three-dimensional point cloud sampling-semantic joint optimization method for holographic communication. In this method, the above step S21 of "the transmitting end inputs the semantic features of multiple adjacent points of each point in the original point cloud into the attention layer together with the semantic features of the point to obtain the V value of the point" can specifically include the following step S41, and the above step S22 of "obtaining the downsampled semantic features based on the semantic information criticality, normalized correlation graph and V value of each point in the original point cloud" can specifically include the following steps S42 to S45: Step S41: For each point in the original point cloud, the sending end inputs the differences between the semantic features of multiple adjacent points of the point and the semantic features of the point into the V linear layer in the attention layer to obtain the V value of the point.
[0042] In this embodiment, the sending end determines the difference between the semantic features of multiple adjacent points of each point in the original point cloud and the semantic features of the point, and inputs the difference between the semantic features of multiple adjacent points of the point and the semantic features of the point into the V linear layer in the attention layer to obtain the V value of the point, and then obtains the V value of each point in the original point cloud.
[0043] Step S42: the sending end sorts the points in the original point cloud in descending order of semantic information criticality, and determines the indexes of the first M points.
[0044] In this embodiment, the sending end uses the semantic sampling layer in the three-dimensional point cloud sampling-semantic joint encoder to globally sort the points in the original point cloud in descending order of the semantic information criticality based on the semantic information criticality of each point in the original point cloud, determine the top M points, and determine the indexes of the top M points.
[0045] Step S43: The sending end filters out the V values of the first M points from the V values of each point in the original point cloud based on the indexes of the first M points.
[0046] In this embodiment, the sending end filters out the V values corresponding to the first M points from the V values of each point in the original point cloud based on the indexes of the first M points through the semantic sampling layer in the three-dimensional point cloud sampling-semantic joint encoder.
[0047] Step S44: The transmitting end selects the normalized correlation graphs of the first M points from the normalized correlation graphs of each point in the original point cloud based on the indexes of the first M points.
[0048] In this embodiment, the sending end uses the semantic sampling layer in the three-dimensional point cloud sampling-semantic joint encoder to filter out the normalized correlation graphs corresponding to the first M points from the normalized correlation graphs of each point in the original point cloud based on the indexes of the first M points.
[0049] Step S45: The transmitting end obtains the downsampling feature based on the V values of the first M points and the normalized correlation graph of the first M points.
[0050] In this embodiment, the sending end processes the V values of the first M points and the normalized correlation graphs of the first M points through the semantic sampling layer in the three-dimensional point cloud sampling-semantic joint encoder to obtain the downsampling features corresponding to the original point cloud.
[0051] In conjunction with any of the above embodiments, in one embodiment, the present invention further provides a 3D point cloud sampling and semantic joint optimization method for holographic communication. In this method, the above step S22 of "determining the semantic information criticality of each point in the original point cloud based on the normalized correlation graph of each point in the original point cloud" can specifically include step S51: Step S51: The sending end calculates the standard deviation of each point in the original point cloud based on T elements corresponding to the point as the semantic information criticality of the point.
[0052] In this embodiment, the number of multiple neighboring points for each point in the original point cloud at the transmitting end is T, where T is an integer greater than 0. In an optional embodiment, for each point in the original point cloud, the kNN method (K nearest neighbor method) can be used to find T points adjacent to the point, thereby obtaining the T neighboring points of the point. The normalized correlation graph of each point in the original point cloud is: a vector representation including T elements, that is, the normalized correlation graph (vector representation) corresponding to the point includes T elements, and the tth element represents the correlation between the semantic features of the tth neighboring point of the point and the semantic features of the point, where t∈[1,T].
[0053] In this embodiment, the sending end calculates the standard deviation of each point in the original point cloud based on the T elements corresponding to the point through the semantic sampling layer in the three-dimensional point cloud sampling-semantic joint encoder, obtains the standard deviation corresponding to the point, and uses the standard deviation corresponding to the point as the semantic information criticality of the point.
[0054] In this embodiment, considering that the difference between the point cloud at the edge of the original point cloud and its surrounding point cloud is large, the standard deviation of its normalized correlation graph is larger, which has richer semantic information and importance (such as Figure 2 As shown, Figure 2 This is an example diagram showing a comparison of standard deviations between edge points and non-edge points in a point cloud according to an embodiment of the present invention. Figure 2In the example, point C is an edge point and point D is a non-edge point. It can be seen that edge points have greater differences and larger standard deviations than non-edge points. Thus, the standard deviation corresponding to the point is used as the semantic information criticality of the point. Based on the order of semantic information criticality from large to small, the points in the original point cloud are sorted to determine the top M points. This can maximize the retention of edge points in the original point cloud and remove non-edge points in the original point cloud, ensuring that the "main body" content with high semantic importance or rich semantic information is semantically sampled, thereby achieving semantic downsampling.
[0055] In one embodiment, if Figure 3 As shown, Figure 3 This is a schematic diagram showing the processing of the attention layer and the semantic sampling layer in a three-dimensional point cloud sampling-semantic joint encoder according to an embodiment of the present invention. Figure 3 In the input, for the semantic features of each point in the original point cloud (dimension is: N xd, N represents a total of N points in the original point cloud), the attention layer finds the k adjacent points (dimension is: N xkxd) adjacent to each point through the (K nearest neighbor method) kNN, and processes the semantic features of each point in the original point cloud (dimension is: N x 1 xd) through the Q linear layer to obtain the Q value of each point (dimension is: N x 1 xd); the semantic features of the k adjacent points of each point in the original point cloud (dimension is: N xkxd) are respectively input into the K linear layer in the attention layer to obtain the K value of each point (dimension is: N xkxd); and the semantic features of the k adjacent points of each point in the original point cloud (dimension is: N xk xd) are respectively input into the V linear layer in the attention layer to obtain the V value of each point (dimension is: N Then, the attention layer determines the local semantic features of each point in the original point cloud based on the Q value and K value of the point.
[0056] The local semantic features and V values of each point in the original point cloud are input into the semantic sampling layer. The semantic sampling layer processes the local semantic features of each point through the softmax operation to obtain the normalized correlation map of each point (dimension is N x 1 x K). Then, the standard deviation is calculated based on the normalized correlation map of each point. , get the semantic information criticality of each point, that is, there are N semantic information criticalities in total; then sort the points in the original point cloud in descending order of the N semantic information criticalities, determine the top M points, and based on the indexes of the top M points, filter out the normalized correlation graphs (dimension Mx 1 x K) of the top M points in the original point cloud; and, based on the indexes of the top M points, filter out the V values (dimension M x 1 x d) of the top M points from the V values (dimension N xk xd) of the top M points in the original point cloud; finally, the semantic sampling layer outputs the downsampled features corresponding to the original point cloud based on the V values of the top M points and the normalized correlation graphs of the top M points.
[0057] In combination with any of the above embodiments, in one implementation, the present invention further provides a 3D point cloud sampling-semantics joint optimization method for holographic communication. In this method, the method is implemented using a pre-trained 3D point cloud sampling-semantics joint model. The transmitting end is deployed with a 3D point cloud sampling-semantics joint encoder in the 3D point cloud sampling-semantics joint model, and the receiving end is deployed with a task network corresponding to the target task and a decoder in the 3D point cloud sampling-semantics joint model. The training process of the 3D point cloud sampling-semantics joint model includes a first stage and a second stage.
[0058] Among them, through the training of the first stage, the intermediate three-dimensional point cloud sampling-semantic joint encoder, intermediate decoder and trained task network trained in the first stage are obtained.
[0059] In this embodiment, by training the first stage of the three-dimensional point cloud sampling-semantic joint model, an intermediate three-dimensional point cloud sampling-semantic joint encoder, an intermediate decoder and a trained task network that have been trained in the first stage can be obtained.
[0060] Furthermore, the second stage may include steps S61 to S65: Step S61: The sending end processes the sample original point cloud through the intermediate three-dimensional point cloud sampling-semantic joint encoder to obtain the sample downsampling semantic features.
[0061] In this embodiment, the sample original point cloud is the original point cloud used for training the 3D point cloud sampling and semantics joint model, and the sample downsampled semantic features are the downsampled semantic features used during the training of the 3D point cloud sampling and semantics joint model. The transmitter can process the sample original point cloud using an intermediate 3D point cloud sampling and semantics joint encoder to obtain the sample downsampled semantic features.
[0062] In an optional embodiment, the transmitting end may extract semantic features of the sample original point cloud through an intermediate three-dimensional point cloud sampling-semantic joint encoder; for each sample point in the sample original point cloud, the sample local semantic features of the sample point are determined based on the correlation between the semantic features of multiple adjacent sample points of the sample point and the semantic features of the sample point; based on the sample local semantic features of each sample point in the sample original point cloud, the semantic information criticality of the sample point is determined; the sample points in the sample original point cloud are sorted in descending order of semantic information criticality, and semantic sampling is performed on the first M sample points to obtain sample downsampled semantic features. The method for obtaining the sample downsampled semantic features is similar to the method for obtaining the downsampled semantic features in steps S11 to S14, and can be implemented with reference to steps S11 to S14.
[0063] Step S62: The sending end sends the sample downsampling semantic features to the receiving end.
[0064] In this embodiment, after obtaining the sample downsampling semantic feature, the transmitting end may send the sample downsampling semantic feature to the receiving end through a wireless channel.
[0065] Step S63: The receiving end restores the sample downsampling point cloud through the intermediate decoder based on the received sample downsampling semantic features.
[0066] In this embodiment, the receiving end can receive the sample downsampling semantic features sent by the transmitting end through a wireless channel to obtain the received sample downsampling semantic features. The receiving end can perform semantic decoding based on the received sample downsampling semantic features through an intermediate decoder deployed at the receiving end to obtain a restored sample downsampling point cloud, that is, to restore the sample downsampling point cloud.
[0067] Step S64: The receiving end performs the target task through the trained task network based on the restored sample downsampling point cloud to obtain a first task execution result.
[0068] In this embodiment, the receiving end can perform the target task based on the restored sample downsampling point cloud through the trained task network deployed in the receiving end, and obtain the first task execution result output by the trained task network.
[0069] Step S65: Based at least on the first task execution result and the label carried by the sample original point cloud, fine-tune the parameters of the intermediate three-dimensional point cloud sampling-semantic joint encoder and the intermediate decoder to obtain the three-dimensional point cloud sampling-semantic joint model.
[0070] In this embodiment, the parameters of the intermediate three-dimensional point cloud sampling-semantic joint encoder and the intermediate decoder can be fine-tuned at least based on the execution results of the first task and the labels carried by the original point cloud of the sample, and finally a trained three-dimensional point cloud sampling-semantic joint encoder and a trained decoder are obtained, and then a three-dimensional point cloud sampling-semantic joint model composed of the trained three-dimensional point cloud sampling-semantic joint encoder and the trained decoder is obtained.
[0071] Among them, when the target task is a classification task, the execution result of the first task is the first classification task result, and the label carried by the sample original point cloud is the category label of the sample original point cloud; when the target task is a semantic segmentation task, the execution result of the first task is the first semantic segmentation task result, and the label carried by the sample original point cloud is the semantic label of the sample original point cloud.
[0072] In combination with any of the above embodiments, in one embodiment, the present invention further provides a three-dimensional point cloud sampling-semantic joint optimization method for holographic communication. In this method, the second stage may include steps S71 to S74: Step S71: the sending end extracts the semantic features of the sample original point cloud through the three-dimensional point cloud sampling-semantic joint encoder to be trained; the sending end sends the semantic features of the sample original point cloud to the receiving end.
[0073] In this embodiment, the sending end can perform semantic encoding on the sample original point cloud through the three-dimensional point cloud sampling-semantic joint encoder to be trained, extract the semantic features of the sample original point cloud, and send the semantic features of the sample original point cloud to the receiving end through a wireless channel.
[0074] Step S72: The receiving end restores the sample original point cloud through the decoder to be trained based on the semantic features of the received sample original point cloud.
[0075] In this embodiment, the receiving end can receive the semantic features of the sample original point cloud sent by the transmitting end through a wireless channel to obtain the semantic features of the received sample original point cloud. The receiving end can perform semantic decoding based on the semantic features of the received sample original point cloud through a decoder to be trained at the receiving end to obtain a restored sample original point cloud, that is, to restore the sample original point cloud.
[0076] Step S73: The receiving end performs the target task through the task network to be trained based on the restored sample original point cloud to obtain a second task execution result.
[0077] In this embodiment, the receiving end can perform the target task based on the restored sample original point cloud through the task network to be trained deployed in the receiving end, and obtain the second task execution result output by the task network to be trained.
[0078] Step S74: Based on the execution result of the second task and the label carried by the sample original point cloud, the parameters of the three-dimensional point cloud sampling-semantic joint encoder to be trained, the decoder to be trained and the task network to be trained are updated to obtain the intermediate three-dimensional point cloud sampling-semantic joint encoder, the intermediate decoder and the trained task network.
[0079] In this embodiment, based on the execution results of the second task and the labels carried by the original point cloud of the sample, the parameters of the three-dimensional point cloud sampling-semantic joint encoder to be trained, the decoder to be trained, and the task network to be trained can be updated to obtain the intermediate three-dimensional point cloud sampling-semantic joint encoder, the intermediate decoder, and the trained task network.
[0080] Among them, when the target task is a classification task, the execution result of the second task is the second classification task result, and the label carried by the sample original point cloud is the category label of the sample original point cloud; when the target task is a semantic segmentation task, the execution result of the second task is the second semantic segmentation task result, and the label carried by the sample original point cloud is the semantic label of the sample original point cloud.
[0081] In one embodiment, if Figure 4 As shown, Figure 4 3D point cloud sampling-semantics joint encoder and decoder in a 3D point cloud sampling-semantics joint model according to an embodiment of the present invention. Figure 4 The 3D point cloud sampling-semantic joint encoder in the 3D point cloud sampling-semantic joint model includes at least: feature embedding layer, two attention layers and two semantic sampling layers; the decoder in the 3D point cloud sampling-semantic joint model includes at least: attention layer, multi-layer perceptron, maximum pooling layer and activation function layer (such as Softmax). At the 3D point cloud sampling-semantic joint encoder end, the original point cloud is first converted into Embed into the feature space (i.e. input feature embedding layer) to extract its semantic information The global attention layer (i.e., attention layer) is then input to guide the subsequent semantic sampling layer to perform semantic sampling and obtain downsampled features. Finally, at the decoder side, the sampled encoded information is restored to the downsampled point cloud , and perform downstream tasks such as classification.
[0082] In combination with any of the above embodiments, in one embodiment, the present invention further provides a three-dimensional point cloud sampling-semantic joint optimization method for holographic communication. In this method, the above step S65 may specifically include steps S81 to S83: Step S81: The receiving end determines the sampling loss based on the difference between the restored sample downsampled point cloud and the sample original point cloud.
[0083] In this embodiment, the receiving end can determine the difference between the restored sample downsampled point cloud and the sample original point cloud based on the restored sample downsampled point cloud and the sample original point cloud, and determine the sampling loss based on the difference between the restored sample downsampled point cloud and the sample original point cloud. .
[0084] Step S82: The receiving end determines the task loss based on the difference between the first task execution result and the second task execution result, and the difference between the first task execution result and the label carried by the sample original point cloud.
[0085] In this embodiment, the receiving end can also determine the difference between the first task execution result and the second task execution result based on the first task execution result and the second task execution result, and determine the difference between the first task execution result and the label carried by the sample original point cloud based on the first task execution result and the label carried by the sample original point cloud; and then determine the task loss based on the difference between the first task execution result and the second task execution result and the difference between the first task execution result and the label carried by the sample original point cloud. .
[0086] Step S83: Based on the sampling loss and the task loss, fine-tune the parameters of the intermediate three-dimensional point cloud sampling-semantic joint encoder and the intermediate decoder to obtain the three-dimensional point cloud sampling-semantic joint model.
[0087] In this embodiment, the total loss can be obtained based on the obtained sampling loss and task loss, and the parameters of the intermediate three-dimensional point cloud sampling-semantic joint encoder and the intermediate decoder can be fine-tuned based on the total loss until the total loss converges. The parameters of the intermediate three-dimensional point cloud sampling-semantic joint encoder and the intermediate decoder are fixed when the total loss converges to obtain the trained three-dimensional point cloud sampling-semantic joint encoder and the trained decoder, and then the three-dimensional point cloud sampling-semantic joint model is obtained.
[0088] In one embodiment, if Figure 5 As shown, Figure 5 FIG1 is a schematic diagram of the architecture of a semantic-oriented point cloud sampling network according to an embodiment of the present invention. Figure 5 The semantic-oriented point cloud sampling network includes a 3D point cloud sampling-semantics joint encoder, a decoder, and a task network. The 3D point cloud sampling-semantics joint encoder and decoder form the 3D point cloud sampling-semantics joint model. Training of the 3D point cloud sampling-semantics joint model involves a first-stage training and a second-stage training.
[0089] First, for the first stage of training, the original point cloud of the sample is input during training End-to-end semantic encoding and decoding are performed through the three-dimensional point cloud sampling-semantic joint encoder to be trained and the decoder to be trained, and then classification, semantic segmentation and other target tasks are performed on the task network to be trained to obtain the second task execution result. Based on the second task execution result and the label carried by the original point cloud of the sample, the parameters of the three-dimensional point cloud sampling-semantic joint encoder to be trained, the decoder to be trained and the task network to be trained are updated to obtain the intermediate three-dimensional point cloud sampling-semantic joint encoder and intermediate decoder, and the parameters of the task network are frozen to obtain the trained task network.
[0090] For the second stage of training, the original point cloud of the sample is trained Input the intermediate 3D point cloud sampling-semantic joint encoder for downsampling, pass through the intermediate 3D point cloud sampling-semantic joint encoder, intermediate decoder and the trained task network, after semantic encoding and semantic decoding, and then perform the task to obtain the restored sample downsampled point cloud And the first task execution result output by the trained task network.
[0091] In the second stage of training, the sampling loss is determined by using the restored sample downsampled point cloud output by the intermediate decoder and the sample original point cloud The task loss is determined by the difference between the first task execution result and the second task execution result, and the difference between the first task execution result and the label carried by the sample original point cloud. The total loss function is: ; Among them, α is the weight, which can be set according to experience; 、 are the parameters to be fine-tuned for the intermediate 3D point cloud sampling-semantic joint encoder and the intermediate decoder, are the parameters of the trained task network. Based on the total loss function, the parameters of the intermediate 3D point cloud sampling-semantic joint encoder and intermediate decoder are fine-tuned to obtain the 3D point cloud sampling-semantic joint model.
[0092] It should be noted that for the sake of simplicity, the method embodiments are described as a series of actions. However, those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.
[0093] Based on the same inventive concept, another embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the three-dimensional point cloud sampling-semantic joint optimization method for holographic communication as described in any of the above embodiments of the present invention are implemented.
[0094] Based on the same inventive concept, another embodiment of the present invention provides an electronic device, such as Figure 6 As shown, Figure 6 This is a schematic diagram of an electronic device according to an embodiment of the present invention. The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When executed by the processor, the computer program implements the steps of the three-dimensional point cloud sampling and semantic joint optimization method for holographic communication described in any of the above embodiments of the present invention.
[0095] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0096] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, apparatus, or computer program products. Thus, embodiments of the present invention may take the form of a fully hardware embodiment, a fully software embodiment, or an embodiment combining software and hardware. Furthermore, embodiments of the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0097] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0098] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0099] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0100] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they are aware of the basic creative concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.
[0101] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or terminal device that includes the element.
[0102] The above is a detailed introduction to the three-dimensional point cloud sampling-semantic joint optimization method, device and medium for holographic communication provided by the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for general technical personnel in this field, according to the ideas of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
Claims
1. A 3D point cloud sampling and semantic joint optimization method for holographic communication, characterized by: The method comprises: The sender in the semantic communication system extracts the semantic features of the original point cloud; The transmitting end determines, for each point in the original point cloud, a local semantic feature of the point based on correlations between the semantic features of a plurality of adjacent points of the point and the semantic feature of the point; The sending end determines the semantic information criticality of each point in the original point cloud based on the local semantic features of the point; Sort the points in the original point cloud in descending order of semantic information criticality, perform semantic sampling on the first M points, and obtain downsampled semantic features; M is an integer greater than 1; The sending end sends the downsampled semantic features to the receiving end in the semantic communication system; The receiving end restores the downsampled point cloud based on the received downsampled semantic features; The receiving end performs a target task based on the downsampled point cloud, where the target task includes at least any one of the following: a classification task and a semantic segmentation task.
2. The 3D point cloud sampling and semantic joint optimization method for holographic communication according to claim 1 is characterized in that: The method is implemented by a pre-trained 3D point cloud sampling-semantic joint model. The transmitting end is deployed with a 3D point cloud sampling-semantic joint encoder in the 3D point cloud sampling-semantic joint model. The 3D point cloud sampling-semantic joint encoder includes an attention layer and a semantic sampling layer. The down-sampled semantic features are obtained according to the following steps: The sending end inputs the semantic features of multiple adjacent points of each point in the original point cloud and the semantic feature of the point into the attention layer to obtain the local semantic feature of the point and the V value of the point; The sending end inputs the local semantic features of each point in the original point cloud and the V value of each point into the semantic sampling layer to obtain a normalized correlation graph of each point in the original point cloud, and determines the semantic information criticality of each point in the original point cloud based on the normalized correlation graph of each point in the original point cloud, and obtains the downsampled semantic features based on the semantic information criticality, the normalized correlation graph and the V value of each point in the original point cloud.
3. The 3D point cloud sampling and semantic joint optimization method for holographic communication according to claim 2 is characterized in that: The sending end inputs the semantic features of multiple adjacent points of each point in the original point cloud and the semantic features of the point into the attention layer to obtain the local semantic features of the point, including: The sending end inputs the difference between the semantic features of multiple adjacent points of each point in the original point cloud and the semantic features of the point into the K linear layers in the attention layer to obtain the K value of the point; The sending end inputs the semantic feature of each point in the original point cloud into the Q linear layer in the attention layer to obtain the Q value of the point; The sending end determines, for each point in the original point cloud, a correlation measure between the point and a plurality of adjacent points based on the Q value and the K value of the point, and uses the correlation measure as a local semantic feature of the point.
4. The 3D point cloud sampling and semantic joint optimization method for holographic communication according to claim 2 is characterized in that: The sending end inputs the semantic features of multiple adjacent points of each point in the original point cloud and the semantic features of the point into the attention layer to obtain the V value of the point, including: The transmitting end inputs the difference between the semantic features of multiple adjacent points of each point in the original point cloud and the semantic features of the point into the V linear layer in the attention layer to obtain the V value of the point; The downsampled semantic features are obtained based on the semantic information criticality, normalized correlation graph and V value of each point in the original point cloud, including: The sending end sorts the points in the original point cloud in descending order of semantic information criticality, and determines the indexes of the first M points; The transmitting end filters out the V values of the first M points from the V values of each point in the original point cloud based on the indexes of the first M points; The transmitting end filters out the normalized correlation graphs of the first M points from the normalized correlation graphs of each point in the original point cloud based on the indexes of the first M points; The transmitting end obtains the downsampling feature based on the V values of the first M points and the normalized correlation graph of the first M points.
5. The 3D point cloud sampling and semantic joint optimization method for holographic communication according to claim 2 is characterized in that: The number of the plurality of neighboring points of each point in the original point cloud at the transmitting end is T, and the normalized correlation graph of the point is a vector representation including T elements, where the t-th element represents the correlation between the semantic feature of the t-th neighboring point of the point and the semantic feature of the point; Determining the semantic information criticality of each point in the original point cloud based on a normalized correlation graph of each point in the original point cloud includes: The sending end calculates the standard deviation of each point in the original point cloud based on T elements corresponding to the point as the semantic information criticality of the point; wherein T is an integer greater than 0, t∈[1,T].
6. The 3D point cloud sampling and semantic joint optimization method for holographic communication according to claim 1 is characterized in that: The method is implemented by a pre-trained 3D point cloud sampling-semantic joint model, wherein the transmitting end is deployed with a 3D point cloud sampling-semantic joint encoder in the 3D point cloud sampling-semantic joint model, and the receiving end is deployed with a task network corresponding to the target task and a decoder in the 3D point cloud sampling-semantic joint model; the training process of the 3D point cloud sampling-semantic joint model includes a first stage and a second stage; Through the first stage of training, an intermediate three-dimensional point cloud sampling-semantic joint encoder, an intermediate decoder, and a trained task network are obtained; The second stage is: The sending end processes the sample original point cloud through the intermediate three-dimensional point cloud sampling-semantic joint encoder to obtain the sample downsampling semantic features; The sending end sends the sample downsampling semantic feature to the receiving end; The receiving end restores the sample downsampling point cloud through the intermediate decoder based on the received sample downsampling semantic features; The receiving end performs the target task through the trained task network based on the restored sample downsampling point cloud to obtain a first task execution result; Based at least on the first task execution result and the label carried by the sample original point cloud, the parameters of the intermediate three-dimensional point cloud sampling-semantic joint encoder and the intermediate decoder are fine-tuned to obtain the three-dimensional point cloud sampling-semantic joint model.
7. The 3D point cloud sampling and semantic joint optimization method for holographic communication according to claim 6 is characterized in that: The first stage is: The transmitting end extracts the semantic features of the sample original point cloud through the three-dimensional point cloud sampling-semantic joint encoder to be trained; the transmitting end sends the semantic features of the sample original point cloud to the receiving end; The receiving end restores the sample original point cloud through the decoder to be trained based on the semantic features of the received sample original point cloud; The receiving end performs the target task through the task network to be trained based on the restored sample original point cloud to obtain a second task execution result; Based on the execution results of the second task and the labels carried by the sample original point cloud, the parameters of the three-dimensional point cloud sampling-semantic joint encoder to be trained, the decoder to be trained and the task network to be trained are updated to obtain the intermediate three-dimensional point cloud sampling-semantic joint encoder, the intermediate decoder and the trained task network.
8. The 3D point cloud sampling and semantic joint optimization method for holographic communication according to claim 7 is characterized in that: Fine-tuning parameters of the intermediate 3D point cloud sampling-semantics joint encoder and the intermediate decoder based at least on the first task execution result and the label carried by the sample original point cloud to obtain the 3D point cloud sampling-semantics joint model, including: The receiving end determines a sampling loss based on a difference between the restored sample downsampled point cloud and the sample original point cloud; The receiving end determines the task loss based on a difference between the first task execution result and the second task execution result, and a difference between the first task execution result and a label carried by the sample original point cloud; Based on the sampling loss and the task loss, parameters of the intermediate three-dimensional point cloud sampling-semantic joint encoder and the intermediate decoder are fine-tuned to obtain the three-dimensional point cloud sampling-semantic joint model.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the computer program is executed by the processor, the three-dimensional point cloud sampling-semantic joint optimization method for holographic communication is implemented as described in any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the three-dimensional point cloud sampling-semantic joint optimization method for holographic communication as described in any one of claims 1 to 8 is implemented.