A Point Cloud Quality Assessment Method Based on Deep Ordinal Regression
Through a deep ordinal regression method, combined with the symmetric cross-modal attention mechanism and deep orderly loss, the problem of point cloud quality evaluation in high-dimensional and without reference conditions is solved, and high accuracy and efficient point cloud quality evaluation is achieved.
Patent Information
- Application Number
- CN202411639402.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-11
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-12-11
AI Technical Summary
It is difficult for the prior art to effectively evaluate point cloud quality, especially in point cloud quality evaluation under high-dimensional and without reference conditions, and traditional methods are difficult to accurately evaluate point cloud geometric accuracy and noise levels.
The point cloud quality evaluation method based on deep ordinal regression is adopted, and the multimodal features of point clouds and pictures are extracted through a symmetric cross-modal attention mechanism and expressed as an orderly regression problem. The predicted probability is converted into continuous point cloud mass fractions using deep order loss and soft order inference.
It realizes high accuracy assessment of point cloud quality, which consumes less time, and can conduct simple and time-consuming quality assessment without reference conditions, improving the effectiveness and efficiency of point cloud quality assessment.
Smart Images

Figure FT_1 
Figure QLYQS_1 
Figure QLYQS_8
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data analysis, and particularly relates to a point cloud quality assessment method based on deep ordinal regression. Background Art
[0002] Compared with traditional 2D images, 3D sensors can often provide geometric, shape, and scale information captured in the real world in a more comprehensive and intuitive way. Point cloud is a common format in 3D data, which refers to a set containing a large number of three-dimensional points and can preserve the most primitive information unchanged in three-dimensional space, ensuring a sense of reality and three-dimensionality.
[0003] With the rapid development of 3D acquisition technology, point clouds have been widely used in fields such as autonomous driving, architecture and surveying, virtual reality and augmented reality, and medical imaging. Since each dense point cloud contains tens of thousands of points, during data storage and transmission, its media content will be compressed, and different degrees and types of distortion will inevitably occur during the compression process. Low-quality point clouds may lead to errors or performance issues during the calculation process. We need to evaluate the degree of distortion, discover defects in the point cloud (such as sparse or missing points), check the geometric accuracy and noise level of the point cloud to ensure it meets the requirements, optimize the data acquisition or processing process, and improve reliability and accuracy. Therefore, point cloud quality assessment is an important part of ensuring the accuracy, integrity, and reliability of three-dimensional data.
[0004] At present, point cloud quality assessment still faces huge challenges. Since point cloud is an irregular data set in three-dimensional space, with high data dimensions, a large number of data, and a lack of topological structure, it is difficult to directly apply traditional image quality evaluation methods, and it is more difficult to directly evaluate compared with two-dimensional images. Traditional point cloud quality assessment methods are divided into three categories according to whether the original point cloud is available: full-reference, semi-reference, and no-reference. In the initial stage of development, MPEG proposed some simple point-based full-reference methods, such as p2point and p2plane. Due to the difficulty of collecting the original point cloud, thanks to the rapid development of deep learning, no-reference point cloud quality assessment methods have been significantly improved. For example, Chetouani et al. extracted handcrafted features block by block and used a classic CNN model for quality regression; Liu et al. used multi-view projection for feature extraction; Zhang et al. used several statistical distributions to estimate quality perception parameters from the distributions of geometric and color attributes; Fan et al. inferred the visual quality of the point cloud through the captured video sequence; Liu et al. adopted an end-to-end sparse CNN for quality prediction; Yang et al. further transmitted the quality information from natural images to understand the quality of point cloud rendering images through domain adaptation.
[0005] Most of the existing deep learning-based reference-free point cloud quality assessment models are formulated as regression problems and trained by minimizing the mean squared error. However, the error measurement does not consider the relative ranking between different ratings on the quality scale, thus affecting the performance of the model. To solve this problem, the present invention reformulates the reference-free point cloud quality assessment learning as an ordinal regression problem, combines multi-modal feature extraction, and converts the prediction probability into a continuous variable of image quality through a deep ordinal loss and using soft ordinal inference to obtain the final point cloud quality score. Summary of the Invention
[0006] To solve the above problems, the present invention proposes a point cloud quality assessment method based on deep ordinal regression, which can effectively evaluate the quality of the input point cloud in point cloud quality assessment, with high accuracy of the assessment result and less time consumption.
[0007] The specific solution is as follows:
[0008] A point cloud quality assessment method based on deep ordinal regression, comprising:
[0009] A feature fusion step based on a symmetric cross-modal attention mechanism, combining two modalities of point cloud and picture, and extracting point cloud visual quality features and picture modality features through a feature encoder; integrating the extracted point cloud visual quality features and picture modality features, analyzing the interaction between modalities and strengthening the feature representation to generate a final feature vector;
[0010] The point cloud quality assessment is expressed as an ordinal regression step. Based on the final feature vector, by simplifying multiple classification problems into a series of binary classification problems, the prediction probability is converted into a continuous point cloud quality score, and an adaptive factor is introduced to make the prediction score adapt to the sparsity of the discretized interval, solving the step effect that is prone to occur when using hard threshold inference.
[0011] Further, the feature fusion step based on the symmetric cross-modal attention mechanism specifically includes:
[0012] A point cloud modality feature extraction step: dividing the input point cloud into local regions and extracting multiple patches, inputting the point cloud feature encoder and extracting point cloud local features;
[0013] An image modality feature extraction step: projecting the point cloud onto a fixed viewpoint to generate a projection picture, and inputting it into an image encoder to extract picture features;
[0014] A step of generating a fused global feature representation: based on the cross-modal symmetric attention mechanism, fusing the extracted point cloud features and picture features to generate a fused global feature representation.
[0015] Further, the base point cloud modality feature extraction step specifically includes:
[0016] Given a normalized point cloud Use farthest point sampling to obtain N δ anchor points For each anchor point, use the K-nearest neighbor algorithm to find N s adjacent points around the anchor point, and use these points to form the corresponding point cloud sub-model. The calculation formula is as follows:
[0017]
[0018] where S is the set of sub-models, KNN(·) represents the KNN operation, and N δ represents the number of anchor points; use the point cloud feature encoder θ P (·) to map the obtained sub-model into the quality-aware embedding space. The calculation formula is as follows:
[0019]
[0020] where represents the quality-aware embedding of the i-th sub-model S l , C P represents the number of output channels of the point cloud encoder θ P (·), is the merged result after average fusion.
[0021] Furthermore, the image modality feature extraction step specifically includes:
[0022] At interval, limit a circular path to rotate around the distorted color point cloud for rendering, and capture the projection N I while keeping the texture consistent. The calculation formula is as follows:
[0023]
[0024] where X and Y represent the abscissa and ordinate of each point in the point cloud on the projection plane, Z represents the coordinate of each point in the point cloud in the vertical direction, the path is centered on the geometric center of the point cloud, and R represents a fixed observation distance; use Open3D for rendering, and use a 2D image encoder to embed the rendered 2D image into the quality-aware space. The calculation formula is as follows:
[0025]
[0026] where represents the quality-aware embedding of the t-th image projection I t , C I represents the number of output channels of the image encoder θ I (·), is the merged result after average fusion.
[0027] Furthermore, the step of generating the fused global feature representation specifically includes:
[0028] Given the features from the point cloud and image modalities Use a linear layer to unify them to the same size:
[0029]
[0030] where is the adjusted feature, W P and W I are learnable linear mappings, C′ is the adjusted number of channels. To further distinguish the components between modalities, use the multi-head attention module:
[0031]
[0032] where Γ(·) represents the multi-head attention operation, β(·) represents the attention function, j μ represents the μ-th head, W, W Q 、W K 、W V represent learnable linear mappings. The final quality embedding is obtained by concatenating the in-modal features and the guided multi-modal features obtained from the symmetric cross-modal attention module. The calculation formula is as follows:
[0033]
[0034] where represents the concatenation operation, Ψ(·) represents the symmetric cross-modal attention operation, represents the final quality-aware feature.
[0035] Furthermore, the point cloud quality assessment is expressed as an ordered regression step, which specifically includes:
[0036] Ordinal loss optimization step: Design a deep ordinal loss function to learn network parameters, and utilize the ordered information in the discrete labels of quality scores to make the network more sensitive to predictions with inconsistent order attributes of quality scores;
[0037] Soft ordinal inference step: Convert the prediction probability to a continuous point cloud quality score through soft ordinal inference, and utilize the probability or confidence predicted by the network to prevent the step effect easily generated by ordinal operations, making the quality score more adaptable to the sparsity of the discretization interval, so as to achieve a better approximation to the ground truth.
[0038] Furthermore, the ordinal loss optimization step specifically includes:
[0039] Discretize the entire range of the existing image quality score database into K subintervals of equal size, and the calculation formula is as follows:
[0040]
[0041] Among them, represents the floor function, d * ∈ {0, 1,..., k - 1} represents the discrete label of the ground truth, s * represents the original continuous value of the image quality score, s min and s max represent the minimum and maximum scores in the given database respectively;
[0042] Unify the continuous quality scores into multiple bins in the logarithmic space, mark the bins according to the range of quality score values covered by each bin, and the ordinal threshold c k ∈ {0, 1,..., c k-1} The calculation formula is as follows:
[0043]
[0044] For each instance c k ∈ {0, 1,..., c k-1}, apply a binary classifier to predict whether the ordinal value of the sample is greater than c k , and predict the hidden sample ordinal value based on the classification results of K - 1 binary classifiers. Define the output of the model as where Y belongs to a 2K-dimensional vector, and the depth ordinal loss calculation formula is as follows:
[0045]
[0046] Among them, d represents the predicted discrete label, represents the ordinal probability when d is greater than k.
[0047] Furthermore, the soft ordinal inference step specifically includes:
[0048] Convert the output probability into a continuous variable of image quality, and the calculation formula is as follows:
[0049]
[0050] Among them represents the introduced adaptive factor, and d between 0 and 1 represents the degree to which the predicted category is close to d + 1.
[0051] The present invention adopts the above technical solutions and has beneficial effects:
[0052] (1) The present invention adopts a symmetric cross-modal attention mechanism to force the model to associate with the synthetic quality perception mode, mobilize the complementary information between the image and point cloud modalities, and jointly utilize the intra-modal and cross-modal features to optimize the quality representation;
[0053] (2) The present invention proposes a deep ordinal loss function for regularization, which imposes a greater loss on the predictions with inconsistent order attributes of the quality labels, making the confidence in learning the quality categories more precisely represented;
[0054] (3) Compared with the same type of methods (reference-free point cloud quality assessment methods), the present invention considers the relative order between quality ratings, making the prediction probability of quality scores more accurate and improving the overall evaluation effect; compared with similar methods (full-reference point cloud quality assessment methods, semi-reference point cloud quality methods), it can perform simpler, less time-consuming and more accurate quality assessment without using the original point cloud. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 It is a schematic diagram of the framework of the point cloud quality assessment method based on deep ordinal regression according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0056] The present invention will be further described in detail below in conjunction with the embodiments and the accompanying drawings, but the embodiments of the present invention are not limited thereto.
[0057] As Figure 1 shown, a point cloud quality assessment method based on deep ordinal regression of the present invention includes the following steps.
[0058] S101, a feature fusion step based on a symmetric cross-modal attention mechanism, combines the point cloud and image modalities, and extracts the point cloud visual quality features and image modality features through a feature encoder; integrates the extracted point cloud visual quality features and image modality features, analyzes the interaction between modalities and strengthens the feature representation, generates a final feature vector, and solves the problem that the information covered by a single modality is not complete enough.
[0059] Specifically, given a colored point cloud where represents the geometric coordinates, represents the additional RGB color information, and the corresponding point cloud modality is obtained by normalizing the original point cloud coordinates. The image modality I is obtained by rendering the colored point cloud Ρ into a 2D projection (the projection does not contain color information). In order to make the segmented point cloud sub-model retain important local geometric structure patterns, it is necessary to repair the segmented point cloud.
[0060] Specifically, the method for repairing the segmented point cloud is the feature fusion step based on the symmetric cross-modal attention mechanism, which specifically includes:
[0061] S1011, Point cloud modal feature extraction step: Divide the input point cloud into local regions and extract multiple patches, input the point cloud into the feature encoder and extract the local features of the point cloud;
[0062] Given a normalized point cloud Use farthest point sampling (FPS) to obtain N δ anchor points For each anchor point, use the k-nearest neighbor algorithm (KNN) to find N s adjacent points around the anchor point, and use these points to form the corresponding point cloud sub-model:
[0063]
[0064] where S is the set of sub-models, and KNN(·) represents the KNN operation. To make the local sub-models maintain the corresponding patterns, use the point cloud feature encoder θ P (·) to map the obtained sub-models into the quality-aware embedding space:
[0065]
[0066] where represents the quality-aware embedding of the i-th sub-model S I , C P represents the number of output channels of the point cloud encoder θ P (·), is the combined result after average fusion, obtaining the final point cloud modal feature.
[0067] S1012, Image modal feature extraction step: Project the point cloud onto a fixed viewing point to generate a projection image, and input it into the image encoder to extract the image features;
[0068] The camera rotates around the distorted color point cloud along a defined circular path at a fixed viewing distance at intervals of for rendering, and captures the projection N I .
[0069]
[0070] where the path is centered on the geometric center of the point cloud, R represents the fixed viewing distance, and Open3D is used for rendering. Then, use the 2D image encoder to embed the rendered 2D image into the quality-aware space:
[0071]
[0072] where represents the quality-aware embedding of the t-th image projection I t , CI Denote the image encoder as θ I The number of output channels of (·), is the merged result after average fusion to obtain the final image modality feature.
[0073] S1013, Generation step of the fused global feature representation: Based on the cross-modal symmetric attention mechanism, fuse the extracted point cloud feature and the picture feature to generate the fused global feature representation;
[0074] Given the features from the point cloud and image modalities Use a linear layer to unify them to the same size:
[0075]
[0076] Among them, is the adjusted feature, W P and W I are learnable linear mappings, and C′ is the adjusted number of channels. To further distinguish the components between modalities, use the multi-head attention module:
[0077]
[0078] Among them, Γ(·) represents the multi-head attention operation, β(·) represents the attention function, and h μ is the μ-th head, and W, W Q , W K , W V are learnable linear mappings. The final quality embedding can be obtained by concatenating the in-modal feature and the guided multi-modal feature obtained from the symmetric cross-modal attention module:
[0079]
[0080] S102, The point cloud quality assessment is formulated as an ordered regression step. Based on the final feature vector, by reducing multiple classification problems to a series of binary classification problems, convert the prediction probability into a continuous point cloud quality score, and introduce an adaptive factor to make the prediction score adapt to the sparsity of the discretized interval, so as to solve the step effect that is likely to occur when using a hard threshold for inference.
[0081] Based on the overall perception of the given image space, place the image into one of the quality categories (for example, in the interval representing "poor" to "excellent" in the continuous range from 0 to 10), and then refine the quality score by comparing the current image with other images belonging to the same local perception quality interval (that is, generate a decimal number), and formulate the model as an ordered regression problem to make it more in line with the human preference level.
[0082] Specifically, the point cloud quality assessment is formulated as an ordered regression step, which specifically includes:
[0083] S1021 Ordinal Loss Optimization Step: Design a deep ordinal loss function to learn network parameters, utilize the ordered information in the quality score discrete labels, and make the network more sensitive to predictions with inconsistent order attributes of quality scores;
[0084] First, discretize the entire range of the existing image quality score database into K equal-sized subintervals:
[0085]
[0086] where, represents the floor function, d * ∈ {0, 1,..., k - 1} is the discrete label of the ground truth, s * is the original continuous value of the image quality score, s min , s max respectively represent the minimum and maximum scores in the given database; discretize the continuous quality scores into multiple bins in the logarithmic space; each bin covers the range of quality score values, and mark the bins according to this range (i.e., the label index of the feature indicates its distance); the ordinal threshold c k ∈ {0, 1,..., c k-1} has the following calculation formula,
[0087]
[0088] Typical multi-class classification losses ignore the ordered information between discrete labels, while quality scores have strong ordered correlations as they form an ordered set; the quality assessment problem can be transformed into an ordered regression problem, and a deep ordered loss (DO-loss) can be designed for the network to learn network parameters, converting the multi-class classification problem into a set of binary classification sub-problems: for each instance c k ∈ {0, 1,..., c k-1}, apply a binary classifier to predict whether the ordinal value of the sample is greater than c k ; and predict the hidden sample ordinal value based on the classification results of K - 1 binary classifiers, and define the output of the model as where Y belongs to a 2K-dimensional vector, and the calculation formula for the deep ordinal loss (DO-loss) can be obtained as follows:
[0089]
[0090]
[0091] where, d represents the predicted discrete label, It is the ordinal probability when d is greater than k; compared with training a classification network using cross-entropy loss conventionally, the deep ordinal loss (DO-loss) can update network parameters more effectively. Due to the inherent order property included in the quality rating, the order loss is more sensitive to predictions that are inconsistent with the order property of the ground truth.
[0092] Based on the output probabilities of K binary classification instances, the calculation formula for predicting the image quality score s is as follows.
[0093]
[0094] where η(·) represents the indicator function, where η(true)=1 and η(false)=0. The rounding operation of hard inference ignores the probability (or confidence) predicted by the network, which may lead to predictions in the transition region that are difficult for the network to distinguish.
[0095] S1022 Soft ordinal inference step: Convert the predicted probability into a continuous point cloud quality score through soft ordinal inference. Utilize the probability or confidence predicted by the network to prevent the step effect easily generated by ordinal operations, make the quality score more adaptable to the sparsity of the discretization interval, and thus achieve an effect closer to the ground truth.
[0096] Precisely utilize the probability (or confidence) predicted by the network to prevent the step effect easily generated by ordinal operations, make the quality score more adaptable to the sparsity of the discretization interval, and thus achieve an effect closer to the ground truth; Using classification instead of regression for quality assessment can naturally obtain the confidence of the quality distribution; The elements in the confidence map of each class only focus on specific quality score intervals, which simplifies network learning. However, it introduces a discretized mean square error that is sensitive to the number of quality score intervals; The inference strategy based on hard thresholds ignores the obtained probability distribution and will lead to a step effect in the feature map; While generalizing the naive hard inference to a soft version, called soft ordinal inference, to solve the above problems; Soft ordinal inference makes full use of the predicted confidence and has strong classification ability for the transition region between objects, and can convert the output probability into a continuous variable of image quality. The specific calculation formula is as follows:
[0097]
[0098] where d is between 0 and 1, indicating the degree to which the predicted class is close to d + 1; In this soft inference, due to the introduction of an adaptive factor the predicted score will adapt to the sparsity of the discretization interval, which makes the inferred score closer to the ground truth.
[0099] Although the present invention has been specifically shown and described in connection with preferred embodiments, those skilled in the art should understand that various changes can be made to the present invention in form and detail without departing from the spirit and scope of the present invention as defined by the appended claims, and all such changes are within the scope of protection of the present invention.
Claims
1. A point cloud quality assessment method based on deep ordinal regression, characterized in that: include: Based on the feature fusion step of the symmetric cross-modal attention mechanism, the two modalities of point cloud and image are combined to extract the point cloud visual quality features and image modal features through the feature encoder; the extracted point cloud visual quality features and image modal features are integrated to analyze the interaction between the modalities and strengthen the feature representation to generate the final feature vector; Point cloud quality assessment is expressed as an ordered regression step. Based on the final feature vector, the prediction probability is converted into a continuous point cloud quality score by simplifying multiple classification problems into a series of binary classification problems. An adaptive factor is introduced to make the prediction score adapt to the sparsity of the discretization interval to solve the step effect that is easy to occur when using hard threshold inference. The point cloud quality assessment is expressed as an ordered regression step, specifically including: Ordinal loss optimization step: Design a deep ordinal loss function to learn network parameters, and use the ordered information in the discrete labels of quality scores to make the network more sensitive to predictions with inconsistent quality score order attributes; Soft ordinal reasoning step: The predicted probability is converted into a continuous point cloud quality score through soft ordinal reasoning. The probability or confidence of network prediction is used to prevent the step effect that is easy to be produced by ordinal operation, so that the quality score is more adapted to the sparsity of the discretization interval; The ordinal loss optimization step specifically includes: The entire range of the existing image quality score database is discretized into K sub-intervals of equal size. The calculation formula is as follows: in, The bottom function, Represents the discrete label of the ground truth. Represents the raw continuous value of the image quality score, , denote the minimum and maximum scores in a given database, respectively; The continuous quality scores are uniformly discretized into multiple bins in the logarithmic space, and the bins are marked according to the range of quality score values covered by each bin, and the ordinal threshold is obtained. The calculation formula is as follows: For each instance , apply the binary classifier to predict whether the ordinal value of the sample is greater than , and based on The classification results of the binary classifier predict the hidden sample ordinal value, and the output of the model is defined as , where Y is a 2K-dimensional vector, the deep ordinal loss calculation formula is as follows: in, represents the predicted discrete label, express Greater than The ordinal probability when ; The soft ordinal number reasoning step specifically includes: The output probability is converted into a continuous variable of image quality, and the calculation formula is as follows: in = , =hd, , Between 0 and 1, the predicted category is close degree.
2. The point cloud quality assessment method based on deep ordinal regression according to claim 1 is characterized in that: The feature fusion step based on the symmetric cross-modal attention mechanism specifically includes: Point cloud modal feature extraction step: divide the input point cloud into local areas and extract multiple patches, input the point cloud feature encoder and extract the local features of the point cloud; Image modality feature extraction step: project the point cloud to a fixed viewpoint to generate a projection image, and input it into the image encoder to extract image features; The steps of generating the fused global feature representation are as follows: Based on the cross-modal symmetric attention mechanism, the extracted point cloud features and image features are fused to generate the fused global feature representation.
3. The point cloud quality assessment method based on deep ordinal regression according to claim 2, characterized in that: The point cloud modal feature extraction step specifically includes: Given a normalized point cloud , obtained by sampling at the farthest point Anchor point For each anchor point, the K nearest neighbor algorithm is used to find Adjacent points are used to form the corresponding point cloud sub-model. The calculation formula is as follows: Among them, where S is the set of sub-models, represents the KNN operation, Indicates the number of anchor points; using point cloud feature encoder The obtained sub-model is mapped into the quality-aware embedding space, and the calculation formula is as follows: in, Represents the Ith sub-model Quality-aware embedding of Represents a point cloud encoder The number of output channels, is the merged result after average fusion.
4. The point cloud quality assessment method based on deep ordinal regression according to claim 2 is characterized in that: The image modality feature extraction step specifically includes: by Rendering a circular path defined by intervals rotating around the distorted color point cloud captures the projection while maintaining texture consistency , the calculation formula is as follows: Among them, X and Y represent the horizontal and vertical coordinates of each point in the point cloud on the projection plane, Z represents the vertical coordinate of each point in the point cloud, the path is centered on the geometric center of the point cloud, and R represents a fixed observation distance; Open3D is used for rendering, and the rendered 2D image is embedded into the quality perception space using a 2D image encoder. The calculation formula is as follows: in, represents the tth image projection Quality-aware embedding of Represents an image encoder The number of output channels, is the merged result after average fusion.
5. The point cloud quality assessment method based on deep ordinal regression according to claim 2 is characterized in that: The fused global feature representation generation step specifically includes: Given features from point cloud and image modalities and , use a linear layer to unify them to the same size: in, and is the adjusted feature, and is a learnable linear mapping, is the adjusted number of channels; through the multi-head attention module: in, It means that long positions should pay attention to the operation. represents the attention function, Q represents the query, K represents the key, V represents the value, Indicates head, , , , Represents a learnable linear mapping, which is finally weighted and summed to obtain the output of the attention head. The final quality embedding is cascaded with the guided multimodal features obtained by the intra-modal features and the symmetric cross-modal attention module. The calculation formula is as follows: in, Indicates cascade operation, represents a symmetric cross-modal attention operation, Represents the final quality-aware features
Citation Information
Patent Citations
Point cloud quality evaluation method and device based on regularization representation
CN118351423A
Point cloud quality evaluation method based on optimal view angle selection and multi-modal fusion
CN118967591A