Multi-dimensional labeling credibility evaluation method based on visual features and perceptual alignment
By constructing a multi-dimensional annotation credibility evaluation method based on visual features and perception alignment, a multi-dimensional feature fusion framework is built, which solves the problem of annotation credibility evaluation in complex scenarios, realizes efficient and automated annotation quality evaluation, and improves the efficiency of annotation tasks and the credibility of evaluation results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-15
- Publication Date
- 2026-04-10
AI Technical Summary
Existing annotation quality assessment methods struggle to achieve high-reliability assessments in complex scenarios, especially in assessing whether annotations have sufficient discriminative evidence. Furthermore, traditional methods are inefficient or lack the ability to judge semantic rationality.
A multi-dimensional annotation credibility evaluation method based on visual features and perception alignment is adopted. By integrating visual features and perception mechanisms, a multi-dimensional interpretable analysis framework is constructed, including analysis of three dimensions: target spatial distribution, image quality, and image saliency. Multiple features are extracted, normalized, and weighted, and finally a comprehensive evaluation score is generated.
It enables intelligent and fully automated evaluation of labeled data, significantly improving the efficiency and scalability of labeling tasks, reducing the frequency of manual review, resolving the problem of inconsistent labeling results, and enhancing the credibility and transparency of evaluation results.
Smart Images

Figure CN121837892A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision, and more specifically to a method for evaluating the credibility of multidimensional annotations based on visual features and perceptual alignment. Background Technology
[0002] With the widespread application of intelligent vision systems in fields such as autonomous driving, industrial quality inspection, and remote sensing monitoring, image annotation data serves as the core foundation for model training and validation, and its quality directly affects the reliability and generalization ability of algorithm performance. High-quality annotation not only requires accurate geometric positioning but also semantic integrity and contextual consistency. Achieving a highly reliable data supply in complex scenarios has become a key bottleneck restricting the implementation of artificial intelligence systems.
[0003] Currently, commonly used methods for evaluating annotation quality mainly include manual review, rule-based validation, and automated detection based on Intersection over Union (IoU). Manual review relies on expert experience, offering high accuracy but low efficiency, making it unsuitable for handling massive amounts of data. Rule-based validation can only identify explicit errors such as boundary overflow and missing labels, lacking the ability to judge semantic reasonableness. IoU-based methods rely on standard truth boxes, suitable for comparing annotation results, but unable to assess whether annotations in a single sample possess sufficient discriminative basis. All of these methods fail to comprehensively reflect the credibility of annotations in real-world application scenarios.
[0004] Against this backdrop, intelligent evaluation technologies based on image content understanding have gradually attracted attention. Research shows that factors such as the visual salience of a target in an image, the degree of background interference, spatial crowding, and geometric deformation directly affect human efficiency in target recognition and annotation confidence. By modeling these perceptually relevant features, an automatic evaluation system that more closely resembles human cognitive mechanisms can be constructed. Specifically, visual feature analysis can capture the distribution characteristics and structural clarity of targets in an image, while the perceptual alignment mechanism simulates the distribution of human eye attention and semantic understanding processes to achieve fine-grained judgments on the reasonableness of annotations. Compared to traditional methods, this type of technology does not rely on standard truth values, possesses stronger universality and interpretability, and is particularly suitable for pre-evaluation of annotation confidence under no-reference conditions.
[0005] Current research has attempted to introduce saliency detection or fuzzy classifiers for coarse quality grading, but these methods still fall short in terms of feature dimensionality completeness, multi-factor coupled modeling, and quantification of evaluation results, making it difficult to accurately characterize the credibility of annotations in complex scenarios. Therefore, there is an urgent need to develop a multi-dimensional annotation credibility evaluation method that integrates visual features and perceptual alignment to improve the quality control capabilities of intelligent systems at the data source. Summary of the Invention
[0006] To address the aforementioned technical issues, this invention provides a multi-dimensional annotation credibility evaluation method based on visual features and perception alignment. By integrating visual features and perception mechanisms, it proposes a multi-dimensional interpretable annotation credibility evaluation framework, aiming to improve annotation efficiency and accuracy, and serve the construction and governance of high-quality visual datasets.
[0007] To achieve the above objectives, the present invention adopts the following technical solution:
[0008] A multidimensional annotation credibility evaluation method based on visual features and perceptual alignment includes the following steps:
[0009] Step 1: Input the image data to be evaluated into three analysis dimensions in parallel, and perform annotation quality analysis based on target spatial distribution, annotation usability analysis based on image quality, and annotation usability analysis based on image saliency respectively;
[0010] Step 2: In the dimension of annotation quality analysis based on target spatial distribution, it is further divided into two sub-dimensions: target region analysis and target dense region analysis. The other two analysis dimensions maintain a single hierarchical structure.
[0011] Step 3: Extract the corresponding visual features for each analysis dimension and its sub-dimensions. Specifically, target region analysis extracts three types of features: target region saliency, target region feature distribution, and target region semantic integrity. Target dense region analysis extracts two types of features: dense region depth density and dense region confusion. Image quality-based annotation usability analysis extracts two types of features: visual recognition and geometric fidelity. Image saliency-based annotation usability analysis extracts two types of features: foreground crowding and background interference.
[0012] Step 4: Normalize and weighted aggregate the features extracted from each dimension to generate four local quality scores. Then, merge the scores of target region analysis and target dense region analysis into a label quality score, and merge the scores of image quality and image saliency analysis into a label usability score.
[0013] Step 5: Based on the annotation quality score and annotation usability score, generate the final comprehensive annotation evaluation score through weighted fusion or rule-based decision-making mechanism.
[0014] Beneficial effects:
[0015] 1. This invention provides a highly automated method for evaluating annotation quality. By constructing a quantitative evaluation system that integrates multi-dimensional features, it achieves intelligent and fully automated evaluation of annotated data, significantly reducing the frequency and cost of manual review. Compared to traditional methods relying on manual sampling checks, this solution greatly improves the overall efficiency and scalability of annotation tasks, and is particularly suitable for rapid iteration scenarios involving large-scale datasets.
[0016] 2. This invention proposes a standardized evaluation framework based on the combination of low-level image features and high-level semantics, which effectively solves the problem of inconsistent annotation results caused by differences in subjective judgment and experience levels of annotators.
[0017] 3. By introducing an interpretability analysis mechanism and a multi-level confidence assessment model, this invention significantly improves the credibility and transparency of automated annotation evaluation results. Attached Figure Description
[0018] Figure 1 This is a flowchart of a multi-dimensional annotation credibility evaluation method based on visual features and perceptual alignment according to the present invention. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.
[0020] The present invention provides a multi-dimensional annotation credibility evaluation method based on visual features and perceptual alignment, which includes three evaluation dimensions: annotation quality analysis based on target spatial distribution, annotation usability analysis based on image quality, and annotation usability analysis based on image saliency. The annotation quality analysis based on target spatial distribution includes two sub-dimensions: target region analysis and target dense region analysis.
[0021] Annotation quality analysis based on target spatial distribution aims to evaluate the distribution characteristics within the target region, whether the target is a sample within a dense region, and to fit the influence of density on overall quality; annotation usability analysis based on image quality aims to evaluate the overall visual discriminability of the image; annotation usability analysis based on image saliency aims to extract the semantic confusion of the image through saliency analysis and to fit the influence of the saliency of this feature on annotation.
[0022] Annotation quality analysis based on target spatial distribution aims to systematically evaluate the distribution characteristics of targets in image space, identify whether targets are concentrated in high-density areas, and quantify the impact of density on annotation quality. Annotation usability analysis based on image quality focuses on evaluating the overall visual recognizability of an image. Through quantitative analysis of the image's low-level visual features, it determines whether the image has sufficient information fidelity to support high-quality manual or automatic annotation, thereby filtering out samples with unreliable annotations due to poor image quality. Annotation usability analysis based on image saliency calculates a visual saliency map, extracts the distribution features of semantic information in the image, and evaluates its semantic complexity and clutter level. This method aims to identify whether there are too many interfering regions, cluttered backgrounds, or distracting elements in the image, and then fits the degree of interference of saliency features on the annotator's attention allocation and annotation consistency.
[0023] Specifically, such as Figure 1 As shown, the present invention provides a multi-dimensional annotation credibility evaluation method based on visual features and perceptual alignment, comprising the following steps:
[0024] Step 1: Perform parallel analysis on the data to be evaluated in three dimensions: annotation quality analysis based on target spatial distribution, annotation usability analysis based on image quality, and annotation usability analysis based on image saliency.
[0025] Step 2: Based on the above analysis, the dimension of annotation quality analysis based on target spatial distribution is further subdivided into two sub-dimensions: (1) target region analysis; (2) target dense region analysis. The other two dimensions (annotation usability analysis based on image quality and annotation usability analysis based on image saliency) maintain a single-level structure and are not further subdivided.
[0026] Step 3: Multi-dimensional labeled feature extraction:
[0027] Based on the preliminary analysis, feature extraction is systematically carried out for each analysis dimension and its sub-dimensions, as follows:
[0028] The target region analysis (sub-dimension) extracts three types of features: target region saliency assessment, target region feature distribution assessment, and target region semantic integrity assessment.
[0029] The target dense region analysis (sub-dimension) extracts two types of features: dense region depth density estimation and dense region confusion degree estimation.
[0030] Image quality-based annotation usability analysis (main dimension) extracts two types of features: visual recognition assessment and geometric fidelity assessment.
[0031] Image saliency-based annotation availability analysis (main dimension) extracts two types of features: forward crowding assessment and background interference assessment.
[0032] Step 4: Generation and fusion of multi-dimensional quality scores:
[0033] Based on the feature extraction of each dimension, a quantitative evaluation is performed on the four analysis dimensions (including two sub-dimensions and two main dimensions) to generate corresponding local quality scores. These scores are then fused hierarchically based on their semantic relevance. The specific process is as follows:
[0034] (1) Each analysis dimension (including two sub-dimensions and two main dimensions) is independently normalized and weighted, generating a quality assessment score for that dimension:
[0035] Target region analysis outputs a target region quality score;
[0036] The analysis of dense target areas outputs a quality score for the dense areas.
[0037] Image quality-based annotation usability analysis outputs an image quality usability score;
[0038] The annotation usability analysis based on image saliency outputs a saliency usability score.
[0039] (2) Perform hierarchical score fusion:
[0040] The quality scores generated from the first two sub-dimensions (target area analysis and target dense area analysis) are integrated using learnable or predefined fusion strategies (such as weighted averaging, rule-based discrimination, etc.) to form a unified annotation quality score, which reflects the rationality and completeness of the annotation content in terms of spatial distribution.
[0041] The usability scores of the latter two main dimensions (image quality and image saliency analysis) are combined to generate a label usability score, which measures the credibility of the labeling results and the strength of visual support under the current observation conditions.
[0042] (3) Finally, two high-level evaluation indicators are output:
[0043] The annotation quality score represents the intrinsic quality of the geometric and semantic structure of the annotation; the annotation usability score represents the extrinsic support capability of the image conditions on which the annotation depends.
[0044] Step 5: Based on the two core evaluation indicators, annotation quality score and annotation usability score, a multi-dimensional comprehensive evaluation model is further constructed to generate the final comprehensive annotation evaluation score. This process uses weighted fusion or rule-based decision-making mechanisms (such as linear combination, nonlinear integration, fuzzy inference systems, etc.) to combine the relative importance of the two indicators in practical application scenarios, thereby achieving a unified quantitative judgment on the overall rationality of the annotations.
[0045] Preferably, in step 1, based on the inherent laws and perceptual preferences of the human visual system, image features are aligned with perception using three dimensions: annotation quality analysis based on target spatial distribution, annotation usability analysis based on image quality, and annotation usability analysis based on image saliency.
[0046] Example:
[0047] This embodiment of a multi-dimensional annotation credibility evaluation method based on visual features and perceptual alignment includes the following steps:
[0048] Step 1: The data to be evaluated is input in parallel and analyzed from three dimensions: annotation quality analysis based on target spatial distribution, annotation usability analysis based on image quality, and annotation usability analysis based on image saliency.
[0049] Step 2: Based on the above analysis, the annotation quality analysis based on target spatial distribution is further subdivided into two sub-dimensions: (1) target region analysis; (2) target dense region analysis. The other two dimensions (annotation usability analysis based on image quality and annotation usability analysis based on image saliency) maintain a single-level structure and are not further subdivided.
[0050] Step 3: Multi-dimensional labeled feature extraction:
[0051] Based on the completed preliminary analysis framework, feature extraction is systematically carried out for each analysis dimension and its sub-dimensions, as follows:
[0052] Target region analysis (sub-dimension) extracts three types of features:
[0053] 1. Target area salience assessment:
[0054] First, based on the human visual system's priority attention mechanism for salient objects, saliency priors are introduced as one of the cues. Advanced saliency detection models (such as...) are then used. Output the full-image saliency. , Indicates the image height (in pixels). This represents the image width (in pixels), where each pixel value indicates the probability that it is perceived as a foreground object. Then, within a given detection box... Extract the corresponding saliency subgraph from the region. Calculate the average significance corresponding ,in Represents the detection box Variance of a single pixel location within the saliency response , This represents the variance operation, used to measure the dispersion of intra-pixel saliency values. The first quality score is defined as a weighted combination of the two. : , , Weighting coefficients set manually.
[0055] 2. Target area feature distribution assessment:
[0056] Considering that saliency relies solely on visual comparison and is insufficient to characterize semantic-level structural organization, this invention introduces a visual language pre-training model based on contrastive learning (such as CLIP-ViT-L / 14 or DINOv2-ViT-giant) to perform non-overlapping slicing (patch-level segmentation) of the image and extract high-level feature vectors for each image patch. , Represents the space of real numbers. It represents the dimension of the feature vector and builds three complementary feature distribution measures on this basis: feature compactness, boundary abruptness and region discrimination, to comprehensively characterize the internal consistency and external contrast of the target.
[0057] Examine the degree of clustering among features of different image patches within the target region. Let... This represents the set of indices of blocks that fall within the labeled area (bbox / mask); This represents a set of image patches within the target region. This represents the total number of image patches after the entire image has been segmented. , Indicates that the index is The feature vector of the image patch.
[0058] If the characteristic mean is used, the energy form of the characteristic covariance is used to measure the degree of dispersion, and the sensitivity to small differences is enhanced by a negative logarithmic transformation, resulting in the logarithmic exponent of characteristic dispersion. :
[0059] ;
[0060] in, express Norm.
[0061] Analyze the degree of jump between the target and its neighboring background in the feature space:
[0062] definition Given a set of image patch pairs representing the target boundary entropy, calculate the sum of the cosine similarity complements between each pair of feature vectors crossing the boundary to obtain the normalized boundary change intensity. , express The set of image patch indexes within, , These represent the feature vectors of the two image blocks in the image block pair, respectively.
[0063] To further quantify the overall distributional difference between the target and its surrounding context, the maximum mean difference (MMD) is used to measure the difference between the target region and the surrounding background region (ring-shaped region). The difference in the feature distribution of the annular region surrounding the target bounding box is denoted as... The feature set is , Let k be the feature vector within the annular region, and k be the index value; then the linear approximation formula for squared MMD is:
[0064] ;
[0065] in, For the target region feature set The number of elements; Background annular region feature set The number of elements; For the target region feature set Elements in; Background annular region feature set The elements; Ф(·) is the feature mapping function corresponding to the Gaussian kernel.
[0066] Based on the above three characteristics, the second quality score is constructed. This is used to characterize the structural rationality and separability of a target in the deep semantic space:
[0067] .
[0068] Target semantic integrity assessment:
[0069] First, semantic coverage measures the proportion of valid semantic pixels belonging to the target category within the detection box, reflecting the model's actual coverage of the target semantic region. Specifically, a high-precision semantic segmentation model (such as the DeepLab series) is used to infer from the input image, generating a pixel-level category prediction map. Given a category and its corresponding detection area Count the number of pixels belonging to a category. ; It is an indicator function, that is, when the condition within the parentheses (i.e., the pixel) is met. semantic prediction results When the condition equals category c), the function value is 1; when the condition is not met, the function value is 0. Let the area of the target region corresponding to the detection box be... Then the semantic coverage rate is:
[0070] ;
[0071] Secondly, considering that targets are often partially invisible in real-world scenarios due to occlusion by other objects, relying solely on semantic coverage may overestimate completeness. Therefore, instance-level occlusion estimation is further introduced. The occlusion ratio for each instance can be output by training or using a pre-trained instance-level occlusion estimation network (e.g., combining a Mask R-CNN convolutional neural network with an occlusion estimation head). .
[0072] Finally, the combined score for the semantic integrity assessment of the target region is obtained by combining these two features. :
[0073] ;
[0074] in, For weight hyperparameters.
[0075] Target-dense region analysis (sub-dimension) extracts the following two types of features:
[0076] 1. Depth density estimation in dense areas:
[0077] To obtain reliable and semantically consistent depth information, advanced monocular depth estimation models (such as the DepthAnything-large model or the DPT-Hybrid depth estimation model based on the hybrid architecture visual Transformer) are employed. A prompt-guided depth estimation mechanism based on the contrastive learning-based visual language pre-trained model CLIP is introduced. Natural language prompts are used to enhance attention to specific categories or regions of interest, thereby improving the accuracy of depth reasoning for key targets in complex backgrounds and obtaining the resulting depth map. Focusing on the target region to be analyzed, the corresponding 3D point set is voxelized or kernel density estimated (KDE) to construct a continuous point cloud density thermal field. On this density field, a set of complementary statistical features are further extracted to describe the dense properties of the region from different perspectives.
[0078] , representing the spatial density mean, where It is the first in the target area The density values corresponding to each sampling point; This is the total number of sampling points within the target area;
[0079] , representing the density concentration, where It is the standard deviation of the density values of all sampling points within the target area. It is the maximum density value within the target area;
[0080] Indicates the number of targets that are not obscured;
[0081] , representing the density distribution entropy, where The index within the target area is The probability corresponding to the density interval.
[0082] The following mass function is finally obtained through synthesis:
[0083] ;
[0084] 2. Estimation of the degree of confusion in densely populated areas:
[0085] At the semantic level, entropy considers the degree of visual confusion between the target and its surrounding sub-regions. It utilizes a pre-trained CNN (such as ResNet-50) to extract global average pooling features from multiple sub-regions (e.g., a 4×4 grid) within the detection box. , The dimension of the features is [value], which capture high-level semantic information of the local region. The mean of the cosine similarity across all pairs is further calculated as a measure of feature consistency within the region. , . The total number of sub-regions. , This is a feature index for the sub-region.
[0086] At the category discrimination level, modern detectors (YOLOv8 / DETR, etc.) are used to obtain the category probability vector of the detection box. ,in For the detection box to belong to the first Predicted probability of class Let this be the number of categories. Its Shannon entropy is introduced to measure its disorder: To eliminate the number of categories The effect of entropy range is normalized to obtain the degree of class confusion: .
[0087] Local entropy is used in terms of texture and structure. Local binary model LBP variance As parameters for quantifying texture features, both reflect the randomness and repetitiveness of pixel intensity variations within a region; higher values indicate richer or more chaotic textures. Therefore, texture activity is defined as the larger of the two: Gradient mean is used. Weak edge ratio Edge alignment Contour curvature As parameters for quantifying structural characteristics, these factors are combined to construct a structural fuzzy term: .
[0088] To address the reliability of model predictions, two uncertainty measures are introduced: confidence variance. Multi-scale prediction inconsistency Combining category confusion, we construct the total uncertainty component: .
[0089] The final mass function is:
[0090] ;
[0091] in, For parameters, This represents an exponential function.
[0092] Image quality-based annotation usability analysis (main dimension) extracts two types of features:
[0093] 1. Visual recognition assessment:
[0094] First, we focus on visual recognizability, that is, whether the image content is clear enough to be clearly identified and accurately delineated. To this end, we introduce two key indicators: sharpness and illumination uniformity. Sharpness The sharpness of image edges is measured by the variance of the Laplacian operator. Illumination uniformity. Then, through the lightness channel in the HSV (Hue-Saturation-Lightness) color space ( The histogram entropy of the channel is modeled, and its standard deviation is further calculated.
[0095] 2. Geometric fidelity evaluation, i.e., whether the spatial structure in the image faithfully reflects the geometric relationships of the real world. To this end, Hough transform is used to detect the main linear structures in the image, and the average offset distance of each detected line relative to the ideal linear model is calculated. Therefore, the linear distortion index is defined. ,in This represents the total number of lines detected.
[0096] To further transform the aforementioned raw features into intuitive quality scores, facilitating cross-sample comparison and fusion decision-making, three types of normalized quality response functions are designed:
[0097] ;in, These are the parameters of the mass function;
[0098] ;in, These are the parameters of the mass function;
[0099] ;in, This represents the standard deviation of the average offset distance.
[0100] Image saliency-based annotation usability analysis (main dimension) extracts two types of features:
[0101] 1. Foreground Crowding Assessment:
[0102] First, we introduce the foreground fill rate. This is used to measure the proportion of the total image area covered by all detected target bounding boxes. The calculation formula is as follows:
[0103] ;
[0104] in, This indicates the area of the corresponding label box. Let W represent the bounding box of the q-th detected target, and let W and H be the width and height of the image, respectively.
[0105] Further propose boundary ambiguity This is used to measure the distinguishability of boundaries between adjacent objects. Specifically, Canny or Sobel edge detection is performed on the original image to obtain a global edge map. Calculate the gap region for each pair of adjacent BBox target bounding boxes. Extract the edge intensity within the region and calculate its average gradient intensity:
[0106] ;
[0107] in, Representing an image At pixel gradient at; For set The number of pixels within.
[0108] Considering that annotation is essentially a selective attention process, this invention introduces a crowding factor based on visual saliency. Saliency maps are generated using advanced saliency models such as Segment All Model (SAM), Deep Gaze III, and Deep Saliency Network (DSS). Extract the following sub-features: number of salient regions. Significant area average Euclidean distance Significant Shannon entropy Image diagonal length Threshold for the number of significant regions The fusion is used to form a significant crowding factor, as shown in the following formula:
[0109] ;
[0110] in, This is the Sigmoid function.
[0111] The quality score for foreground crowding is defined as:
[0112] .
[0113] 2. Background interference level assessment:
[0114] First, consider the complexity of the background texture. The gray-level co-occurrence matrix (GLCM) was used to analyze the second-order statistical properties of the segmented background region. GLCMs were constructed under multiple directions and displacement distances, and two key indicators were extracted from them:
[0115] Texture entropy: defined as It is used to measure the spatial uncertainty or randomness of texture; among which Indicates the grayscale value in the image and The probability of both occurring simultaneously.
[0116] Contrast ratio: defined as It reflects the intensity of grayscale difference between adjacent pixels.
[0117] These two metrics together characterize the structural activity of the background in local space, and the texture complexity score is finally obtained by weighted combination:
[0118] ;
[0119] in, These are the weight parameters.
[0120] Introducing contextual semantic diversity This method aims to capture the information redundancy of the background at a high semantic level. A semantic segmentation model (such as DeepLabV3+) is used to obtain the category label map of background pixels, and the category of the foreground object is filtered out. The frequency of occurrence of each category in the background region is calculated. ,in If the number of background categories is , then its Shannon entropy is . Then normalize the value to get the score:
[0121] ;
[0122] in, It is the total number of all categories in the dataset.
[0123] Analyzing the energy distribution characteristics of the background from a frequency domain perspective, we propose a high-frequency energy proportion... Based on the two-dimensional Fourier transform of the image, the background region image is transformed into the frequency domain space to obtain the amplitude spectrum. Polar coordinate resampling is used to create radial bins, separating low-frequency and high-frequency components. The high-frequency energy percentage is defined as follows:
[0124] ;
[0125] in, The frequency domain amplitude spectrum of the background region image after two-dimensional Fourier transform. f0 is the frequency domain coordinate and the threshold for distinguishing between high and low frequencies.
[0126] In summary, the three indicators mentioned above reveal the essential mechanisms of background interference from three orthogonal dimensions: spatial domain structure, semantic layer cognition, and frequency domain energy distribution. To form a unified background interference quality score, these indicators are linearly fused as follows:
[0127] ;
[0128] Step 4: Generation and fusion of multi-dimensional quality scores, including:
[0129] Based on the feature extraction of each dimension, the four analysis units are quantitatively evaluated to generate corresponding local quality scores, and then fused hierarchically according to their semantic relevance. The specific process is as follows:
[0130] (1) Each analysis dimension (including two sub-dimensions and two main dimensions) is independently normalized and weighted, generating a quality assessment score for that dimension:
[0131] Target region analysis outputs target region quality score:
[0132] ;
[0133] Target dense area analysis outputs dense area quality score:
[0134] ;
[0135] Image quality-based annotation usability analysis outputs an image quality usability score:
[0136] ;
[0137] in, These are the weight parameters.
[0138] Image saliency-based annotation usability analysis → Output "Saliency Usability Score":
[0139] ;
[0140] in These are the weight parameters.
[0141] (2) Perform hierarchical score fusion:
[0142] The quality scores generated for the first two sub-dimensions (target region analysis and target dense region analysis) are obtained by using learnable or predefined fusion strategies (such as weighted averaging, rule-based discrimination, etc.). The annotations are integrated to form a unified annotation quality score, which is used to reflect the rationality and completeness of the spatial distribution of the annotation content.
[0143] The usability scores of the latter two main dimensions (image quality and image saliency analysis) The data is then fused to generate an "annotation usability score," which measures the credibility of the annotation results and the strength of visual support under the current observation conditions.
[0144] (3) Finally, two high-level evaluation indicators are output:
[0145] Mark quality score : Characterizes the intrinsic quality of the geometric and semantic structure of the annotation;
[0146] Usability score : Represents the external support capability of the image conditions on which the annotation depends.
[0147] Step 5: Obtain the labeled quality score Usability score of annotation Based on the two core evaluation indicators, a multi-dimensional comprehensive evaluation model is further constructed to generate the final comprehensive evaluation score for annotation. This process, through weighted fusion or rule-based decision-making mechanisms (such as linear combination, nonlinear integration, fuzzy inference systems, etc.), combines the relative importance of the two in practical application scenarios to achieve a unified quantitative judgment on the overall rationality of the annotation.
[0148] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A multi-dimensional annotation credibility evaluation method based on visual features and perceptual alignment, characterized in that, Includes the following steps: Step 1: Input the image data to be evaluated into three analysis dimensions in parallel, and perform annotation quality analysis based on target spatial distribution, annotation usability analysis based on image quality, and annotation usability analysis based on image saliency respectively; Step 2: In the dimension of annotation quality analysis based on target spatial distribution, it is further divided into two sub-dimensions: target region analysis and target dense region analysis. The other two analysis dimensions maintain a single hierarchical structure. Step 3: Extract the corresponding visual features for each analysis dimension and its sub-dimensions. Specifically, target region analysis extracts three types of features: target region saliency, target region feature distribution, and target region semantic integrity. Target dense region analysis extracts two types of features: dense region depth density and dense region confusion. Image quality-based annotation usability analysis extracts two types of features: visual recognition and geometric fidelity. Image saliency-based annotation usability analysis extracts two types of features: foreground crowding and background interference. Step 4: Normalize and weighted aggregate the features extracted from each dimension to generate four local quality scores. Then, merge the scores of target region analysis and target dense region analysis into a label quality score, and merge the scores of image quality and image saliency analysis into a label usability score. Step 5: Based on the annotation quality score and annotation usability score, generate the final comprehensive annotation evaluation score through weighted fusion or rule-based decision-making mechanism.
2. The multi-dimensional annotation credibility evaluation method based on visual features and perceptual alignment as described in claim 1, characterized in that, The target region saliency assessment is based on the saliency map output by the saliency detection model. The average saliency response and the variance of the saliency response are calculated within the labeled region to characterize the likelihood of the target being noticed.
3. The multi-dimensional annotation credibility evaluation method based on visual features and perceptual alignment as described in claim 1, characterized in that, The target region feature distribution evaluation uses a visual Transformer to extract high-level features of image patches, and measures the internal consistency and external contrast of the target through three indicators: feature compactness, boundary abruptness, and regional discrimination.
4. The multi-dimensional annotation credibility evaluation method based on visual features and perceptual alignment as described in claim 1, characterized in that, The semantic integrity assessment of the target region combines the category prediction map output by the semantic segmentation model to calculate the semantic coverage rate, and introduces the occlusion ratio output by the instance-level occlusion estimation network to correct the integrity judgment.
5. The multi-dimensional annotation credibility evaluation method based on visual features and perceptual alignment as described in claim 1, characterized in that, The depth density estimation of the dense area adopts a monocular depth estimation model combined with a natural language prompting mechanism to generate a depth map, and performs kernel density estimation on the target's corresponding 3D point set to extract spatial density statistical features.
6. The multi-dimensional annotation credibility evaluation method based on visual features and perceptual alignment as described in claim 1, characterized in that, The estimation of the degree of confusion in the dense region includes calculating the feature consistency of sub-regions within the target, the normalized Shannon entropy of the detection box category probability vector, texture activity, structural ambiguity term, and model prediction uncertainty component.
7. The multi-dimensional annotation credibility evaluation method based on visual features and perceptual alignment as described in claim 1, characterized in that, The visual recognition assessment measures image sharpness using the Laplacian operator variance and illumination uniformity using the HSV color space luminance channel histogram entropy and its standard deviation.
8. The multi-dimensional annotation credibility evaluation method based on visual features and perceptual alignment as described in claim 1, characterized in that, The geometric fidelity assessment uses Hough transform to detect the main straight line structure of the image and calculates the average offset distance of the detected straight line relative to the ideal straight line model to construct the straight line distortion index.
9. The multi-dimensional annotation credibility evaluation method based on visual features and perceptual alignment as described in claim 1, characterized in that, The foreground crowding assessment includes calculating the proportion of the total image area covered by all detection boxes, the mean edge intensity of the gap region between adjacent bounding boxes, and the number of salient regions extracted based on the salientity map, the mean Euclidean distance of the salient regions, and the salient Shannon entropy.
10. The multidimensional annotation credibility evaluation method based on visual features and perceptual alignment as described in claim 1, characterized in that, The background interference assessment includes calculating the background texture entropy and contrast based on the gray-level co-occurrence matrix, calculating the normalized Shannon entropy of background semantic diversity based on the semantic segmentation results, and calculating the proportion of high-frequency energy in the background based on Fourier transform.