Image data annotation method and system

By constructing a dynamic edge weight matrix and a quantum state circuit to optimize the image data annotation method, the labeling error problem caused by dynamic changes in data distribution in the prior art is solved, efficient and accurate image data annotation is achieved, and the adaptability and robustness of the model are enhanced.

CN120256836AActive Publication Date: 2025-07-04CHINA NAT INST OF STANDARDIZATION +1

Patent Information

Application Number
CN202510747994.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-07-04
Estimated Expiration
2045-06-05

AI Technical Summary

Technical Problem

The existing image data annotation methods rely on static graph structures or fixed weight matrix, which are difficult to adapt to the dynamic changes in data distribution, resulting in the accumulation of label propagation errors. The existing singular value decomposition methods ignore local topological structures, resulting in the loss of semantic information. The clustering algorithm based on classic machine learning has limited processing capabilities for high-dimensional sparse data, and lacks a quantitative evaluation mechanism for potential conflict labels. The collaborative optimization of quantum computing and NeRF models and label propagation cannot be effectively solved, resulting in the labeling results being prone to deviations.

Method used

By collecting image and text data, preprocessing hash codes and building a dynamic edge weight matrix, combining singular value decomposition to generate a low-rank approximation matrix, updating it using kernel density estimation and dynamic equations, Euler discretization combines CLP propagation and EA optimization, topological weighted entanglement entropy is calculated through quantum state circuits and partial traces, potential conflict pairs are generated and priority is calculated, regional features are output using CNN and NeRF models, and extracted and updated in combination with CLIP models, and finally a verification fusion label distribution matrix is ​​generated.

Benefits of technology

It improves the accuracy and consistency of image data labeling, improves the efficiency and accuracy of labeling, enhances the adaptability and robustness of the model to dynamic scenes, significantly reduces the labeling deviation, and improves the ability of feature extraction and semantic representation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256836A_ABST
    Figure CN120256836A_ABST
Patent Text Reader

Abstract

The invention discloses an image data annotation method and system, and relates to the technical field of computer vision, and the method comprises the steps: collecting image and text data, carrying out the updating through kernel density estimation and a kinetic equation, carrying out the optimization through employing Euler discretization in combination with CLP propagation and EA, carrying out the conversion through amplitude coding, and carrying out the construction through a quantum state circuit. The method comprises the following steps: carrying out calculation by using partial traces, generating topological weighted entanglement entropy through weighted fusion, carrying out updating based on an attention mechanism, generating potential conflict pairs, calculating priorities, generating final conflict pairs through a CNN model, outputting region features through a NeRF model, carrying out extraction by using a CLIP model, and carrying out updating through a loss function. According to the method, VAE is combined with GLP propagation and EA for optimization, the precision and consistency of labeling are improved, and through joint verification of quantum entanglement entropy calculation and the NeBF model, the efficiency and accuracy of labeling are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to an image data annotation method and system. Background Art

[0002] With the rapid development of deep learning, multi-modal modeling, and big data processing technologies, image data annotation, as one of the core technologies in the field of computer vision, has made great progress. Early image data annotation mainly relied on manual annotation or rule-based automated methods, which were costly and prone to introducing subjective biases. To improve efficiency, technologies such as VAE, GAN, and NeRF have begun to be gradually used, and clustering algorithms and dimensionality reduction technologies have also been widely used for data preprocessing and feature extraction.

[0003] Existing image data annotation methods still have deficiencies. They rely on static graph structures or fixed weight matrices, making it difficult to adapt to the dynamic changes in data distribution, resulting in the accumulation of annotation propagation errors. Existing singular value decomposition methods usually ignore local topological structures, causing semantic information loss. Clustering algorithms based on classical machine learning have limited processing capabilities for high-dimensional sparse data and lack a quantitative evaluation mechanism for potential conflicting labels. The collaborative optimization of quantum computing and NeRF models with label propagation has not been effectively solved, leading to prone biases in annotation results. Summary of the Invention

[0004] In view of the above existing problems, the present invention is proposed.

[0005] Therefore, the present invention provides an image data annotation method and system to solve the problems of relying on static graph structures or fixed weight matrices, being difficult to adapt to the dynamic changes in data distribution, resulting in the accumulation of annotation propagation errors, existing singular value decomposition methods usually ignoring local topological structures, causing semantic information loss, clustering algorithms based on classical machine learning having limited processing capabilities for high-dimensional sparse data, lacking a quantitative evaluation mechanism for potential conflicting labels, and the collaborative optimization of quantum computing and NeRF models with label propagation not being effectively solved, leading to prone biases in annotation results.

[0006] To solve the above technical problems, the present invention provides the following technical solutions: In a first aspect, the present invention provides an image data annotation method, which includes collecting image and text data, preprocessing the two collected data, generating hash codes through VAE and constructing a dynamic edge weight matrix, generating a low-rank approximation matrix by combining singular value decomposition, updating through kernel density estimation and kinetic equations, using Euler discretization combined with CLP propagation and EA for optimization, based on the optimization results, performing conversion using amplitude encoding, and constructing through a quantum state circuit, performing calculations using partial trace, generating topological weighted entanglement entropy through weighted fusion, updating based on the attention mechanism, generating potential conflict pairs, and calculating priorities, generating final conflict pairs through a CNN model, outputting regional features through a NeRF model, extracting using a CLIP model, and updating through a loss function, generating a verification fusion label distribution matrix through an MLP model, generating a heat map through ECharts, and storing the two collected data.

[0007] As a preferred solution of the image data annotation method of the present invention, wherein: the collecting of image and text data, and using Euler discretization combined with CLP propagation and EA for optimization includes: Based on the multimodal feature vector and the attention region feature, calculate the global similarity and the local attention similarity respectively through the cosine formula; Use VAE to map the multimodal feature vector to a hash code, and calculate the hash similarity through a Gaussian kernel function; The global similarity, the local attention similarity, and the hash similarity are weighted and fused to obtain a comprehensive similarity, and the comprehensive similarity is defined as an edge, and the image ID is a node to construct a dynamic edge weight matrix; Based on the elbow method, set the number of categories K, use K-Means clustering to perform K clusters on the multimodal feature vector, and map to string labels to generate initial category labels, and generate an initial probability distribution through maximum likelihood estimation; Map the hash code to the initial probability distribution through kernel density estimation, and update the initial probability distribution; Based on each node, define a kinetic equation, use the Euler method for discretization, iterate on the updated initial probability distribution, and perform GLP propagation using a low-rank approximation matrix; Perform weighted fusion on the propagated initial probability distribution and the updated initial probability distribution to generate a fusion label distribution matrix; Based on the dynamic edge weight matrix, update the fusion label distribution matrix through EA; Based on the updated fusion distribution label matrix, define an EA objective function and optimize the two weight parameters in the kinetic equation; Based on each node, screen the maximum value in the optimized fusion distribution label matrix to generate the final label of each node.

[0008] As a preferred solution of the image data annotation method described in the present invention, wherein: based on the optimization result, through the CNN model, generate the final conflict pairs, including: Convert the optimized fusion label distribution matrix into a quantum state vector through amplitude encoding, define an edge set based on the dynamic edge weight matrix, construct an entangled state with the quantum state vector through a quantum state circuit, and generate a topological feature vector for each node through TDA; Based on the entangled state, calculate the reduced density matrix through partial trace; Normalize the topological feature vector using the L2 norm and map it using the Sigmoid function to generate the weight of the topological feature vector; Calculate the entanglement entropy of the reduced density matrix through Qiskit, perform weighted fusion by combining the weight of the topological feature vector, generate the topological weighted entanglement entropy, and perform adjustment to generate the adjusted topological weighted entanglement entropy; Based on the multi-modal feature vector, obtain the modal weight through the attention mechanism, perform a multiplication operation on the adjusted topological weighted entanglement entropy to generate the latest topological weighted entanglement entropy; If the final label of each node is not equal to the neighbor label and the edge weight between the two is greater than the edge weight threshold, then mark this node and the neighbor node as potential conflict pairs and calculate the priority; Sort the priorities in descending order, select the top E potential conflict pairs, where E is the number of potential conflict pairs, to generate a set of potential conflict pairs; Construct a GNN model, input the topological feature vector and the edge weight matrix of the set of potential conflict pairs into the GNN model, output the visual weight and the text weight, and calculate the cross-modal similarity of the set of potential conflict pairs through the CLIP model; Calculate the difference between the cross-modal similarity and the comprehensive similarity, select the potential conflict pairs greater than the difference threshold to generate the final conflict pairs.

[0009] As a preferred solution of the image data annotation method described in the present invention, wherein: output the region features through the NeRF model, extract them using the CLIP model, and update them through the loss function, including: Based on the final conflict pairs, initialize the NeRF model to generate a high-dimensional feature map, and use CBAM to generate spatial and channel attention weights to output the initial region features; Use the CLIP model to extract the region semantic embedding of the final conflict pairs to obtain a semantic feature vector, map it to the CLIP dimension using a linear layer, and define the loss function of the model parameters; Output the updated region features through the Adam optimizer; Perform weighted fusion on the region semantic embedding and the mapping result to generate a fusion feature; Extract the feature vectors of each initial class label, and obtain the class average features through summation and averaging. Calculate the weighted Euclidean distance between the fused features and the class average features; Sort the weighted Euclidean distances in descending order, select the fused features with the smallest weighted Euclidean distance, update the attention region features, and generate updated attention region features.

[0010] As a preferred solution of the image data annotation method described in the present invention, wherein: generating the verification fusion label distribution matrix through the MLP model and generating the heat map through ECharts includes: Construct an MLP model, input the updated attention region features and the class mean features into the MLP model, map them to class probabilities, calculate the updated label distribution through the Softmax function, and update the fusion label distribution matrix of the final conflict pairs to generate the verification fusion label distribution matrix; Convert the verification fusion label distribution matrix into a metadata table and generate a heat map through ECharts.

[0011] As a preferred solution of the image data annotation method described in the present invention, wherein: preprocessing the two types of collected data includes: Preprocess the text data, including tokenizing through the BPE algorithm, extracting semantic features from the tokenization results through the BERT model, performing dimensionality reduction processing on the semantic features using PCA, and performing normalization processing using the L2 norm to generate text features; Use the VGG-16 model to extract the convolutional features of the image data, extract the fully connected layer features through the FC7 layer to generate the FC7 vector, and encode through the BOW model to generate the BOW vector; Use the Conv5-3 layer of VGG-16 to extract the features of the image data to generate a feature map, and perform normalization processing using the L2 norm to generate multi-scale spatial features; Concatenate the text features, the normalized FC7 vector, the normalized BOW vector, and the multi-scale spatial features, perform dimensionality reduction through the fully connected layer, and perform normalization processing using the L2 norm to generate a multi-modal feature vector; Use the Selective Search algorithm to segment the feature map to generate candidate regions, and extract the features of each candidate region through the ResNet-50 model; Sort the discriminative scores of all candidate regions in descending order, select the region with the highest discriminative score, extract the features of the corresponding candidate region, and concatenate them with the multi-modal feature vector to generate the attention region features.

[0012] As a preferred solution of the image data annotation method described in the present invention, wherein: storing the two types of collected data includes: The two types of collected data and the generated heat map from the analysis are transmitted to the central database for storage via the SFTP protocol and maintained and updated regularly.

[0013] In a second aspect, the present invention provides an image data annotation system, including, A processing optimization module for collecting image and text data, preprocessing the two types of collected data, generating hash codes through VAE and constructing a dynamic edge weight matrix, generating a low-rank approximation matrix by combining singular value decomposition, updating through kernel density estimation and kinetic equations, and optimizing using Euler discretization in combination with CLP propagation and EA; A calculation update module for, based on the optimization result, performing conversion using amplitude encoding, constructing through a quantum state circuit, performing calculations using partial trace, generating topological weighted entanglement entropy through weighted fusion, updating based on the attention mechanism, generating potential conflict pairs, calculating priorities, generating final conflict pairs through a CNN model, outputting regional features through a NeRF model, extracting using a CLIP model, and updating through a loss function; A verification upload module for generating a verification fusion label distribution matrix through an MLP model and generating a heat map through ECharts; An encryption storage module for storing the two types of collected data.

[0014] In a third aspect, the present invention provides a computer device including a memory and a processor, where the memory stores a computer program, and: when the computer program is executed by the processor, any step of the image data annotation method described in the first aspect of the present invention is implemented.

[0015] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, and: when the computer program is executed by the processor, any step of the image data annotation method described in the first aspect of the present invention is implemented.

[0016] The beneficial effects of the present invention are as follows: The present invention collects image and text data, preprocesses the two collected data, generates hash codes through VAE and constructs a dynamic edge weight matrix, combines singular value decomposition to generate a low-rank approximation matrix, updates through kernel density estimation and kinetic equations, uses Euler discretization combined with CLP propagation and EA for optimization, based on the optimization results, performs conversion using amplitude encoding, constructs through a quantum state circuit, performs calculations using partial trace, generates topological weighted entanglement entropy through weighted fusion, updates based on the attention mechanism, generates potential conflict pairs, calculates priorities, generates final conflict pairs through a CNN model, outputs regional features through a NeRF model, extracts using a CLIP model, and updates through a loss function; improving the accuracy and consistency of annotation, and enhancing the efficiency and accuracy of annotation. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for description in the embodiments. Obviously, the following described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0018] Figure 1 It is a flowchart of the image data annotation method in Embodiment 1; Figure 2 It is a schematic diagram of the image data annotation system in Embodiment 1; Figure 3 It is a schematic diagram of conflict pair identification in Embodiment 1; Figure 4 It is a schematic diagram of topological weighted entanglement entropy generation in Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will provide a detailed description of the specific embodiments of the present invention with reference to the drawings in the specification.

[0020] In the following description, many specific details are set forth to fully understand the present invention. However, the present invention can also be implemented in other ways different from the described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0021] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that can be included in at least one implementation of the present invention. The "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor is it an embodiment that is separate or selectively mutually exclusive with other embodiments.

[0022] Example 1, referring to Figures 1 to 4 , which is the first embodiment of the present invention. This embodiment provides an image data annotation method, including the following steps: S1. Collect image and text data, preprocess the two collected data, generate hash codes through VAE and construct a dynamic edge weight matrix, generate a low-rank approximation matrix by combining singular value decomposition, update through kernel density estimation and kinetic equations, and optimize using Euler discretization combined with CLP propagation and EA; Specifically, preprocessing the two collected data includes: Using a camera and social software to collect image data and text data; Preprocess the text data, including tokenizing through the BPE algorithm, extracting semantic features from the tokenization results through the BERT model, reducing the dimensionality of the semantic features using PCA, and normalizing using the L2 norm to generate text features; Use the VGG-16 model to extract convolutional features of the image data and perform global average pooling; For the pooling result, use the FC7 layer of the VGG-16 model to extract fully connected layer features, reduce the dimensionality through PCA to generate an FC7 vector, and encode through the BOW model to generate a BOW vector; Normalize the BOW vector and the FC7 vector respectively using the L2 norm; Use the Conv5-3 layer of VGG-16 to extract features of the image data to generate a feature map, perform pooling through SPP, reduce the dimensionality through a fully connected layer, filter through the ReLU activation function, and normalize using the L2 norm to generate multi-scale spatial features; Concatenate the text features, the normalized FC7 vector, the normalized BOW vector, and the multi-scale spatial features, reduce the dimensionality through a fully connected layer, and normalize using the L2 norm to generate a multi-modal feature vector; Use the Selective Search algorithm to segment the feature map to generate candidate regions, and extract the features of each candidate region through the ResNet-50 model; Calculate the variance between the features of each candidate region and the class average feature (obtained based on the initial class label) through the within-class variance, and normalize through Min-Max normalization to generate a discriminative score; Sort the discriminative scores of all candidate regions in descending order, select the region with the highest discriminative score, extract the features of the corresponding candidate region, and concatenate them with the multi-modal feature vector to generate attention region features.

[0023] Through multi-source data collection, the data diversity and representativeness are enhanced, the bias problems that may be caused by a single data source are avoided, the requirements of dynamic scenarios can be adapted, data support is provided for real-time annotation tasks, thereby enhancing the generality and robustness of the system. Through the combination of BPE and BERT, in-depth mining of text semantics is achieved, and complex semantic associations are retained. CA dimensionality reduction and L2 normalization effectively reduce redundant information and noise interference, improving the computational efficiency and stability of features, laying a high-quality text feature foundation for subsequent multi-modal feature fusion. Through the convolutional feature extraction of VGG-16, accurate representation of complex image textures and semantic structures is realized. The combination of FC7 vectors and BOW vectors achieves complementary expression of global semantics and local pattern features, enhancing the semantic richness of features. Through SPP pooling, adaptive feature extraction of images with different sizes and resolutions is realized, solving the sensitivity problem of traditional fixed-size pooling to image scale changes. The Selective Search algorithm combined with the ResNet-50 model realizes accurate positioning and in-depth feature extraction of key regions of images, improving the semantic relevance and feature quality of regions.

[0024] Furthermore, image and text data are collected and optimized using Euler discretization combined with CLP propagation and EA, including: Based on multi-modal feature vectors and attention region features, calculate the global similarity and local attention similarity respectively through the cosine formula; Use VAE to map multi-modal feature vectors to hash codes and calculate hash similarity through the Gaussian kernel function; Fuse the global similarity, local attention similarity, and hash similarity through weighted fusion to obtain the comprehensive similarity; Define the comprehensive similarity as the edge, with the image ID as the node, and obtain the edge weight based on the hash similarity to construct a dynamic edge weight matrix. The formula is: , where, is the edge weight between node and node at time , is the exponential function, is the scalar parameter of the time decay factor (obtained through experimental optimization), is the dynamic time constant between node and node (obtained through dynamic adjustment based on simulated annealing), is the comprehensive similarity; Extract the similarities of all nodes, sort them in descending order, filter the top N maximum similarities (set based on empirical rules), where N is the number of nodes, and perform a sum average to obtain the local density; Based on the local density, calculate the number of neighbors, with the formula: , where, is the proportionality coefficient (set based on statistical methods of data distribution), is the number of neighbors of node , is the node 's local density; Based on the number of neighbors, construct a sparse adjacency matrix for each node; Based on the sparse adjacency matrix, calculate the degree matrix and Laplacian matrix for each node, with the formula: ,

[0025] where, is an element of the degree matrix, representing the degree of node at time , is the Laplacian matrix at time , is the sparse adjacency matrix at time ; Perform singular value decomposition on the Laplacian matrix to generate a low-rank approximation matrix; Based on the elbow method, set the number of classes K, use K-Means clustering to perform K clusters on the multi-modal feature vectors, and map them to string labels (such as K classes like people, cats... etc.), generating initial class labels; Based on the initial class labels, obtain the probability distribution through maximum likelihood estimation and perform normalization through softmax normalization to generate the initial probability distribution; Based on the multi-modal feature vectors, attention region features (mapped through a fully connected layer), and timestamps (encoded through sine), calculate the LTC time constant through Laplacian eigenvalue decomposition; Map the hash code to the initial probability distribution through kernel density estimation and update the initial probability distribution, with the formula: , where, is the hash prior weight (set based on cross-validation), is the updated initial probability distribution of node , is the initial probability distribution of the hash code of node ; is the number of class labels; Based on each node, define the dynamic equation, the formula is: , where, is the influence neighbor weight (set based on Laplacian regularization), is the attention weight (set based on adversarial training), is the node 's attention guiding vector (sort the Laplacian matrix in descending order, and select the eigenvector corresponding to the minimum value as the attention guiding vector), is the node 's initial probability distribution of dynamic update, is the node 's dynamic time constant; Use the Euler method for discretization, iterate the initial probability distribution of update, and combine with the low-rank approximation matrix for GLP propagation, stop after reaching the maximum number of iterations (determined based on mathematical derivation), and output the initial probability distribution of iteration, the formula is: , where, is the retention weight factor of the initial label (obtained based on the label propagation algorithm), is the propagation initial probability distribution of the node , is the time 's initial probability distribution of the node , is the node and the node 's initial probability distribution of update at time ; For the initial probability distribution of update, generate the hash code of the initial probability distribution of update through VME, and calculate the hash loss as the constraint of the multi-modal feature vector, the formula is: , where, is the hash loss, is the hash code of the node ; Perform weighted fusion on the propagation initial probability distribution and the initial probability distribution of update to generate a fused label distribution matrix; Based on the initial probability distribution of update and the propagation initial probability distribution, calculate the consistency loss of the initial probability distribution of update and the initial label deviation of the propagation initial probability respectively through the mean squared error loss function, and generate the total loss through weighted fusion; Based on the dynamic edge weight matrix, calculate the voting weights of the neighbors of each node and the voting distribution of the node through EA, and update the fused label distribution matrix; Based on the updated fused distribution label matrix, define the EA objective function and optimize the two weight parameters in the dynamic equation. The formula is: , where, is the value of the objective function, is the total loss, is the variance weight (static setting based on empirical ratio), is the label distribution variance (obtained by global variance calculation based on the updated fused label distribution matrix), is the node 's updated fused distribution label matrix; Based on each node, screen the maximum value in the optimized fused distribution label matrix to generate the final label for each node.

[0026] Generating low-dimensional hash codes through VAE reduces the computational complexity of multi-modal features. Introducing a time decay factor into the dynamic edge weight matrix enhances the model's adaptability to dynamic scenarios. The weighted fusion of comprehensive similarities improves semantic consistency, significantly enhancing the annotation accuracy and robustness. The low-rank approximation matrix effectively filters out the noise in high-dimensional data and retains key semantic information. The introduction of the sparse adjacency matrix and Laplacian matrix enhances the modeling ability of the data topological structure, enabling better capture of potential correlations between data in complex scenarios, improving the annotation consistency and accuracy. K-Means clustering combined with kernel density estimation significantly improves the semantic discrimination of initial labels. Kernel density estimation introduces non-parametric probability modeling, enhancing the model's adaptability to uneven data distributions and effectively reducing annotation bias in dynamic scenarios. The efficient discretization of the Euler method reduces the complexity of iterative calculations. GLP propagation utilizes the topological information of the low-rank approximation matrix, significantly enhancing the stability and convergence speed of label propagation, and better balancing local and global semantics in complex network structures, reducing annotation errors. EA optimizes weight parameters through global search, avoiding the risk of traditional gradient optimization methods falling into local optima. The voting mechanism combined with the dynamic edge weight matrix enhances the model's ability to handle potential data conflicts, significantly improving the robustness and adaptability of the annotation results.

[0027] S2. Based on the optimization results, perform conversion using amplitude encoding, construct through a quantum state circuit, perform calculations using partial trace, generate topological weighted entanglement entropy through weighted fusion, update based on the attention mechanism, generate potential conflict pairs, calculate priorities, generate final conflict pairs through a CNN model, output regional features through a NeRF model, extract using a CLIP model, and update through a loss function; Specifically, based on the optimization results, the final conflict pairs are generated through the CNN model, including: The optimized fusion label distribution matrix is converted into a quantum state vector through amplitude encoding. An edge set is defined based on the dynamic edge weight matrix. The quantum state vector is used to construct an entangled state through a quantum state circuit, and through TDA, a topological feature vector (including Betti numbers and persistent diagram features) of each node is generated; Based on the entangled state, the reduced density matrix is calculated through partial trace; The L2 norm is used to normalize the topological feature vector, and the Sigmoid function is used for mapping to generate the weight of the topological feature vector; The entanglement entropy of the reduced density matrix is calculated through Qiskit, and weighted fusion is performed in combination with the weight of the topological feature vector to generate the topological weighted entanglement entropy; The final label is adjusted using temperature scaling, the adjustment result is normalized through the Sigmoid function, the adjustment factor is calculated, and the topological weighted entanglement entropy is multiplied to generate the adjusted topological weighted entanglement entropy; Based on the multi-modal feature vector, the modal weight is obtained through the attention mechanism, and the adjusted topological weighted entanglement entropy is multiplied to generate the latest topological weighted entanglement entropy; Based on the latest topological weighted entanglement entropy, the mean value and standard deviation value are calculated through the arithmetic mean and sample standard deviation respectively, and a dynamic adjustment threshold is generated through the adaptive statistical method; Nodes with the latest topological entanglement entropy greater than the dynamic adjustment threshold are screened and defined as agents (including the final label, the latest topological weighted entanglement entropy, the topological feature vector, and the fusion distribution matrix); Based on the final label, the actions of the agent (adjusting the label to the neighbor node label set) are defined, and a reward function is set; The Q-table is initialized, the Q value is updated through the Bellman equation, and the update stops when the maximum number of updates is reached (set based on the early stopping method), and the final Q value of each node is output to generate the agent Q-table; If the final label of each node is not equal to the neighbor label and the edge weight between the two is greater than the edge weight threshold (set based on the fixed threshold method), then this node and the neighbor node are marked as potential conflict pairs, and the priority is calculated. The formula is: , where, is the latest topological weighted entanglement entropy of node , is the Q value of node , is the final label of node , is the potential conflict pair (node And node ); Sort the priorities in descending order, filter the top E potential conflict pairs (Top-K filtering based on priority), where E is the number of potential conflict pairs, and generate a set of potential conflict pairs; Construct a GNN model, including an input layer, a convolutional layer, a fully connected layer, and an output layer; Use the Cora dataset to train the GNN model; Input the topological feature vectors and edge weight matrices of the set of potential conflict pairs into the GNN model, output the visual weight and text weight, and calculate the cross-modal similarity of the set of potential conflict pairs through the CLIP model; Calculate the difference between the cross-modal similarity and the comprehensive similarity, filter the potential conflict pairs greater than the difference threshold, and generate the final conflict pairs.

[0028] Through amplitude encoding and quantum state circuits, an efficient conversion from classical data to quantum states is achieved, significantly reducing the data dimension while retaining key topological features. The computational efficiency is improved through the parallelism of quantum computing. At the same time, the intrinsic geometric structure of the data is captured by TDA, enhancing the ability to recognize complex data patterns in the annotation process. The quantum state information of the subsystem is accurately extracted by calculating the reduced density matrix through partial trace. Based on the principles of quantum mechanics, the modeling accuracy of the complex relationships between multi-modal data is significantly improved, providing reliable data support for subsequent entanglement entropy generation. Through the combination of the L2 norm and the Sigmoid function, the standardization and non-linear weighting of the feature vectors are realized, enhancing the robustness of the model to sparse or high-dimensional data. Through the fusion of quantum entanglement entropy and topological features, an accurate modeling of the intrinsic correlation of high-dimensional data is achieved, significantly enhancing the representational ability of feature correlation. The introduction of temperature scaling and adjustment factors further optimizes the dynamic range of the entropy value, enhancing the adaptability of the model to dynamic datasets. Through the attention mechanism, the adaptive weighting of modal features is realized, significantly improving the ability of the model to capture the semantic correlation between image and text modalities. Through Q-learning and priority ranking, the intelligent screening of conflict pairs is realized, improving the accuracy and efficiency of conflict recognition. The introduction of the priority formula further optimizes resource allocation, ensuring the priority processing of high-impact conflict pairs. Through the synergistic effect of the GNN and CLIP models, the refined screening of conflict pairs is realized, significantly improving the semantic consistency of the annotation results.

[0029] Furthermore, output the regional features through the NeRF model, extract them using the CLIP model, and update them through the loss function, including: Based on the final conflict pairs, initialize the NeRF model, render the regional features of the image data to generate a high-dimensional feature map, and use CBAM to generate spatial and channel attention weights, and output the initial regional features; Use the CLIP model to extract the regional semantic embeddings of the final conflict pairs, obtain the semantic feature vectors, and map them to the CLIP dimension using a linear layer. Define the loss function of the model parameters. The formula is as follows: , where, is the loss function value of the model parameters, is the expectation operation on the parameter (NeRF model) distribution , is the regional feature generating the image of the final conflict node the logarithm of the probability, is the parameter posterior distribution (approximated by variational inference), is the parameter prior distribution (obtained through the standard normal distribution), is the divergence, is the hyperparameter of the semantic embedding (set based on the heuristic ratio method), is the region semantic feature vector, is the hyperparameter of the label loss (dynamically adjusted based on the number of training epochs), is the node final label fusion label distribution matrix; Minimize the loss function value of the model parameters through the Adam optimizer, stop when the preset number of times is reached (set based on the learning rate scheduling method), update the parameters of the NeBF model, and output the updated regional features; Perform weighted fusion on the regional semantic embeddings and mapping results to generate fused features; Extract the feature vectors of each initial class label, obtain the class average features through summation and averaging, and calculate the weighted Euclidean distance between the fused features and the class average features. The formula is as follows: , where, is the weighted Euclidean distance of the region , is the region fused feature vector (the regional feature and the semantic feature vector are obtained through weighted fusion), is the Sigmoid activation function, is the node topological feature vector, is the label class; Sort the weighted Euclidean distances in descending order, select the fusion features with the smallest weighted Euclidean distance, update the attention region features, and generate updated attention region features.

[0030] Through the region feature rendering and attention mechanism of the NeRF model combined with CBAM, the adaptive enhancement of the high-dimensional feature map is realized, which significantly improves the pertinence and robustness of feature extraction, can retain more geometric information, reduce feature loss, provides a high-quality feature basis for subsequent semantic embedding and label generation, improves the accuracy and consistency of image data annotation. Through the semantic embedding of the CLIP model and the optimization of the comprehensive loss function, the deep alignment of region features and semantic information is realized, the semantic consistency of the annotation task is improved, the deep correlation between images and texts can be effectively captured, and the annotation deviation is reduced. The Adam optimizer combined with the learning rate scheduling method further ensures the stability of model convergence, thus providing reliable support for real-time annotation in dynamic scenarios. By weighted fusion of region semantic embedding and mapping results, the effective integration of visual features and semantic features is realized, the information integrity and expression ability of feature vectors are significantly enhanced. Compared with single-modal features, the fusion features can better capture semantic associations in complex scenes, reduce information redundancy, and provide a more reliable feature basis for subsequent distance calculation and feature update. Through the calculation and selection of weighted Euclidean distance, the accurate update of attention region features is realized, and the focusing ability of the model on key regions is significantly improved.

[0031] S3. Generate a verification fusion label distribution matrix through the MLP model and generate a heat map through ECharts; Specifically, generating a verification fusion label distribution matrix through the MLP model and generating a heat map through ECharts includes: Construct an MLP model, including an input layer, a hidden layer, and an output layer; Use the COCO multi-modal dataset to train the MLP model; Input the updated attention region features and class mean features into the MLP model, map them to class probabilities, calculate the updated label distribution through the Softmax function, update the fusion label distribution matrix of the final conflict pairs, and generate a verification fusion label distribution matrix; Convert the verification fusion label distribution matrix into a metadata table and generate a heat map through ECharts.

[0032] By constructing an MLP model, the processing ability of high-dimensional multimodal features is significantly improved, thus enhancing the accuracy and robustness of label distribution. Training the MLP model using the MNIST dataset not only improves the model's adaptability to image data but also reduces the training cost on specific task data. By combining the region-of-interest features and class mean features and mapping them to the probability space, the conflict problem in multimodal data fusion is alleviated, significantly enhancing the semantic consistency and accuracy of label distribution and better adapting to dynamic data changes in complex scenarios. The generated fused label distribution matrix has higher credibility and application value. By generating a heatmap using ECharts, the interpretability and user-friendliness of the annotation results are significantly improved.

[0033] S4. Store the two types of collected data; Specifically, storing the two types of collected data and the heatmap includes: Transfer the two types of collected data and the generated heatmap through the SFTP protocol to the central database for storage, and perform maintenance and updates regularly.

[0034] By transferring the two types of collected data and the generated heatmap through the SFTP protocol to the central database for storage, the data security risk is significantly reduced, the credibility and compliance of the system are enhanced. By establishing a secure communication channel, the reliability and integrity of data transmission are improved, ensuring the secure migration of data from the local to the central database and reducing the risk of data loss or damage caused by network vulnerabilities. By performing maintenance and updates regularly, the stability and scalability of the data storage system are enhanced. At the same time, through regular security updates, the security risk caused by system aging or vulnerabilities is further reduced.

[0035] This embodiment also provides an image data annotation system, including: A processing optimization module for collecting image and text data, preprocessing the two types of collected data, generating hash codes through VAE and constructing a dynamic edge weight matrix, generating a low-rank approximation matrix by combining singular value decomposition, updating through kernel density estimation and kinetic equations, and optimizing using Euler discretization combined with CLP propagation and EA; A calculation and update module for performing transformation using amplitude encoding based on the optimization result, constructing through a quantum state circuit, performing calculations using partial trace, generating topological weighted entanglement entropy through weighted fusion, updating based on the attention mechanism, generating potential conflict pairs, calculating priorities, generating final conflict pairs through a CNN model, outputting region features through a NeRF model, extracting using a CLIP model, and updating through a loss function; A verification and upload module for generating a verification fused label distribution matrix through an MLP model and generating a heatmap through ECharts; An encryption storage module for storing the two types of collected data.

[0036] This embodiment also provides a computer device applicable to the case of an image data annotation method, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the image data annotation method proposed in the above embodiment.

[0037] The computer device may be a terminal. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the outer shell of the computer device, or an external keyboard, a touchpad, or a mouse, etc.

[0038] This embodiment also provides a storage medium with a computer program stored thereon. When the program is executed by a processor, it implements the image data annotation method proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (abbreviated as SRAM), an electrically erasable programmable read-only memory (abbreviated as EEPROM), an erasable programmable read-only memory (abbreviated as EPROM), a programmable read-only memory (abbreviated as PROM), a read-only memory (abbreviated as ROM), a magnetic memory, a flash memory, a magnetic disk, or an optical disc.

[0039] In summary, the present invention collects image and text data, preprocesses the two collected data, generates hash codes through VAE and constructs a dynamic edge weight matrix, combines singular value decomposition to generate a low-rank approximation matrix, updates through kernel density estimation and kinetic equations, uses Euler discretization combined with CLP propagation and EA for optimization, based on the optimization results, performs conversion using amplitude encoding, constructs through a quantum state circuit, performs calculations using partial trace, generates topological weighted entanglement entropy through weighted fusion, updates based on the attention mechanism, generates potential conflict pairs, calculates priorities, generates final conflict pairs through a CNN model, outputs regional features through a NeRF model, extracts using a CLIP model, and updates through a loss function; improving the accuracy and consistency of annotation and enhancing the efficiency and accuracy of annotation.

[0040] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.

Claims

1. An image data annotation method, characterized in that: Including, Collecting image and text data, preprocessing the two collected data, generating hash codes through VAE and constructing a dynamic edge weight matrix, generating a low-rank approximation matrix by combining singular value decomposition, updating through kernel density estimation and kinetic equations, and optimizing using Euler discretization combined with CLP propagation and EA; Based on the optimization results, performing conversion using amplitude encoding, constructing through a quantum state circuit, performing calculations using partial trace, generating topological weighted entanglement entropy through weighted fusion, updating based on the attention mechanism, generating potential conflict pairs, calculating priorities, generating final conflict pairs through a CNN model, outputting regional features through a NeRF model, extracting using a CLIP model, and updating through a loss function; Generating a verification fusion label distribution matrix through an MLP model and generating a heat map through ECharts; Storing the two collected data.

2. The image data annotation method according to claim 1, characterized in that: The collecting of image and text data and the optimization using Euler discretization combined with CLP propagation and EA include: Based on the multimodal feature vector and the attention region feature, calculating the global similarity and the local attention similarity respectively through the cosine formula; Using VAE to map the multimodal feature vector to a hash code and calculating the hash similarity through a Gaussian kernel function; Fusing the global similarity, the local attention similarity, and the hash similarity through weighting to obtain a comprehensive similarity, defining the comprehensive similarity as an edge, using the image ID as a node, and constructing a dynamic edge weight matrix; Setting the number of categories K based on the elbow method, performing K-Means clustering on the multimodal feature vector into K clusters, mapping to string labels, generating initial category labels, and generating an initial probability distribution through maximum likelihood estimation; Mapping the hash code to the initial probability distribution through kernel density estimation and updating the initial probability distribution; Based on each node, defining a kinetic equation, discretizing using the Euler method, iterating on the updated initial probability distribution, and performing GLP propagation using a low-rank approximation matrix; Performing weighted fusion on the propagated initial probability distribution and the updated initial probability distribution to generate a fusion label distribution matrix; Based on the dynamic edge weight matrix, updating the fusion label distribution matrix through EA; Based on the updated fusion distribution label matrix, defining an EA objective function and optimizing the two weight parameters in the kinetic equation; Based on each node, screening the maximum value in the optimized fusion distribution label matrix to generate the final label for each node.

3. The image data annotation method according to claim 2, wherein: The generating of the final conflict pairs through a CNN model based on the optimization results includes: Converting the optimized fusion label distribution matrix into a quantum state vector through amplitude encoding, defining an edge set based on the dynamic edge weight matrix, constructing an entangled state from the quantum state vector through a quantum state circuit, and generating a topological feature vector for each node through TDA; Calculating the reduced density matrix through partial trace based on the entangled state; Normalizing the topological feature vector using the L2 norm and mapping using the Sigmoid function to generate the weight of the topological feature vector; Calculate the entanglement entropy of the reduced density matrix through Qiskit, perform weighted fusion by combining the weights of the topological eigenvectors, generate the topological weighted entanglement entropy, and make adjustments to generate the adjusted topological weighted entanglement entropy; Based on the multi-modal eigenvector, obtain the modal weights through the attention mechanism, perform multiplication on the adjusted topological weighted entanglement entropy to generate the latest topological weighted entanglement entropy; If the final label of each node is not equal to the neighbor label and the edge weight between the two is greater than the edge weight threshold, then mark this node and the neighbor node as a potential conflict pair and calculate the priority; Sort the priorities in descending order, select the top E potential conflict pairs, where E is the number of potential conflict pairs, to generate a set of potential conflict pairs; Construct a GNN model, input the topological eigenvector and edge weight matrix of the set of potential conflict pairs into the GNN model, output the visual weight and text weight, and calculate the cross-modal similarity of the set of potential conflict pairs through the CLIP model; Calculate the difference between the cross-modal similarity and the comprehensive similarity, select the potential conflict pairs greater than the difference threshold to generate the final conflict pairs.

4. The image data annotation method according to claim 3, wherein: The output of the regional features through the NeRF model, extraction using the CLIP model, and update through the loss function include: Based on the final conflict pairs, initialize the NeRF model to generate a high-dimensional feature map, and use CBAM to generate spatial and channel attention weights to output the initial regional features; Use the CLIP model to extract the regional semantic embedding of the final conflict pairs to obtain the semantic feature vector, and use a linear layer to map it to the CLIP dimension, and define the loss function of the model parameters; Output the updated regional features through the Adam optimizer; Perform weighted fusion on the regional semantic embedding and the mapping result to generate a fused feature; Extract the feature vectors of each initial class label, obtain the class average feature through summation and averaging, and calculate the weighted Euclidean distance between the fused feature and the class average feature; Sort the weighted Euclidean distances in descending order, select the fused feature with the smallest weighted Euclidean distance, and update the attention regional features to generate the updated attention regional features.

5. The image data annotation method according to claim 4, characterized in that: The generation of the verification fusion label distribution matrix through the MLP model and the generation of the heat map through ECharts include: Construct an MLP model, input the updated attention regional features and class mean features into the MLP model, map them to class probabilities, calculate the updated label distribution through the Softmax function, and update the fusion label distribution matrix of the final conflict pairs to generate the verification fusion label distribution matrix; Convert the verification fusion label distribution matrix into a metadata table and generate a heat map through ECharts.

6. The image data annotation method according to claim 1, characterized in that: The preprocessing of the two types of data collected includes: Preprocess the text data, including tokenization through the BPE algorithm, extraction of semantic features from the tokenization results through the BERT model, dimensionality reduction processing of the semantic features using PCA, and normalization processing using the L2 norm to generate text features; Use the VGG-16 model to extract the convolutional features of the image data, extract the fully connected layer features through the FC7 layer to generate the FC7 vector, and encode it through the BOW model to generate the BOW vector; Extract the features of the image data using the Conv5-3 layer of VGG-16 to generate a feature map, perform normalization processing using the L2 norm, and generate multi-scale spatial features; Concatenate the text features, normalized FC7 vector, normalized BOW vector, and multi-scale spatial features, perform dimensionality reduction through a fully connected layer, and perform normalization processing using the L2 norm to generate a multi-modal feature vector; Use the Selective Search algorithm to segment the feature map to generate candidate regions, and extract the features of each candidate region through the ResNet-50 model; Sort the discriminative scores of all candidate regions in descending order, select the region with the highest discriminative score, extract the features of the corresponding candidate region, and concatenate them with the multi-modal feature vector to generate attention region features.

7. The image data annotation method according to claim 5, wherein: The storage of the two types of collected data includes: Transfer the two types of collected data and the generated heatmap through the SFTP protocol to the central database for storage, and perform maintenance and updates regularly.

8. An image data annotation system, based on the image data annotation method according to any one of claims 1 to 7, characterized in that: Including, A processing optimization module for collecting image and text data, preprocessing the two types of collected data, generating hash codes through VAE and constructing a dynamic edge weight matrix, generating a low-rank approximation matrix by combining singular value decomposition, updating through kernel density estimation and kinetic equations, and optimizing using Euler discretization combined with CLP propagation and EA; A calculation update module for performing conversion using amplitude encoding based on the optimization results, constructing through a quantum state circuit, performing calculations using partial trace, generating topological weighted entanglement entropy through weighted fusion, updating based on the attention mechanism, generating potential conflict pairs, calculating priorities, generating final conflict pairs through a CNN model, outputting region features through a NeRF model, extracting using a CLIP model, and updating through a loss function; A verification upload module for generating a verification fusion label distribution matrix through an MLP model and generating a heatmap through ECharts; An encryption storage module for storing the two types of collected data.

9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that: When the processor executes the computer program, it implements the steps of an image data annotation method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of an image data annotation method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Image-text retrieval method and system based on cross-modal semantic analysis

    CN118132677A

  • Unsupervised cross-modal hash retrieval method based on CLIP and attention fusion mechanism

    CN118861327A

  • Comparative learning unsupervised cross-modal hash retrieval algorithm based on graph attention mechanism

    CN119377462A

  • Interactive behavior understanding method for posture reconstruction based on features of skeleton and image

    US20250022165A1

Cited By

  • Material dynamic coding method, system and equipment

    CN120597838A

  • Dynamic management method and system for social group members

    CN120821878A

  • Multi-modal power data retrieval method and device and medium

    CN120929662A

  • A multimodal power data retrieval method, device, and medium

    CN120929662B

  • Automatic processing method and system for graphic annotation

    CN121236344A