An Automatic Image Data Annotation Method and System Based on Deep Learning

By adopting a combination of deep learning and reinforcement learning in image data annotation, the labeling strategy is dynamically adjusted, and the problems of low labeling efficiency, low accuracy and low intelligence in the existing technology are solved, and efficient, accurate and intelligent automatic labeling of image data is achieved.

CN119741706BActive Publication Date: 2025-07-01YUNHAI SPACETIME (BEIJING) TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510251726.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-07-01
Estimated Expiration
2045-03-05

AI Technical Summary

Technical Problem

In the prior art, image data labeling is low efficiency, low accuracy and low intelligence, making it difficult to meet the needs of fast labeling and dynamic adjustment of labeling strategies.

Method used

The automatic labeling method of image data based on deep learning is adopted, and the automatic labeling strategy adjustment model is constructed in combination with reinforcement learning algorithms. The labeling strategy is dynamically adjusted by real-time performance data to improve labeling efficiency and accuracy.

Benefits of technology

It significantly improves the efficiency and accuracy of image data labeling, reduces artificial errors and subjective inconsistencies, improves the degree of intelligence, and is suitable for large-scale image data sets and fast response application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119741706B_ABST
    Figure CN119741706B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of computer vision, and discloses an automatic image data annotation method and system based on deep learning. The method includes the following steps: using a deep learning algorithm to construct an automatic image data annotation model, and using a reinforcement learning algorithm to construct an automatic annotation strategy adjustment model; collecting real-time performance data of the automatic image data annotation model, and inputting the real-time performance data into the automatic annotation strategy adjustment model; according to the real-time performance data, using the automatic annotation strategy adjustment model to adjust the automatic annotation strategy of the automatic annotation strategy adjustment model, so as to obtain an adjusted automatic image data annotation model; collecting real-time image data, and using the adjusted automatic image data annotation model to automatically annotate the real-time image data, so as to obtain automatically annotated real-time image data. The present invention solves the problems of low annotation efficiency, low annotation accuracy and low intelligence level existing in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and particularly relates to an automatic image data annotation method and system based on deep learning. Background Art

[0002] Computer vision is an important branch of artificial intelligence, which aims to enable computers to "see" and "understand" the content in images and videos. With the development of computer vision technology, how to efficiently and automatically annotate image data has become an important development direction for computer vision applications. The need for image data annotation has also promoted the further innovation and application of computer vision technology. With the continuous progress of technology, image annotation technology will be more efficient, accurate, and better meet the needs of different fields and scenarios.

[0003] The existing image data annotation technologies have the following defects:

[0004] 1) Low annotation efficiency: The existing image annotation methods often rely on manual operation, which is not only time-consuming and laborious, but also inefficient on large-scale image datasets and difficult to meet the need for rapid annotation;

[0005] 2) Low annotation accuracy: Manual annotation is easily affected by subjective factors, such as the experience and fatigue degree of annotators, resulting in errors and inconsistencies in the annotation results;

[0006] 3) Low intelligence level: The existing image annotation technologies lack the ability to dynamically adjust annotation strategies. Once encountering a situation where the annotation performance deteriorates, they are unable to timely adjust the annotation strategies to adapt to the new data distribution. Summary of the Invention

[0007] In order to solve the problems of low annotation efficiency, low annotation accuracy, and low intelligence level existing in the prior art, the purpose of the present invention is to provide an automatic image data annotation method and system based on deep learning.

[0008] The technical solution adopted by the present invention is as follows:

[0009] An automatic image data annotation method based on deep learning, comprising the following steps:

[0010] Using a deep learning algorithm to construct an automatic image data annotation model, and using a reinforcement learning algorithm to construct an automatic annotation strategy adjustment model;

[0011] According to a preset model adjustment mechanism, collecting real-time performance data of the automatic image data annotation model, and inputting the real-time performance data into the automatic annotation strategy adjustment model;

[0012] Adjust the model using an automatic annotation strategy according to real-time performance data, and adjust the automatic annotation strategy of the automatic annotation strategy adjustment model to obtain an adjusted automatic annotation model for image data;

[0013] Collect real-time image data, and use the adjusted automatic annotation model for image data to automatically annotate the real-time image data to obtain automatically annotated real-time image data.

[0014] Furthermore, use deep learning algorithms to construct an automatic annotation model for image data, and use reinforcement learning algorithms to construct an automatic annotation strategy adjustment model, including the following steps:

[0015] Collect a number of historical image data, and preprocess the number of historical image data to obtain a number of preprocessed historical image data;

[0016] According to a number of preprocessed historical image data, use deep learning algorithms to construct an automatic annotation model for image data;

[0017] Set a model adjustment mechanism, and according to the model adjustment mechanism, collect a number of historical performance data of the automatic annotation model for image data, and preprocess it to obtain a number of preprocessed historical performance data;

[0018] According to a number of preprocessed historical performance data, use reinforcement learning algorithms to construct an automatic annotation strategy adjustment model.

[0019] Furthermore, the automatic annotation model for image data is constructed based on the FPN-Faster R-CNN-DANN-DBN algorithm, and the automatic annotation model for image data includes a graph feature extraction module constructed based on the FPN algorithm, a target region localization module constructed based on the Faster R-CNN algorithm, a domain adversarial training module constructed based on the DANN algorithm, and an automatic annotation module constructed based on the DBN algorithm. The graph feature extraction module, the target region localization module, and the automatic annotation module are connected in sequence. The domain adversarial training module is connected to the target region localization module, and the domain adversarial training module includes a label predictor and a domain classifier connected in sequence.

[0020] Furthermore, according to a number of preprocessed historical image data, use deep learning algorithms to construct an automatic annotation model for image data, including the following steps:

[0021] Use the CNN algorithm to construct the basic network architecture of the graph feature extraction module; the basic network architecture includes alternately connected convolutional layers and pooling layers;

[0022] Use the FPN algorithm to construct a feature pyramid, and horizontally connect the outputs of all convolutional layers to the feature pyramid in a top-down order to obtain an initial graph feature extraction module;

[0023] Use the Faster R-CNN algorithm to construct an initial target region localization module, and use the DANN algorithm to construct an initial domain adversarial training module;

[0024] Use the DBN algorithm to construct an initial automatic annotation module;

[0025] Integrate the initial graph feature extraction module, the initial target region localization module, the initial domain adversarial training module, and the initial automatic annotation module to obtain an initial automatic annotation model for image data;

[0026] Combine the loss functions of the initial graph feature extraction module, the initial target region localization module, the initial domain adversarial training module, and the initial automatic annotation module to obtain a comprehensive loss function;

[0027] Optimize the initial model parameters of the initial automatic annotation model for image data to obtain an optimized automatic annotation model for image data;

[0028] Based on the comprehensive loss function, use a number of preprocessed historical image data to optimize and train the optimized automatic annotation model for image data to obtain a final automatic annotation model for image data.

[0029] Furthermore, with the goal of minimizing the model error, use the swarm intelligence optimization algorithm to optimize the initial model parameters of the initial automatic annotation model for image data to obtain an optimized automatic annotation model for image data.

[0030] Furthermore, the automatic annotation strategy adjustment model is constructed based on the MPO-PPO algorithm, and the automatic annotation strategy adjustment model includes a meta-strategy optimization module constructed based on the MPO algorithm and a reinforcement learning module constructed based on the PPO algorithm. The reinforcement learning module is provided with an agent, a policy network, and an experience replay pool. The agent is respectively connected to the policy network, the experience replay pool, and the meta-strategy optimization module, and the meta-strategy optimization module is connected to the experience replay pool.

[0031] Furthermore, according to a number of preprocessed historical performance data, use the reinforcement learning algorithm to construct an automatic annotation strategy adjustment model, including the following steps:

[0032] Use the FCM clustering algorithm to cluster a number of preprocessed historical performance data to obtain a number of cluster centers and corresponding cluster clusters, and set historical strategy adjustment categories for each cluster cluster;

[0033] Use the experience replay mechanism to initialize the experience replay pool, construct the agent of the reinforcement learning module, and define the state space, action space, and reward function of the agent;

[0034] Taking the historical automatic annotation strategy adjustment problem as a simulation environment, using the PPO algorithm, constructing a policy network, and combining an experience replay pool and an agent to obtain an initial reinforcement learning module;

[0035] Setting several policy adjustment categories as several sub-scenarios for meta-policy optimization, and using the MPO algorithm to construct an initial meta-policy optimization module;

[0036] According to several preprocessed historical performance data, training and optimizing the initial reinforcement learning module, extracting the historical policy network parameters of the policy network of the reinforcement learning module during training and optimization to obtain a final reinforcement learning module, and generating several historical reinforcement learning experiences;

[0037] Based on several sub-scenarios, according to several historical policy network parameters, training and optimizing the initial meta-policy optimization module to obtain a final meta-policy optimization module, and generating several historical meta-policy optimization experiences;

[0038] Storing several historical meta-policy optimization experiences and several historical reinforcement learning experiences into the experience replay pool of the final reinforcement learning module, and integrating the final meta-policy optimization module and the final reinforcement learning module to obtain an automatic annotation strategy adjustment model.

[0039] Furthermore, according to real-time performance data, using the automatic annotation strategy adjustment model to adjust the automatic annotation strategy of the automatic annotation strategy adjustment model to obtain an adjusted image data automatic annotation model, including the following steps:

[0040] Obtaining the similarity between the real-time performance data and several clustering centers, and taking the historical policy adjustment category of the clustering center with the highest similarity as the real-time policy adjustment category of the real-time performance data;

[0041] According to the real-time policy adjustment category, extracting the corresponding real-time meta-policy optimization experience from several historical meta-policy optimization experiences in the experience replay pool of the automatic annotation strategy adjustment model;

[0042] According to the real-time meta-policy optimization experience, using the meta-policy optimization module to initialize the real-time policy network parameters of the policy network of the reinforcement learning module to obtain an updated policy network;

[0043] Randomly extracting several real-time reinforcement learning experiences from several historical reinforcement learning experiences in the experience replay pool, and updating the action space of the reinforcement learning module according to the several real-time reinforcement learning experiences to obtain an updated action space;

[0044] According to the real-time performance data, updating the state space of the reinforcement learning module to obtain an updated state space;

[0045] Using an agent, control the updated policy network to generate the probability distribution of all possible actions in the action space corresponding to each state in the state space;

[0046] Take the possible action with the highest probability distribution in the action space as the execution action of the state, and integrate the execution actions of all states in the state space to obtain several automatic annotation strategy adjustment actions;

[0047] According to several automatic annotation strategy adjustment actions, adjust the automatic annotation strategy of the automatic annotation strategy adjustment model to obtain an adjusted image data automatic annotation model.

[0048] Furthermore, collect real-time image data, and use the adjusted image data automatic annotation model to automatically annotate the real-time image data to obtain automatically annotated real-time image data, including the following steps:

[0049] Collect real-time image data and preprocess the real-time image data to obtain preprocessed real-time image data;

[0050] Use the graph feature extraction module of the adjusted image data automatic annotation model to extract several real-time feature graphs of the preprocessed real-time image data and perform feature graph fusion to obtain real-time fusion features;

[0051] Use the target region localization module of the adjusted image data automatic annotation model to perform target region localization according to the real-time fusion features to obtain several real-time target regions;

[0052] Use the automatic annotation module of the adjusted image data automatic annotation model to automatically annotate several real-time target regions according to the real-time fusion features to obtain automatically annotated real-time image data.

[0053] An image data automatic annotation system based on deep learning is used to implement the image data automatic annotation method. The system includes a model construction unit, a performance data collection unit, a model adjustment unit, and an automatic annotation unit that are connected in sequence.

[0054] The beneficial effects of the present invention are:

[0055] An image data automatic annotation method and system based on deep learning provided by the present invention can automatically process image data, significantly reducing manual participation, thereby greatly improving the annotation efficiency, especially when dealing with large-scale image data sets; the automatic annotation model of image data constructed using deep learning algorithms can more accurately identify and annotate the targets in the images, reducing human errors and subjective inconsistencies, improving the annotation accuracy, and increasing the processing speed of image data, which is applicable to application scenarios that require quick responses; the automatic annotation strategy adjustment model constructed using reinforcement learning algorithms improves the degree of intelligence, dynamically adjusts the annotation strategy according to real-time performance data, making the model update and iteration simpler and more efficient. By combining reinforcement learning and meta-strategy optimization, the automatic annotation model can better generalize to new data sets, improving the performance of the automatic annotation model in different scenarios. The domain adversarial training module realizes the transfer of cross-domain knowledge, improving the performance of the automatic annotation model in cross-domain image annotation tasks.

[0056] Other beneficial effects of the present invention will be further described in the specific implementation manners. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 is a flowchart of the image data automatic annotation method based on deep learning in the present invention.

[0058] Figure 2 is a structural block diagram of the image data automatic annotation system based on deep learning in the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0059] The present invention will be further explained below with reference to the accompanying drawings and specific embodiments.

[0060] Embodiment 1:

[0061] As Figure 1 shown, this embodiment provides an image data automatic annotation method based on deep learning, including the following steps:

[0062] S1: Using deep learning algorithms, construct an automatic annotation model for image data, and using reinforcement learning algorithms, construct an automatic annotation strategy adjustment model, including the following steps:

[0063] S1-1: Collect a number of historical image data, and preprocess the number of historical image data to obtain a number of preprocessed historical image data;

[0064] The preprocessing includes data cleaning, Gaussian denoising, image enhancement, and normalization processing of the historical image data, improving the quality of the data and providing data support for subsequent model construction;

[0065] S1-2: Using deep learning algorithms, construct an automatic image data annotation model based on a number of preprocessed historical image data;

[0066] The automatic image data annotation model is constructed based on the Feature Pyramid Networks (FPN)-Faster R-CNN-Domain-Adversarial Neural Network (DANN)-Deep Belief Network (DBN) algorithms. The automatic image data annotation model includes a graph feature extraction module constructed based on the FPN algorithm, an object region localization module constructed based on the Faster R-CNN algorithm, a domain adversarial training module constructed based on the DANN algorithm, and an automatic annotation module constructed based on the DBN algorithm. The graph feature extraction module, the object region localization module, and the automatic annotation module are connected in sequence. The domain adversarial training module is connected to the object region localization module, and the domain adversarial training module includes a label predictor and a domain classifier connected in sequence;

[0067] The graph feature extraction module includes a basic network architecture constructed based on the Convolutional Neural Networks (CNN) algorithm. Feature maps are extracted on feature maps of different scales, and through the feature pyramid network structure, skip connections are used to fuse feature maps of different levels extracted by the CNN. Through upsampling and lateral connections, a feature pyramid is constructed to transfer high-level semantic information to the low level and enhance the semantic expression ability of low-level features. In this way, low-level detailed features and high-level semantic features can be effectively combined. The object region localization module uses the Faster R-CNN object detection algorithm, including a region proposal network and a classifier, for locating object regions in the image. The domain adversarial training module is used to distinguish features of the source domain and the target domain, and tries to distinguish which domain these features come from. Through this adversarial training, the object region localization module gradually learns to ignore domain-related information and only retain task-related information, improving the prediction efficiency. The gradient reversal layer is located in the label predictor, and its function is to reverse the gradient during backpropagation, so that the domain classifier can learn to distinguish features of the source domain and the target domain during training, while the feature extraction part of the object region localization module is trained to extract and learn domain-invariant features, thus realizing adversarial training. The domain classifier is used to predict whether the feature belongs to the source domain or the target domain according to the reversed features output by the gradient reversal layer. The automatic annotation module generates annotation labels based on the fused features;

[0068] Using deep learning algorithms, construct an automatic image data annotation model based on a number of preprocessed historical image data, including the following steps:

[0069] S1-2-1: Use the CNN algorithm to construct the basic network architecture of the graph feature extraction module; the basic network architecture includes alternately connected convolutional layers and pooling layers; through the alternate connection of convolutional layers and pooling layers, multi-scale feature maps can be effectively extracted from images, local features of images can be captured, and the spatial dimension of features can be reduced by the pooling layer while maintaining important feature information;

[0070] S1-2-2: Use the FPN algorithm to construct a feature pyramid, and horizontally connect the outputs of all convolutional layers to the feature pyramid in a top-down order to obtain the initial graph feature extraction module; the feature pyramid helps to improve the accuracy of object detection, especially in object detection at different scales;

[0071] S1-2-3: Use the Faster R-CNN algorithm to construct the initial object region localization module, and use the DANN algorithm to construct the initial domain adversarial training module; the Faster R-CNN algorithm improves the efficiency and accuracy of object detection and realizes object region localization; the DANN algorithm enables the model to share feature representations between the source domain and the target domain through adversarial training. The label predictor and the domain classifier are respectively used to predict the target category and the domain label, which helps the model to transfer learning between different data distributions and improve the generalization ability of the model on unseen data;

[0072] S1-2-4: Use the DBN algorithm to construct the initial automatic annotation module; pre-train the network in an unsupervised manner and then perform supervised fine-tuning, which helps to improve the accuracy of annotation;

[0073] S1-2-5: Integrate the initial graph feature extraction module, the initial object region localization module, the initial domain adversarial training module, and the initial automatic annotation module to obtain the initial image data automatic annotation model;

[0074] S1-2-6: Combine the loss functions of the initial graph feature extraction module, the initial object region localization module, the initial domain adversarial training module, and the initial automatic annotation module to obtain the comprehensive loss function; the comprehensive loss function can balance the importance of different modules and optimize the performance of the entire model;

[0075] S1-2-7: With the goal of minimizing the model error, use the Improved Sparrow Search Algorithm (ISSA) to optimize the initial model parameters of the initial image data automatic annotation model to obtain the optimized image data automatic annotation model, including the following steps:

[0076] S1-2-7-1: Minimize the model error as the optimization goal of the ISSA algorithm, and encode the initial model parameters of the initial image data automatic annotation model as the individual vector of the ISSA algorithm;

[0077] S1-2-7-2: Set the algorithm parameters of the ISSA algorithm, and set the fitness function of the ISSA algorithm according to the optimization goal;

[0078] The algorithm parameters include ISSA population parameters, the maximum number of iterations, and search space limits, etc. In this embodiment, the ISSA population parameters include the search space with dimensions, determine the search space food as and the sparrow position as ; where, is the search space food matrix, are all elements of the search space food matrix, is the sparrow position matrix, are all elements of the sparrow position matrix, is the ISSA population size, is the dimension of the model optimization problem; h is the number of ISSA individuals; the search space limits include the search space upper limit and the search space lower limit is ; T is the transpose symbol;

[0079] The formula of the fitness function is:

[0080]

[0081] In the formula, is the fitness value of the ISSA individual ; is the model error function; is the ISSA individual variable; c is the ISSA individual indicator;

[0082] S1-2-7-3: Based on the algorithm parameters and the individual vector, use the Circle chaotic mapping sequence to perform population initialization to obtain an initial ISSA population including several initial ISSA individuals;

[0083] The formula is:

[0084]

[0085] In the formula, is the initial ISSA individual of the Circle chaotic mapping; is the randomly generated initial ISSA individual; cIt is the individual indication quantity of ISSA;

[0086] S1-2-7-4: According to the fitness function, use the ISSA algorithm to perform identity assignment and update on the initial ISSA population, obtain an updated ISSA population including several updated ISSA individuals, and retain the optimal individual, including the following steps:

[0087] S1-2-7-4-1: Use the fitness function to obtain the fitness value of each initial ISSA individual in the initial ISSA population;

[0088] S1-2-7-4-2: Sort the initial ISSA individuals according to the fitness values of the initial ISSA individuals to obtain the initial discoverers, initial joiners, and initial predators;

[0089] S1-2-7-4-3: Update the initial ISSA population to obtain an updated ISSA population; the updated ISSA population includes updated discoverers, updated joiners, and updated predators;

[0090] The update formula for the discoverer is:

[0091]

[0092] In the formula, are respectively the t +1, t th iteration of the c th discoverer ISSA individual; is the maximum number of iterations; is a random number between 0 and 1; is a normally distributed random number; is matrix, all of whose elements are 1; is the warning value; is the safety threshold;

[0093] The update formula for the joiner is:

[0094]

[0095] In the formula, are respectively the t +1, t th iteration of the c th joiner ISSA individual; is the best position occupied by the exposed individual; is the current worst position; is a random number between 0 and 1; is A matrix with all elements being 1 or -1; h is the number of ISSA individuals; A + is the update parameter;

[0096] The update formula for the predator is:

[0097]

[0098] In the formula, are respectively the t +1, t th iteration of the c th predator ISSA individual; is the step size control parameter, and , is the convergence factor, is a non-zero positive real number for step size control; is the current best position; are respectively the current, best, and worst fitness of the ISSA individual; is the minimum constant to prevent the denominator from being zero;

[0099]

[0100] In the formula, is the convergence factor; tanh(.) is the hyperbolic tangent function; , are respectively the maximum and minimum values of the convergence factor; λ is the decreasing rate parameter, is the decreasing period parameter, λ = -2 π , = π ;

[0101] S1-2-7-4-5: Introduce the Gaussian mutation mechanism to perform Gaussian mutation on the updated ISSA population, obtain the Gaussian mutated ISSA population including several Gaussian mutated ISSA individuals, and retain the optimal individual;

[0102] The formula is:

[0103]

[0104] In the formula, is the Gaussian mutated ISSA individual; G (1,1) is the Gaussian mutation parameter; is the updated ISSA individual;

[0105] S1-2-7-6: Introduce a dynamic reverse mechanism to perform dynamic reversal on the updated ISSA population, obtaining a dynamically reversed ISSA population including a number of dynamically reversed ISSA individuals, and retaining the optimal individual;

[0106] The formula is:

[0107]

[0108] In the formula, is the dynamically reversed ISSA individual; is the decreasing inertia coefficient; is the upper limit of the search space; is the lower limit of the search space; is the updated ISSA individual;

[0109] S1-2-7-7: If the current iteration number reaches the maximum iteration number or the fitness value of the optimal individual is lower than the fitness value threshold, output the optimal individual and proceed to the next step;

[0110] S1-2-7-8: Decode the individual vector of the optimal individual to obtain the optimal initial model parameters of the initial image data automatic annotation model;

[0111] S1-2-7-9: Optimize the initial image data automatic annotation model according to the optimal initial model parameters to obtain an optimized image data automatic annotation model;

[0112] The ISSA algorithm optimizes the image data automatic annotation model, avoids the sensitivity of the image data automatic annotation model to the initial value, improves the training efficiency and accuracy of the image data automatic annotation model, and prevents the image data automatic annotation model from falling into premature convergence;

[0113] S1-2-8: Based on the comprehensive loss function, use a number of preprocessed historical image data to optimize and train the optimized image data automatic annotation model to obtain the final image data automatic annotation model;

[0114] S1-3: Set a model adjustment mechanism, and according to the model adjustment mechanism, collect a number of historical performance data of the image data automatic annotation model and perform preprocessing to obtain a number of preprocessed historical performance data;

[0115] S1-4: According to a number of preprocessed historical performance data, use a reinforcement learning algorithm to construct an automatic annotation strategy adjustment model;

[0116] The automatic annotation strategy adjustment model is constructed based on the Meta-Policy Optimization (MPO)-Proximal Policy Optimization (PPO) algorithm. The automatic annotation strategy adjustment model includes a meta-policy optimization module constructed based on the MPO algorithm and a reinforcement learning module constructed based on the PPO algorithm. The reinforcement learning module is provided with an agent, a policy network, and an experience replay pool. The agent is respectively connected to the policy network, the experience replay pool, and the meta-policy optimization module, and the meta-policy optimization module is connected to the experience replay pool;

[0117] According to a number of preprocessed historical performance data, using a reinforcement learning algorithm, construct an automatic annotation strategy adjustment model, including the following steps:

[0118] S1-4-1: Use the Fuzzy-C Means (FCM) clustering algorithm to cluster a number of preprocessed historical performance data to obtain a number of cluster centers and corresponding cluster clusters, and set historical policy adjustment categories for each cluster cluster; Clustering helps to identify different policy adjustment categories in the performance data and provides a basis for subsequent meta-policy optimization;

[0119] S1-4-2: Use the experience replay mechanism to initialize the experience replay pool, construct the agent of the reinforcement learning module, and define the state space, action space, and reward function of the agent; The experience replay pool can break the correlation between data and improve the stability and efficiency of learning;

[0120] The experience replay pool is used to store meta-policy optimization experiences and reinforcement learning experiences. The reinforcement learning experiences include states, actions, rewards, and next states. The meta-policy optimization experiences include the network parameters of the policy network in the reinforcement learning module and the corresponding policy adjustment categories;

[0121] S1-4-3: Take the historical automatic annotation strategy adjustment problem as a simulation environment, use the PPO algorithm to construct a policy network, and combine the experience replay pool and the agent to obtain an initial reinforcement learning module; PPO is a reinforcement learning algorithm that maximizes the cumulative reward by optimizing the parameters of the policy network, which can effectively balance exploration and exploitation while maintaining the stability of training;

[0122] S1-4-4: Set a number of policy adjustment categories as a number of sub-scenarios for meta-policy optimization, and use the MPO algorithm to construct an initial meta-policy optimization module; Allowing the model to learn different policies in different scenarios can improve the adaptability and generalization ability of the policies. Meta-policy optimization can enable the model to quickly adapt to different tasks or environments without having to learn from scratch;

[0123] S1-4-5: Train and optimize the initial reinforcement learning module based on a number of preprocessed historical performance data, extract the historical policy network parameters of the policy network of the reinforcement learning module during the training and optimization, obtain the final reinforcement learning module, and generate a number of historical reinforcement learning experiences;

[0124] S1-4-6: Based on a number of sub-scenarios, train and optimize the initial meta-policy optimization module according to a number of historical policy network parameters, obtain the final meta-policy optimization module, and generate a number of historical meta-policy optimization experiences;

[0125] S1-4-7: Store a number of historical meta-policy optimization experiences and a number of historical reinforcement learning experiences in the experience replay pool of the final reinforcement learning module, and integrate the final meta-policy optimization module and the final reinforcement learning module to obtain an automatic annotation policy adjustment model;

[0126] S2: According to the preset model adjustment mechanism, collect the real-time performance data of the image data automatic annotation model, and input the real-time performance data into the automatic annotation policy adjustment model;

[0127] The model adjustment mechanism includes the time period for collecting the performance data of the image data automatic annotation model, and the performance metrics included in the performance data, such as accuracy, recall, F1 score, etc.;

[0128] S3: According to the real-time performance data, use the automatic annotation policy adjustment model to adjust the automatic annotation policy of the automatic annotation policy adjustment model to obtain an adjusted image data automatic annotation model, including the following steps:

[0129] S3-1: Obtain the similarity between the real-time performance data and a number of clustering centers, and use the historical policy adjustment category of the clustering center with the highest similarity as the real-time policy adjustment category of the real-time performance data; This helps to identify the policy adjustment category to which the real-time data belongs, so as to select the most appropriate meta-optimization strategy;

[0130] S3-2: According to the real-time policy adjustment category, extract the corresponding real-time meta-policy optimization experience from a number of historical meta-policy optimization experiences in the experience replay pool of the automatic annotation policy adjustment model; This allows the model to utilize past learning experiences in similar situations and quickly adapt to new annotation scenarios;

[0131] S3-3: According to the real-time meta-policy optimization experience, use the meta-policy optimization module to initialize the real-time policy network parameters of the policy network of the reinforcement learning module to obtain an updated policy network; This ensures that the policy network can quickly adjust its behavior according to the current annotation task requirements;

[0132] S3-4: Randomly select several real-time reinforcement learning experiences from a number of historical reinforcement learning experiences in the experience replay pool, and update the action space of the reinforcement learning module according to the several real-time reinforcement learning experiences to obtain an updated action space; the updated action space can reflect the possible actions of the current annotation task, improving the adaptability and flexibility of the automatic annotation strategy;

[0133] S3-5: Update the state space of the reinforcement learning module according to the real-time performance data to obtain an updated state space; this helps the model better understand the current annotation task, thereby making more appropriate decisions;

[0134] S3-6: Use an agent to control the updated policy network to generate the probability distribution of all possible actions in the action space corresponding to each state in the state space; this allows the model to select the optimal action among multiple possible actions to maximize the expected reward;

[0135] S3-7: Take the possible action with the highest probability distribution in the action space as the execution action of the state, and integrate the execution actions of all states in the state space to obtain several automatic annotation strategy adjustment actions; this helps the model make more accurate and effective decisions during the automatic annotation process;

[0136] The automatic annotation strategy adjustment actions include:

[0137] Weight reallocation action: Adjust the weights of different parts in the comprehensive loss function of the image data automatic annotation model. For example, if there are many errors in the localization of the detected target regions, the weight of the localization loss in the total loss can be increased;

[0138] Feature selection optimization action: According to the specific situation of performance decline, reselect or optimize the features used for annotation, which may include adding new features or discarding irrelevant features;

[0139] Sample relabeling action: For the image data mislabeled by the model, relabel it and use these samples as part of the training data to improve the model;

[0140] S3-8: Adjust the automatic annotation strategy of the automatic annotation strategy adjustment model according to several automatic annotation strategy adjustment actions to obtain an adjusted image data automatic annotation model;

[0141] S4: Collect real-time image data and use the adjusted image data automatic annotation model to automatically annotate the real-time image data to obtain automatically annotated real-time image data, including the following steps:

[0142] S4-1: Collect real-time image data and preprocess the real-time image data to obtain preprocessed real-time image data. The preprocessing may include operations such as scaling, cropping, normalization, denoising, etc., to ensure that the image data is suitable for input into the adjusted automatic image data annotation model, which can improve the quality of the image data, reduce the complexity of model processing, and contribute to improving the accuracy and efficiency of subsequent annotation.

[0143] S4-2: Use the graph feature extraction module of the adjusted automatic image data annotation model to extract several real-time feature maps of the preprocessed real-time image data and perform feature map fusion to obtain real-time fused features. Through feature map fusion, the model can capture detailed information and high-level structures in the image, which helps to more accurately locate and identify the target.

[0144] S4-3: Use the target area localization module of the adjusted automatic image data annotation model to perform the target area localization module based on the real-time fused features to obtain several real-time target areas. Accurate target area localization is the key to automatic annotation, ensuring that the automatic annotation module can annotate the correct image area.

[0145] S4-4: Use the automatic annotation module of the adjusted automatic image data annotation model to automatically annotate several real-time target areas based on the real-time fused features to obtain the real-time image data after automatic annotation. The output of the automatic annotation module is the image data with annotation information, which directly provides an interpretation of the image content and is crucial for image analysis and understanding.

[0146] Embodiment 2:

[0147] As Figure 2 shown, this embodiment provides an automatic image data annotation system based on deep learning for implementing the automatic image data annotation method. The system includes a model construction unit, a performance data collection unit, a model adjustment unit, and an automatic annotation unit that are connected in sequence.

[0148] The model construction unit is used to construct an automatic image data annotation model using a deep learning algorithm and construct an automatic annotation strategy adjustment model using a reinforcement learning algorithm.

[0149] The performance data collection unit is used to collect the real-time performance data of the automatic image data annotation model according to a preset model adjustment mechanism and input the real-time performance data into the automatic annotation strategy adjustment model.

[0150] The model adjustment unit is used to adjust the automatic annotation strategy of the automatic annotation strategy adjustment model according to the real-time performance data using the automatic annotation strategy adjustment model to obtain an adjusted automatic image data annotation model.

[0151] An automatic annotation unit, which is used to collect real-time image data and automatically annotate the real-time image data by using an adjusted image data automatic annotation model to obtain automatically annotated real-time image data.

[0152] A method and system for automatically annotating image data based on deep learning provided by the present invention can automatically process image data, significantly reduce manual participation, and thus greatly improve the annotation efficiency, especially when dealing with large-scale image data sets; the image data automatic annotation model constructed by using deep learning algorithms can more accurately identify and annotate the targets in the images, reduce human errors and subjective inconsistencies, improve the accuracy of annotation, and increase the processing speed of image data, which is applicable to application scenarios that require quick response; the automatic annotation strategy adjustment model constructed by using reinforcement learning algorithms improves the degree of intelligence, dynamically adjusts the annotation strategy according to real-time performance data, makes the model update and iteration more simple and efficient, and through the combination of reinforcement learning and meta-strategy optimization, enables the automatic annotation model to better generalize to new data sets, improving the performance of the automatic annotation model in different scenarios; the domain adversarial training module realizes the transfer of cross-domain knowledge and improves the performance of the automatic annotation model in cross-domain image annotation tasks.

[0153] The present invention is not limited to the above optional embodiments, and anyone can obtain other various forms of products under the inspiration of the present invention. The above specific embodiments should not be construed as limiting the protection scope of the present invention, and the protection scope of the present invention should be defined by the claims, and the description can be used to interpret the claims.

Claims

1. A method for automatic annotation of image data based on deep learning, characterized in that: The steps include: Use deep learning algorithms to build an automatic image data annotation model, and use reinforcement learning algorithms to build an automatic annotation strategy adjustment model; According to the preset model adjustment mechanism, real-time performance data of the image data automatic annotation model is collected, and the real-time performance data is input into the automatic annotation strategy adjustment model; According to the real-time performance data, the automatic annotation strategy of the automatic annotation strategy adjustment model is adjusted using the automatic annotation strategy adjustment model to obtain an adjusted image data automatic annotation model; Collecting real-time image data, and using the adjusted image data automatic annotation model to automatically annotate the real-time image data to obtain automatically annotated real-time image data; The automatic labeling strategy adjustment model is constructed based on the MPO-PPO algorithm, and the automatic labeling strategy adjustment model includes a meta-strategy optimization module constructed based on the MPO algorithm and a reinforcement learning module constructed based on the PPO algorithm. The reinforcement learning module is provided with an intelligent agent, a strategy network and an experience replay pool. The intelligent agent is respectively connected to the strategy network, the experience replay pool and the meta-strategy optimization module, and the meta-strategy optimization module is connected to the experience replay pool. Based on some pre-processed historical performance data, a reinforcement learning algorithm is used to build an automatic labeling strategy adjustment model, which includes the following steps: Use the FCM clustering algorithm to cluster a number of pre-processed historical performance data to obtain a number of cluster centers and corresponding cluster clusters, and set a historical strategy adjustment category for each cluster cluster; Use the experience replay mechanism to initialize the experience replay pool, build the agent of the reinforcement learning module, and define the state space, action space and reward function of the agent; The historical automatic labeling strategy adjustment problem is used as a simulation environment, the PPO algorithm is used to build a strategy network, and the initial reinforcement learning module is obtained by combining the experience replay pool and the intelligent agent. Set several policy adjustment categories as several sub-scenarios of meta-policy optimization, and use the MPO algorithm to build an initial meta-policy optimization module; According to some pre-processed historical performance data, the initial reinforcement learning module is trained and optimized, and the historical policy network parameters of the policy network of the reinforcement learning module in the training optimization are extracted to obtain the final reinforcement learning module, and generate some historical reinforcement learning experiences; Based on several sub-scenarios and several historical policy network parameters, the initial meta-policy optimization module is trained and optimized to obtain the final meta-policy optimization module, and several historical meta-policy optimization experiences are generated; Several historical meta-strategy optimization experiences and several historical reinforcement learning experiences are stored in the experience replay pool of the final reinforcement learning module, and the final meta-strategy optimization module and the final reinforcement learning module are integrated to obtain an automatic labeling strategy adjustment model.

2. The method for automatic annotation of image data based on deep learning according to claim 1, characterized in that: Use deep learning algorithms to build an automatic image data annotation model, and use reinforcement learning algorithms to build an automatic annotation strategy adjustment model, including the following steps: Collecting a number of historical image data, and preprocessing the number of historical image data to obtain a number of preprocessed historical image data; Based on some pre-processed historical image data, a deep learning algorithm is used to build an automatic image data annotation model; A model adjustment mechanism is set, and according to the model adjustment mechanism, image data is collected to automatically annotate a number of historical performance data of the model, and preprocessing is performed to obtain a number of preprocessed historical performance data; Based on some preprocessed historical performance data, a reinforcement learning algorithm is used to build an automatic labeling strategy adjustment model.

3. The method for automatic annotation of image data based on deep learning according to claim 2, characterized in that: The image data automatic annotation model is constructed based on the FPN-Faster R-CNN-DANN-DBN algorithm, and the image data automatic annotation model includes a graph feature extraction module constructed based on the FPN algorithm, a target area positioning module constructed based on the Faster R-CNN algorithm, a domain adversarial training module constructed based on the DANN algorithm, and an automatic annotation module constructed based on the DBN algorithm. The graph feature extraction module, the target area positioning module, and the automatic annotation module are connected in sequence, the domain adversarial training module is connected to the target area positioning module, and the domain adversarial training module includes a label predictor and a domain classifier connected in sequence.

4. The method for automatic annotation of image data based on deep learning according to claim 3, characterized in that: Based on some pre-processed historical image data, a deep learning algorithm is used to build an automatic image data annotation model, including the following steps: Using the CNN algorithm, a basic network architecture of a graph feature extraction module is constructed; the basic network architecture includes alternately connected convolutional layers and pooling layers; Use the FPN algorithm to build a feature pyramid, and connect the outputs of all convolutional layers to the feature pyramid horizontally in a top-down order to obtain the initial graph feature extraction module; Use the Faster R-CNN algorithm to build the initial target area positioning module, and use the DANN algorithm to build the initial domain adversarial training module; Use the DBN algorithm to build the initial automatic annotation module; Integrate the initial graph feature extraction module, the initial target area positioning module, the initial domain adversarial training module and the initial automatic annotation module to obtain the initial image data automatic annotation model; Combining the loss functions of the initial graph feature extraction module, the initial target region localization module, the initial domain adversarial training module, and the initial automatic labeling module, a comprehensive loss function is obtained; Optimizing initial model parameters of the initial image data automatic annotation model to obtain an optimized image data automatic annotation model; Based on the comprehensive loss function, several preprocessed historical image data are used to optimize the training of the optimized image data automatic annotation model to obtain the final image data automatic annotation model.

5. The method for automatic annotation of image data based on deep learning according to claim 4, characterized in that: Taking minimizing the model error as the optimization goal, a swarm intelligence optimization algorithm is used to optimize the initial model parameters of the initial image data automatic annotation model to obtain an optimized image data automatic annotation model.

6. The method for automatic annotation of image data based on deep learning according to claim 1, characterized in that: According to the real-time performance data, the automatic annotation strategy of the automatic annotation strategy adjustment model is adjusted using the automatic annotation strategy adjustment model to obtain an adjusted image data automatic annotation model, including the following steps: Obtaining similarities between the real-time performance data and a number of cluster centers, and using the historical strategy adjustment category of the cluster center with the highest similarity as the real-time strategy adjustment category of the real-time performance data; According to the real-time strategy adjustment category, the corresponding real-time meta-strategy optimization experience is extracted from a number of historical meta-strategy optimization experiences in the experience playback pool of the automatic labeling strategy adjustment model; Based on the experience of real-time meta-strategy optimization, the meta-strategy optimization module is used to initialize the real-time policy network parameters of the policy network of the reinforcement learning module to obtain an updated policy network. Randomly extract a number of real-time reinforcement learning experiences from a number of historical reinforcement learning experiences in the experience replay pool, and update the action space of the reinforcement learning module according to the number of real-time reinforcement learning experiences to obtain an updated action space; According to the real-time performance data, the state space of the reinforcement learning module is updated to obtain an updated state space; Use the agent to control the updated policy network to generate the probability distribution of all possible actions in the action space corresponding to each state in the state space; The possible action with the highest probability distribution in the action space is taken as the execution action of the state, and the execution actions of all states in the state space are integrated to obtain several automatic labeling strategy adjustment actions; According to a number of automatic labeling strategy adjustment actions, the automatic labeling strategy of the automatic labeling strategy adjustment model is adjusted to obtain an adjusted image data automatic labeling model.

7. The method for automatic annotation of image data based on deep learning according to claim 6, characterized in that: Collecting real-time image data, using the adjusted image data automatic annotation model to automatically annotate the real-time image data, and obtaining the automatically annotated real-time image data, includes the following steps: Collecting real-time image data, and preprocessing the real-time image data to obtain preprocessed real-time image data; Using the graph feature extraction module of the adjusted image data automatic annotation model, extract several real-time feature graphs of the preprocessed real-time image data, and perform feature graph fusion to obtain real-time fusion features; Using the adjusted image data to automatically annotate the model’s target region positioning module, perform the target region positioning module based on the real-time fusion features to obtain several real-time target regions; The automatic annotation module of the adjusted image data automatic annotation model is used to automatically annotate several real-time target areas according to the real-time fusion features to obtain the automatically annotated real-time image data.

8. A deep learning-based automatic image data annotation system, used to implement the automatic image data annotation method according to any one of claims 1 to 7, characterized in that: The system comprises a model building unit, a performance data collection unit, a model adjustment unit and an automatic marking unit which are connected in sequence.

Citation Information

Patent Citations

  • Model training method based on NLP large model

    CN118734924A

  • Decision management method and system based on machine learning

    CN119150050A