Demonstration learning method and demonstration learning system based on visual segmentation large model
By minimizing the labeling investment, combining feature search models and visual segmentation large models, automatic labeling of unlabeled pictures has been solved, and the problem of difficulty in obtaining defect labeling data in industrial intelligent manufacturing is realized, and an efficient and accurate labeling process is achieved, reducing costs.
Patent Information
- Application Number
- CN202510139070.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-05-30
AI Technical Summary
In the field of industrial intelligent manufacturing, the process of obtaining sufficient and reliable product appearance defect labeling data is time-consuming and labor-intensive, and is easily limited by differences in subjective judgments, which affects the consistency and accuracy of labeling.
By minimizing the annotation input, comparing the similarity between the marked and unlabeled images in combination with the feature search model, the distance between the two in the feature space is calculated to infer the category of the unlabeled image. The visual segmentation model and Swin Transformer are used as the backbone network of the feature retrieval model, and the adapter fine-tuning and Triplet comparison function are used to achieve automated annotation.
It significantly improves the labeling efficiency and accuracy of industrial quality inspection defect samples, reduces manual intervention, reduces labeling costs, improves the consistency and accuracy of labeling, and provides higher quality data support for deep learning models in defect detection applications.
Smart Images

Figure CN120070975A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of artificial intelligence and data annotation, and particularly relates to a teaching learning method and a teaching learning system based on a large visual segmentation model. Background Art
[0002] In the field of industrial intelligent manufacturing, using deep learning models to accurately detect product appearance defects has become an indispensable technical means.
[0003] However, in the actual application process, a significant difficulty lies in obtaining sufficient and reliable defect annotation data, which is often accompanied by high human and material costs. Traditional annotation methods, such as relying on visual inspection and manually annotating defect areas, are not only time-consuming and laborious but also easily limited by subjective judgment differences, affecting the consistency and accuracy of annotation. To address this problem, based on the large visual segmentation model SAM or similar models, by selecting the target area, it can automatically capture the outer contour points of defects to a certain extent, thus reducing the burden of manual annotation. Although this method has made significant progress compared to pure manual annotation, it still relies on a considerable degree of manual intervention. Summary of the Invention
[0004] To overcome the deficiencies of the prior art, the purpose of the present invention is to provide a sample processing method and a processing system that can solve the above problems.
[0005] Design principle: This application realizes efficient automation of defect detection annotation with minimal annotation input, and combines a feature retrieval model (i.e., an embedding model) to compare the similarity of annotated and unannotated images or instances, calculates the distance between the two in the feature space, and thus obtains the class attribution of the unannotated images or instances.
[0006] A teaching learning method based on a large visual segmentation model includes the following steps: constructing a large visual segmentation model, the large visual segmentation model adopts the YOLO-World model structure, and obtains a large visual segmentation model for industrial quality inspection through the adapter fine-tuning method for the backbone network; the adapter network consists of 3 cascaded convolutional layers, activation layers, and normalization layers; constructing a feature retrieval model, the feature retrieval model uses the Swin Transformer as the backbone network and the Triplet comparison function in contrastive learning as the loss function; calculating the feature vector embedding of the image, and calculating or extracting the feature vector embedding of each image or instance through the feature retrieval model; classifying and teaching learning, and classifying the target image through image classification teaching learning.
[0007] Preferably, the image classification teaching and learning method includes the following steps: S11. Annotate images, annotate several pictures, check the number of pictures in each annotated category, and ensure that there are 3 to 5 pictures in each category; if the statistics do not meet the requirements, teaching and learning cannot start, and if they meet the requirements, proceed to the next step; S12. Calculate the feature vectors of the images, and calculate the feature vector embeddings of the annotated and unannotated pictures through the feature retrieval model; S13. Calculate the similarity of the unannotated pictures, and confirm the similarity of the unannotated pictures compared with the annotated pictures; S14. Confirm the image category attribution, and obtain the category attribution of the unannotated pictures through nearest neighbor or Top5 voting.
[0008] Preferably, in the confirmation of image category attribution, the method for confirming the category of the nearest neighbor picture is: calculate the picture that is most similar to the unannotated picture among the annotated pictures, and assign the category of the most similar annotated picture to the unannotated picture.
[0009] Preferably, in the confirmation of image category attribution, the method for confirming the category by Top5 voting is: screen the top 5 annotated images for the unannotated picture, conduct weighted voting according to the selected categories, and determine the category of the unannotated image according to the voting result.
[0010] Preferably, the instance segmentation teaching and learning method includes: S21. Annotate images, perform instance segmentation annotation on the images, and for each category, at least annotate 3 to 5 instances as examples; S22. Calculate the feature vectors of each instance in the annotated images, and calculate the feature vector embeddings of each instance in the annotated images through the feature retrieval model; S23. Obtain the prediction targets, detect the unannotated pictures through the YOLO-World model, and obtain all the prediction targets in the unannotated images; S24. Calculate the feature vectors of the prediction targets, and calculate the feature vector embeddings of all the prediction targets in the unannotated images through the feature retrieval model; S25. Confirm the image category attribution, compare the feature vector embeddings of all the prediction targets in the image set to be taught and learned with the feature vector embeddings of the annotated instances. If the similarity between the two is within the set range, it is considered detected, and the category attribution of the prediction target is obtained. If it is not within this set range, continue to compare with the feature vector embeddings of other instance targets until the category attribution of the prediction target is detected or it does not match all the annotated instances and then discard the prediction target.
[0011] Preferably, the inter-class similarity and intra-class similarity of different class targets of the image are combined for statistics to obtain a similarity threshold, and then the category attribution of the target is performed. The specific steps are as follows: S31. Calculate each one according to all the annotated images Class target The average intra-class similarity of the class, average inter-class similarity; S32. Take the best threshold between the two average similarities, so that each target has the highest accuracy in class attribution through this threshold。
[0012] The present invention also provides a teaching learning system based on a large visual segmentation model. The teaching learning system includes: a large visual segmentation model module for constructing and training a large visual segmentation model; a feature retrieval model module for constructing and training a feature retrieval model; an image classification teaching learning module for executing an image classification teaching learning method; an instance segmentation teaching learning module for executing an instance segmentation teaching learning method; a data management module for storing and managing labeled and unlabeled picture data; and a user interaction module for a user to input teaching learning parameters and view annotation results.
[0013] Preferably, the large visual segmentation model module includes a pre-training data input end at the physical level, a trained backbone network model, a trained physical layer adapter model, a splicing unit, and a physical level spliced feature map output end. Among them, the trained backbone network model outputs a first feature map trained in four stages, and the trained physical layer adapter model outputs a third feature map processed by a three-stage adaptation network; during the training process, the parameters of the backbone network remain unchanged, and the parameters of the adapter network are fine-tuned. The adapter network includes 3 cascaded convolutional layers, activation layers, and normalization layers.
[0014] Preferably, both the image classification teaching learning module and the instance segmentation teaching learning module support selecting static criteria or dynamic expansion for the labeled images.
[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: Through the present invention, the annotation efficiency and accuracy of industrial quality inspection defect samples have been greatly improved, and the annotation efficiency has been significantly enhanced. By automatically comparing the similarity of labeled and unlabeled pictures in the feature space, this method can quickly infer the category of unlabeled pictures, thus greatly reducing the workload of manual annotation and improving the overall annotation efficiency. Specifically, it is manifested in the following points:
[0016] (1) Reduced annotation costs: Due to the reduction of manual intervention, this method reduces the human and material costs required to obtain sufficient and reliable annotation data, making the application of deep learning models in defect detection in the industrial manufacturing field more economically feasible.
[0017] (2) Improved annotation consistency and accuracy: Compared with the traditional manual annotation method, this method reduces the differences in subjective judgments, improves the annotation consistency and accuracy, and provides higher-quality data support for the subsequent training of defect detection models.
[0018] (3) Promoted the process of industrial intelligence: The successful application of this method further promotes the process of industrial intelligence in the industrial manufacturing field, providing strong technical support for achieving more efficient and accurate product quality control. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 Schematic diagram of a large visual segmentation model based on adapter fine-tuning; Figure 2 Schematic diagram of the process of the image classification teaching learning method; Figure 3 Schematic diagram of the process of the instance segmentation teaching learning method. Specific implementation manners
[0020] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0021] Learning from demonstration method A teaching learning method based on a large visual segmentation model, see Figure 1 - Figure 3 , the teaching learning method includes the following steps (the steps are not necessarily distinguished in sequence).
[0022] Construct a large visual segmentation model. The large visual segmentation model adopts the YOLO-World model structure, and a large visual segmentation model for industrial quality inspection is obtained by fine-tuning the backbone network through an adapter method; the adapter network consists of 3 cascaded convolutional layers, activation layers and normalization layers.
[0023] Construct a feature retrieval model. The feature retrieval model uses Swin Transformer as the backbone network and the Triplet comparison function in contrastive learning as the loss function.
[0024] Calculate the feature vector embedding of the image. Through the feature retrieval model, calculate or extract the feature vector embedding of each image or instance.
[0025] Classification teaching learning, classify the target image through image classification teaching learning or instance segmentation teaching learning.
[0026] The core of the solution is: calculate the similarity comparison between the labeled and unlabeled pictures, calculate the distance between the two in the feature space, so as to obtain the class attribution of the unlabeled picture; design the image classification teaching learning method and the instance segmentation teaching learning method, and select the corresponding teaching learning method according to different labeling requirements.
[0027] Among them, see Figure 2, the image classification teaching learning method includes the following steps (the order is not necessarily distinguished).
[0028] S11. Label the images, label several pictures, and check the number of labeled pictures in each category to ensure that there are 3 to 5 pictures in each category; if the statistics do not meet the requirements, the teaching learning cannot start, and if the requirements are met, proceed to the next step.
[0029] S12. Calculate the feature vectors of the images, and calculate the feature vector embeddings of the labeled and unlabeled pictures through the feature retrieval model.
[0030] S13. Calculate the similarity of the unlabeled pictures, and confirm the similarity of the unlabeled pictures compared with the labeled pictures.
[0031] S14. Confirm the image category attribution, and obtain the category attribution of the unlabeled pictures through nearest neighbor or Top5 voting.
[0032] Further, in the confirmation of image category attribution, the method for confirming the category of the nearest neighbor picture is: calculate the picture that is most similar to the unlabeled picture among the labeled pictures, and assign the category of the most similar labeled picture to the unlabeled picture. This process can select static criteria or dynamic expansion for the labeled images.
[0033] Further, in the confirmation of image category attribution, the method for confirming the category by Top5 voting is: screen the top 5 labeled images for the unlabeled pictures, perform weighted voting according to the selected categories (the weighting rule can be customized as a method of decreasing coefficients), and determine the category of the unlabeled image according to the voting results. This process can select static criteria or dynamic expansion for the labeled images.
[0034] Among them, see Figure 3 , the instance segmentation teaching learning method includes the following steps (the order is not necessarily distinguished).
[0035] S21. Perform instance segmentation labeling on the images, and for each category, label at least 3 to 5 instances as examples.
[0036] S22. Calculate the feature vectors of each instance in the labeled images, and calculate the feature vector embeddings of the instances in the labeled images through the feature retrieval model.
[0037] S23. Obtain the prediction targets, detect the unlabeled pictures through the YOLO-World model, and obtain all the prediction targets in the unlabeled images.
[0038] S24. Calculate the feature vectors of the prediction targets, and calculate the feature vector embeddings of the prediction targets through the feature retrieval model.
[0039] S25. Image category attribution confirmation: Compare the feature vector embedding of each predicted target in the image set to be taught and learned with the feature vector embedding of the labeled target. If the similarity between the two is within the set range, it is considered detected, and the category attribution of the predicted target is obtained. If it is not within the set range, continue to compare with the feature vector embeddings of other labeled targets until the category attribution of the predicted target is detected or the predicted target is discarded after not matching all labeled instances.
[0040] Among them, the inter-class similarity and intra-class similarity of different classes of images are combined for statistics to obtain a similarity threshold, and then the category attribution of the target is carried out. The specific steps are as follows.
[0041] S31. Calculate for each The average intra-class similarity of the class target, inter-class average similarity based on all labeled images.
[0042] S32. Take the best threshold between the two average similarities, so that each target passes this threshold for this class attribution with the highest accuracy .
[0043] In this method, the visual segmentation large model adopts the YOLO-World model structure, and the visual segmentation large model for industrial quality inspection is obtained by fine-tuning the backbone network through an adapter. The specific model structure is shown in Figure 1 , and during training, the parameters of the backbone network are kept unchanged, and the parameters of the adapter network are fine-tuned. The adapter network consists of 3 cascaded convolutional layers, activation layers, and normalization layers. The Embedding model uses the Swin Transformer as the backbone network and the Triplet comparison function in contrastive learning as the loss function. Based on the visual large model, the present invention designs two teaching and learning methods, namely image classification teaching and learning and instance segmentation teaching and learning.
[0044] The image classification teaching and learning proposed by the present invention is as Figure 2 shown. When using this method, it is required that each class has at least 3 - 5 images. If the statistics do not meet the requirements, teaching and learning cannot start. Then, the embedding of each image is calculated through a unified embedding model, the similarity of the unlabeled images is calculated, and the nearest neighbor or Top5 voting (optional) is used to determine the category attribution. The specific steps are as follows: 1) Calculate the most similar image of the unlabeled image to the labeled image, and assign its category to the unlabeled image. This process can select a static standard or dynamic expansion for the labeled images. 2) Screen the top 5 labeled images for the unlabeled image, and conduct weighted voting according to the selected categories (the weighting rule can be customized as a method of decreasing coefficients), and determine the category of the unlabeled image according to the voting result. This process can select a static standard or dynamic expansion for the labeled images.
[0045] The instance segmentation teaching learning method proposed by the present invention is as follows Figure 3 as shown. When using this method, one or more targets are selected as examples, and then the embeddings of the labeled instances in each image are extracted through a unified embedding model; finally, YOLO-World is used to perform relatively strict detection on the image set to be taught and learned, calculate the embeddings of each predicted target, and compare them with the embeddings of existing targets. If the similarity is within a certain range, the target is detected. The similarity can be statistically analyzed by combining the inter-class similarity and intra-class similarity of different types of targets in the image, and a reasonable threshold is given, and then the category attribution of the target is performed. The specific steps are as follows: 1) Calculate the intra-class average similarity and inter-class average similarity of each type of target according to all the labeled images; 2) Take the best threshold between the two average similarities so that the accuracy of each target's category attribution through this threshold is the highest.
[0046] Learning from demonstration system A teaching learning system based on a large visual segmentation model, see Figure 1 , the teaching learning system includes the following modules.
[0047] Large visual segmentation model module: used to construct and train a large visual segmentation model.
[0048] Feature retrieval model module: used to construct and train a feature retrieval model.
[0049] Image classification teaching learning module: used to execute the image classification teaching learning method.
[0050] Instance segmentation teaching learning module: used to execute the instance segmentation teaching learning method.
[0051] Data management module: used to store and manage labeled and unlabeled picture data.
[0052] User interaction module: used for users to input the parameters of teaching learning and view the annotation results.
[0053] Among them, the large visual segmentation model module includes a pre-training data input end at the physical level, a trained backbone network model, a trained physical layer adapter model, a splicing unit, and a physical level spliced feature map output end.
[0054] Data is input from the pre-training data input end as the first feature map and processed or trained through the four stages of the trained backbone network model. The trained backbone network model outputs the first feature map trained through the four stages and outputs it at the end.
[0055] The first feature map processed by the first stage of the backbone network model is input into the first adapter network of the trained physical layer adapter model for processing. The first feature map processed by the second stage of the backbone network model is input into the second adapter network of the trained physical layer adapter model for processing. The first feature map processed by the third stage of the backbone network model is input into the third adapter network of the trained physical layer adapter model for processing, and finally the third special function map is output, that is, the third feature map processed by the three-level adaptation network of the trained physical layer adapter model. During the training process, the parameters of the backbone network remain unchanged, and the parameters of the adapter network are fine-tuned. The adapter network includes 3 cascaded convolutional layers, activation layers, and normalization layers. The first feature map and the third feature map are concatenated in the concatenation unit to output a concatenated feature map.
[0056] Among them, both the image classification teaching and learning module and the instance segmentation teaching and learning module support selecting static standards or dynamic expansion for the annotated images.
[0057] Computer-readable storage medium The present invention also provides a computer-readable storage medium, on which computer instructions are stored, and when the computer instructions run, they execute the steps of the foregoing method. Among them, for the method, please refer to the detailed introduction in the foregoing part, and details will not be repeated here.
[0058] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program. This program can be stored in a computer-readable storage medium. The computer-readable medium includes permanent and non-permanent, removable and non-removable media, and information storage can be achieved by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media, such as modulated data signals and carrier waves.
[0059] Terminal The present invention also provides a terminal, including a memory and a processor. The memory stores data provider information and computer instructions capable of running on the processor. When the processor runs the computer instructions, it executes the steps of the foregoing method. Among them, for the foregoing method, please refer to the detailed introduction in the foregoing part, and details are not described herein again.
[0060] In addition, those skilled in the art can understand that various aspects of the present application can be described and illustrated by several patentable types or situations, including any new and useful process, machine, product, or combination of substances, or any new and useful improvement thereof. Accordingly, various aspects of the present application can be executed entirely by hardware, entirely by software (including firmware, resident software, microcode, etc.), or by a combination of hardware and software. The above-mentioned hardware or software can all be referred to as "data blocks", "modules", "engines", "units", "components" or "systems". In addition, various aspects of the present application may be embodied as a computer product located in one or more computer-readable media, and the product includes computer-readable program codes.
[0061] A computer storage medium may include a propagated data signal containing computer program codes therein, such as on a baseband or as part of a carrier wave. The propagated signal may have various forms of representation, including electromagnetic form, optical form, etc., or a suitable combination thereof. A computer storage medium may be any computer-readable medium other than a computer-readable storage medium, and this medium can be connected to an instruction execution system, device, or equipment to implement communication, propagation, or transmission of a program for use. The program codes located on the computer storage medium can be propagated through any suitable medium, including radio, cable, fiber optic cable, RF, or similar media, or any combination of the above media.
[0062] The computer program codes required for the operations of various parts of this application can be written in any one or more programming languages, including object-oriented programming languages such as Java, Scala, Smalltalk, Eiffel, JADE, Emerald, C++, C#, VB.NET, Python, etc., conventional procedural programming languages such as C language, VisualBasic, Fortran2003, Perl, COBOL2002, PHP, ABAP, dynamic programming languages such as Python, Ruby, and Groovy, or other programming languages. The program codes can run entirely on the user's computer, or run on the user's computer as an independent software package, or partially run on the user's computer and partially on a remote computer, or run entirely on a remote computer or processing device. In the latter case, the remote computer can be connected to the user's computer in any network form, such as a local area network (LAN) or a wide area network (WAN), or connected to an external computer (e.g., via the Internet), or in a cloud computing environment, or used as a service such as software as a service (SaaS).
[0063] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A teaching learning method based on a large visual segmentation model, characterized in that: The following steps are involved: Construct a large visual segmentation model. The large visual segmentation model adopts the YOLO-World model structure. The large visual segmentation model for industrial quality inspection is obtained by fine-tuning the backbone network with an adapter. Constructing a feature retrieval model, wherein the feature retrieval model uses Swin Transformer as a backbone network and a Triplet comparison function in contrastive learning as a loss function; Calculate the feature vector embedding of the image, and calculate or extract the feature vector embedding of each image or instance through the feature retrieval model; Classification teaching learning, classify the target image through image classification teaching learning.
2. The method according to claim 1, characterized in that The image classification teaching learning method comprises the following steps: S11, label images, label several pictures, check the number of labeled pictures of each category, and ensure that there are 3 to 5 pictures of each category; S12, calculating the feature vector of the image, and calculating the feature vector embedding of the labeled image and the unlabeled image through the feature retrieval model; S13, calculating the similarity of the unlabeled images, and confirming the similarity of the unlabeled images to the labeled images; S14: Confirm the image category attribution, and obtain the category attribution of the unlabeled image through nearest neighbor or Top5 voting.
3. The method according to claim 2, characterized in that In the image category attribution confirmation, the category confirmation method of the nearest neighbor image is: calculate the image that is most similar to the unlabeled image and the labeled image, and assign the category of the most similar labeled image to the unlabeled image.
4. The method according to claim 2, characterized in that: In the image category confirmation, the category confirmation method of Top5 voting is as follows: filter the TOP5 labeled images for the unlabeled images, perform weighted voting based on the filtered categories, and determine the unlabeled image category based on the voting results.
5. The method according to claim 1, characterized in that: The instance segmentation teaching learning method comprises: S21, perform instance segmentation and annotation on the image. For each category, at least 3-5 instances are annotated as examples; S22, calculating the feature vector of each instance in the labeled image, and calculating the feature vector embedding of the instance in the labeled image through the feature retrieval model; S23, obtaining predicted targets, detecting unlabeled images through the YOLO-World model, and obtaining all predicted targets in the unlabeled images; S24, calculating the feature vector of the predicted target, and calculating the feature vector embedding of the predicted target through the feature retrieval model; S25. Confirm the image category attribution. Compare the feature vector embedding of each predicted target in the image set that needs to be taught and learned with the feature vector embedding of the labeled target. If the similarity between the two is within the set range, it is detected and the category attribution of the predicted target is obtained. If it is not within the set range, it is compared with the feature vector embedding of other labeled targets until the category attribution of the predicted target is detected or the predicted target is discarded after it does not match all labeled instances.
6. The method according to claim 5, characterized in that Combine the inter-class similarity and intra-class similarity of different classes of objects in the image to obtain the similarity threshold, and then classify the objects. The specific steps are as follows: S31, calculate each according to all the labeled images The average similarity within the class of the target, Average similarity between classes; S32、 Take the best threshold between the two average similarities so that each target can be classified into the category through this threshold The highest accuracy 。 7. A teaching learning system based on a large visual segmentation model, characterized in that: The teaching learning system includes: Visual segmentation large model module: used to build and train a visual segmentation large model; Feature retrieval model module: used to build and train feature retrieval models; Image classification teaching learning module: used to execute image classification teaching learning method; Instance segmentation teaching learning module: used to execute instance segmentation teaching learning method; Data management module: used to store and manage labeled and unlabeled image data; User interaction module: used for users to input teaching learning parameters and view annotation results.
8. The teaching learning system according to claim 7, characterized in that: The visual segmentation large model module includes a pre-trained data input terminal of the physical level, a trained backbone network model, a trained physical layer adapter model, a splicing unit and a physical level splicing feature map output terminal, wherein the trained backbone network model outputs a first feature map trained in four stages, and the trained physical layer adapter model outputs a third feature map processed by a three-level adaptation network; during the training process, the parameters of the backbone network are kept unchanged, and the parameters of the adapter network are fine-tuned, and the adapter network includes 3 cascaded convolutional layers, an activation layer and a normalization layer.
9. The teaching learning system according to claim 7, characterized in that: The image classification teaching learning module and the instance segmentation teaching learning module both support the selection of static standards or dynamic extensions for labeled images.
Citation Information
Patent Citations
Active learning self-iteration image classification method and system
CN112560971A
Clustering method, clustering device and computer readable storage medium
CN113255841A
Zero sample instance segmentation method and system, readable storage medium and computer
CN117407557A
Intelligent data labeling method and device
CN118298250A
Target detection method, target detection model training method, target detection model training device and related equipment
CN118657930A