Sample processing method and processing system

By combining multiple technical means based on rules-based detection, model-assisted evaluation and manual review, a sample processing method was designed, which solved the problem of low data integration efficiency in the existing technology and achieved a significant improvement in the performance of deep learning models.

CN120071053AActive Publication Date: 2025-05-30TZTEK TECHNOLOGY CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510139073.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-05-30
Estimated Expiration
2045-02-08

AI Technical Summary

Technical Problem

The existing technology is inefficient when building large-scale high-quality labeled data sets, which affects the training efficiency, performance and generalization capabilities of deep learning models.

Method used

By combining rule-based detection, model-assisted evaluation and manual review, a sample processing method is designed, including high-value sample mining, low-value sample filtering and semi-automatic interactive inspection to improve data integration efficiency.

Benefits of technology

While greatly reducing the workload of data labeling, the most informative and representative samples are accurately screened out to provide high-quality data support for model training, thereby significantly improving the performance of deep learning models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071053A_ABST
    Figure CN120071053A_ABST
Patent Text Reader

Abstract

The invention provides a sample processing method and system, and belongs to the field of artificial intelligence and data processing, and the sample processing method comprises the steps of unlabeled high-value sample mining, low-value sample filtering and semi-automatic interactive inspection. The data processing system comprises a high-value sample mining module, a low-value sample filtering module, a semi-automatic interaction checking module, a data labeling module, a data storage module and a feedback optimization module. According to the method, the data labeling workload can be greatly reduced, samples with the most information content and representativeness can be accurately screened out, high-quality data support is provided for model training, the deep learning model performance is remarkably improved, and application and popularization in the field of data processing are facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of artificial intelligence and data processing, and particularly relates to a sample processing method and a processing system. Background Art

[0002] The successful application of deep learning models in actual projects depends on three key prerequisites: advanced deep network architectures, powerful computing platforms, and large-scale datasets with high-quality labels. Among them, the first two conditions are usually ready during the project preparation stage. Therefore, during the project implementation process, the focus is on constructing a large-scale high-quality labeled dataset.

[0003] In fact, the quality and quantity of data are often the key factors determining the success or failure of a project. Data integration requires not only data collection, classification, and annotation, but also data inspection, handling inconsistencies, evaluating, and correcting annotation errors. Common quality inspection methods include rule-based methods, model-assisted methods, and manual evaluation. Rule-based methods detect obvious errors through preset conditions. Model-assisted methods use various task models or pre-trained models to automatically verify labels and discover potential abnormal annotations. Manual evaluation relies on annotators for review. Relying solely on manual evaluation is less efficient. Therefore, data integration generally adopts a semi-automated approach, checking most data through automated methods, and manually evaluating samples with low confidence in model processing. Summary of the Invention

[0004] In order to overcome the deficiencies of the prior art, the purpose of the present invention is to provide a sample processing method and a processing system, which can solve the above problems.

[0005] Design Principle: To solve the problem of the integration efficiency of large-scale datasets and improve the training efficiency, performance, and generalization ability of deep learning models, the present invention combines various technical means such as rule-based detection, model-assisted evaluation, and manual review. The efficient data integration method mainly includes high-value sample mining, low-value sample filtering, and semi-automatic interactive inspection.

[0006] A sample processing method, the sample processing method includes the following steps: mining unlabeled high-value samples, selecting samples that contribute the most to improving the model performance through minimum confidence and maximum value sample mining; filtering low-value samples, initially screening similarities using embedding feature vectors and hash algorithms, and confirming the similarity image library through manual inspection to obtain a self-deduplicated dataset; semi-automatic interactive inspection: including consistency inspection and boundary confirmation, comparing the automatic segmentation with the manual annotation results, screening out the data with the largest inconsistencies, and dynamically screening boundary samples.

[0007] Preferably, the mining of unlabeled high-value samples includes: S11, using the minimum confidence method to mine uncertain samples, selecting samples with model output probability values ​​close to 0.5, and samples with the smallest difference between the model's largest and second largest probability predictions; S12, through maximum value sample mining, using DBScan clustering and model parameter updating, selecting samples that contribute most to improving model performance.

[0008] Preferably, low-value sample filtering includes: S21, cutting the original data set into standard images, and using embedding to obtain data set features; S22, taking images with close feature distribution as suspected image pairs, and further using a hash algorithm to perform preliminary screening for similarity to obtain a batch of similar data sets; S23, manually checking similar data sets, moving images with higher similarity to a similarity image library, and using dissimilar images as self-deduplication data sets; S24, the self-deduplication data sets enter the search system, and perform similarity search on the data sets to be stored with the data sets already existing in the standard image library to obtain suspected image groups; S25, manually checking and screening suspected image groups, moving similar images to the similarity image library, and adding dissimilar images to the processed standard image library.

[0009] Preferably, the semi-automatic interactive check includes: S31, consistency check: the model obtained by training part of the data is used to make predictions on other data, and the data is screened for manual review based on the missed and over-detected targets; the results are compared using the automatic segmentation method and manual labeling, that is, the intersection and union ratio of the automatic segmentation and manually labeled targets is compared, and sorting and screening are performed according to the maximum inconsistency; S32, boundary confirmation: boundary samples are dynamically screened based on the embedding distribution, and the deviation is adjusted.

[0010] Preferably, for boundary confirmation, dynamic sample boundary screening is adopted, and the implementation steps include: S41. Given an image or target dataset containing two categories, calculate the first feature vector set (embeddings_1) and the second feature vector set (embeddings_2) of all images or targets in the two categories; S42. Search using the principle of maximum and minimum distances. Find the first farthest data in the first feature vector set relative to the overall second feature vector set, denoted as embedding_fps_1. Similarly, find the second farthest data in the second feature vector set relative to the overall first feature vector set, denoted as embedding_fps_2; S43. Calculate the similarity of all feature vector sets embeddings (embeddings_1 and embeddings_2) relative to the farthest data (embedding_fps_1 and embedding_fps_2), and sort the similarities; S44. Given similarity thresholds for the first and second farthest data, and perform nearest neighbor retrieval on the data near the similarity thresholds according to the similarity sorting results, and select a fixed number of samples not greater than a certain value on both sides for display or download.

[0011] The present invention also provides a data processing system, which includes: a high-value sample mining module that selects samples that contribute the most to improving the model performance through the minimum confidence method and the maximum value sample mining technology; a low-value sample filtering module that performs preliminary similarity screening using feature vector embedding feature extraction and a hashing algorithm, and confirms images with relatively high similarity through manual inspection; a semi-automatic interactive inspection module that screens data through consistency inspection and boundary confirmation and dynamically adjusts samples.

[0012] Preferably, the data processing system further includes: a data annotation module for annotating high-value samples; a data storage module for storing the screened samples and annotation results; a feedback optimization module for optimizing the similarity calculation and automatic duplicate removal algorithm according to user feedback.

[0013] Preferably, the low-value sample filtering module includes: a cropping set duplicate removal unit and a pre-warehouse comparison with the gallery set duplicate removal unit to achieve complete data cleaning. Among them, the cropping set duplicate removal unit includes input of data to be cleaned and feature extraction, similarity calculation, automatic duplicate removal and manual inspection, result storage and feedback optimization, stores the results confirmed by manual inspection in the similarity image library, and collects user feedback to continuously optimize the algorithm model for similarity calculation and automatic duplicate removal; among them, the pre-warehouse comparison with the gallery set duplicate removal unit includes feature extraction and search matching, manual confirmation, and updating the standard gallery.

[0014] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention can greatly reduce the workload of data annotation while accurately screening out the most informative and representative samples, providing high-quality data support for the training of the model, thereby significantly improving the performance of the deep learning model. Description of the Drawings

[0015] Figure 1 It is a schematic diagram of the partial process framework for sample processing or data cleaning of the present invention. Detailed Embodiments

[0016] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0017] Sample processing method A sample processing method, the sample processing method includes the following steps.

[0018] Mining unlabeled high-value samples, selecting the samples that contribute the most to the improvement of the model performance through minimum confidence and maximum value sample mining; the role is to find high-value samples.

[0019] Filtering low-value samples, using feature vector embedding (feature vector, which represents an image target here and is well-known in the field) and hash algorithm for preliminary similarity screening, and obtaining a self-duplicate-removing data set through manual inspection to confirm the similarity image library; the role is to remove low-value samples.

[0020] Semi-automatic interactive inspection: including consistency inspection and boundary confirmation, screening out the data with the largest inconsistency by comparing the automatic segmentation with the manual annotation results, and dynamically screening boundary samples. The role is to manually screen the obtained samples to remove unreasonable annotation samples, etc. In addition, evaluate whether the consistency of the annotation results is reasonable.

[0021] Among them, for mining unlabeled high-value samples, the contribution degrees of different data samples to the model performance are different. By carefully selecting the most representative and informative samples for annotation, the workload of annotation can be reduced while training an efficient and superior-performance model. Here, multiple methods are used for mining high-value samples. Specific examples are as follows.

[0022] S11. Use the minimum confidence method to mine uncertain samples, and select samples with model output probability values close to 0.5 and samples with the smallest difference between the maximum and the second-largest predicted probabilities of the model. Specifically, select those samples that make the model uncertain (i.e., several samples with model output probability values closest to 0.5) and samples with the smallest difference between the maximum and the second-largest predicted probabilities of the model.

[0023] S12. Through the mining of the most valuable samples, use DBScan clustering and model parameter update to select the samples that contribute the most to the improvement of the model performance. Specifically, by calculating the contribution to labeling a certain sample under the current model, select those samples that are most likely to improve the model performance for labeling. The specific method is as follows:

[0024] a) Use the DBScan method to cluster the unlabeled data to obtain an unlabeled clustered sample set {x}; b) Use the current model M to predict each sample x in the unlabeled clustered sample set x to obtain a predicted label y; c) Assume that the true label of x is y, retrain the model M or update the model parameters to obtain a new model M'; d) Extract one sample from each class in the clustered sample set, compare the output differences between M and M', and the cosine similarity can be used for measurement, and select the sample set x that causes the largest change in the model.

[0025] For the filtering of low-value samples, extremely similar data samples have low value. If there are a large number of them, it will hinder the learning efficiency and expand the impact of sample imbalance. In a specific example, see Figure 1 , including the following steps.

[0026] S21. Cut the original data set into standard images, and use embedding (feature vector, which represents the image target here and is well-known in the field) to obtain the data set features.

[0027] S22. Take the images with relatively close feature distributions as suspected image pairs, and further use the hash algorithm to conduct a preliminary screening of the similarity to obtain a batch of similar data sets.

[0028] S23. Manually check the similar data sets, move the images with higher similarity to the similarity image library, and use the dissimilar images as the self-duplicate-removing data set.

[0029] S24. The self-duplicate-removing data set enters the search system, and the data set to be stored in the library is searched for similarity with the data sets already existing in the standard image library to obtain a suspected picture group.

[0030] S25. Manually check and screen the suspected picture group, move the similar images to the similarity image library, and add the dissimilar images to the processed standard image library.

[0031] Among them, semi-automatic interactive inspection includes the following steps.

[0032] S31. Consistency check: The model trained with part of the data is used to make predictions on other data, and data is screened for manual review based on missed detections and over-detections of the target; the results are compared using an automatic segmentation method and manual annotation, that is, the intersection over union of the automatically segmented and manually annotated targets is compared, and sorting and screening are performed according to the maximum inconsistency.

[0033] S32. Boundary confirmation: Dynamically screen boundary samples based on the embedding distribution and adjust the deviation.

[0034] For boundary confirmation, when adjusting a certain class, boundary samples are dynamically displayed, and the main reference for the boundary samples is the deviation from the class center. Here, the deviation is mainly adjusted according to the embedding distribution of a certain two classes or a certain class of NG / OK targets, and boundary samples are dynamically screened. The boundary confirmation adopts dynamic screening of sample boundaries, and the steps are as follows.

[0035] S41. Given an image or target dataset containing two classes (the target dataset is the dataset after cropping, so both can be regarded as classification datasets), calculate the embeddings of all images or targets of the two classes: the first feature vector set (embeddings_1) and the second feature vector set (embeddings_2).

[0036] S42. Search using the principle of maximum and minimum distance. Find the first farthest data in the first feature vector set relative to the overall second feature vector set, denoted as embedding_fps_1. Similarly, find the second farthest data in the second feature vector set relative to the overall first feature vector set, denoted as embedding_fps_2; specifically, find the data that is the farthest from embeddings_2 as a whole in embeddings_1, and search using the principle of maximum and minimum distance, denoted as embedding_fps_1; find the data that is the farthest from embeddings_1 as a whole in embeddings_2, and search using the principle of maximum and minimum distance, denoted as embedding_fps_2.

[0037] S43. Calculate the similarity of all feature vector sets embeddings (embeddings_1 and embeddings_2) with respect to the farthest data (embedding_fps_1 and embedding_fps_2), and sort the similarities. Specifically, calculate the cosine similarity of all embeddings with respect to embedding_fps_1, denoted as sim_1; calculate the cosine similarity of all embeddings with respect to embedding_fps_2, denoted as sim_2. Calculate sim_1 - sim_2, sort it, and adjust the range through linear transformation to -0.5 to 0.5.

[0038] S44. Given the similarity threshold for the first and second farthest data, and perform nearest neighbor retrieval on the data near the similarity threshold according to the similarity sorting result, and take a fixed number of samples not greater than it on both sides for display or download.

[0039] Data processing system A data processing system, which includes: a high-value sample mining module, a low-value sample filtering module, a semi-automatic interaction inspection module, a data annotation module, a data storage module, and a feedback optimization module.

[0040] The high-value sample mining module selects the samples that contribute the most to the improvement of the model performance through the minimum confidence method and the maximum value sample mining technology; it is used to find high-value samples.

[0041] The low-value sample filtering module, see Figure 1 , uses the embedding feature extraction of the feature vector and the hash algorithm to perform a preliminary screening of similarity, and confirms the images with relatively high similarity through manual inspection; it is used to remove low-value samples. Specifically, the low-value sample filtering module includes: a duplicate removal unit for the cropping set and a duplicate removal unit for comparing with the gallery set before storage to achieve complete data cleaning; the duplicate removal unit for the cropping set includes the input of data to be cleaned and feature extraction, similarity calculation, automatic duplicate removal and manual inspection, result storage and feedback optimization, stores the results confirmed by manual inspection in the similarity image library, and collects user feedback to continuously optimize the algorithm models of similarity calculation and automatic duplicate removal. The duplicate removal unit for the cropping set screens suspected duplicate image pairs through embedding feature extraction and the hash algorithm, and combines manual review to optimize the cleaning effect. The duplicate removal unit for comparing with the gallery set before storage includes feature extraction and search matching, manual confirmation, and updating the standard gallery. The duplicate removal unit for comparing with the gallery before storage ensures the uniqueness of the data to be stored through the embedding feature comparison between the standard image library and the dataset to be stored.

[0042] The cropping set deduplication unit mainly targets the dataset to be cleaned and is implemented through the following steps: (1)Data input and feature extraction: Input the data to be cleaned into the system, and extract the feature vectors of the images through the embedding algorithm, which serves as the basis for subsequent processing.

[0043] (2)Similarity calculation: Use the hashing algorithm to perform fast similarity calculation on the image features, generate suspected duplicate image pairs, and reduce the computational complexity.

[0044] (3)Automatic deduplication and manual inspection: Use an automated algorithm to automatically screen the suspected duplicate image pairs, delete most of the duplicate data; for the part with low confidence, conduct a review through manual inspection to ensure the accuracy of the results.

[0045] (4)Result storage and feedback optimization: Store the results confirmed by manual inspection in the similarity image library, collect user feedback, and continuously optimize the algorithm models for similarity calculation and automatic deduplication.

[0046] Summary: The cropping set deduplication unit targets the dataset to be cleaned, generates feature vectors through embedding feature extraction, combines the hashing algorithm to calculate image similarity, screens out possible duplicate image pairs, uses automatic deduplication to process most of the duplicate data, and hands over the part with low confidence to manual inspection for review. Finally, store the confirmed results in the similarity image library, and optimize the cleaning effect through the user feedback mechanism.

[0047] The deduplication unit for comparison with the image library before storage, before the data is stored in the library, performs deduplication by comparison using the following steps: (1)Feature extraction and search matching: Extract the feature vectors of the data to be stored and the data in the standard image library through the embedding algorithm, and use the search system to quickly match the two to screen out suspected duplicate image groups.

[0048] (2)Manual confirmation: For the suspected duplicate groups marked by the search system, confirm through manual inspection whether they are duplicate data, so as to ensure the uniqueness and integrity of the data stored in the library.

[0049] (3)Update the standard image library: Update the non-duplicate data confirmed by manual inspection to the standard image library to improve the coverage of the library and provide a more perfect benchmark for subsequent data cleaning.

[0050] Summary: The deduplication unit for comparison with the image library set before storage, before the data is stored in the library, through the embedding feature comparison between the standard image library and the dataset to be stored, uses the search system to screen out suspected duplicate image groups, ensures the uniqueness of the data stored in the library after manual inspection confirmation, and updates the new data to the standard image library to improve the data coverage and quality.

[0051] Semi-automatic interactive inspection module, which screens data and dynamically adjusts samples through consistency check and boundary confirmation. Specifically, the semi-automatic interactive inspection module includes a consistency check unit and a boundary confirmation unit. Among them, the consistency check unit is used to predict other data by a model trained with part of the data, and screen data for manual review according to the missed detection and over-detection of the target; among them, the boundary confirmation unit is used to dynamically screen boundary samples based on the embedding distribution and adjust the deviation degree.

[0052] By combining automated algorithms with manual review, the present invention realizes an efficient and accurate semi-automated data cleaning process, which can effectively improve the quality and reliability of the training dataset of the deep learning model.

[0053] Data annotation module, used to annotate high-value samples.

[0054] Data storage module, used to store the screened samples and annotation results, including a deduplicated dataset storage unit, a standard image library storage unit, a dataset to be warehoused storage unit, a suspected picture group storage unit, a similarity image library storage unit, etc.

[0055] Feedback optimization module, used to optimize the similarity calculation and automatic deduplication algorithm according to user feedback.

[0056] Through the above methods and systems, an efficient semi-automated data cleaning process can be realized, while ensuring the high quality and high reliability of the data. Compared with the traditional method that relies on manual evaluation, the present invention improves the efficiency while significantly reducing the workload of manual intervention. In addition, by combining user feedback and dynamically updating the standard image library, the present invention has the ability to continuously optimize and adapt to the characteristics of different datasets, and can be widely applied to the preprocessing stage of deep learning model training data.

[0057] After actual tests, the data annotation category accuracy of the present invention is greater than 99%, the edge consistency is greater than 98%, and the data integration efficiency is increased by 50%.

[0058] Computer-readable storage medium The present invention also provides a computer-readable storage medium, on which computer instructions are stored, and when the computer instructions run, they execute the steps of the foregoing processing method. Among them, for the method, please refer to the detailed introduction in the foregoing part, and will not be elaborated here.

[0059] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium. Computer-readable media include permanent and non-permanent, removable and non-removable media and can be implemented by any method or technology for information storage. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media do not include transitory media such as modulated data signals and carrier waves.

[0060] Terminal The present invention also provides a terminal, including a memory and a processor. The memory stores data provider information and computer instructions that can run on the processor. When the processor runs the computer instructions, it executes the steps of the foregoing method. Among them, for the foregoing method, please refer to the detailed introduction in the foregoing part and will not be elaborated herein.

[0061] In addition, those skilled in the art can understand that various aspects of the present application can be described and illustrated by several patentable types or situations, including any new and useful process, machine, product, or composition of matter, or any new and useful improvement thereof. Accordingly, various aspects of the present application can be executed entirely by hardware, entirely by software (including firmware, resident software, microcode, etc.), or by a combination of hardware and software. The above hardware or software can all be referred to as "data blocks", "modules", "engines", "units", "components", or "systems". In addition, various aspects of the present application may be embodied as a computer product located in one or more computer-readable media, which includes computer-readable program codes.

[0062] A computer storage medium may include a propagated data signal that contains computer program code, for example, on a baseband or as part of a carrier wave. The propagated signal may have various forms of manifestation, including electromagnetic form, optical form, etc., or a suitable combination of forms. A computer storage medium can be any computer-readable medium other than a computer-readable storage medium, and this medium can be connected to an instruction execution system, device, or equipment to implement communication, propagation, or transmission of a program for use. The program code located on the computer storage medium can be propagated through any suitable medium, including radio, cable, fiber optic cable, RF, or similar media, or any combination of the above media.

[0063] The computer program code required for the operation of each part of this application can be written in any one or more programming languages, including object-oriented programming languages such as Java, Scala, Smalltalk, Eiffel, JADE, Emerald, C++, C#, VB.NET, Python, etc., conventional procedural programming languages such as C language, VisualBasic, Fortran2003, Perl, COBOL2002, PHP, ABAP, dynamic programming languages such as Python, Ruby, and Groovy, or other programming languages. The program code can run entirely on the user's computer, or run on the user's computer as an independent software package, or run partially on the user's computer and partially on a remote computer, or run entirely on a remote computer or processing device. In the latter case, the remote computer can be connected to the user's computer through any network form, such as a local area network (LAN) or a wide area network (WAN), or connected to an external computer (for example, through the Internet), or in a cloud computing environment, or used as a service such as software as a service (SaaS).

[0064] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A sample processing method, characterized in that: The sample processing method includes the following steps: Unlabeled high-value sample mining: by mining the minimum confidence and maximum value samples, we select the samples that contribute most to improving model performance. Low-value samples are filtered, using embedding feature vectors and hash algorithms for initial similarity screening, and similar image libraries are confirmed through manual inspection to obtain a self-deduplication dataset; Semi-automatic interactive checking: including consistency checking and boundary confirmation. By comparing automatic segmentation with manual annotation results, the data with the greatest inconsistency can be screened out, and boundary samples can be dynamically screened.

2. The sample processing method according to claim 1, characterized in that: The unlabeled high-value sample mining includes: S11. Use the minimum confidence method to mine uncertain samples, select samples whose model output probability value is close to 0.5, and samples whose model prediction has the largest probability and the smallest difference between the second largest probability; S12. Through maximum value sample mining, using DBScan clustering and model parameter updating, select the samples that contribute most to improving model performance.

3. The sample processing method according to claim 1, characterized in that: The low-value sample filtering includes: S21, the original data set is cut into standard images, and the data set features are obtained using embedding; S22, taking images with close feature distribution as suspected image pairs, and further using a hash algorithm to preliminarily screen the similarity to obtain a batch of similar data sets; S23, manually check the similar data sets, move the images with higher similarity to the similarity image library, and use the dissimilar images as the self-deduplication data sets; S24, the self-deduplicated data set enters the search system, and the data set to be stored is searched for similarity with the data set already existing in the standard image library to obtain a suspected image group; S25, manually checking and screening the suspected image group, moving similar images to the similarity image library, and adding dissimilar images to the standard image library.

4. The sample processing method according to claim 1, characterized in that: Semi-automatic interactive checks include: S31, consistency check: The model obtained by training part of the data is used to make predictions on other data, and the data is screened for manual review based on the missed and over-detected targets; the results are compared using the automatic segmentation method and the manual annotation method, that is, the intersection and union ratio of the automatic segmentation and the manual annotation targets are compared, and sorting and screening are performed based on the maximum inconsistency; S32, boundary confirmation: dynamically filter boundary samples based on embedding distribution and adjust the deviation.

5. The sample processing method according to claim 4, characterized in that: Boundary confirmation uses dynamic screening of sample boundaries, and the execution steps include: S41, given an image or target data set containing two categories, calculating a first feature vector set and a second feature vector set of all images or targets of the two categories; S42, searching by using the principle of maximum and minimum distance, searching in the first feature vector set for the first farthest data relative to the second feature vector set as a whole, recorded as embedding_fps_1, and similarly searching in the second feature vector set for the second farthest data relative to the first feature vector set as a whole, recorded as embedding_fps_2; S43, calculating the similarity of all feature vector sets embeddings relative to the first farthest data and the second farthest data, and sorting the similarities; S44, a similarity threshold is given relative to the first farthest data and the second farthest data, and a nearest neighbor search is performed on the data near the similarity threshold according to the similarity sorting result, and samples on both sides that are not greater than a fixed number are selected for display or download.

6. A data processing system, characterized in that: The data processing system includes: High-value sample mining module, which selects samples that contribute most to model performance improvement through the minimum confidence method and maximum value sample mining technology; used to find high-value samples; The low-value sample filtering module uses feature vector embedding feature extraction and hashing algorithm to perform a preliminary similarity screening, and confirms images with high similarity through manual inspection; it is used to remove low-value samples; Semi-automatic interactive checking module that screens data and dynamically adjusts samples through consistency checking and boundary confirmation.

7. The data processing system according to claim 6, characterized in that: The data processing system also includes: Data annotation module, used to annotate high-value samples; Data storage module, used to store screened samples and annotation results; Feedback optimization module, used to optimize similarity calculation and automatic deduplication algorithm based on user feedback.

8. The data processing system according to claim 6, characterized in that: The low-value sample filtering module includes: the deduplication unit of the cropping set and the deduplication unit of the comparison with the gallery set before storage to achieve complete data cleaning. The cut set deduplication unit includes the input and feature extraction of the data to be cleaned, similarity calculation, automatic deduplication and manual inspection, result storage and feedback optimization. The results confirmed by manual inspection are stored in the similarity image library, and user feedback is collected to continuously optimize the algorithm model of similarity calculation and automatic deduplication. Among them, the deduplication unit compared with the gallery set before storage includes feature extraction and search matching, manual confirmation, and updating of the standard gallery.

Citation Information

Patent Citations

  • SVM active learning classification algorithm for large-scale training data

    CN104331716A

  • Labelled sample maintaining method and two-way learning interactive classification method

    CN104657745A

  • Data processing method and device and computer equipment

    CN110163265A

  • Intelligent weld defect identification model evolution method based on active learning

    CN116258175A

  • Kitchen garbage target detection method based on active learning selection strategy

    CN118429625A