Sample labeling method, classification model training method, object classification method and object searching method
By obtaining the classification prediction results of the samples to be marked, selecting target samples with high similarity and using a large language model for annotation, the confusion problem in the automated annotation method is solved, the accuracy and consistency of the labeling are improved, sample selection is optimized, and the adaptability and reliability of the system are enhanced.
Patent Information
- Application Number
- CN202510492635.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-08-01
AI Technical Summary
In the prior art, the automated labeling method has confusion and labeling errors when selecting samples, resulting in insufficient accuracy of model labeling, and the traditional manual labeling process is costly and inefficient, making it difficult to quickly respond to changes in business needs.
By obtaining the classification prediction results of the samples to be marked, selecting the target samples based on similarity, and using a large language model for annotation, combining sample pool and labeling instructions, the sample selection and labeling process is optimized.
It improves the accuracy and consistency of labeling, reduces confusion caused by sample complexity and diversity, enhances the adaptability and reliability of the system, and reduces the need for manual intervention.
Smart Images

Figure CN120408272A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of machine learning technology, and particularly to a method for sample annotation, classification model training, object classification, and object search. Background Art
[0002] In the related field of automatically annotating data, the traditional manual annotation is inefficient and costly. The few-shot prompting technology has become the mainstream choice. By embedding a small number of labeled examples in the prompt words, it guides the large model to follow the established rules for accurate annotation.
[0003] To ensure the effectiveness of the annotation process, it is crucial to select appropriate examples. Appropriate examples can not only help the model understand the annotation rules but also improve its annotation accuracy for new samples. However, in the related technologies, due to the complexity and diversity of the example content, even if seemingly appropriate examples are selected, it is still possible for the model to be confused and make annotation errors.
[0004] Therefore, there is an urgent need for a more accurate method for sample annotation. Summary of the Invention
[0005] In view of this, the embodiments of this specification provide a method for sample annotation. One or more embodiments of this specification also relate to a method for training a classification model, a method for object classification, a method for object search, a sample annotation device, a classification model training device, an object classification device, an object classification device, an object search system, a computing device, a computer-readable storage medium, and a computer program product to solve the technical defects existing in the prior art.
[0006] According to the first aspect of the embodiments of this specification, a method for sample annotation is provided, including:
[0007] Obtain a sample to be annotated;
[0008] Perform classification prediction on the sample to be annotated to obtain a first classification result corresponding to the sample to be annotated;
[0009] Obtain second classification results corresponding to multiple labeled samples, and determine a target sample from the multiple labeled samples based on the similarity between the first classification result and each second classification result;
[0010] Input the target sample and the sample to be annotated into an annotation model to obtain an annotation result of the sample to be annotated.
[0011] According to the second aspect of the embodiments of this specification, a method for training a classification model is provided, including:
[0012] Obtain an initial classification model, multiple training samples, and label information of the training samples. Among them, the label information of the multiple training samples includes the annotation results obtained by using the sample annotation method described in the first aspect;
[0013] Train the initial classification model based on the multiple training samples and the label information to obtain a target classification model.
[0014] According to the third aspect of the embodiments of this specification, an object classification method is provided, including:
[0015] Obtain an object classification request, where the object classification request includes object data of a target object;
[0016] Input the object data into the target classification model to determine the classification result of the target object, where the target classification model is trained based on the object classification model training method described in the second aspect.
[0017] According to the fourth aspect of the embodiments of this specification, an object search method is provided, including:
[0018] Obtain an object search request, where the object search request includes a search statement;
[0019] Input the search statement into the target classification model to determine the search type corresponding to the search statement, where the target classification model is trained based on the object classification model training method described in the second aspect;
[0020] Based on the search type, call the target search service to obtain a search result.
[0021] According to the fifth aspect of the embodiments of this specification, a sample annotation device is provided, including:
[0022] A first acquisition module configured to acquire samples to be annotated;
[0023] A prediction module configured to perform classification prediction on the samples to be annotated to obtain a first classification result corresponding to the samples to be annotated;
[0024] A determination module configured to acquire second classification results corresponding to multiple annotated samples, and determine a target sample from the multiple annotated samples based on the similarity between the first classification result and each second classification result;
[0025] An annotation module configured to input the target sample and the samples to be annotated into an annotation model to obtain an annotation result of the samples to be annotated.
[0026] According to the sixth aspect of the embodiments of this specification, a classification model training device is provided, including:
[0027] A second acquisition module, configured to acquire an initial classification model, a plurality of training samples, and label information of the training samples, wherein the label information of the plurality of training samples includes an annotation result obtained by using the sample annotation method described in the first aspect;
[0028] A training module, configured to train the initial classification model based on the plurality of training samples and the label information to obtain a target classification model.
[0029] According to a seventh aspect of the embodiments of the present specification, there is provided an object classification device, including:
[0030] A third acquisition module, configured to acquire an object classification request, wherein the object classification request includes object data of a target object;
[0031] A first classification module, configured to input the object data into the target classification model to determine a classification result of the target object, wherein the target classification model is trained based on the object classification model training method described in the second aspect.
[0032] According to an eighth aspect of the embodiments of the present specification, there is provided an object search device, including:
[0033] A fourth acquisition module, configured to acquire an object search request, wherein the object search request includes a search statement;
[0034] A second classification module, configured to input the search statement into the target classification model to determine a search type corresponding to the search statement, wherein the target classification model is trained based on the object classification model training method described in the second aspect;
[0035] A search module, configured to call a target search service based on the search type to obtain a search result.
[0036] According to a ninth aspect of the embodiments of the present specification, there is provided an object search system, including:
[0037] A request interface, configured to acquire an object search request, wherein the object search request includes a search statement;
[0038] A classification unit, configured to input the search statement into the target classification model to determine a search type corresponding to the search statement, wherein the target classification model is trained based on the object classification model training method described in the second aspect;
[0039] A search unit, configured to call a target search service based on the search type to obtain a search result;
[0040] The request interface is further configured to feedback the search result.
[0041] According to a tenth aspect of the embodiments of the present specification, there is provided a computing device, including:
[0042] A memory and a processor;
[0043] The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the sample annotation method described in the first aspect above, or the classification model training method described in the second aspect above, or the object classification method described in the third aspect above, or the object search method described in the fourth aspect above are implemented.
[0044] According to the eleventh aspect of the embodiments of the present specification, a computer-readable storage medium is provided, which stores computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps of the sample annotation method described in the first aspect above, or the classification model training method described in the second aspect above, or the object classification method described in the third aspect above, or the object search method described in the fourth aspect above are implemented.
[0045] According to the twelfth aspect of the embodiments of the present specification, a computer program product is provided, including computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps of the sample annotation method described in the first aspect above, or the classification model training method described in the second aspect above, or the object classification method described in the third aspect above, or the object search method described in the fourth aspect above are implemented.
[0046] An embodiment of the present specification realizes obtaining a sample to be annotated; performing classification prediction on the sample to be annotated to obtain a first classification result corresponding to the sample to be annotated; obtaining second classification results corresponding to multiple annotated samples, and determining a target sample from the multiple annotated samples based on the similarity between the first classification result and each second classification result; inputting the target sample and the sample to be annotated into an annotation model to obtain an annotation result of the sample to be annotated. By using a classification model to obtain the classification prediction result of the sample to be annotated and screening according to the similarity between the sample to be annotated and the annotated samples, it is ensured that the selected target sample has a high similarity with the sample to be annotated at the classification prediction level, thereby reducing the confusion caused by the complexity and diversity of the sample content. On this basis, by further processing the target sample and the sample to be annotated through the annotation model, a final annotation result is generated, significantly improving the accuracy and consistency of the annotation. This method not only optimizes the example selection but also enhances the adaptability and reliability of the system. Description of the Drawings
[0047] Figure 1 is a flowchart of a sample annotation method provided by an embodiment of the present specification;
[0048] Figure 2 is a flowchart of a classification model training method provided by an embodiment of the present specification;
[0049] Figure 3 It is a flowchart of an object classification method provided by an embodiment of this specification;
[0050] Figure 4 It is a flowchart of an object search method provided by an embodiment of this specification;
[0051] Figure 5 It is a process architecture diagram of a sample annotation method provided by an embodiment of this specification;
[0052] Figure 6 It is a processing flowchart of a sample annotation method provided by an embodiment of this specification;
[0053] Figure 7 It is a front-end schematic diagram of an object search system provided by an embodiment of this specification;
[0054] Figure 8 It is a structural schematic diagram of an object search system provided by an embodiment of this specification;
[0055] Figure 9 It is a structural schematic diagram of a sample annotation device provided by an embodiment of this specification;
[0056] Figure 10 It is a structural schematic diagram of a classification model training device provided by an embodiment of this specification;
[0057] Figure 11 It is a structural schematic diagram of an object classification device provided by an embodiment of this specification;
[0058] Figure 12 It is a structural schematic diagram of an object search device provided by an embodiment of this specification;
[0059] Figure 13 It is a structural block diagram of a computing device provided by an embodiment of this specification. Detailed implementation manners
[0060] In the following description, many specific details are set forth in order to provide a thorough understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of this specification. Therefore, this specification is not limited by the specific implementations disclosed below.
[0061] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and encompasses any or all possible combinations of one or more of the associated listed items.
[0062] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0063] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for the user to choose to authorize or refuse.
[0064] First, the noun terms involved in one or more embodiments of this specification are explained.
[0065] Perplexity: It is an indicator to measure the prediction ability of a model and is often used to evaluate the performance of a language model or a classification model. Essentially, it is a measure of the uncertainty of the model's prediction distribution. The lower the perplexity, the more accurate the model's prediction of the data; the higher the perplexity, the greater the uncertainty of the model and the worse the prediction effect. In a broader application, such as a classification task, perplexity can be understood as a form of the entropy of the probability distribution output by the model. For a classification problem, if the model assigns a high confidence to the correct category for each sample (i.e., the predicted probability distribution is concentrated on the correct label), then its perplexity will be lower. On the contrary, if the predicted probability distribution of the model is relatively uniform, indicating that the model lacks confidence in classifying the sample, the perplexity will be higher at this time.
[0066] In the context of AI search business scenarios, accurately identifying and classifying users' query intents is a crucial step in providing high-quality search services. This classification not only determines whether to initiate the AI search service but also affects subsequent answer generation strategies, directly related to the improvement of the user search experience and the optimization of search efficiency. With the development of AI technology, understanding and classifying users' query intents have become one of the core challenges in improving the quality of search services.
[0067] Traditional approaches rely on large-scale manually annotated data to train classification models. This process includes formulating detailed annotation specifications, recruiting and training professional annotators, performing multiple rounds of annotation and quality inspection, and handling annotation disputes and boundary cases. In contrast, using large language models (LLMs) for automated annotation demonstrates significant advantages: rapid deployment, high annotation efficiency, controllable costs, good standard consistency, and easy scalability to adapt to new annotation requirements. Currently, mainstream large model automated annotation methods mostly adopt few-shot prompting techniques, that is, by embedding a small number of annotated examples in the prompt words, expecting the model to refer to these examples and follow established rules for accurate annotation.
[0068] Although large language models perform well in automated annotation, there is still a certain gap in accuracy compared to manual annotation. Most existing example selection methods are based on the semantic similarity between the samples to be annotated and the samples in the example pool. The core idea is to select the examples with the closest semantics to assist in annotation. However, this method has limitations: samples with similar semantics may correspond to different annotation results, leading to model confusion and annotation errors. In addition, the traditional manual annotation process has a long cycle and high costs, making it difficult to quickly respond to changes in business requirements. Although the automatic annotation of large models improves efficiency, there is still room for improvement in maintaining annotation quality.
[0069] To address the above problems, in this specification, a sample annotation method is provided. One or more embodiments of this specification are also related to a classification model training method, an object classification method, an object search method, a sample annotation device, a classification model training device, an object classification device, an object classification device, an object search system, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail one by one in the following embodiments.
[0070] See Figure 1 , Figure 1 is a flowchart of a sample annotation method provided by an embodiment of this specification, specifically including the following steps.
[0071] Step 102: Obtain the sample to be annotated.
[0072] A sample to be labeled refers to a data instance that the system has not yet classified or tagged. These data can be in various forms such as text, images, audio, etc., depending on the application scenario. For example, in a text classification task, the samples to be labeled can be a series of unclassified user comments; in an image recognition task, they can be pictures without labeled categories.
[0073] In practical applications, the process of obtaining samples to be labeled involves collecting and preparing data from various sources to ensure that the data is suitable for subsequent labeling and training processes. The system first determines the data type and format to be labeled, and then extracts data from the corresponding data sources. These data sources can include, but are not limited to, databases, file systems, web services, API interfaces, etc. To ensure the effectiveness and accuracy of subsequent labeling work, the system also needs to preprocess the collected data, such as cleaning, deduplication, format conversion, etc., to eliminate noise and inconsistencies and ensure data quality.
[0074] For the step of obtaining samples to be labeled, an optional method is to extract a portion of unlabeled data from an existing large-scale dataset as samples to be labeled; another optional implementation method is to capture newly generated data in real time and perform preliminary processing immediately to ensure the freshness and relevance of the data. Whichever method is used, the system needs to ensure that the obtained samples can represent the feature distribution of the target domain, thereby improving the effect of labeling and training.
[0075] In the embodiments of this specification, through a carefully designed acquisition mechanism, the system can efficiently collect high-quality samples to be labeled from multiple channels. This process not only ensures the diversity and representativeness of the data, but also lays a solid foundation for subsequent labeling work, helping to improve the generalization ability and prediction accuracy of the final model.
[0076] Exemplarily, assume that the system aims to develop a social media content review tool for automatically identifying inappropriate content. To obtain samples to be labeled, the system extracts a large number of un-reviewed posts from the logs of the social platform as the initial dataset. These posts cover a variety of languages and topics, ensuring broad representativeness of the data. Next, the system preprocesses these posts, removes duplicates and irrelevant information, and at the same time converts the text content into a unified format. In addition, the system also filters out the posts that are most likely to contain sensitive content based on the timestamp and interaction volume of the posts as the objects to be labeled first. In this way, the system not only obtains a large number of samples to be labeled, but also ensures the quality and relevance of these samples, providing strong support for subsequent labeling and model training.
[0077] Step 104: Perform classification prediction on the samples to be labeled to obtain the first classification result corresponding to the samples to be labeled.
[0078] Classification prediction refers to the process of using a trained machine learning model to analyze new data (i.e., samples to be labeled) to determine the category or label to which it belongs. The result of classification prediction is that the system assigns one or more possible category labels to each sample to be labeled based on the model's learning and understanding. For example, in a text classification task, classification prediction can categorize a blog post into different topics such as technology, entertainment, and health; in an image recognition task, it can identify the types of objects in an image, such as cats, dogs, cars, etc.
[0079] The first classification result is the preliminary classification label generated by the system for the sample to be labeled through the classification prediction process. This result reflects the model's current understanding and judgment of the sample, but has not yet been finalized or manually reviewed. The first classification result is the basis for subsequent sample selection and further processing, so its accuracy and reliability are crucial.
[0080] In practical applications, classifying and predicting the classification of unlabeled samples involves several key steps. First, the system loads a pre-trained initial classification model, which has the ability to understand the specific task. Then, the system inputs the unlabeled sample into the classification model, which calculates the probability distribution of each possible category based on the sample's feature vector. Finally, the system selects the most likely category as the first classification result based on these probability distributions. To ensure the accuracy of the classification prediction, the system can also adopt ensemble learning methods to combine the prediction results of multiple models, or use cross-validation techniques to optimize model performance.
[0081] For the step of classifying and predicting the unlabeled samples, one option is to directly use pre-trained deep learning models, such as convolutional neural networks (CNNs) or recurrent neural networks (RNNs). These models have been fully trained on large-scale datasets and have strong generalization capabilities. Another option is to use lightweight models, such as decision trees or support vector machines (SVMs). Although these models are less complex, they can also provide fast and accurate prediction results for specific tasks. Regardless of which method is used, the system must ensure that the classification prediction process is efficient and reliable to meet the needs of actual application scenarios.
[0082] In the examples of this specification, by performing classification predictions on the samples to be labeled, the system can generate a first classification result for each sample. This not only provides a preliminary category label but also lays the foundation for subsequent sample selection and labeling optimization. This process significantly improves the efficiency and quality of labeling, reduces the need for manual intervention, and enhances the system's automation.
[0083] Exemplarily, assume that the system aims to develop a news recommendation engine for pushing personalized news according to user interests. To classify and predict the samples to be labeled, the system first loads a pre-trained text classification model that has been fully trained on a large corpus containing various news categories. Next, the system takes the latest news articles collected from social platforms, news websites, etc. as samples to be labeled and inputs them one by one into the classification model. Based on the content features of each article, such as keywords, sentence structure, sentiment tendency, etc., the model calculates the probability distribution of each news category and assigns a preliminary category label to each article accordingly. For example, an article about the progress of artificial intelligence may be labeled as "Technology", while an article describing tourist attractions will be labeled as "Travel". In this way, the system not only quickly generates the first classification results for a large number of news articles, but also provides accurate data support for subsequent personalized recommendation services, ensuring the relevance and diversity of the recommended content.
[0084] Furthermore, classifying and predicting the samples to be labeled to obtain the first classification results corresponding to the samples to be labeled includes: inputting the samples to be labeled into the target classification model to obtain the first classification results corresponding to the samples to be labeled output by the target classification model, where the target classification model is pre-trained based on multiple training samples and the label information of the training samples.
[0085] In practical applications, the process of classifying and predicting the samples to be labeled and obtaining the first classification results involves several key steps. First, the system inputs the samples to be labeled into the target classification model. The target classification model calculates the probability distribution of each possible category based on the feature vectors of the samples and assigns a preliminary category label to each sample to be labeled as the first classification result. To ensure the accuracy of the classification prediction, the system can also adopt ensemble learning methods, combine the prediction results of multiple models, or use cross-validation techniques to optimize the model performance. In addition, for some complex tasks, the system may introduce additional feature engineering to extract richer sample features, thereby improving the quality of the classification prediction.
[0086] In the embodiments of this specification, by inputting the samples to be labeled into the target classification model, the system can generate the corresponding first classification results for each sample to be labeled. This process not only provides preliminary category labels, but also lays a foundation for subsequent example selection and annotation optimization. By making full use of the pre-trained target classification model, the system significantly improves the efficiency and quality of the annotation work, reduces the need for manual intervention, and enhances the automation and reliability of the system.
[0087] Exemplarily, during the development of a news recommendation engine, assume that the system needs to perform classification prediction on the latest news articles to be annotated. First, the system loads a pre-trained text classification model that has been fully trained on a large corpus containing various news categories. Next, the system takes the latest news articles collected from social platforms, news websites, etc. as samples to be annotated and inputs them one by one into the classification model. The model calculates the probability distribution of each news category based on the content features of each article (such as keywords, sentence structure, sentiment tendency, etc.) and assigns a preliminary category label to each article as the first classification result. For example, an article about the progress of artificial intelligence may be labeled as "Technology", while an article describing tourist attractions will be labeled as "Travel". In this way, the system not only quickly generates the first classification results of a large number of news articles, but also provides accurate data support for subsequent personalized recommendation services, ensuring the relevance and diversity of the recommended content.
[0088] Step 106: Obtain the second classification results corresponding to multiple annotated samples, and determine the target sample from the multiple annotated samples based on the similarity between the first classification result and each second classification result.
[0089] An annotated sample refers to a data instance that has been given a correct category label by manual or automatic methods. These samples usually go through strict review and verification to ensure the accuracy and reliability of their labels. For example, in a text classification task, the annotated samples can be a series of articles with clear topic labels; in an image recognition task, they can be pictures attached with object category labels.
[0090] The second classification result refers to the classification prediction label generated by the system for each annotated sample. This process is similar to performing classification prediction on samples to be annotated, but since the annotated samples already have true labels, the second classification results are mainly used to evaluate model performance, calculate similarity, and assist in example selection.
[0091] Similarity refers to the degree of matching between two classification results, usually measured by a certain mathematical formula or algorithm. Common similarity metrics include cosine similarity, Euclidean distance, etc. In this step, similarity is used to compare the first classification result of the sample to be annotated and the second classification result of the annotated sample to determine the most suitable reference example.
[0092] In practical applications, the process of obtaining the second classification results corresponding to multiple labeled samples and determining the target sample involves several key steps. First, the system loads a batch of labeled samples and their true labels, and uses the classification model to predict these samples to obtain the second classification results. Next, the system compares the first classification result of the sample to be labeled with the second classification results of all labeled samples and calculates the similarity between them. Then, based on the similarity scores, it selects a group of labeled samples that are closest to the sample to be labeled as the target sample. This process not only considers the matching of classification labels, but may also combine other features (such as text content, image pixels, etc.) to improve the accuracy of similarity calculation.
[0093] For the step of obtaining the second classification results corresponding to multiple labeled samples and determining the target sample, an optional approach is to directly use a pre-trained classification model to predict the labeled samples, and then use a metric method such as cosine similarity to calculate the similarity between the classification results; another optional implementation is to introduce additional feature engineering, extract richer sample features (such as TF-IDF values, word vectors, etc.), and calculate the comprehensive similarity based on this to more accurately select the target sample. Whichever approach is taken, the system needs to ensure that the entire process is efficient and reliable to provide a solid foundation for subsequent annotation optimization.
[0094] In the embodiments of this specification, by obtaining the second classification results corresponding to multiple labeled samples and determining the target sample based on similarity, the system can find the most suitable reference example for each sample to be labeled. This process not only improves the accuracy of example selection, reduces the confusion caused by sample complexity and diversity, but also significantly improves the accuracy and consistency of the final annotation results. This method optimizes the example selection mechanism and enhances the adaptability and reliability of the system.
[0095] Exemplarily, in a document classification system, assume that the system needs to find a suitable reference example for a newly submitted scientific and technological article (a sample to be labeled). First, the system loads a batch of labeled scientific and technological articles and uses a pre-trained text classification model to predict these articles to obtain the second classification results. Next, the system compares the first classification result (possibly the "science and technology" category) of the new article with the second classification results of the labeled articles and calculates the similarity between them. To ensure the accuracy of the similarity calculation, the system not only considers the matching of classification labels but also combines features such as the subject keywords and sentence structures of the articles. Finally, the system sorts according to the similarity scores and selects the five most similar new articles as target samples. In this way, the system not only finds a suitable reference example for the new article but also provides strong support for subsequent labeling work, ensuring the accuracy and consistency of the classification results. This not only improves the labeling efficiency but also provides a more accurate document classification service for users.
[0096] Further, obtaining the second classification results corresponding to multiple labeled samples includes: based on the sample information of the multiple labeled samples, searching in the example pool for the second classification results corresponding to the multiple labeled samples.
[0097] The example pool refers to a database or dataset that stores a large number of labeled samples and their related information. This information includes not only the content features of the samples (such as text content, image pixels, etc.) but also the true labels of the samples and other metadata (such as source, timestamp, etc.). The example pool provides the basic data support for subsequent classification prediction, model training, and example selection.
[0098] In practical applications, the process of obtaining the second classification results corresponding to multiple labeled samples involves several key steps. First, the system searches in the example pool and extracts the corresponding labeled samples based on the sample information of the multiple labeled samples. The sample information can include but is not limited to the unique identifier of the sample, content features, true labels, etc. Next, the system uses the classification model to classify and predict these labeled samples to generate the second classification results. To ensure the accuracy of the classification prediction, the system can also adopt an ensemble learning method, combine the prediction results of multiple models, or use cross-validation techniques to optimize the model performance. In addition, for some complex tasks, the system may introduce additional feature engineering to extract richer sample features, thereby improving the quality of the classification prediction.
[0099] For the step of obtaining the second classification results corresponding to multiple labeled samples, one optional approach is to directly batch-load the pre-labeled data from the sample pool and use a pre-trained deep learning model for classification prediction; another optional implementation is to grab the latest data from an online platform in real time, perform annotation and classification prediction immediately to ensure the freshness and relevance of the data. Whichever approach is adopted, the system must ensure that the entire process is efficient and reliable to provide a solid foundation for subsequent sample selection and model training.
[0100] In the embodiments of this specification, by looking up and extracting the second classification results corresponding to multiple labeled samples from the sample pool based on the sample information of the multiple labeled samples, the system can provide high-quality data support for subsequent sample selection and model evaluation. This process not only improves the accuracy of sample selection, reduces the confusion caused by the complexity and diversity of samples, but also significantly enhances the generalization ability and prediction accuracy of the final model. This method optimizes the sample selection mechanism and enhances the adaptability and reliability of the system.
[0101] Exemplarily, in a news classification system, assume that the system needs to obtain a batch of labeled technology news articles as reference samples. First, the system looks up and extracts a series of news articles labeled "technology" from the sample pool based on the sample information of the labeled samples (such as article title, publication date, author, etc.). Next, the system uses a pre-trained text classification model to perform classification prediction on these articles to generate the second classification results. To ensure the accuracy of the classification prediction, the system not only considers the topic labels of the articles, but also conducts a comprehensive analysis in combination with the content features of the articles (such as keywords, sentence structure, etc.). Finally, the system selects a group of labeled samples closest to the samples to be labeled as the target samples according to the second classification results. In this way, the system not only obtains high-quality reference samples, but also provides strong support for subsequent annotation work, ensuring the accuracy and consistency of the classification results.
[0102] Further, before looking up the second classification results corresponding to multiple labeled samples from the sample pool based on the sample information of the multiple labeled samples, it further includes: obtaining multiple labeled samples; inputting the multiple labeled samples into a target classification model to obtain the second classification results corresponding to each of the labeled samples output by the target classification model, where the target classification model is pre-trained based on the multiple labeled samples; and storing the second classification results corresponding to the multiple labeled samples in the sample pool.
[0103] The target classification model refers to the final model that can efficiently and accurately classify new data after training. Through the patterns and features learned from the training samples and label information, this model can provide reliable classification services in practical applications. The target classification model in this step is pre-trained based on multiple labeled samples, ensuring its generalization ability and accuracy.
[0104] In practical applications, the process of obtaining multiple labeled samples and generating the second classification results involves several key steps. First, the system obtains multiple labeled samples, which can come from various sources such as databases, file systems, web services, API interfaces, etc. To ensure the quality and consistency of the data, the system also needs to preprocess the collected samples, such as cleaning, deduplication, format conversion, and other operations. Next, the system inputs these labeled samples into the pre-trained target classification model to obtain the second classification results corresponding to each labeled sample output by the model. This process utilizes the powerful classification ability of the target classification model to ensure the accuracy and reliability of the classification results. Then, the system stores the second classification results corresponding to the multiple labeled samples in the example pool for convenient search and use of these classification results in subsequent steps.
[0105] For the step of obtaining multiple labeled samples and generating the second classification results, one optional method is to directly extract a part of the labeled data from an existing large-scale dataset as the samples to be processed; another optional implementation method is to capture newly generated data in real-time and immediately perform preliminary processing and annotation to ensure the freshness and relevance of the data. Whichever method is adopted, the system needs to ensure that the entire process is efficient and reliable to meet the requirements of the actual application scenario.
[0106] In the embodiments of this specification, by obtaining multiple labeled samples and inputting them into the target classification model, the system can obtain the second classification results corresponding to each labeled sample and store these results in the example pool. This process not only improves the accuracy and consistency of the classification results but also provides high-quality data support for subsequent example selection and model evaluation. Through the pre-trained target classification model, the system further enhances the reliability of classification prediction, reduces the need for manual intervention, and improves the automation level and efficiency of the system.
[0107] Exemplarily, in a document classification system, assume that the system needs to update the classification information in its example pool. First, the system obtains a batch of labeled technology articles from multiple channels as new labeled samples. To ensure the quality of these samples, the system preprocesses them, including removing duplicates, correcting formatting errors, etc. Next, the system inputs this batch of labeled samples into a pre-trained target classification model, which has been fully trained on a large corpus containing various news categories. The model generates a second classification result for each article based on the content features of each article, such as keywords, sentence structure, sentiment tendency, etc. For example, an article about the progress of artificial intelligence may be labeled as "Technology - Artificial Intelligence", while an article describing a tourist attraction will be labeled as "Travel". Finally, the system stores these second classification results together with the samples in the example pool to ensure that these classification information can be conveniently searched and used in subsequent steps. In this way, the system not only enriches the classification information in the example pool, but also provides more accurate and comprehensive data support for future classification tasks, ensuring the accuracy and consistency of the classification results.
[0108] Furthermore, before inputting multiple labeled samples into the target classification model, it also includes: obtaining an initial classification model, multiple training samples, and the label information of the training samples; training the initial classification model based on the multiple training samples and the label information to obtain the target classification model.
[0109] The initial classification model refers to a machine learning model before training on specific task data. Such a model can be implemented by any type of classification algorithm, such as decision tree, support vector machine (SVM), neural network, etc. The initial classification model usually already has certain basic knowledge or pre-trained weights, but has not been adjusted for the dataset of a specific task.
[0110] Multiple training samples refer to the data set used to train the classification model, and each sample represents an object instance that the system attempts to classify. These instances can come from various different sources and formats, such as text files, pictures, audio clips, etc. For example, in an image recognition application, the training samples can be a series of pictures labeled with different category labels, such as pictures of cats, dogs, cars, etc.
[0111] The label information of the training samples refers to the correct answer or target value associated with each training sample. The label information is crucial for supervised learning because it guides the learning process of the model. The label information in this step particularly includes the annotation results obtained by using the above sample annotation method, ensuring the accuracy and consistency of the labels.
[0112] In practical applications, the process of obtaining an initial classification model, multiple training samples, and their label information involves several key steps. First, the system loads an initial classification model suitable for the target task. This model may have been pre-trained on a general dataset and has a certain generalization ability. Next, the system collects a large number of training samples from multiple sources and ensures that these samples cover all aspects of the target task to improve the robustness of the model. Then, the system uses the above sample annotation method to generate high-quality label information for these training samples, ensuring the accuracy and consistency of the labels. Finally, the system inputs the training samples and label information into the initial classification model and adjusts the model parameters through an iterative optimization process (such as gradient descent) to enable the model to gradually learn to extract useful features from the data, thereby obtaining the target classification model.
[0113] For the steps of obtaining an initial classification model, multiple training samples and label information and conducting training, an optional approach is to directly use a pre-trained large language model or other deep learning models as the initial classification model and then fine-tune it in combination with domain-specific datasets; another optional implementation method is to train a brand-new classification model from scratch, which is applicable when there is sufficient large-scale and high-quality labeled data. Regardless of which method is adopted, the system needs to ensure that the entire training process is efficient and reliable to provide solid support for subsequent applications.
[0114] In the embodiments of this specification, by obtaining an initial classification model, multiple training samples, and their label information and training the initial classification model based on this information, the system can obtain an efficient and accurate target classification model. This process not only improves the generalization ability and accuracy of the model but also ensures its performance in dealing with complex and diverse practical problems. By making full use of high-quality annotation results, the system further enhances the effectiveness of the training data and provides strong guarantee for the performance of the final model.
[0115] Exemplarily, during the development of a social media content moderation tool, assume that the system needs to train a target classification model capable of automatically identifying inappropriate content. First, the system selects a pre-trained convolutional neural network (CNN) as the initial classification model, which has been fully trained on a large-scale image dataset. Next, the system extracts a large number of posts from the logs of the social platform as training samples and generates detailed label information for these posts using the above sample annotation method, including categories such as "normal", "advertisement", "technology", etc. Then, the system inputs these training samples and label information into the initial classification model and gradually adjusts the model parameters through multiple iterations of training, enabling it to learn how to identify inappropriate content based on the content features of the posts. Finally, the system obtains a fully trained target classification model that can not only quickly and accurately classify new posts but also adapt to the changing network environment, providing users with a safe and reliable social media experience.
[0116] Furthermore, the first classification result includes a first predicted probability distribution, and the second classification result includes a second predicted probability distribution; determining a target sample from multiple labeled samples based on the similarity between the first classification result and each second classification result includes: for any labeled sample, calculating the predicted probability distribution similarity between the to-be-labeled sample and any labeled sample according to the first predicted probability distribution and the second predicted probability distribution of any labeled sample, where the first predicted probability distribution represents the probability distribution of the to-be-labeled sample among multiple candidate labels, and the second predicted probability distribution represents the probability distribution of multiple labeled samples among multiple candidate labels; determining at least one target sample from multiple labeled samples based on the predicted probability distribution similarity.
[0117] The predicted probability distribution similarity refers to the matching degree between two predicted probability distributions, usually measured by a certain mathematical formula or algorithm. Common similarity metrics include cosine similarity, Euclidean distance, etc. In this step, the similarity is used to compare the first predicted probability distribution of the to-be-labeled sample and the second predicted probability distribution of the labeled sample to determine the most suitable reference example.
[0118] In practical applications, the process of determining a target sample from multiple labeled samples based on the similarity between the first classification result and each second classification result involves several key steps. First, for any labeled sample, the system calculates the similarity of the prediction probability distributions between the first prediction probability distribution of the sample to be labeled and the second prediction probability distribution of this labeled sample. This calculation can use various similarity measurement methods, such as cosine similarity, Kullback-Leibler divergence, etc., depending on the application scenario and requirements. Then, the system filters out at least one target sample from multiple labeled samples whose prediction probability distribution similarity reaches a preset threshold. The preset threshold can be adjusted according to the requirements of a specific task to ensure that the selected target samples have sufficient similarity, thereby improving the accuracy and efficiency of subsequent labeling work.
[0119] For the step of calculating the prediction probability distribution similarity and determining the target sample, an optional way is to directly use cosine similarity as the measurement standard. This method is simple and efficient and applicable to most scenarios; another optional implementation method is to introduce more complex similarity measurement methods, such as Kullback-Leibler divergence or Jensen-Shannon divergence. Although these methods have a higher computational cost, they can provide more accurate results in some specific tasks. No matter which method is adopted, the system needs to ensure that the whole process is efficient and reliable to provide a solid foundation for subsequent labeling optimization.
[0120] In the embodiments of this specification, based on the similarity between the first classification result and each second classification result, the system can determine the target sample from multiple labeled samples. This process not only improves the accuracy of sample selection, reduces the confusion caused by sample complexity and diversity, but also significantly improves the accuracy and consistency of the final labeling result. By using the prediction probability distribution similarity, the system optimizes the sample selection mechanism and enhances the adaptability and reliability of the system.
[0121] Exemplarily, in a document classification system, assume that the system needs to find a suitable reference example for a newly submitted scientific and technological article (a sample to be labeled). First, the system uses a pre-trained text classification model to generate a first classification result for this new article, including a high probability value for the "science and technology" category and lower probability values for other categories. Next, the system calculates the similarity of the predicted probability distributions between the second classification results of these articles and the first classification result of the new article one by one from a batch of labeled scientific and technological articles. To ensure the accuracy of the similarity calculation, the system adopts the cosine similarity metric method and sets a reasonable preset threshold. Finally, the system selects five labeled articles whose predicted probability distribution similarity reaches the preset threshold as target samples. In this way, the system not only finds a suitable reference example for the new article, but also provides strong support for the subsequent labeling work, ensuring the accuracy and consistency of the classification results. This not only improves the labeling efficiency, but also provides users with a more accurate document classification service.
[0122] Step 108: Input the target sample and the sample to be labeled into the labeling model to obtain the labeling result of the sample to be labeled.
[0123] The target sample refers to the data instance that is closest to the sample to be labeled selected based on the similarity of the classification prediction results from multiple labeled samples. These samples have high reference value and can provide reliable labeling guidance for the sample to be labeled.
[0124] The sample to be labeled refers to the data instance that the system has not yet classified or labeled. These data can be in various forms such as text, images, audio, etc., depending on the application scenario.
[0125] The labeling model (large language model) is a powerful natural language processing tool. After being trained on a large corpus, it has the ability to understand complex contexts and generate high-quality text. In this step, the model is used to receive the target sample and the sample to be labeled as inputs and generate the final labeling result based on the knowledge and patterns learned internally.
[0126] In practical applications, the process of inputting the target sample and the sample to be labeled into the labeling model involves several key steps. First, the system prepares the data formats of the target sample and the sample to be labeled to ensure that they meet the input requirements of the model. Then, the system passes these samples to the labeling model, and the model will comprehensively consider the context information provided by the target sample and the content features of the sample to be labeled for in-depth analysis. Next, the labeling model uses its powerful natural language understanding and generation capabilities to generate accurate and reasonable labeling results for the sample to be labeled. To ensure the labeling quality, the model can also combine the prompt words or rules in the context to further optimize the output results.
[0127] For the step of inputting the target sample and the sample to be labeled into the labeling model, an optional approach is to directly use pre-trained large language models such as BERT. These models have been fully trained on a vast amount of text data and can understand and process various types of text content. Another optional implementation is to fine-tune the large language model for a specific task. By introducing domain-specific datasets, the model can better adapt to specific application scenarios, thereby improving the accuracy of labeling. Regardless of which approach is adopted, the system needs to ensure that the model can efficiently process the input and generate high-quality labeling results to meet the requirements of practical applications.
[0128] In the embodiments of this specification, by inputting the target sample and the sample to be labeled into the labeling model, the system can make full use of the powerful capabilities of the large language model to generate accurate labeling results for each sample to be labeled. This process not only improves the efficiency of the labeling work, reduces the need for manual intervention, but also significantly enhances the quality and consistency of the labeling results, and strengthens the automation and reliability of the system.
[0129] Exemplarily, in a document classification system, assume that the system needs to generate the final category label for a newly submitted scientific and technological article (the sample to be labeled). First, the system has selected five most similar new articles as target samples based on the similarity of the classification prediction results. Next, the system inputs these five target samples and the newly submitted article into a fine-tuned large language model. The model not only understands the category information and content features of the target samples, but also combines the specific content of the newly submitted article for in-depth analysis. Finally, the large language model generates a label of the "science and technology" category for the newly submitted article, along with an explanation of why this article belongs to this category. For example, the model may point out that the key technologies and research findings mentioned in the article are highly relevant to other scientific and technological articles in the target samples. In this way, the system not only provides an accurate classification label for the new article, but also provides detailed classification basis for users, ensuring the transparency and credibility of the classification results. This not only improves the automation level of the system, but also brings a more intelligent and personalized service experience for users.
[0130] Furthermore, before inputting the target sample and the sample to be labeled into the labeling model to obtain the labeling result of the sample to be labeled, it further includes: obtaining labeling indication information, where the labeling indication information is used to indicate the labeling task of the labeling model; constructing labeling prompt information based on the labeling indication information and the target sample; inputting the target sample and the sample to be labeled into the labeling model to obtain the labeling result of the sample to be labeled, including: inputting the labeling prompt information and the sample to be labeled into the labeling model to obtain the labeling result of the sample to be labeled.
[0131] Annotation instruction information refers to the specific instructions or parameters received by the system on how to perform the annotation task. Such information can include, but is not limited to, classification categories, annotation rules in specific domains, priority settings, etc. The annotation instruction information is used to guide the annotation model to perform the specific annotation task and ensure that its output meets the expected requirements.
[0132] Annotation prompt information refers to a series of prompts or auxiliary information constructed based on the annotation instruction information and the target sample. These prompt messages are designed to provide additional context and support for the annotation model, helping the model to more accurately understand the content of the sample to be annotated and generate high-quality annotation results. For example, in a text classification task, the annotation prompt information may include domain-related keywords, background knowledge, or other information helpful for classification.
[0133] The annotation model is a large language model, a powerful natural language processing tool. Trained on a large corpus, it has the ability to understand complex contexts and generate high-quality text. In this step, the model is used to receive the target sample, the sample to be annotated, and the annotation prompt information as input, and based on the knowledge and patterns learned internally, generate the final annotation result.
[0134] In practical applications, the process of obtaining annotation instruction information and constructing annotation prompt information involves several key steps. First, the system obtains the annotation instruction information from the user terminal or other sources, which clarifies the specific requirements of this annotation task. Next, based on these annotation instruction information and the selected target sample, the system constructs detailed annotation prompt information. This process may involve operations such as extracting key features from the target sample, integrating domain knowledge, defining classification criteria, etc., to ensure that the annotation prompt information can effectively guide the work of the annotation model. Then, the system inputs the constructed annotation prompt information together with the sample to be annotated into the annotation model, and the model will comprehensively consider the prompt information and the sample content, conduct in-depth analysis, and finally generate accurate annotation results.
[0135] For the step of obtaining annotation instruction information and constructing annotation prompt information, an optional method is to directly receive the annotation instruction information from the user through the API interface and use a predefined template to quickly construct the annotation prompt information; another optional implementation method is to integrate the front-end interface, allowing users to customize the annotation rules and prompt information, and the system automatically parses and processes this data, and then calls the annotation model for prediction. Whichever method is adopted, the system needs to ensure that the entire process is efficient and reliable to provide accurate annotation services for users.
[0136] In the embodiments of this specification, by obtaining annotation indication information and constructing annotation prompt information, the system can provide detailed context support for each sample to be annotated, thereby significantly improving the quality and consistency of the annotation results. Utilizing the powerful capabilities of the annotation model, the system not only improves the efficiency of the annotation work, reduces the need for manual intervention, but also enhances the automation and reliability of the system. This method optimizes the annotation process, ensures the transparency and credibility of the annotation results, and brings a more intelligent and personalized service experience to users.
[0137] Exemplarily, in a document classification system, assume that the system needs to generate a final category label for a newly submitted scientific and technological article (sample to be annotated). First, the system receives the annotation indication information provided by the user, clarifying that the goal of this task is to identify articles in the field of "artificial intelligence". Next, based on these indication information and five selected target samples (all confirmed "artificial intelligence" articles), the system constructs detailed annotation prompt information, including domain keywords (such as "machine learning", "deep learning"), research directions (such as "natural language processing", "computer vision"), etc. Then, the system inputs these annotation prompt information together with the newly submitted article into a pre-trained large language model. The model not only understands the domain background in the prompt information but also conducts in-depth analysis by combining the specific content of the new article. Finally, the large language model generates a label of the "artificial intelligence" category for the new article and attaches an explanation of why this article belongs to this category. For example, the model may point out that the key technologies and research findings mentioned in the article are highly relevant to the domain characteristics in the prompt information. In this way, the system not only provides an accurate classification label for the new article but also provides detailed classification basis for users, ensuring the transparency and credibility of the classification results. This not only improves the intelligence level of the system but also brings a more intelligent and personalized service experience to users.
[0138] Further, after inputting the target sample and the sample to be annotated into the annotation model to obtain the annotation result of the sample to be annotated, it further includes: receiving feedback information regarding the annotation result; and adjusting the annotation model based on the feedback information.
[0139] Feedback information refers to the opinions or suggestions provided by users regarding the annotation result. Such information may include, but is not limited to, confirming the correctness of the annotation, pointing out errors, providing directions for improvement, etc. Feedback information is crucial for optimizing the annotation model because it directly reflects the performance of the model in actual applications and the expectations of users.
[0140] In practical applications, after inputting the target samples and samples to be labeled into the labeling model and obtaining the labeling results of the samples to be labeled, the system still needs to perform several key steps. First, the system outputs the labeling results to the user side and presents them to the user in an intuitive and understandable way. This step can be achieved in various ways, such as displaying on a web page, sending via email, or updating the content in a mobile application. Next, the system receives the feedback information on the labeling results from the user side. To ensure the effectiveness and timeliness of the feedback information, the system can provide a simple user interface to allow users to easily submit their opinions and suggestions. Finally, the system adjusts the labeling model based on the received feedback information. This process may involve operations such as retraining the model, adjusting parameters, adding new training data, etc., depending on the content and nature of the feedback information.
[0141] For the steps of outputting the labeling results to the user side and receiving feedback information, an optional way is to display the labeling results on the user interface through an API interface and collect the users' feedback in real time; another optional implementation method is to integrate the front-end interface, allowing users to view the labeling results and submit feedback on the same interface, and the system automatically parses and processes this data. Whichever way is adopted, the system needs to ensure that the whole process is efficient and reliable, so as to quickly respond to the users' query needs and provide accurate labeling services.
[0142] In the embodiments of this specification, by outputting the labeling results to the user side and receiving feedback information, the system can collect valuable user opinions, and thus adjust the labeling model in a targeted manner. This process not only improves the accuracy and consistency of the labeling results, but also enhances the adaptive ability and user experience of the system. Using the users' feedback information, the system can continuously optimize the model performance to ensure its performance in dealing with complex and diverse practical problems. This method forms a closed-loop optimization process, enabling the system to continuously improve and better meet the users' needs.
[0143] Exemplarily, in a document classification system, assume that the system has generated a labeling result of the "Artificial Intelligence" category for a newly submitted scientific and technological article (sample to be labeled). First, the system displays this labeling result on the user interface, and the user can clearly see that the article is classified as "Artificial Intelligence". If the user believes that this classification is inaccurate, they can select the "Feedback" button on the same interface and provide specific improvement suggestions, such as pointing out that this article actually leans more towards the sub-field of "Machine Learning". After receiving the user's feedback information, the system will take it into account for subsequent model adjustment. For example, the system may increase more training data on the "Machine Learning" field or adjust the model parameters to improve the recognition ability for specific fields. In this way, the system not only provides users with a transparent and interactive labeling experience but also continuously optimizes its own performance using the user's feedback information to ensure the accuracy of future classification tasks.
[0144] Further, after inputting the target sample and the sample to be labeled into the labeling model to obtain the labeling result of the sample to be labeled, it further includes: generating an audit task based on the sample to be labeled and the labeling result, and sending the audit task to the audit end; in response to the audit approval instruction sent by the audit end, taking the sample to be labeled as a labeled sample.
[0145] An audit task refers to a task generated by the system based on the sample to be labeled and its corresponding labeling result to ensure the accuracy and reliability of the labeling result. An audit task usually includes the sample content to be audited, the preliminary labeling result, and any relevant context information for the auditor to check and confirm.
[0146] The audit end refers to the platform or interface responsible for the audit task, which can be a dedicated audit tool, a web interface, a mobile application, etc. The audit end is used to display the audit task to the auditor for viewing and allows them to submit audit opinions or instructions (such as "Audit passed", "Audit failed").
[0147] In practical applications, after inputting the target sample and the sample to be labeled into the labeling model and obtaining the labeling result of the sample to be labeled, the system also needs to perform several key steps. First, the system generates an audit task based on the sample to be labeled and its corresponding labeling result. This process may involve extracting the key features of the sample, integrating relevant background information, defining audit criteria, etc., to ensure that the audit task can provide sufficient context support to help the auditor make an accurate judgment. Then, the system sends the generated audit task to the audit end for the auditor to review. The audit task can be delivered to the auditor in various ways, such as through internal network distribution, email notification, or integrated into a dedicated audit management system.
[0148] Next, in response to the approval instruction sent by the review terminal, the system regards the sample to be labeled as a labeled sample. If the reviewers confirm that the labeling result is correct, they will send an approval instruction through the review terminal. After receiving this instruction, the system will officially include the sample to be labeled and its labeling result in the set of labeled samples, which means that these samples have undergone manual review and have high credibility and reliability. For samples that fail the review, the system can adjust the labeling process according to the feedback information, such as retraining the model, adjusting parameters, adding new training data, etc., to improve the accuracy of subsequent labeling.
[0149] For the step of generating a review task and sending the review task to the review terminal, an optional method is to directly push the review task to the reviewer's workstation through the API interface and receive the review feedback in real time; another optional implementation method is to integrate the front-end interface, allowing reviewers to view and process multiple review tasks on a unified platform, and the system automatically parses and processes this data. Regardless of which method is adopted, the system needs to ensure that the entire process is efficient and reliable, so as to quickly respond to review requirements and ensure the labeling quality.
[0150] In the embodiments of this specification, by generating a review task based on the sample to be labeled and the labeling result and sending the review task to the review terminal, the system can ensure that each labeling result undergoes a strict review process, thereby improving the accuracy and consistency of the labeling work. Utilizing the professional knowledge and experience of the reviewers, the system not only enhances the credibility of the labeling result but also provides high-quality data support for subsequent applications. This method optimizes the labeling process, forms a closed-loop quality control mechanism, enables the system to continuously improve, and better meets the needs of users.
[0151] Exemplarily, in a document classification system, assume that the system has generated a labeling result of the "artificial intelligence" category for a newly submitted scientific and technological article (sample to be labeled). First, the system generates a detailed review task based on this article and its labeling result, which includes the full text of the article, the preliminary category label, and some auxiliary information (such as domain keywords, research directions). Then, the system sends this review task to the review terminal, and the reviewers can view the task details on the platform and decide whether to agree with this classification. If the reviewers confirm that the labeling is correct, they will send an approval instruction through the review terminal. After receiving the instruction, the system officially includes this article and its labeling result in the set of labeled samples, which means that it has undergone manual review and has high credibility. For cases where the review fails, the system can adjust the labeling process according to the improvement suggestions provided by the reviewers to improve the future classification accuracy. In this way, the system not only provides users with a transparent and interactive labeling experience but also continuously optimizes its own performance using the review mechanism to ensure the accuracy of future classification tasks.
[0152] One embodiment of this specification realizes obtaining samples to be labeled; classifying and predicting the samples to be labeled to obtain a first classification result corresponding to the samples to be labeled; obtaining second classification results corresponding to multiple labeled samples, and determining a target sample from the multiple labeled samples based on the similarity between the first classification result and each second classification result; inputting the target sample and the samples to be labeled into a labeling model to obtain a labeling result for the samples to be labeled. By using a classification model to obtain the classification prediction result of the samples to be labeled and screening according to the similarity between the samples to be labeled and the labeled samples, it is ensured that the selected target sample is highly similar to the samples to be labeled at the classification prediction level, thereby reducing the confusion caused by the complexity and diversity of the sample content. On this basis, by further processing the target sample and the samples to be labeled through the labeling model, the final labeling result is generated, significantly improving the accuracy and consistency of the labeling. This method not only optimizes the example selection but also enhances the adaptability and reliability of the system.
[0153] See Figure 2 , Figure 2 is a flowchart of a classification model training method provided by one embodiment of this specification, which specifically includes the following steps.
[0154] Step 202: Obtain an initial classification model, multiple training samples, and label information of the training samples. Among them, the label information of the multiple training samples includes the labeling results obtained by using the above sample labeling method.
[0155] Step 204: Train the initial classification model based on the multiple training samples and the label information to obtain a target classification model.
[0156] The initial classification model refers to a machine learning model before training on specific task data. This model can be implemented by any type of classification algorithm, such as decision tree, support vector machine (SVM), neural network, etc. The initial classification model usually already has certain basic knowledge or pre-trained weights, but has not been adjusted for the dataset of a specific task.
[0157] The multiple training samples refer to the data set used to train the classification model, and each sample represents an object instance that the system attempts to classify. These instances can come from various different sources and formats, such as text files, pictures, audio clips, etc. For example, in an image recognition application, the training samples can be a series of pictures marked with different category labels, such as pictures of cats, dogs, cars, etc.
[0158] The label information of training samples refers to the correct answers or target values associated with each training sample. Label information is crucial for supervised learning as it guides the learning process of the model. The label information in this step specifically includes the annotation results obtained using the above sample annotation method, ensuring the accuracy and consistency of the labels.
[0159] The target classification model refers to the final model that, after training, can efficiently and accurately classify new data. The model can provide reliable classification services in practical applications by learning patterns and features from training samples and label information.
[0160] In practical applications, the process of obtaining the initial classification model, multiple training samples, and their label information involves several key steps. First, the system loads an initial classification model suitable for the target task, which may have been pre-trained on a general dataset and has certain generalization capabilities. Next, the system collects a large number of training samples from multiple sources and ensures that these samples cover all aspects of the target task to improve the robustness of the model. Then, the system uses the above sample annotation method to generate high-quality label information for these training samples, ensuring the accuracy and consistency of the labels. Finally, the system inputs the training samples and label information into the initial classification model and adjusts the model parameters through an iterative optimization process (such as gradient descent) to enable the model to gradually learn to extract useful features from the data, thereby obtaining the target classification model.
[0161] For the step of obtaining the initial classification model, multiple training samples, and label information and conducting training, an optional approach is to directly use a pre-trained large language model or other deep learning models as the initial classification model and then fine-tune it with domain-specific datasets; another optional implementation method is to train a brand-new classification model from scratch, which is applicable when there is sufficient large-scale and high-quality labeled data. Regardless of which method is adopted, the system needs to ensure that the entire training process is efficient and reliable to provide solid support for subsequent applications.
[0162] In the embodiments of this specification, by obtaining the initial classification model, multiple training samples, and their label information and training the initial classification model based on this information, the system can obtain an efficient and accurate target classification model. This process not only improves the generalization ability and accuracy of the model but also ensures its performance in dealing with complex and diverse practical problems. By making full use of high-quality annotation results, the system further enhances the effectiveness of the training data and provides strong guarantee for the performance of the final model.
[0163] Exemplarily, during the development of a social media content moderation tool, assume that the system needs to train a target classification model capable of automatically identifying inappropriate content. First, the system selects a pre-trained convolutional neural network (CNN) as the initial classification model, which has been fully trained on a large-scale image dataset. Next, the system extracts a large number of posts from the logs of the social platform as training samples, and uses the above sample annotation method to generate detailed label information for these posts, including categories such as "normal", "advertisement", "technology", etc. Then, the system inputs these training samples and label information into the initial classification model, and through multiple iterative trainings, gradually adjusts the model parameters to enable it to learn how to identify inappropriate content based on the content features of the posts. Finally, the system obtains a fully trained target classification model, which can not only quickly and accurately classify new posts, but also adapt to the changing network environment to provide users with a safe and reliable social media experience.
[0164] By obtaining the initial classification model, multiple training samples and their label information (generated by the sample annotation method described in steps 102 - 108), this method can efficiently train the target classification model. Using high-quality annotation results, the training process not only improves the generalization ability of the model, but also ensures its accuracy and stability when dealing with complex and diverse samples. The finally obtained target classification model can provide more accurate classification services in practical applications, significantly improving the overall performance of the classification task.
[0165] See Figure 3 , Figure 3 is a flowchart of an object classification method provided by an embodiment of this specification, specifically including the following steps.
[0166] Step 302: Obtain an object classification request, where the object classification request includes object data of the target object.
[0167] Step 304: Input the object data into the target classification model to determine the classification result of the target object, where the target classification model is trained based on the above classification model training method.
[0168] An object classification request refers to a request received by the system to classify a specific object. This request usually contains object data of the target object, and these data can be in various forms such as text, image, audio, etc., depending on the application scenario. For example, in an image recognition task, the object classification request may be a picture; in a text classification task, it may be a document or a paragraph of text.
[0169] The object data of the target object refers to the specific content and characteristic information of the object to be classified. These data are the basis for the system to perform classification prediction and determine the accuracy and reliability of the final classification result. For example, for a news article, its object data may include metadata such as the title, text, author, and publication time.
[0170] The target classification model refers to the final model that can efficiently and accurately classify new data after training. Based on the patterns and features learned from the training samples and label information, this model can provide reliable classification services in practical applications. The target classification model in this step is trained based on the above classification model training method, ensuring the generalization ability and accuracy of the model.
[0171] In practical applications, the process of obtaining an object classification request and determining the classification result of the target object involves several key steps. First, the system receives an object classification request from a user terminal or other systems, and this request contains the object data of the target object. Next, the system preprocesses these object data, such as format conversion, feature extraction, etc., to ensure that they meet the input requirements of the target classification model. Then, the system inputs the processed object data into the target classification model, and the model will generate a classification result for each target object according to the knowledge and patterns learned internally. To ensure the accuracy of the classification result, the system can also combine context information or other auxiliary data to further optimize the output result.
[0172] For the step of obtaining an object classification request and inputting the object data into the target classification model, an optional way is to directly use the API interface to receive the object classification request and transfer the object data to the classification model through the RESTful service; another optional implementation method is to integrate the front-end interface, allowing users to upload files or input text, and the system automatically parses and processes these data, and then calls the classification model for prediction. No matter which method is adopted, the system needs to ensure that the whole process is efficient and reliable, so as to quickly respond to the user's query needs and provide accurate classification services.
[0173] In the embodiments of this specification, by obtaining an object classification request and inputting the object data into the target classification model, the system can generate accurate classification results for each target object. This process not only improves the efficiency of the classification work, reduces the need for manual intervention, but also significantly improves the quality and consistency of the classification results, enhancing the automation degree and reliability of the system. Using the well-trained target classification model, the system can provide high-quality classification services in a variety of application scenarios to meet the needs of different users.
[0174] Exemplarily, in an online product recommendation system, assume that the system needs to classify the products browsed by users in order to provide personalized recommendations. First, the system receives an object classification request, which contains the details of the products currently browsed by the user (such as name, description, price, category, etc.). Next, the system preprocesses these product details, extracts key features (such as keywords, category labels, etc.), and converts them into a format suitable for input to the classification model. Then, the system inputs the processed product data into the target classification model, which has been fully trained with a large amount of labeled product data and can accurately identify various product categories. Finally, the target classification model generates a label of the "electronic product" category for the product, along with an explanation of why this product belongs to this category. For example, the model may point out that the technical specifications and functional features in the product description highly match those of known electronic products. In this way, the system not only provides accurate product classification for users, but also lays a solid foundation for subsequent personalized recommendation services, ensuring the relevance and diversity of the recommended content. This not only improves the intelligence level of the system, but also brings a more convenient and personalized shopping experience to users.
[0175] The object classification method realizes fast and accurate object classification by receiving a classification request containing target object data and inputting this data into the target classification model trained through Step 202 - Step 204. This method makes full use of the well-trained classification model, can determine the classification result of the target object in a short time, and provides high-precision classification services. This not only enhances the response speed of the system, but also ensures the consistency and reliability of the classification results, and is applicable to a variety of application scenarios.
[0176] See Figure 4 , Figure 4 is a flowchart of an object search method provided by an embodiment of this specification, specifically including the following steps.
[0177] Step 402: Obtain an object search request, where the object search request includes a search statement.
[0178] Step 404: Input the search statement into the target classification model to determine the search type corresponding to the search statement, where the target classification model is trained based on the above classification model training method.
[0179] Step 406: Based on the search type, call the target search service to obtain the search result.
[0180] An object search request refers to a request received by the system that demands the search for a specific object. This request usually contains search statements provided by the user, which can be text queries, keywords, or other forms of natural language expressions. For example, on an online shopping platform, an object search request might be a query about "red dresses".
[0181] A search statement refers to the specific query content provided by the user in an object search request. These statements express the user's search intention and are the basis for the system to understand the user's needs and provide relevant search results. For example, "How to make pasta" or "Latest popular tech news".
[0182] The search type refers to the specific search category determined after classifying the search statement according to the target classification model. Different types of searches may correspond to different search strategies and services. For example, simple information queries, product searches, navigation instructions, etc. Each type may require invoking different search services to obtain the most relevant search results.
[0183] The target search service refers to a service module designed specifically for a certain type of search. It is responsible for performing specific search operations according to the classified search type and returning the corresponding search results. For example, for the product search type, the target search service might be a commodity database query interface; for the information query type, it might be a knowledge base or a web search engine.
[0184] In practical applications, the process of obtaining an object search request and processing the search statement involves several key steps. First, the system receives an object search request from the user terminal, which contains the search statement provided by the user. Next, the system inputs the search statement into the target classification model, and the model will determine the corresponding search type for each search statement based on the knowledge and patterns learned internally. To ensure the accuracy of classification, the system can also combine context information or other auxiliary data to further optimize the output results. Finally, the system invokes the corresponding target search service based on the determined search type to obtain the most relevant search results and feedback them to the user.
[0185] For the step of obtaining an object search request and inputting the search statement into the target classification model, an optional method is to directly use the API interface to receive the object search request and pass the search statement to the classification model through the RESTful service; another optional implementation method is to integrate the front-end interface, allowing the user to input text queries, and the system automatically parses and processes this data, and then invokes the classification model for prediction. Regardless of which method is adopted, the system needs to ensure that the entire process is efficient and reliable, so as to quickly respond to the user's query needs and provide accurate search services.
[0186] In the embodiments of this specification, by obtaining an object search request and inputting the search statement into the target classification model, the system can accurately determine the search type corresponding to the search statement, and based on the search type, call the appropriate target search service to obtain the most relevant search results. This process not only improves the efficiency of the search work, reduces the need for manual intervention, but also significantly improves the quality and consistency of the search results, enhancing the automation and reliability of the system. Using a well-trained target classification model, the system can provide high-quality search services in a variety of application scenarios to meet the needs of different users.
[0187] Exemplarily, on a comprehensive online service platform, assume that the system needs to process a search statement about "how to make pizza at home". First, the system receives the object search request, which contains this search statement. Next, the system inputs the search statement into the target classification model, which has been fully trained with a large number of labeled search statements and can accurately identify various search types. After analyzing the search statement, the target classification model determines that it belongs to the "recipe query" type. Then, the system calls the target search service specifically for recipe queries based on this search type. This service queries the recipe database on the platform and finds a series of recipes and tutorials related to "how to make pizza at home". Finally, the system presents these search results to the user, providing not only detailed production steps but also information such as required materials and tool suggestions, ensuring that the user can find the information they need. This not only improves the intelligence level of the system but also brings a more convenient and personalized search experience to the user.
[0188] The object search method determines the corresponding search type by parsing the search statement submitted by the user and using the target classification model trained in steps 202 - 204, and then calls the corresponding target search service to obtain accurate search results. This method combines advanced classification techniques and efficient search algorithms to ensure the relevance and accuracy of the search results. It can not only quickly respond to the user's query needs but also provide a personalized search experience according to the search type, significantly improving the user's satisfaction and search efficiency.
[0189] The following combines the attached Figure 5 、 Figure 6 , taking the application of the sample annotation method provided in this specification in a note sharing application as an example, to further illustrate the sample annotation method. Among them, Figure 5 is a flowchart of a sample annotation method provided by an embodiment of this specification, mainly including the following key components:
[0190] Samples to be annotated: It can be the unannotated note content uploaded by the user, which needs to be classified and annotated to optimize subsequent search and recommendation functions.
[0191] Classification model: A classifier trained based on labeled notes, used to perform preliminary classification prediction on newly uploaded notes.
[0192] Example pool: Contains standard labeled notes verified manually, providing high-quality data support for subsequent annotation optimization and model training.
[0193] Large language model: A large language model that performs automatic annotation tasks, with powerful natural language processing capabilities, capable of generating labels that match the note content under the prompt of high-quality labeled sample data tags.
[0194] Figure 6 It is a process flow chart of a sample annotation method provided by an embodiment of this specification, specifically including the following steps:
[0195] Step 602: Train a classification model using the standard labeled data in the example pool.
[0196] Specifically, train the classification model P c (·|θ) using the standard labeled data in the sample pool (such as high-quality notes reviewed by the community). This model receives the sample input x and outputs the probability distribution P c (x|θ) = [p1, p2, …, p n , where n is the number of labels.
[0197] Step 604: Calculate the predicted probability distribution for all samples in the example pool using the trained model, and establish an index structure with the probability distribution as the key and the example as the value.
[0198] Specifically, use the trained model to calculate the predicted probability distribution for all note samples in the example pool and establish an index structure with the labeled probability distribution as the key and the note sample as the value. as the value.
[0199] Step 606: For each sample to be labeled, first obtain its predicted probability distribution (i.e., the model perplexity) through the classification model.
[0200] Specifically, the model perplexity is expressed as P c (x a |θ).
[0201] Step 608: Calculate the cosine similarity between this distribution and the perplexity distribution of all samples in the example pool.
[0202] Specifically, the calculation formula is as follows:
[0203]
[0204] Step 610: Based on the similarity ranking, select the top 10 most similar samples.
[0205] Step 612: Input the selected samples together with the labeling rules into the large language model to guide it to label the target samples.
[0206] Step 614: Evaluate the annotation results.
[0207] Specifically, the annotation results were evaluated. In order to verify the effectiveness of the method, an evaluation was conducted in a note-sharing application, and the evaluation task was to accurately determine the topic type of the note. For example, 2,000 correct notes were manually annotated according to the rules, a preliminary classification model was trained, and an annotation sample pool was constructed. Next, in order to evaluate the annotation results, a set of test notes was selected from the sample pool, and the annotation results based on the above steps were obtained. The accuracy between these results and the manually annotated benchmark answers was calculated. The accuracy was quantitatively analyzed by calculating indicators such as precision, recall, and F1 score, thereby comprehensively evaluating the annotation effect of the sample annotation method on the same topic type.
[0208] Take the note sharing application as an example. Figure 7 This is a front-end schematic diagram of an object search system provided by an embodiment of this specification:
[0209] After the user logs in to the note sharing app, at the top of the homepage, there are a details button, a "Follow" channel, a "Discover" channel, a "Nearby" channel and a query button.
[0210] Click the query button on the note-sharing app's homepage to open an input box and receive the user's search query. The entered search query is fed into the trained target classification model, which determines the search type and calls the corresponding target search service to obtain accurate search results. The search results are displayed in the center of the homepage.
[0211] In the middle of the homepage, the note covers of four searched target notes are displayed in order, namely: "Create your ever-changing autumn OOTD", "Explore the hidden food corners in the city", "How to take cinematic photos with a mobile phone", and "Those domestic niche travel destinations that cannot be missed".
[0212] A label can be added under each note cover to indicate the category to which the note belongs (such as fashion, food, photography, travel). This helps users quickly understand the content of the note and supports further filtering based on category.
[0213] The bottom of the homepage includes a navigation bar, which contains a "Home" button, a "Video" button, a "Post Notes" button, a "Message" button, and a "Me" button.
[0214] Corresponding to the above method embodiments, this specification also provides an object search system embodiment. Figure 8 FIG. is a schematic structural diagram of an object search system provided by an embodiment of this specification. As Figure 8 shown, the object search system 800 includes a request interface 802, a classification unit 804, a search unit 806, a target classification model 808, and a target search service 810;
[0215] The request interface 802 is configured to obtain an object search request, where the object search request includes a search statement;
[0216] The classification unit 804 is configured to input the search statement into the target classification model 808 to determine the search type corresponding to the search statement, where the target classification model is trained based on the foregoing object classification model training method;
[0217] The search unit 806 is configured to obtain a search result by invoking the target search service 810 based on the search type;
[0218] The request interface is further configured to feedback the search result.
[0219] The object search system 800 is a search system integrating multiple functional modules, aiming to provide efficient and personalized search services. Its core components include a request interface 802, a classification unit 804, a search unit 806, a target classification model 808, and a target search service 810.
[0220] The request interface 802 is responsible for receiving an object search request from a user terminal. When a user inputs a search statement through their device (such as a smartphone, tablet, etc.), the request interface 802 captures this request and passes it to other components inside the system for processing. After the search is completed, the request interface 802 is responsible for feeding back the search result to the user terminal to ensure that the user can obtain the required information in a timely manner.
[0221] The classification unit 804 is an important part of the object search system 800, and it is responsible for processing the received search statement. Specifically, the classification unit 804 inputs the search statement into the trained target classification model 808 to determine the search type corresponding to the search statement. The target classification model 808 is trained based on the foregoing object classification model training method and can accurately identify and classify different types of queries, laying a foundation for subsequent precise searches.
[0222] The search unit 806 is a core part of the object search system 800. It is responsible for calling the appropriate target search service 810 according to the search type determined by the classification unit 804 to obtain search results. The search unit 806 not only considers the text content of the search statement, but also combines the user's attribute characteristics and historical interaction behaviors to generate more personalized and relevant search results. In addition, the search unit 806 also sorts the search results to ensure the quality and relevance of the recommended content and improve the user experience.
[0223] The target classification model 808 is a machine learning model trained by the aforementioned classification model training method, and is used to parse and classify the user's search statement. This model is trained based on the labeled data obtained by the aforementioned sample annotation method, and has the ability to understand complex queries and can accurately map the search statement to the appropriate search type, so as to guide the search unit 806 to select the correct search strategy and service.
[0224] The target search service 810 is a service module that performs actual search operations. According to the instructions of the search unit 806, the target search service 810 retrieves relevant information from the content database or external resources and generates the final search results. After these results are optimized, they are fed back to the user terminal through the request interface 802 to ensure that the user obtains the most relevant and valuable information
[0225] Corresponding to the above method embodiments, this specification also provides an embodiment of a sample annotation device. Figure 9 It is a schematic structural diagram of a sample annotation device provided by an embodiment of this specification. As Figure 9 shown, the device includes:
[0226] The first acquisition module 902 is configured to acquire samples to be annotated;
[0227] The prediction module 904 is configured to perform classification prediction on the samples to be annotated to obtain the first classification result corresponding to the samples to be annotated;
[0228] The determination module 906 is configured to obtain the second classification results corresponding to multiple labeled samples, and determine the target sample from the multiple labeled samples based on the similarity between the first classification result and each second classification result;
[0229] The annotation module 908 is configured to input the target sample and the samples to be annotated into the annotation model to obtain the annotation result of the samples to be annotated.
[0230] Optionally, the determination module 906 is further configured to search for the second classification results corresponding to the multiple labeled samples from the sample pool based on the sample information of the multiple labeled samples.
[0231] Optionally, the sample annotation device further includes a classification module, configured to obtain a plurality of annotated samples; input the plurality of annotated samples into a target classification model to obtain second classification results corresponding to each of the annotated samples output by the target classification model, where the target classification model is pre-trained based on the plurality of annotated samples; store the second classification results corresponding to the plurality of annotated samples in an example pool.
[0232] Optionally, the sample annotation device further includes an initial training module, configured to obtain an initial classification model, a plurality of training samples, and label information of the training samples; train the initial classification model based on the plurality of training samples and the label information to obtain a target classification model.
[0233] Optionally, the prediction module is further configured to input the sample to be annotated into the target classification model to obtain a first classification result corresponding to the sample to be annotated output by the target classification model, where the target classification model is pre-trained based on a plurality of annotated samples.
[0234] Optionally, the first classification result includes a first predicted probability distribution, and the second classification result includes a second predicted probability distribution; correspondingly, the determination module 906 is further configured to calculate the similarity of the predicted probability distributions between the sample to be annotated and any one of the annotated samples according to the first predicted probability distribution and the second predicted probability distribution of any one of the annotated samples for any one of the annotated samples; determine at least one target sample from the plurality of annotated samples based on the similarity of the predicted probability distributions.
[0235] Optionally, the sample annotation device further includes an indication information construction module, configured to obtain annotation indication information, where the annotation indication information is used to indicate the annotation task of the annotation model; construct annotation prompt information based on the annotation indication information and the target sample; correspondingly, the annotation module 908 is further configured to input the annotation prompt information and the sample to be annotated into the annotation model to obtain the annotation result of the sample to be annotated.
[0236] Optionally, the sample annotation device further includes a feedback adjustment module, configured to output the annotation result to the user terminal; receive feedback information on the annotation result fed back by the user terminal; adjust the annotation model based on the feedback information.
[0237] Optionally, the sample annotation device further includes a sample conversion module, configured to generate an audit task based on the sample to be annotated and the annotation result, and send the audit task to the audit terminal; in response to an audit pass instruction sent by the audit terminal, use the sample to be annotated as an annotated sample.
[0238] The above is a schematic solution of a sample annotation device according to this embodiment. It should be noted that the technical solution of this sample annotation device and the technical solution of the above sample annotation method belong to the same concept. For the details not described in the technical solution of the sample annotation device, reference can be made to the description of the technical solution of the above sample annotation method.
[0239] Corresponding to the above method embodiment, this specification also provides an embodiment of a classification model training device. Figure 10 It is a schematic structural diagram of a classification model training device provided by an embodiment of this specification. As Figure 10 shown, the device includes:
[0240] A second acquisition module 1002, configured to acquire an initial classification model, a plurality of training samples, and label information of the training samples, where the label information of the plurality of training samples includes the annotation results obtained by using the sample annotation method described in the first aspect;
[0241] A training module 1004, configured to train the initial classification model based on the plurality of training samples and the label information to obtain a target classification model.
[0242] Applied to this classification model training device, the classification model training device acquires an initial classification model, a plurality of training samples and their label information through the second acquisition module 1002, where the label information is generated by the foregoing sample annotation method. Subsequently, the training module 1004 trains the initial classification model based on these high-quality annotation results, and finally obtains a target classification model. This process not only improves the generalization ability of the model and the accuracy of processing complex and diverse samples, but also ensures its stability and reliability in practical applications, thus significantly improving the overall performance of the classification task.
[0243] The above is a schematic solution of a classification model training device according to this embodiment. It should be noted that the technical solution of this classification model training device and the technical solution of the above classification model training method belong to the same concept. For the details not described in the technical solution of the classification model training device, reference can be made to the description of the technical solution of the above classification model training method.
[0244] Corresponding to the above method embodiment, this specification also provides an embodiment of an object classification device. Figure 11 It is a schematic structural diagram of an object classification device provided by an embodiment of this specification. As Figure 11 shown, the device includes:
[0245] A third acquisition module 1102, configured to acquire an object classification request, where the object classification request includes object data of a target object;
[0246] The first classification module 1104 is configured to input object data into a target classification model and determine the classification result of a target object, where the target classification model is trained based on the object classification model training method described in the second aspect.
[0247] Applied to the object classification device, the object classification device uses the third acquisition module 1102 to receive a classification request containing target object data, and inputs this data into the target classification model trained by the foregoing classification model training method through the first classification module 1104. This method can determine the classification result of the target object in a short time, provide high-precision and consistent classification services, enhance the response speed of the system and the reliability of the classification result, and is applicable to a variety of application scenarios.
[0248] The above is a schematic solution of an object classification device in this embodiment. It should be noted that the technical solution of this object classification device and the technical solution of the above object classification method belong to the same concept. For the details not described in the technical solution of this object classification device, reference can be made to the description of the technical solution of the above object classification method.
[0249] Corresponding to the above method embodiment, this specification also provides an embodiment of an object search device. Figure 12 It is a schematic structural diagram of an object search device provided by an embodiment of this specification. As Figure 12 shown, the device includes:
[0250] The fourth acquisition module 1202 is configured to acquire an object search request, where the object search request includes a search statement.
[0251] The second classification module 1204 is configured to input the search statement into a target classification model and determine the search type corresponding to the search statement, where the target classification model is trained based on the object classification model training method described in the second aspect.
[0252] The search module 1206 is configured to obtain a search result by calling a target search service based on the search type.
[0253] Applied to the object search device, the object search device parses the search statement submitted by the user through the fourth acquisition module 1202, and inputs the search statement into the target classification model trained by the foregoing classification model training method with the help of the second classification module 1204 to determine the search type. Then, the search module 1206 calls the corresponding target search service to obtain an accurate search result. This device combines advanced classification technology and efficient search algorithms, ensures the relevance and accuracy of the search result, provides a personalized search experience at the same time, and significantly improves the user satisfaction and search efficiency.
[0254] The above is a schematic solution of an object search device according to this embodiment. It should be noted that the technical solution of the object search device and the technical solution of the above object search method belong to the same concept. For the details not described in the technical solution of the object search device, reference can be made to the description of the technical solution of the above object search method.
[0255] Figure 13 FIG. 1300 shows a structural block diagram of a computing device 1300 according to an embodiment of the present specification. The components of the computing device 1300 include, but are not limited to, a memory 1310 and a processor 1320. The processor 1320 is connected to the memory 1310 through a bus 1330, and a database 1350 is used to store data.
[0256] The computing device 1300 further includes an access device 1340, which enables the computing device 1300 to communicate via one or more networks 1360. Examples of these networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 1340 may include one or more of any type of wired or wireless network interfaces (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (WiMAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC).
[0257] In an embodiment of the present specification, the above components of the computing device 1300 and Figure 13 other components not shown in FIG. may also be connected to each other, for example, through a bus. It should be understood that Figure 13 the structural block diagram of the computing device shown in FIG. is for illustrative purposes only and is not a limitation on the scope of the present specification. Those skilled in the art can add or replace other components as needed.
[0258] The computing device 1300 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smart phone), a wearable computing device (e.g., a smart watch, smart glasses, etc.) or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 1300 can also be a mobile or stationary server.
[0259] The processor 1320 is configured to execute the following computer program / instructions, and when the computer program / instructions are executed by the processor, the steps of the above sample annotation method, or classification model training method, or object classification method, or object search method are implemented.
[0260] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiment of the computing device, since it is basically similar to the embodiment of the sample annotation method, or classification model training method, or object classification method, or object search method, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the embodiment of the sample annotation method, or classification model training method, or object classification method, or object search method.
[0261] An embodiment of this specification also provides a computer-readable storage medium, which stores computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the above sample annotation method, or classification model training method, or object classification method, or object search method are implemented.
[0262] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiment of the computer-readable storage medium, since it is basically similar to the embodiment of the sample annotation method, or classification model training method, or object classification method, or object search method, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the embodiment of the sample annotation method, or classification model training method, or object classification method, or object search method.
[0263] An embodiment of this specification also provides a computer program product, including computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the above sample annotation method, or classification model training method, or object classification method, or object search method are implemented.
[0264] The above is a schematic solution of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solutions of the above sample annotation method, or classification model training method, or object classification method, or object search method belong to the same concept. For the details not described in the technical solution of the computer program product, reference can be made to the descriptions of the technical solutions of the above sample annotation method, or classification model training method, or object classification method, or object search method.
[0265] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0266] The computer instructions include computer program code, which can be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, external hard drives, magnetic disks, optical discs, computer memories, read-only memories (ROMs), random access memories (RAMs), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0267] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that the embodiments of this specification are not limited by the described order of actions, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this specification.
[0268] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments. [[ID=!4]]
[0269] The preferred embodiments of the present specification disclosed above are only used to help explain the present specification. The alternative embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of the embodiments of the present specification. The present specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of the present specification, so that those skilled in the art can well understand and utilize the present specification. The present specification is only limited by the claims and their full scope and equivalents.
Claims
1. A method for sample annotation, characterized in that, Including: Obtain a sample to be labeled; Perform classification prediction on the sample to be labeled to obtain a first classification result corresponding to the sample to be labeled; Obtain second classification results corresponding to multiple labeled samples, and determine a target sample from the multiple labeled samples based on the similarity between the first classification result and each second classification result; Input the target sample and the sample to be labeled into a labeling model to obtain a labeling result of the sample to be labeled.
2. The method according to claim 1, wherein Before searching for the second classification results corresponding to the multiple labeled samples from the sample pool based on the sample information of the multiple labeled samples, it further includes: Obtain multiple labeled samples; Input the multiple labeled samples into a target classification model to obtain second classification results corresponding to each of the labeled samples output by the target classification model, where the target classification model is pre-trained based on the multiple labeled samples; Store the second classification results corresponding to the multiple labeled samples into the sample pool; The obtaining of the second classification results corresponding to the multiple labeled samples includes: Search for the second classification results corresponding to the multiple labeled samples from the sample pool based on the sample information of the multiple labeled samples.
3. The method according to claim 2, wherein Before inputting the multiple labeled samples into the target classification model, it further includes: Obtain an initial classification model, multiple training samples, and label information of the training samples; Train the initial classification model based on the multiple training samples and the label information to obtain a target classification model.
4. The method according to any one of claims 1-3, characterized in that, The performing of classification prediction on the sample to be labeled to obtain a first classification result corresponding to the sample to be labeled includes: Input the sample to be labeled into the target classification model to obtain a first classification result corresponding to the sample to be labeled output by the target classification model, where the target classification model is pre-trained based on multiple training samples and the label information of the training samples.
5. The method according to claim 1, wherein The first classification result includes a first predicted probability distribution, and the second classification result includes a second predicted probability distribution. The first predicted probability distribution represents the probability distribution of the sample to be labeled among multiple candidate labels, and the second predicted probability distribution represents the probability distribution of the multiple labeled samples among the multiple candidate labels; The determining of the target sample from the multiple labeled samples based on the similarity between the first classification result and each second classification result includes: For any labeled sample, calculate the similarity of the predicted probability distribution between the sample to be labeled and the any labeled sample according to the first predicted probability distribution and the second predicted probability distribution of the any labeled sample; Determine at least one target sample from the multiple labeled samples based on the similarity of the predicted probability distribution.
6. The method according to claim 1, wherein Before inputting the target sample and the sample to be labeled into the labeling model to obtain a labeling result of the sample to be labeled, it further includes: Obtain labeling indication information, where the labeling indication information is used to indicate the labeling task of the labeling model; Construct labeling prompt information based on the labeling indication information and the target sample; Inputting the target sample and the sample to be labeled into a labeling model to obtain a labeling result for the sample to be labeled includes: Inputting the labeling prompt information and the sample to be labeled into the labeling model to obtain a labeling result for the sample to be labeled.
7. The method according to claim 1, wherein After inputting the target sample and the sample to be labeled into the labeling model to obtain a labeling result for the sample to be labeled, it further includes: Receiving feedback information regarding the labeling result; Adjusting the labeling model based on the feedback information; or Generating an audit task based on the sample to be labeled and the labeling result, and sending the audit task to an audit terminal; In response to an audit passed instruction sent by the audit terminal, regarding the sample to be labeled as a labeled sample.
8. A method for training a classification model, characterized in that, Includes: Obtaining an initial classification model, multiple training samples, and label information for the training samples, where the label information for the multiple training samples includes labeling results obtained by using the method according to any one of claims 1 - 7; Training the initial classification model based on the multiple training samples and the label information to obtain a target classification model.
9. A method for object classification, characterized in that, Includes: Obtaining an object classification request, where the object classification request includes object data of a target object; Inputting the object data into the target classification model to determine a classification result for the target object, where the target classification model is trained according to the method of claim 8.
10. An object search method, characterized in that, Includes: Obtaining an object search request, where the object search request includes a search statement; Inputting the search statement into the target classification model to determine a search type corresponding to the search statement, where the target classification model is trained according to the method of claim 8; Based on the search type, invoking a target search service to obtain a search result.
11. A sample annotation device, characterized in that, Includes: A first obtaining module configured to obtain a sample to be labeled; A prediction module configured to perform a classification prediction on the sample to be labeled to obtain a first classification result corresponding to the sample to be labeled; A determination module configured to obtain second classification results corresponding to multiple labeled samples, and determine a target sample from the multiple labeled samples based on the similarity between the first classification result and each second classification result; A labeling module configured to input the target sample and the sample to be labeled into a labeling model to obtain a labeling result for the sample to be labeled.
12. A classification model training device, characterized in that, Includes: A second obtaining module configured to obtain an initial classification model, multiple training samples, and label information for the training samples, where the label information for the multiple training samples includes labeling results obtained by using the method according to any one of claims 1 - 9; A training module configured to train the initial classification model based on the multiple training samples and the label information to obtain a target classification model.
13. An object classification device, characterized in that Includes: A third obtaining module configured to obtain an object classification request, where the object classification request includes object data of a target object; The first classification module is configured to input the object data into a target classification model and determine the classification result of the target object, where the target classification model is trained based on the method described in claim 8.
14. An object search device, characterized in that, It includes: The fourth acquisition module is configured to acquire an object search request, where the object search request includes a search statement; The second classification module is configured to input the search statement into a target classification model and determine the search type corresponding to the search statement, where the target classification model is trained based on the method described in claim 8; The search module is configured to call a target search service to obtain a search result based on the search type.
15. A computing device, characterized in that, It includes: A memory and a processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the sample annotation method described in any one of claims 1-7, or the classification model training method described in claim 8, or the object classification method described in claim 9, or the object search method described in claim 10 are implemented.
16. A computer-readable storage medium, characterized in that, It stores computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the sample annotation method described in any one of claims 1-7, or the classification model training method described in claim 8, or the object classification method described in claim 9, or the object search method described in claim 10 are implemented.
17. A computer program product, characterized in that, It includes computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the sample annotation method described in any one of claims 1-7, or the classification model training method described in claim 8, or the object classification method described in claim 9, or the object search method described in claim 10 are implemented.
Citation Information
Patent Citations
Classifier training method, classifier and sentiment classification system
CN105930411A
Method and device for selecting sample image, storage medium and server
CN111310846A
Data screening method and device, electronic equipment and storage medium
CN112560993A
Foundation cloud image cloud detection method based on model self-training
CN112669298A
Text classification model training and classification method and system and data processing system
CN113177119A