Drift sample acquisition method and classification model optimization method

Drift samples are identified through manual sampling and similarity thresholds, and automatic labeling of large models is used to solve the problem of discovery and optimization of drift samples in the model, improving model performance and reducing manpower consumption.

CN120372278APending Publication Date: 2025-07-25ANT ZHIXIN HANGZHOU INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510361616.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The prior art is difficult to effectively discover and optimize drift samples in the model, resulting in a degradation of model performance.

Method used

Through manual sampling of online sample sets, the first drift sample and the second drift sample are automatically identified using the similarity threshold, and the drift sample label is automatically marked with the large model to generate a drift sample set for optimization training for the classification model.

Benefits of technology

Reliance on manual inspections has been reduced, the model performance and usage effect has been significantly improved, and manpower consumption has been reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372278A_ABST
    Figure CN120372278A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a drift sample acquisition method and a classification model optimization method, and the method comprises the steps: determining a difficult sample with a manual judgment label not consistent with a model prediction label in an online sample set through a manual sampling inspection mode, and then carrying out the calculation of the difficult sample for each difficult sample, if the similarity between a target training sample having the highest similarity with the difficult sample in the training sample set and the difficult sample is less than a first similarity threshold, determining the difficult sample as a first drift sample, and further determining an online sample having the similarity with the first drift sample greater than a second similarity threshold in the online sample set as a second drift sample, and generating a drift sample set containing the first drift sample and the second drift sample, and then performing optimization training on the classification model by using the determined drift sample set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to computer technology, and in particular to a drift sample acquisition method and a classification model optimization method. Background Art

[0002] Since the beginning of machine learning, researchers have discovered the phenomenon of data drift. It refers to the significant changes between the data used for model training and the new data in actual application scenarios over time, and this change will lead to a decline in model performance, manifested as a decrease in output accuracy, a decrease in relevance, or a weakening of the prediction effect.

[0003] Therefore, after data drift occurs, how to discover the drifted data samples and optimize the model training based on the drifted data samples to improve model performance and usage effect is a technical problem that needs to be solved urgently. Summary of the invention

[0004] According to the above-mentioned purpose, the embodiments of this specification provide a drift sample acquisition method and a classification model optimization method. By adopting the above-mentioned drift sample acquisition method, a large number of drift samples with data drift in the online sample set can be expanded based on a small number of drift samples manually sampled, without the need to manually check the drift of all samples in the online sample set, thereby reducing the dependence on manual work. Then, through the classification model optimization method, the acquired drift samples are automatically labeled using a large model without relying on manual labeling, and the classification model is optimized and trained using the labeled drift samples, which can significantly improve the model performance and usage effect.

[0005] This specification embodiment proposes a drift sample acquisition method, the method comprising:

[0006] Selecting some online samples from the online sample set as random inspection samples, and determining difficult samples in which the manually judged labels and the model predicted labels do not match each other in the random inspection samples;

[0007] For each of the difficult samples, if the similarity between the target training sample with the highest similarity to the difficult sample in the training sample set and the difficult sample is less than a first similarity threshold, then determining the difficult sample as a first drift sample;

[0008] An online sample in the online sample set whose similarity to the first drift sample is greater than a second similarity threshold is determined as a second drift sample, and a drift sample set including the first drift sample and the second drift sample is generated.

[0009] Further, in some embodiments, the determining that the difficult sample is a first drift sample if the similarity between the training sample with the highest similarity to the difficult sample in the training sample set and the difficult sample is lower than the first similarity threshold includes:

[0010] Calculate the similarity between each training sample in the training sample set and the difficult sample, and determine the target training sample with the highest similarity to the difficult sample in the training sample set;

[0011] If the similarity between the target training sample and the difficult sample is less than the first similarity threshold, determine that the difficult sample is a first drift sample.

[0012] Further, in some embodiments, the calculating the similarity between each training sample in the training sample set and the difficult sample includes:

[0013] Perform word segmentation on the difficult sample to obtain a set of words corresponding to the difficult sample;

[0014] Calculate the similarity between the set of words and each training sample in the training sample set based on a text matching algorithm.

[0015] Further, in some embodiments, the determining that the online sample in the online sample set with a similarity greater than the second similarity threshold to the first drift sample is a second drift sample includes:

[0016] Calculate the similarity between each online sample in the online sample set and the first drift sample;

[0017] Determine that the online sample with a similarity greater than the second similarity threshold to the first drift sample is a second drift sample.

[0018] Further, in some embodiments, the text matching algorithm includes, but is not limited to, the TF-IDF algorithm and the BM25 algorithm.

[0019] The embodiments of this specification also provide a classification model optimization method, including:

[0020] Obtain a drift sample set by using a drift sample acquisition method;

[0021] Re-label each drift sample in the drift sample set through a large model to obtain the sample label corresponding to each drift sample;

[0022] Add each drift sample and the sample label corresponding to each drift sample to the training sample set to obtain an updated training sample set;

[0023] Optimize and train the classification model based on the updated training sample set to obtain an optimized classification model.

[0024] An embodiment of this specification also proposes a drifting sample acquisition device, including:

[0025] A difficult sample determination module, configured to extract some online samples from the online sample set as spot-check samples, and determine difficult samples in each of the spot-check samples where the manually judged label and the model predicted label do not match;

[0026] A first sample determination module, for each of the difficult samples, if the similarity between the target training sample in the training sample set with the highest similarity to the difficult sample and the difficult sample is less than the first similarity threshold, then determine the difficult sample as the first drifting sample;

[0027] A second sample determination module, configured to determine the online samples in the online sample set with a similarity greater than the second similarity threshold to the first drifting sample as the second drifting samples, and generate a drifting sample set including the first drifting samples and the second drifting samples.

[0028] An embodiment of this specification also proposes a classification model optimization device, including:

[0029] A sample set acquisition module, configured to obtain a drifting sample set by using a drifting sample acquisition method;

[0030] A label qualification module, configured to re-qualify the labels of each drifting sample in the drifting sample set through a large model to obtain the sample labels corresponding to each drifting sample;

[0031] A sample set update module, configured to add each of the drifting samples and the sample labels corresponding to each of the drifting samples to the training sample set to obtain an updated training sample set;

[0032] An optimization training module, configured to optimize and train the classification model based on the updated training sample set to obtain an optimized classification model.

[0033] An embodiment of this specification also provides a computer program product, where the computer program product stores at least one instruction, and the at least one instruction is suitable for being loaded and executed by a processor to perform the above method steps.

[0034] An embodiment of this specification also provides a computer storage medium, where the computer storage medium stores a computer program, and the computer program is suitable for being loaded and executed by a processor to perform the steps of the above method.

[0035] An embodiment of this specification also provides an electronic device, including: a processor and a memory; wherein, the memory stores a computer program, and the computer program is adapted to be loaded and executed by the processor to perform the steps of the above method.

[0036] In the embodiment of this specification, by adopting the above drift sample acquisition method, a large number of drift samples with data drift in the online sample set can be expanded based on a small number of drift samples manually sampled, without the need for manual drift inspection of all samples in the online sample set, reducing the dependence on manual labor. Then, through the classification model optimization method, the large model is used to automatically label the acquired drift samples without relying on manual labeling, and the labeled drift samples are used to optimize and train the classification model, which can significantly improve the model performance and usage effect. Moreover, the drift sample acquisition method and classification model optimization method provided in the embodiment of this specification only require a small amount of manual labor for online sample sampling in the actual application process, can automatically find drift samples, amplify drift samples, and optimize and train the classification model, reducing labor consumption while optimizing the model performance, and having the effect of reducing costs and increasing efficiency. Description of the Drawings

[0037] Figure 1 It is a schematic flowchart of a drift sample acquisition method provided by an embodiment of this specification;

[0038] Figure 2 It is a schematic flowchart of a drift sample acquisition method provided by an embodiment of this specification;

[0039] Figure 3 It is a schematic diagram of the data relationship of determining drift samples provided by an embodiment of this specification;

[0040] Figure 4 It is a schematic flowchart of a classification model optimization method provided by an embodiment of this specification;

[0041] Figure 5 It is a schematic structural diagram of a drift sample acquisition device provided by an embodiment of this specification;

[0042] Figure 6 It is a schematic structural diagram of a classification model optimization device provided by an embodiment of this specification;

[0043] Figure 7 It is a schematic structural diagram of an electronic device provided by an embodiment of this specification. Detailed Embodiments

[0044] To make the purpose, technical solutions, and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by this specification.

[0045] For machine learning, data drift is a common phenomenon over time. Data drift is mainly divided into three types: 1) Concept drift: The definition of the label changes with the data domain; 2) Label drift: The meaning of the label remains unchanged, and only the label distribution changes with the data domain; 3) Feature drift: The labeling function P(Y|X) is fixed, and only the marginal distribution P(X) changes with the data domain, that is, P(Y|X) = Q(Y|X), P(X) ≠ Q(X). Statisticians call this covariate shift because features are also called covariates. For text, data drift means that the overall text expression features have shifted, usually occurring in scenarios where the topic or the scope covered by the data source changes. Generally speaking, it may be that the semantics, language patterns, keywords mentioned, expression formats, etc. have changed.

[0046] Data drift will lead to a decline in model performance, manifested as a decrease in output accuracy, relevance, or a weakening of the prediction effect.

[0047] Based on this, the embodiments of this specification provide a method for obtaining drift samples and a method for optimizing a classification model. By using the above method for obtaining drift samples, a large number of drift samples with data drift in the online sample set can be expanded based on a small number of drift samples randomly selected by humans. There is no need for humans to check for drift in all samples in the online sample set, reducing the dependence on humans. Then, through the method for optimizing the classification model, the large model is used to automatically label the obtained drift samples without relying on manual labeling, and the labeled drift samples are used to optimize and train the classification model, which can significantly improve the model performance and usage effect. Moreover, the method for obtaining drift samples and the method for optimizing the classification model provided by the embodiments of this specification only require a small amount of human labor for random sampling of online samples in the actual application process, can automatically find drift samples, expand drift samples, and optimize and train the classification model, reducing labor consumption while optimizing the model performance, and having the effect of reducing costs and increasing efficiency.

[0048] Please refer to Figure 1 , which is a schematic flowchart of a method for obtaining drift samples provided by the embodiments of this specification. In the embodiments of this specification, the method for obtaining drift samples is applied to a drift sample acquisition device or an electronic device configured with a drift sample acquisition device. The following will be directed at Figure 1The following describes the process shown in detail. The drift sample acquisition method may specifically include the following steps:

[0049] S102: Extract some online samples from the online sample set as spot-check samples, and determine difficult samples in each spot-check sample where the manually judged label does not match the model predicted label.

[0050] Among them, the online sample set is the sample data flowing into the line during the model application period.

[0051] After the model goes online, as time goes by, the data flowing into the line will change, and the phenomenon of data drift will occur. During the model online application period, manual spot-checks will be carried out on the online sample data to determine whether data drift has occurred.

[0052] In one or more embodiments of this specification, an online sample set is generated according to the sample data flowing into the model. Some online samples are extracted from the online sample set by humans as spot-check samples, and humans judge the labels of each spot-check sample. If the manually judged label does not match the model predicted label, then determine that this spot-check sample is a difficult sample.

[0053] S104: For each difficult sample, if the similarity between the target training sample with the highest similarity to the difficult sample in the training sample set and the difficult sample is less than the first similarity threshold, then determine the difficult sample as the first drift sample.

[0054] Among them, the training sample set is the sample set used for model training.

[0055] It is not difficult to understand that the manually judged label has personal subjective awareness, and it is impossible to directly determine that this difficult sample is a drift sample relative to the training sample set through manual judgment.

[0056] In one or more embodiments of this specification, after determining the difficult samples, for each difficult sample, judge whether the difficult sample is a drift sample by comparing the similarity between each training sample in the training sample set and the difficult sample; if the similarity between the target training sample with the highest similarity to the difficult sample in the training sample set and the difficult sample is less than the first similarity threshold, then determine the difficult sample as the first drift sample; if the similarity between the target training sample with the highest similarity to the difficult sample in the training sample set and the difficult sample is greater than or equal to the first similarity threshold, then determine that the difficult sample is not the first drift sample.

[0057] It can be understood that if all the training samples in the training sample set are not similar to the difficult sample, then determine that this difficult sample is the first drift sample; if there are training samples in the training sample set that are similar enough to the difficult sample, then this difficult sample has not deviated from the training sample set, and determine that this difficult sample is not the first drift sample.

[0058] S106. Determine the online samples in the online sample set that are more similar to the first drifting sample than the second similarity threshold as the second drifting samples, and generate a drifting sample set that includes the first drifting sample and the second drifting samples.

[0059] In one or more embodiments of this specification, after determining the first drifting sample, calculate the similarity between each online sample in the online sample set and the difficult samples to determine whether each online sample is a drifting sample; if the similarity between an online sample and the first drifting sample is greater than the second similarity threshold, it indicates that the online sample is similar enough to the first drifting sample, and then determine that the online sample is the second drifting sample. Finally, generate a drifting sample set that includes the first drifting sample and the second drifting samples.

[0060] It can be understood that the first drifting sample is only the drifting sample determined from some of the online samples inspected manually, and there are still a large number of uninspected online samples in the online sample set. By calculating the similarity between each online sample in the online sample set and the difficult samples, and determining the online samples in the online sample set that are more similar to the first drifting sample than the second similarity threshold as the second drifting samples, a large number of drifting samples can be automatically determined from the online sample set, expanding the number of drifting samples.

[0061] In the embodiments of this specification, a large number of drifting samples with data drift in the online sample set can be expanded based on a small number of drifting samples inspected manually, without the need for manual drift inspection of all samples in the online sample set, reducing the dependence on manual labor. The finally obtained drifting samples can be used for the optimized training of the model, improving the model performance and usage effect.

[0062] In one embodiment, please refer to Figure 2 , step S104. For each of the difficult samples, if the similarity between the target training sample with the highest similarity to the difficult sample in the training sample set and the difficult sample is less than the first similarity threshold, then determine the difficult sample as the first drifting sample. Specifically, it may include the following steps:

[0063] Step S1042. Calculate the similarity between each training sample in the training sample set and the difficult sample, and determine the target training sample with the highest similarity to the difficult sample in the training sample set.

[0064] Step S1044. If the similarity between the target training sample and the difficult sample is less than the first similarity threshold, then determine the difficult sample as the first drifting sample.

[0065] In the embodiments of this specification, after obtaining difficult samples through random inspection, for each difficult sample, calculate the similarity between each training sample in the training sample set and the difficult sample, and determine the target training sample in the training sample set with the highest similarity to the difficult sample. After determining the target training sample with the highest similarity to the difficult sample, if the similarity between the target training sample and the difficult sample is less than the first similarity threshold, it indicates that there is no training sample in the training sample set that is similar enough to the difficult sample, and then determine the difficult sample as the first drift sample with data drift. At this time, the detection accuracy of the model for the difficult sample is low and the effect is poor.

[0066] In a feasible implementation, calculating the similarity between each training sample in the training sample set and the difficult sample can specifically be: First, perform word segmentation on the difficult sample to obtain a set of words corresponding to the difficult sample, and then calculate the similarity between the set of words and each training sample in the training sample set based on the text matching algorithm.

[0067] Among them, word segmentation of the difficult sample can be performed based on a pre-trained large model. The text matching algorithm includes but is not limited to the TF-IDF algorithm and the BM25 algorithm.

[0068] Further, in one embodiment, in step S106, determining the online samples in the online sample set with a similarity greater than the second similarity threshold to the first drift sample as the second drift sample can specifically be: Calculate the similarity between each online sample in the online sample set and the first drift sample, and then determine the online samples with a similarity greater than the second similarity threshold to the first drift sample as the second drift sample.

[0069] Among them, the similarity between each online sample and the first drift sample can be calculated based on the text matching algorithm, and then select the online samples with a similarity greater than the second similarity threshold as the second drift sample.

[0070] It is not difficult to understand that the first drift sample is only the drift sample determined from a part of the online samples randomly inspected manually, and there are still a large number of uninspected online samples in the online sample set. By calculating the similarity between each online sample in the online sample set and the difficult sample, and determining the online samples in the online sample set with a similarity greater than the second similarity threshold to the first drift sample as the second drift sample, a large number of drift samples can be automatically determined from the online sample set, expanding the number of drift samples.

[0071] Please refer to Figure 3 , for a schematic diagram of the data relationship for determining drift samples provided by the embodiments of this specification. As Figure 3As shown, the circle 1 represents the training sample set, the circle 2 represents the online sample set, the circle 3 represents the difficult samples randomly selected from the online sample set manually, the circle 4 represents the first drift samples determined from the difficult samples by calculating the similarity between the difficult samples and each training sample in the training sample set, and the circle 5 represents the drift sample set including the first drift samples and the second drift samples obtained by calculating the similarity between the first drift samples and each online sample in the online sample set and amplifying according to the similarity.

[0072] Please refer to Figure 4 , which is a schematic flowchart of a classification model optimization method provided by an embodiment of this specification. In the embodiment of this specification, the classification model optimization method is applied to a classification model optimization device or an electronic device configured with a classification model optimization device. The following will elaborate in detail on Figure 4 the process shown, and the classification model optimization method may specifically include the following steps:

[0073] S202. Obtain a drift sample set based on a drift sample acquisition method;

[0074] Obtain a drift sample set from the online sample set through the drift sample acquisition methods proposed in the above embodiments.

[0075] S204. Re-perform label qualification on each drift sample in the drift sample set through a large model to obtain the sample labels corresponding to each drift sample respectively;

[0076] In the embodiment of this specification, after obtaining the drift sample set, re-perform label qualification on each drift sample in the drift sample set through a pre-trained large model to obtain the sample labels corresponding to each drift sample respectively.

[0077] Among them, the large model can perform label qualification on the drift samples in the ICL manner based on zero-shot or few-shot. Learning can be carried out on a specific task with a small number of labeled samples. By designing a task-related instruction to form a prompt template, using a small number of labeled samples or no samples as prompts to guide the large model to generate prediction results on the drift samples.

[0078] Zero-shot means that the model has not encountered samples of a certain category during the training process, but can still classify unseen categories through category descriptions; Zero-Shot prompt means that the model generates responses only based on the task description without examples. Few-shot allows the model to learn new categories with a limited number of examples.

[0079] S206. Add each drift sample and the sample label corresponding to each drift sample to the training sample set to obtain an updated training sample set;

[0080] S208. Optimize and train the classification model based on the updated training sample set to obtain an optimized classification model.

[0081] In the embodiments of this specification, after qualitatively determining the labels of each drift sample in the drift sample set, add each drift sample and the sample label corresponding to each drift sample to the training sample set to obtain an updated training sample set. Then, use the updated training sample set to optimize and train the classification model to obtain an optimized classification model. This is to improve the model effect and accuracy of the classification model and eliminate the impact of data drift during model use on the application effect of the classification model.

[0082] Please refer to Figure 5 , which is a schematic structural diagram of a drift sample acquisition device provided by the embodiments of this specification. As Figure 5 shown, the drift sample acquisition device 1 can be implemented as all or part of an electronic device through software, hardware, or a combination of both. According to some embodiments, the drift sample acquisition device 1 includes a difficult sample determination module 11, a first sample determination module 12, and a second sample determination module 13, specifically including:

[0083] The difficult sample determination module 11 is used to extract some online samples from the online sample set as spot-check samples, and determine difficult samples in each of the spot-check samples where the manually judged label and the model prediction label do not match.

[0084] The first sample determination module 12 is used to, for each of the difficult samples, if the similarity between the target training sample in the training sample set with the highest similarity to the difficult sample and the difficult sample is less than the first similarity threshold, determine the difficult sample as the first drift sample.

[0085] The second sample determination module 13 is used to determine the online samples in the online sample set with a similarity greater than the second similarity threshold to the first drift sample as the second drift samples, and generate a drift sample set including the first drift samples and the second drift samples.

[0086] Optionally, the first sample determination module 12 is specifically used for:

[0087] Calculate the similarity between each training sample in the training sample set and the difficult sample, and determine the target training sample in the training sample set with the highest similarity to the difficult sample;

[0088] If the similarity between the target training sample and the difficult sample is less than the first similarity threshold, determine the difficult sample as the first drift sample.

[0089] Optionally, when the first sample determination module 12 calculates the similarity between each training sample in the calculated training sample set and the difficult sample, it is specifically configured to:

[0090] Perform word segmentation on the difficult sample to obtain a set of words corresponding to the difficult sample;

[0091] Calculate the similarity between the set of words and each training sample in the training sample set based on a text matching algorithm.

[0092] Optionally, the second sample determination module 13 is specifically configured to:

[0093] Calculate the similarity between each online sample in the online sample set and the first drift sample;

[0094] Determine that the online sample whose similarity to the first drift sample is greater than the second similarity threshold is the second drift sample.

[0095] Optionally, the text matching algorithm includes but is not limited to the TF-IDF algorithm and the BM25 algorithm.

[0096] The above device embodiments correspond to the method embodiments. For specific descriptions, reference can be made to the description in the method embodiment section, which will not be elaborated here. The device embodiments are obtained based on the corresponding method embodiments and have the same technical effects as the corresponding method embodiments. For specific descriptions, reference can be made to the corresponding method embodiments.

[0097] Please refer to Figure 6 , which is a schematic structural diagram of a classification model optimization device provided in an embodiment of this specification. As Figure 6 shown, the classification model optimization device 2 can be implemented as all or part of an electronic device through software, hardware, or a combination of both. According to some embodiments, the classification model optimization device 2 includes a sample set acquisition module 21, a label qualification module 22, a sample set update module 23, and an optimization training module 24, specifically including:

[0098] The sample set acquisition module 21 is configured to acquire a drift sample set by using a drift sample acquisition method;

[0099] The label qualification module 22 is configured to re-qualify the labels of each drift sample in the drift sample set through a large model to obtain sample labels corresponding to each drift sample;

[0100] The sample set update module 23 is configured to add each drift sample and the sample label corresponding to each drift sample to the training sample set to obtain an updated training sample set;

[0101] The optimization training module 24 is configured to perform optimization training on the classification model based on the updated training sample set to obtain an optimized classification model.

[0102] The above device embodiments correspond to the method embodiments. For specific descriptions, reference can be made to the descriptions in the method embodiment section, which will not be elaborated here. The device embodiments are obtained based on the corresponding method embodiments and have the same technical effects as the corresponding method embodiments. For specific descriptions, reference can be made to the corresponding method embodiments.

[0103] An embodiment of this specification also provides a computer storage medium, which can store multiple instructions. The instructions are suitable for being loaded and executed by a processor to perform the drift sample acquisition method and the classification model optimization method as described in the above Figures 1 to 4 shown embodiments. The specific execution process can be referred to Figures 1 to 4 the specific descriptions of the shown embodiments and will not be elaborated here.

[0104] This specification also provides a computer program product. The computer program product stores at least one instruction. The at least one instruction is loaded and executed by the processor to perform the drift sample acquisition method and the classification model optimization method as described in the above Figures 1 to 4 shown embodiments. The specific execution process can be referred to Figures 1 to 4 the specific descriptions of the shown embodiments and will not be elaborated here.

[0105] An embodiment of this specification also provides a schematic structural diagram of the electronic device shown in FIG. 7. As Figure 7 shown, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, other hardware required for other services may also be included. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above drift sample acquisition method and classification model optimization method.

[0106] Of course, in addition to the software implementation, this specification does not exclude other implementation manners, such as logical devices or a combination of software and hardware. That is to say, the execution subject of the following processing flow is not limited to each logical unit and can also be hardware or a logical device.

[0107] In the 1990s, it was obvious to distinguish whether an improvement to a technology was an improvement in hardware (e.g., improvement to circuit structures such as diodes, transistors, switches, etc.) or an improvement in software (improvement to method flows). However, with the development of technology, many improvements to method flows today can be regarded as direct improvements to hardware circuit structures. Almost all designers obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented with a hardware entity module. For example, a Programmable Logic Device (PLD) (e.g., a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can program themselves to "integrate" a digital system onto a single PLD, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL). There is not only one type of HDL, but many types, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing a little logical programming on the method flow with the above-mentioned several hardware description languages and programming it into an integrated circuit, it is easy to obtain a hardware circuit that implements the logical method flow.

[0108] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that, in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. to achieve the same functions. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.

[0109] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0110] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0111] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.

[0112] This specification is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the specification. It should be understood that each flow and / or block in the flowchart and / or block diagram, and combinations of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0113] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0114] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0115] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.

[0116] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.

[0117] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0118] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0119] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0120] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0121] The various embodiments in this specification are described in a progressive manner. For the same or similar parts among the various embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiments.

[0122] The above is only the embodiments of this specification and is not intended to limit this specification. For those skilled in the art, various modifications and changes can be made to this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims of this specification.

Claims

1. A drift sample acquisition method, comprising: Selecting some online samples from the online sample set as random inspection samples, and determining difficult samples in which the manually judged labels and the model predicted labels do not match each other in the random inspection samples; For each of the difficult samples, if the similarity between the target training sample with the highest similarity to the difficult sample in the training sample set and the difficult sample is less than a first similarity threshold, then determining the difficult sample as a first drift sample; An online sample in the online sample set whose similarity to the first drift sample is greater than a second similarity threshold is determined as a second drift sample, and a drift sample set including the first drift sample and the second drift sample is generated.

2. The method according to claim 1, wherein if the similarity between the training sample with the highest similarity to the difficult sample in the training sample set and the difficult sample is lower than a first similarity threshold, determining that the difficult sample is a first drift sample comprises: Calculating the similarity between each training sample in the training sample set and the difficult sample, and determining a target training sample in the training sample set having the highest similarity to the difficult sample; If the similarity between the target training sample and the difficult sample is less than the first similarity threshold, the difficult sample is determined to be a first drift sample.

3. The method according to claim 2, wherein the calculating the similarity between each training sample in the training sample set and the difficult sample comprises: Performing word segmentation processing on the difficult sample to obtain a word set corresponding to the difficult sample; The similarity between the word set and each training sample in the training sample set is calculated based on a text matching algorithm.

4. The method according to claim 1, wherein determining that the online sample in the online sample set whose similarity with the first drift sample is greater than a second similarity threshold is the second drift sample comprises: Calculate the similarity between each online sample in the online sample set and the first drift sample; An online sample whose similarity to the first drift sample is greater than a second similarity threshold is determined as a second drift sample.

5. According to the method of claim 1, the text matching algorithm includes but is not limited to the TF-IDF algorithm and the BM25 algorithm.

6. A classification model optimization method, comprising: Acquire a drift sample set using the drift sample acquisition method as described in claims 1 to 4; Re-labeling and qualitatively analyzing each drift sample in the drift sample set by using a large model to obtain a sample label corresponding to each drift sample; Adding each of the drift samples and the sample labels respectively corresponding to each of the drift samples to the training sample set to obtain an updated training sample set; The classification model is optimized and trained based on the updated training sample set to obtain an optimized classification model.

7. A drift sample acquisition device, comprising: A difficult sample determination module is used to extract some online samples from the online sample set as random inspection samples, and determine difficult samples whose labels determined by manual judgment and labels predicted by the model do not match in each of the random inspection samples; The first sample determination module is configured to, for each of the difficult samples, if the similarity between the target training sample with the highest similarity to the difficult sample in the training sample set and the difficult sample is less than the first similarity threshold, determine the difficult sample as the first drift sample; The second sample determination module is configured to determine the online samples in the online sample set with a similarity greater than the second similarity threshold to the first drift sample as the second drift samples, and generate a drift sample set including the first drift samples and the second drift samples.

8. A classification model optimization device, comprising: A sample set acquisition module for acquiring a drift sample set by using a drift sample acquisition method; A label qualification module for re-qualifying the labels of each drift sample in the drift sample set through a large model to obtain the sample labels corresponding to each drift sample respectively; A sample set update module for adding each of the drift samples and the sample labels corresponding to each of the drift samples to the training sample set to obtain an updated training sample set; An optimization training module for optimizing and training the classification model based on the updated training sample set to obtain an optimized classification model.

9. A storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 or 6 are implemented.

10. An electronic device, comprising: A processor and a memory; wherein, the memory stores a computer program, and the computer program is adapted to be loaded and executed by the processor to perform the steps of the method according to any one of claims 1 to 5 or 6.

11. A computer program product, on which at least one instruction is stored, and when the at least one instruction is executed by a processor, the steps of the method according to any one of claims 1 to 5 or 6 are implemented.