Data classification

By dividing the training samples into classification and inference tasks and independently optimizing the machine learning model, the performance limitations and high resource consumption of CoT response are solved, achieving higher accuracy and reliability.

CN122249821APending Publication Date: 2026-06-19FACE CUTE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
FACE CUTE CO LTD
Filing Date
2025-04-28
Publication Date
2026-06-19

AI Technical Summary

Technical Problem

Existing machine learning models suffer from limitations in predicting CoT responses and their associated tasks, including high resource consumption, sensitivity to noise, and high compound error rates. In particular, when integrating classification and inference tasks, model performance and reliability are affected.

Method used

By dividing the training samples into a first sample for classification and a second sample for inference, the classification and CoT inference capabilities of the machine learning model are optimized independently. The model is trained separately using two cue words, which reduces cognitive burden and lowers the risk of error propagation.

Benefits of technology

It improves the accuracy and reliability of machine learning models in classification and CoT inference tasks, reduces the overall error rate, and optimizes model performance and resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122249821A_ABST
    Figure CN122249821A_ABST
Patent Text Reader

Abstract

A method, apparatus, and computer program product for data classification are provided. In this method, samples are acquired for training a machine learning model. The samples include prompt words and responses to the prompt words; the prompt words include input data, and the responses include the classification of the input data and the reason why the input data belongs to that classification. A first sample is determined based on the input data and its classification; the first sample includes a first prompt word and a first response. A second sample is determined based on the input data, its classification, and the reason; the second sample includes a second prompt word and a second response. The machine learning model is updated based on the first and second samples. Therefore, the machine learning model can be updated in a more reliable and accurate manner.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-references

[0001] This application claims the benefit of U.S. Patent Application 18 / 920,692, filed October 18, 2024, entitled “Data Classification,” the entire contents of which are incorporated herein by reference. Technical Field

[0002] This disclosure relates generally to machine learning, and more specifically to methods, apparatus, and computer program products for data classification. Background Technology

[0003] With the continuous technological advancements in machine learning, particularly the innovation and development of language models, their integration and application have become widespread across various fields and industries. Nevertheless, numerous challenges remain related to the application and utilization of machine learning models, especially when CoT responses are required and their primary predictions for associated tasks are needed. Therefore, improving the performance of machine learning models when CoT responses are required is desirable. Summary of the Invention

[0004] In a first aspect of this disclosure, a method for data classification is provided. In this method, samples for training a machine learning model are acquired. The samples include a cue word and a response to the cue word, the cue word including input data, and the response including a classification of the input data and a reason why the input data belongs to that classification. A first sample is determined based on the input data and its classification, and the first sample includes a first cue word and a first response. A second sample is determined based on the input data, its classification, and the reason why the input data belongs to, the second sample including a second cue word and a second response. The machine learning model is updated based on the first and second samples.

[0005] In a second aspect of this disclosure, an electronic device is provided. The electronic device includes a computer processor coupled to a computer-readable storage unit, the storage unit including instructions that, when executed by the computer processor, implement the method according to a first aspect of this disclosure.

[0006] In a third aspect of this disclosure, a computer program product is provided, comprising a computer-readable storage medium having program instructions embodied therein, which are executed by an electronic device to cause the electronic device to perform the method according to a first aspect of this disclosure.

[0007] The present invention is provided to present in a simplified form the selection of concepts further described in the specific implementations below. The present invention is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description

[0008] The above and other objects, features and advantages of this disclosure will become more apparent from the accompanying drawings, which describe some implementations of this disclosure in more detail, wherein the same reference numerals generally refer to the same parts in the implementations of this disclosure.

[0009] Figure 1 This diagram illustrates the training of a machine learning model using a single prompt word.

[0010] Figure 2 An example diagram illustrating data classification according to an implementation of this disclosure is shown;

[0011] Figure 3 Example diagrams of a sample implementation according to this disclosure are shown;

[0012] Figure 4 An example diagram showing the task division according to the implementation of this disclosure is provided;

[0013] Figure 5A An example diagram of a first template according to an implementation of this disclosure is shown;

[0014] Figure 5B An example diagram of a second template according to an implementation of this disclosure is shown;

[0015] Figure 6 An example diagram is shown illustrating the construction of a dataset for training a machine learning model according to an implementation of this disclosure;

[0016] Figure 7 An example flowchart of a method for data classification according to an implementation of this disclosure is shown; and

[0017] Figure 8 A block diagram of a computing device in which various implementations of the present disclosure may be implemented is shown. Detailed Implementation

[0018] The principles of this disclosure will now be described with reference to some implementations. It should be understood that these implementations are described for illustrative purposes only and to assist those skilled in the art in understanding and implementing this disclosure, and do not imply any limitation on the scope of this disclosure. The disclosure described herein can be implemented in various ways other than those described below.

[0019] In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.

[0020] References to "an implementation," "implementation," "example implementation," etc., in this disclosure indicate that the described implementation may include specific features, structures, or characteristics, but not every implementation necessarily includes such features, structures, or characteristics. Furthermore, such phrases do not necessarily refer to the same implementation. Moreover, when a specific feature, structure, or characteristic is described in conjunction with an example implementation, it can be assumed that, whether explicitly described or not, the influence of such feature, structure, or characteristic on other implementations is within the knowledge of those skilled in the art.

[0021] It should be understood that although the terms “first” and “second” may be used in this document to describe various elements, these elements should not be limited by these terms. These terms are used only to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the example implementation. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.

[0022] The terminology used herein is for the purpose of describing a particular implementation only and is not intended to limit the example implementations. As used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It will be further understood that the terms “comprising,” “including,” “having,” “having,” “containing,” and / or “containing” as used herein specify the presence of the stated features, elements, and / or components, but do not preclude the presence or addition of one or more other features, elements, components, and / or combinations thereof.

[0023] The principles of this disclosure will now be described with reference to some implementations. It should be understood that these implementations are described for illustrative purposes only and to assist those skilled in the art in understanding and implementing this disclosure, without implying any limitation on the scope of this disclosure. The disclosure described herein can be implemented in various ways other than those described below. In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.

[0024] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of relevant laws, regulations and rules.

[0025] It is understood that before using the technical solutions disclosed in the various implementations of this disclosure, users should be notified in an appropriate manner, in accordance with relevant laws and regulations, of the types, scope of use, and usage scenarios of the personal information involved in this disclosure, and user authorization should be obtained.

[0026] For example, in response to receiving a user's active request, a prompt message is sent to the user to explicitly notify the user that the requested operation will require the acquisition and use of the user's personal information. Therefore, the user can independently choose, based on the prompt message, whether to provide personal information to the software or hardware such as electronic devices, applications, servers, or storage media that perform the operations of the technical solutions disclosed herein.

[0027] As an optional but not limited implementation, the method of sending a prompt to the user in response to a user's active request may include, for example, a pop-up window, in which the prompt message may be presented in text form. Furthermore, the pop-up window may also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0028] It is understood that the above-described notification and user authorization process is merely exemplary and does not limit the implementation of this disclosure. Other methods that comply with applicable laws and regulations are also applicable to the implementation of this disclosure.

[0029] As briefly mentioned above, machine learning models have become commonplace across various fields and industries. Machine learning models can include millions or even billions of parameters and require vast amounts of input data for training. However, these parameters are designed and distributed across various hierarchical structures within the model architecture, and can capture complex interrelationships within the input data. For example, machine learning models can identify and map complex relationships within image pixels, morphemes in speech and text. The advantage of machine learning models lies in the inherent depth of their ability to capture relationships. Due to the large number of parameters, machine learning models have an inherent ability to understand and map complex abstract relationships and concepts, not just simple patterns, compared to smaller models. Therefore, machine learning models excel at complex tasks such as image recognition, natural language processing, and speech and audio analysis.

[0030] However, there are many challenges regarding the application and utilization of machine learning models. First, their training and development require significant computational resources, thus their performance and training may be limited by the availability of their hardware and software. Second, the complexity and scale of these machine learning models are related to the increasing optimization and tuning of model parameters. Furthermore, machine learning models may be more sensitive to noise and anomalies in the input data, which can affect the accuracy of their predictions and outputs based on the interpretations and inferences they generate for predictions. These challenges are particularly pronounced when dealing with problems requiring the CoT response and its primary predictions for related tasks.

[0031] CoT refers to the reasoning process a model generates for its primary prediction task. For example, image classification and text classification problems might fall under the CoT category, where the model not only predicts an image / text and classifies it into a specific category / label, but also provides an explanation / reasoning process regarding why the image / text belongs to the predicted category / label. For instance, a machine learning model can predict that an image containing tall buildings and skyscrapers will be classified as "urban," where the model's CoT response is "skyscrapers are more likely to indicate an urban area." Text consisting of movie reviews can be predicted by a machine learning model to be classified as "positive opinions," where the model's CoT response is "this review contains praise for the movie."

[0032] For the above classification task, considering CoT, the following three methods are used to train the machine learning model for classification.

[0033] The first approach directly reuses a pre-trained machine learning model, where the input data (images / text) is formatted as part of a cue word, which is then fed into the machine learning model for predictions about its category / label and CoT inference. This method is concise and requires no supervised fine-tuning, as it directly utilizes the model's existing capabilities and avoids many model training-related problems. Furthermore, this method is practical for scenarios requiring minimal structured responses, and by following a few cue word structures, the model can provide good answers. However, this method may only be suitable for scenarios requiring specifically structured formatted responses. Additionally, if the cue words are highly specialized, the model may face length limitations in its responses, and its comprehension ability will also decrease. This method is only suitable for simple cue words; for very complex cue words without examples, the model struggles to respond accurately.

[0034] In the second approach, the input data is annotated with its actual category / label and CoT inference to form a training dataset. This training dataset is used to supervise and fine-tune the trained machine learning model, thereby improving the effectiveness and accuracy of the machine learning model in its category prediction and subsequent inference. Regarding this method, the structural similarity of responses can be learned into the model parameters through training samples, making the model more accurate and concise in its inference behind responses and its primary category prediction. Furthermore, this method enables the customization of models for multiple applications and tasks in different domains, making them general and flexible to meet various industry requirements. However, training the model requires additional resources, such as GPU units, larger memory storage, and higher power requirements. For efficient model training, the training dataset, along with the CoT interpretation, needs to be of very high quality, and a high data matching ratio is also required.

[0035] In the third approach, the input data is annotated with the actual class / label and CoT inference to form a training dataset, which is used to first supervise and fine-tune the trained machine learning model. Then, the model's output predictions, along with the CoT, are compared with the actual class / label and CoT, and the model is further fine-tuned using human-annotated responses so that its generated results more closely match human-annotated responses. This not only improves the effectiveness and accuracy of the machine learning model in terms of class prediction and subsequent inference, but also improves the structural development of its inference, making it more consistent with human-annotated responses. Regarding this approach, alignment with human preference data allows the model to understand and align with human preferences and inference more deeply, potentially reducing its prediction error rate. By learning from human intervention, the model becomes more robust to anomalous or adversarial inputs. However, this approach is resource-intensive, especially when composed of large datasets, as it requires significant human resources for continuous preference feedback. Models trained using reinforcement learning (RLHF) with human feedback may suffer from problems such as hallucinations and biases, which not only further reduce model accuracy but may also pose potential risks to business security and confidential information. Furthermore, alignment with human feedback reduces the diversity of samples generated by the model, leading to model degradation / collapse as it deviates from the original predictive optimization to human preference optimization.

[0036] As can be seen from the above, although the second method can solve the CoT-based classification problem, it uses supervised fine-tuning to fine-tune the machine learning model. However, a common practice for this fine-tuning scheme is to collect existing business data and then fine-tune that data using a single prompt word. This prompt word requests both primary classification / labeling and CoT inference, with the CoT inference having two objectives: the first is to accurately identify / classify the input text into a predefined category / label, and the second is to provide a detailed explanation / description for predicting the following category / label.

[0037] Figure 1 This illustrates a diagram of training a machine learning model using a single prompt word. (For example...) Figure 1As shown, samples 112 are obtained for training machine learning model 120. Sample 112 includes a cue word 110 and a response 130. The cue word 110 includes the input data (denoted as X), and the response 130 includes the classification of the input data (denoted as Y) and the reason why the input data belongs to the classification (denoted as J). Machine learning model 120 is updated based on sample 112. However, the single cue word approach may have the following drawbacks. First, business data often consists of complex and varied datasets that indicate significant data sparsity. Combining these two tasks within a single cue word increases the cognitive load on the model, making it more challenging to simultaneously optimize both classification and CoT inference, thus impacting overall model performance as it may lead to increased errors in both classification and inference. Furthermore, integrating the classification and CoT-based inference tasks into a single cue word also creates a higher order dependency on model completion and prediction, so if the model makes an error in classification, it will propagate to CoT inference and lead to a compound error effect that may reduce the overall performance and reliability of the model.

[0038] In view of the above, this disclosure refers to Figure 2 A data classification scheme is proposed, illustrated in Figure 200 as an example of data classification according to an implementation of this disclosure. Figure 2 As shown, samples 112 are obtained for training machine learning model 120. Sample 112 includes a cue word 110 and a response 130 to the cue word 110. The cue word includes input data (denoted as X), and the response includes the classification of the input data (denoted as Y) and the reason why the input data belongs to that classification (denoted as J). A first sample 210 is determined based on the input data and the classification of the input data. The first sample 210 includes a first cue word and a first response. A second sample 220 is determined based on the input data, the classification of the input data, and the reason. The second sample 220 includes a second cue word and a second response. Machine learning model 120 is updated based on the first and second samples.

[0039] Using these implementations of the present disclosure, samples used to train a machine learning model can be divided into a first sample for training the model's classification ability and a second sample for training its reasoning ability, with the first and second samples used to train the machine learning model respectively. In this way, the machine learning model can initially focus solely on classification, and then subsequently generate CoT inference. This division reduces the cognitive burden on the model, enabling it to perform each task with higher accuracy and efficiency.

[0040] Reference Figure 3 Example describing sample 112, Figure 3 Example Figure 300 shows a sample 112 according to an implementation of this disclosure. (As shown...) Figure 3As shown, sample 112 may include a cue word and a response to the cue word. The cue word may include input data 310, a potential category of input data 310 (e.g., negative, positive, and neutral opinions), and a sentence instructing machine learning model 120 to provide a reason why input data 310 belongs to that category. The response may include the category 320 to which input data 310 belongs and a reason 330 for input data 310 belonging to category 320.

[0041] In this disclosed implementation, the machine learning model can perform the task of outputting a classification of target data and a response explaining why the target data belongs to that classification. The task can be divided into a first task and a second task implemented after the first task. The first task outputs the classification of the target data, and the second task outputs a response explaining why the target data belongs to that classification. (See reference...) Figure 4 Examples describing task division, Figure 4 Figure 400 shows an example of a task being divided according to an implementation of this disclosure.

[0042] like Figure 4 As shown, task 430 can be configured to output a classification of the target data and a response indicating why the target data belongs to that classification. Task 430 can be divided into a first task 410 that outputs classification and a second task that outputs reasoning. According to the first task 410, a first sample 210 can be obtained based on the input data and its classification. In one example, the first sample 210 may include the input data and its classification. According to the second task, a second sample 220 can be obtained based on the input data, its classification, and the reason. In one example, the second sample 220 may include the input data, its classification, and the reason.

[0043] By utilizing these implementations of this disclosure, the cognitive load on the machine learning model is simplified by dividing the task, with each cue word focusing on a specific aspect. Therefore, the machine learning model can be fine-tuned independently for classification and inference, allowing for targeted optimization to improve the overall performance of the machine learning model. Since the two-cue-word model treats classification and CoT inference as independent tasks, errors in the first task (classification) are less likely to affect the second task (inference), thus minimizing the risk of error propagation and leading to more reliable and accurate predictions in both stages. This is supported by probabilistic analysis and derivation on the proposed two-cue-word model. By independently decomposing the tasks, the composite error rate common in the single-cue-word model is reduced, which is supported by the probabilistic framework derived from both the two-cue-word model and the single-cue-word model provided below.

[0044] In a model with a single cue word, the joint probability of correct classification and interpretation when applying Bayesian rules is: Where X represents the input text (also known as input data), and Y represents the correct label or category. The label or category to be predicted is represented by J, which stands for explanation / description (also known as cause). This indicates the generated explanation / description.

[0045] In the proposed two-prompt-word model, the task is divided into a first task and a second task. The first task is configured to classify the input text, and the probability of correct classification can be formulated as follows: The second task is configured to provide explanations based on classification labels, and the probability of a correct explanation can be formulated as follows: Therefore, based on the probability of the response from each model discussed above, the overall combined error for each model response can be calculated, i.e., the extent to which the model response deviates from the correct label / classification and the original interpretation.

[0046] In a single prompt word model, the combination error can be represented as follows:

[0047] In the two proposed prompt word models, the independent error for the first task can be represented as follows:

[0048] The independent error for the second task can be expressed as follows:

[0049] Given the independence of the two tasks, the inclusion-exclusion principle of probability can be applied. Therefore, the combined error of the two cue word models can be expressed as follows:

[0050] Equation (1) can be derived from Equation (4), but the error for each task (i.e., Equations (2) and (3)) is smaller than that for Equation (1). When comparing the proposed two cue word systems for CoT-based classification tasks with traditional single cue word systems in terms of model accuracy for classifying input text and CoT interpretation, the probability of correctly classifying the input text is significantly reduced. And representing the probability of accurately providing an interpretation of the category label. Greater than represents the joint probability of correct classification and interpretation. And the model comes from the two proposed cue words. and It has a significantly lower error rate when classifying input text and providing more accurate CoT-based inference, thereby improving classification accuracy.

[0051] In this disclosed implementation, the first template corresponding to the first task can be obtained. See reference... Figure 5A To describe an example of the first template, Figure 5A Example Figure 500A shows a first template according to an implementation of this disclosure. (As shown...) Figure 5A As shown, the first template 510 can be represented in natural language format and includes a first position 531 for inserting input data and a second position 532 for inserting classification. The first sample 210 can be obtained by updating the first template with the input data and the classification of the input data. It should be noted that... Figure 5A The first template 510 shown is merely an example, and other first templates suitable for inserting input data and classification may exist. For example, the first template 510 could be: "Clues: Provide a classification of input data {X}. Response: Classify {Y}". Utilizing these implementations of this disclosure, during fine-tuning, by providing initial samples, the machine learning model can better grasp most of the knowledge relevant to the classification task. Therefore, the machine learning model can better solve classification-related problems.

[0052] In this implementation, the first prompt word in the first sample 210 can be obtained by updating the prompt word portion 512 in the first template 510 with input data. By updating the prompt word portion 512, the first prompt word includes the input data inserted at the first position 531. The first response in the first sample 210 can be obtained by updating the response portion 514 in the first template 510 with classification. By updating the response portion 514, the first response includes the classification inserted at the second position 532.

[0053] In the implementation of this disclosure, multiple candidate categories of the input data can be added to the first prompt word based on a length limit for that first prompt word. There may be a length limit for the prompt word input to the machine learning model 120. If it is determined that the length of the first prompt word does not exceed the length limit, candidate categories of the input data (e.g., negative, positive, and neutral) can be added to the first prompt word. Using these implementations of this disclosure, the classification task can be described more accurately, thereby improving the accuracy of the classification output by the machine learning model.

[0054] In this disclosed implementation, the second template corresponding to the second task can be obtained. See reference for further details. Figure 5B Example describing the second template, Figure 5B Example Figure 500B shows a second template according to an implementation of this disclosure. Figure 5BAs shown, the second template 520 can be represented in natural language format and includes a third position 533 for inserting input data, a fourth position 534 for inserting a classification, and a fifth position 535 for inserting a cause. The second sample 520 can be obtained by updating the second template with the input data, the classification of the input data, and the cause. It should be noted that... Figure 5B The second template 520 shown is merely an example, and other second templates suitable for inserting input data, categories, and reasons may exist. For example, the second template could be: "Clues: Explain why input data {X} belongs to category {Y}. Response: Reason: {J}". Utilizing these implementations of this disclosure, during fine-tuning, by providing second samples, the machine learning model can better grasp most of the knowledge relevant to the reasoning task. Therefore, the machine learning model can better solve problems related to providing why input data belongs to a category.

[0055] In this implementation, the second prompt word in the second sample 520 can be obtained by updating the prompt word portion 522 of the second template 520 with input data and a classification. By updating the prompt word portion 522, the second prompt word includes the input data inserted at the third position 533 and the classification inserted at the fourth position 534. The second response in the second sample 520 can be obtained by updating the response portion 524 of the second template 520 with a reason. By updating the response portion 524, the second response includes the reason inserted at the fifth position 535.

[0056] You can refer to this. Figure 6 Describe the construction of the dataset used to train machine learning model 120. Figure 6 An example diagram is shown illustrating the construction of a dataset for training a machine learning model 120 according to an implementation of this disclosure. Figure 6 As shown, the labeled dataset 610 can be used to generate a dataset 620 for classification and a dataset 630 for inference. For the purposes of the machine learning model 120, a ratio can be determined between a first number of first samples (also referred to as dataset 620) and a second number of second samples (also referred to as dataset 630). Based on this ratio, a first number of first samples 210 and a second number of second samples 220 can be obtained. In some examples, the first number of first samples 210 and the second number of second samples 220 can form a merged dataset 640 for training the machine learning model 120.

[0057] In some implementations, the purpose of machine learning model 120 is to expect higher accuracy in the classification performed by machine learning model 120. The first number of the first plurality of first samples 210 can be relatively large, so the ratio can be set to a value greater than 1. In some implementations, the purpose of machine learning model 120 is to expect higher accuracy in the reasons given by machine learning model 120. The second number of the second plurality of second samples 220 can be relatively large, so the ratio can be set to a value less than 1. Using these implementations of the present disclosure, the performance of both tasks can be improved according to preferences.

[0058] In the implementation of this disclosure, a batch of samples can be selected from a first plurality of first samples 210 and a second plurality of second samples 220 based on a predetermined batch number. In the example, the predetermined batch number can be 128. After selecting the batch of samples, the machine learning model can be updated based on the batch of samples. Using these implementations of this disclosure, the accuracy of the machine learning model output can be improved by batch training the machine learning model in each iteration.

[0059] In the implementation of this disclosure, in response to receiving a target cue word including target input data, the machine learning model 120 can provide a target classification of the target input data and a reason why the target input data belongs to the target classification. After the machine learning model 120 is trained, during the inference phase, the machine learning model 120 can output the target classification and reason based on the received target cue word. Using these implementations of this disclosure, the results output by the machine learning model can be more accurate after training on two independent tasks.

[0060] The preceding paragraphs have described the details of data classification. Based on the implementation of this disclosure, a method for data classification is provided. Further details regarding this method will be found in [reference needed]. Figure 7 ,in Figure 7 An example flowchart of a method 700 for data classification according to an implementation of this disclosure is shown. In box 710, samples for training a machine learning model are obtained. The samples include a cue word and a response to the cue word, the cue word including input data, and the response including a classification of the input data and a reason why the input data belongs to that classification. In box 720, a first sample is determined based on the input data and its classification, and the first sample includes a first cue word and a first response. In box 730, a second sample is determined based on the input data, its classification, and the reason why the input data belongs to that classification, the second sample including a second cue word and a second response. In box 740, the machine learning model is updated based on the first and second samples.

[0061] In the implementation of this disclosure, the machine learning model performs a task to output a classification of target data and a response explaining why the target data belongs to that classification. Determining the first sample and the second sample includes: dividing the task into a first task and a second task implemented after the first task, wherein the first task outputs a classification of the target data and the second task outputs a response explaining why the target data belongs to that classification; obtaining the first sample based on the input data and the classification of the input data according to the first task; and obtaining the second sample based on the input data, the classification of the input data, and the reason according to the second task.

[0062] In the implementation of this disclosure, obtaining the first sample includes: obtaining a first template corresponding to the first task, the first template being represented in natural language format and including a first position for inserting input data and a second position for inserting classification; and obtaining the first sample by updating the first template with input data and the classification of the input data.

[0063] In the implementation of this disclosure, obtaining the first sample by updating the first template with input data and the classification of input data includes: obtaining the first prompt word in the first sample by updating the prompt word part in the first template with input data; and obtaining the first response in the first sample by updating the response part in the first template with classification.

[0064] In the implementation disclosed herein, obtaining the first prompt word includes: adding multiple candidate categories of the input data to the first prompt word based on the length limit of the first prompt word.

[0065] In the implementation of this disclosure, obtaining the second sample includes: obtaining a second template corresponding to the second task, the second template being represented in natural language format and including a third position for inserting input data, a fourth position for inserting classification, and a fifth position for inserting reasons; and obtaining the second sample by updating the second template with input data, the classification of the input data, and the reasons for the input data.

[0066] In the implementation of this disclosure, obtaining a second sample by updating the second template with input data, the classification of the input data, and the reason includes: obtaining a second prompt word in the second sample by updating the prompt word part of the second template with input data and classification; and obtaining a second response in the second sample by updating the response part of the second template with the reason.

[0067] In an implementation of this disclosure, method 700 further includes: determining, based on the purpose of the machine learning model, a ratio between the first number of the first plurality of first samples and the second number of the second plurality of second samples; and obtaining the first plurality of first samples and the second plurality of second samples based on the ratio.

[0068] In the implementation of this disclosure, updating the machine learning model based on the first sample and the second sample includes: selecting a batch of samples from a first plurality of first samples and a second plurality of second samples based on a predetermined batch number; and updating the machine learning model based on the batch of samples.

[0069] In an implementation of this disclosure, method 700 further includes: in response to receiving a target cue word including target input data, a machine learning model provides a target classification of the target input data and a reason why the target input data belongs to the target classification.

[0070] According to an implementation of this disclosure, an apparatus for data classification is provided. The apparatus includes: a sample acquisition module configured to acquire samples for training a machine learning model, the samples including prompt words and responses to the prompt words, the prompt words including input data, and the responses including a classification of the input data and a reason why the input data belongs to that classification; a first sample determination module configured to determine a first sample based on the input data and the classification of the input data, the first sample including a first prompt word and a first response; a second sample determination module configured to determine a second sample based on the input data, the classification of the input data, and the reason why the input data belongs to that classification, the second sample including a second prompt word and a second response; and a model update module configured to update the machine learning model based on the first and second samples.

[0071] According to an implementation of this disclosure, an electronic device is provided for implementing method 700. The electronic device includes a computer processor coupled to a computer-readable storage unit, the storage unit including instructions that, when executed by the computer processor, implement a method for data classification. The method includes: acquiring samples for training a machine learning model, the samples including prompt words and responses to prompt words, the prompt words including input data, and the responses including a classification of the input data and a reason why the input data belongs to that classification; determining a first sample based on the input data and the classification of the input data, the first sample including a first prompt and a first response; determining a second sample based on the input data, the classification of the input data, and the reason why the input data belongs to that classification, the second sample including a second prompt word and a second response; and updating the machine learning model based on the first sample and the second sample.

[0072] In the implementation of this disclosure, the machine learning model performs a task to output a classification of target data and a response explaining why the target data belongs to that classification. Determining the first sample and the second sample includes: dividing the task into a first task and a second task implemented after the first task, wherein the first task outputs a classification of the target data and the second task outputs a response explaining why the target data belongs to that classification; obtaining the first sample based on the input data and the classification of the input data according to the first task; and obtaining the second sample based on the input data, the classification of the input data, and the reason according to the second task.

[0073] In the implementation of this disclosure, obtaining the first sample includes: obtaining a first template corresponding to the first task, the first template being represented in natural language format and including a first position for inserting input data and a second position for inserting classification; and obtaining the first sample by updating the first template with input data and the classification of the input data.

[0074] In the implementation of this disclosure, obtaining the first sample by updating the first template with input data and the classification of input data includes: obtaining the first prompt word in the first sample by updating the prompt word part in the first template with input data; and obtaining the first response in the first sample by updating the response part in the first template with classification.

[0075] In the implementation disclosed herein, obtaining the first prompt word includes: adding multiple candidate categories of the input data to the first prompt word based on the length limit of the first prompt word.

[0076] In the implementation of this disclosure, obtaining the second sample includes: obtaining a second template corresponding to the second task, the second template being represented in natural language format and including a third position for inserting input data, a fourth position for inserting classification, and a fifth position for inserting reasons; and obtaining the second sample by updating the second template with input data, the classification of the input data, and the reasons for the input data.

[0077] In the implementation of this disclosure, obtaining a second sample by updating the second template with input data, the classification of the input data, and the reason includes: obtaining a second prompt word in the second sample by updating the prompt word part of the second template with input data and classification; and obtaining a second response in the second sample by updating the response part of the second template with the reason.

[0078] In an implementation of this disclosure, method 700 further includes: determining, based on the purpose of the machine learning model, a ratio between the first number of the first plurality of first samples and the second number of the second plurality of second samples; and obtaining the first plurality of first samples and the second plurality of second samples based on the ratio.

[0079] In the implementation of this disclosure, updating the machine learning model based on the first sample and the second sample includes: selecting a batch of samples from a first plurality of first samples and a second plurality of second samples based on a predetermined batch number; and updating the machine learning model based on the batch of samples.

[0080] In an implementation of this disclosure, method 700 further includes: in response to receiving a target cue word including target input data, a machine learning model provides a target classification of the target input data and a reason why the target input data belongs to the target classification.

[0081] According to an implementation of this disclosure, a computer program product is provided, the computer program product including a computer-readable storage medium having program instructions embodied therein, the program instructions being executed by an electronic device to cause the electronic device to perform method 700.

[0082] Figure 8 A block diagram of a computing device 800 in which various implementations of the present disclosure may be implemented is shown. It should be understood that... Figure 8 The computing device 800 shown is for illustrative purposes only and does not imply any limitation on the functionality and scope of this disclosure. The computing device 800 can be used to implement the method 700 described in the implementation of this disclosure. Figure 8 As shown, the computing device 800 can be a general-purpose computing device. The computing device 800 may include at least one or more processors or processing units 810, memory 820, storage unit 830, one or more communication units 840, one or more input devices 850, and one or more output devices 860.

[0083] Processing unit 810 can be a physical or virtual processor and can implement various processes based on program 825 stored in memory 820. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of computing device 800. Processing unit 810 may also be referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.

[0084] Computing device 800 typically includes various computer storage media. Such media can be any media accessible to computing device 800, including but not limited to volatile and non-volatile media, or removable and non-removable media. Memory 820 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory) or any combination thereof. Storage cell 830 can be any removable or non-removable media and may include machine-readable media such as memory, flash drives, disks, or other media that can be used to store information and / or data and can be accessed within computing device 800.

[0085] The computing device 800 may also include additional removable / non-removable volatile / non-volatile memory media. Although in Figure 8Not shown, but may provide disk drives for reading from and / or writing to removable non-volatile disks, and optical disk drives for reading from and / or writing to removable non-volatile optical disks. In this case, each drive may be connected to a bus (not shown) via one or more data media interfaces.

[0086] The communication unit 840 communicates with another computing device via a communication medium. Furthermore, the functionality of the components in the computing device 800 can be implemented by a single computing cluster or multiple computing machines that can communicate via communication connections. Therefore, the computing device 800 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.

[0087] Input device 850 can be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 860 can be one or more of various output devices, such as a monitor, speaker, printer, etc. With the aid of communication unit 840, computing device 800 can also communicate with one or more external devices (not shown), such as storage devices and display devices, wherein one or more devices enable a user to interact with computing device 800 or any device (such as a network card, modem, etc.), enabling computing device 800 to communicate with one or more other computing devices (if needed). Such communication can be performed via input / output (I / O) interface (not shown).

[0088] In some implementations, some or all components of computing device 800 may be deployed within a cloud computing architecture, rather than integrated into a single device. In a cloud computing architecture, components can be remotely provided and work together to achieve the functionality described in this disclosure. In some implementations, cloud computing provides computing, software, data access, and storage services without requiring end users to know the physical location or configuration of the systems or hardware providing these services. In various implementations, cloud computing provides services via a wide area network (WAN), such as the Internet, using appropriate protocols. For example, a cloud computing provider offers applications via a WAN that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture, along with the corresponding data, may be stored on servers at remote locations. Computing resources in a cloud computing environment may be consolidated or distributed across locations in remote data centers. Cloud computing infrastructure can provide services through shared data centers, although they act as a single access point for users. Therefore, cloud computing architectures can be used to provide the components and functionality described herein from service providers at remote locations. Alternatively, they may be provided from conventional servers or directly installed or otherwise installed on client devices.

[0089] The functions described herein can be performed at least in part by one or more hardware logic components. For example, but not limited to, illustrative types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), etc.

[0090] Program code used to perform the methods described herein can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program code enables the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely or partially on a machine, partially as a standalone software package on a machine, partially on a remote machine, or entirely on a remote machine or server.

[0091] In the context of this disclosure, a machine-readable medium can be any tangible medium that may contain or store a program used by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. More specific examples of machine-readable storage media will include electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable optical disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0092] Furthermore, although operations are shown in a specific order, this should not be construed as requiring that such operations be performed in the specific order shown or sequentially, or that all shown operations be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the topics described herein, but rather as descriptions of features that may be specific to a particular implementation. Certain features described in the context of a single implementation may also be implemented in combination within a single implementation. Conversely, various features described in a single implementation may also be implemented individually or in any suitable sub-combination in multiple implementations.

[0093] Although the subject matter has been described in language specific to structural features and / or methodological actions, it should be understood that the subject matter specified in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are disclosed as exemplary forms for implementing the claims.

[0094] Based on the foregoing, it should be understood that this document has described specific implementations of the currently disclosed technology for illustrative purposes, but various modifications can be made without departing from the scope of this disclosure. Therefore, the technology disclosed herein is not limited except for the appended claims.

[0095] The subject matter and functional operations described in this disclosure can be implemented in various systems, digital electronic circuits, or computer software, firmware, or hardware (including the structures disclosed in this specification and their equivalents), or combinations thereof. Implementations of the subject matter described in this specification can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a tangible and non-transitory computer-readable medium for execution by or control of the operation of a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a combination of substances influencing machine-readable propagation signals, or combinations thereof. The terms "data processing unit" or "data processing apparatus" encompass all means, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for the computer program in question, for example, code constituting processor firmware, a protocol stack, a database management system, an operating system, or combinations thereof.

[0096] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language, including compiled or interpreted languages, and can be deployed in any form, including as standalone programs or as modules, components, subroutines, or other units suited to a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), as a single file dedicated to the program in question, or as multiple coordinating files (e.g., files storing one or more modules, subroutines, or code sections). Computer programs can be deployed to execute on one or more computers located at a single site or distributed across multiple sites and interconnected via a communication network.

[0097] Processors suitable for executing computer programs include, for example, general-purpose and special-purpose microprocessors, and any one or more processors of any type of digital computer. Typically, a processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include, or be operatively coupled to, receiving data from or transferring data to, or both to, one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data. However, a computer does not need to have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices. The processor and memory may be supplemented by or incorporated into special-purpose logic circuitry.

[0098] The specification together with the accompanying drawings is intended to be illustrative only, where illustrative means example. As used herein, the use of "or" is intended to include "and / or" unless the context clearly indicates otherwise.

[0099] While this disclosure contains numerous details, these details should not be construed as limiting the scope of any disclosure or claimable content, but rather as descriptions of features specific to particular implementations of a particular disclosure. Certain features described in this disclosure in the context of a single implementation may also be implemented in combination within a single implementation. Conversely, various features described in the context of a single implementation may also be implemented individually or in any suitable sub-combination in multiple implementations. Furthermore, although features may be described above as functioning in certain combinations and even initially claimed in this way, in some cases, one or more features from the claimed combination may be removed from the combination, and the claimed combination may involve sub-combinations or variations thereof.

[0100] Similarly, although operations are shown in a specific order in the accompanying drawings, this should not be construed as requiring the operations to be performed in the specific order shown or sequentially, or to perform all shown operations to achieve the desired result. Furthermore, the separation of various system components in the implementations described in this disclosure should not be construed as requiring such separation in all implementations. Only some implementations and examples have been described, and other implementations, enhancements, and variations can be made based on what is described and shown in this disclosure.

Claims

1. A method for data classification, comprising: Obtain samples for training a machine learning model, the samples including prompt words and responses to the prompt words, the prompt words including input data, and the responses including a classification of the input data and a reason why the input data belongs to the classification; A first sample is determined based on the input data and the classification of the input data, the first sample including a first prompt word and a first response; A second sample is determined based on the input data, the classification of the input data, and the reason, and the second sample includes a second prompt word and a second response; as well as The machine learning model is updated based on the first sample and the second sample.

2. The method of claim 1, wherein the machine learning model performs the task of outputting a classification of target data and a response on why the target data belongs to the classification, and determining the first sample and the second sample comprises: The task is divided into a first task and a second task implemented after the first task. The first task outputs the classification of the target data, and the second task outputs a response on why the target data belongs to the classification. Based on the first task, the first sample is obtained based on the input data and the classification of the input data; as well as According to the second task, the second sample is obtained based on the input data, the classification of the input data, and the reason.

3. The method according to claim 2, wherein obtaining the first sample comprises: Obtain a first template corresponding to the first task. The first template is represented in natural language format and includes a first position for inserting the input data and a second position for inserting the classification. as well as A first sample is obtained by updating the first template with the input data and the classification of the input data.

4. The method of claim 3, wherein obtaining the first sample by updating the first template with the input data and the classification of the input data comprises: The first prompt word in the first sample is obtained by updating the prompt word portion in the first template with the input data; as well as The first response in the first sample is obtained by updating the response portion of the first template with the classification.

5. The method according to claim 4, wherein obtaining the first prompt word includes: Based on the length limit of the first prompt word, multiple candidate categories of the input data are added to the first prompt word.

6. The method of claim 2, wherein obtaining the second sample comprises: Obtain a second template corresponding to the second task. The second template is represented in natural language format and includes a third position for inserting the input data, a fourth position for inserting the category, and a fifth position for inserting the reason. as well as The second sample is obtained by updating the second template with the input data, the classification of the input data, and the reason.

7. The method of claim 6, wherein obtaining the second sample by updating the second template with the input data, the classification of the input data, and the reason comprises: The second prompt word in the second sample is obtained by updating the prompt word portion in the second template with the input data and the classification; as well as The second response in the second sample is obtained by updating the response portion of the second template with the stated reason.

8. The method according to claim 1, further comprising: Based on the purpose of the machine learning model, determine the ratio between the first number of the first plurality of first samples and the second number of the second plurality of second samples; as well as Based on the ratio, the first plurality of first samples and the second plurality of second samples are obtained.

9. The method of claim 8, wherein updating the machine learning model based on the first sample and the second sample comprises: Based on a predetermined batch number, a batch of samples is selected from the first plurality of first samples and the second plurality of second samples; as well as The machine learning model is updated based on the batch of samples.

10. The method according to claim 1, further comprising: In response to receiving a target prompt word including target input data, the machine learning model provides a target classification of the target input data and a reason why the target input data belongs to the target classification.

11. An electronic device comprising a computer processor coupled to a computer-readable storage unit, the storage unit including instructions that, when executed by the computer processor, implement a method for data classification, the method comprising: Obtain samples for training a machine learning model, the samples including prompt words and responses to the prompt words, the prompt words including input data, and the responses including a classification of the input data and a reason why the input data belongs to the classification; A first sample is determined based on the input data and the classification of the input data, the first sample including a first prompt word and a first response; A second sample is determined based on the input data, the classification of the input data, and the reason, and the second sample includes a second prompt word and a second response; as well as The machine learning model is updated based on the first sample and the second sample.

12. The electronic device of claim 11, wherein the machine learning model performs the task of outputting a classification of target data and a response on why the target data belongs to the classification, and determining the first sample and the second sample comprises: The task is divided into a first task and a second task implemented after the first task. The first task outputs the classification of the target data, and the second task outputs a response on why the target data belongs to the classification. Based on the first task, the first sample is obtained based on the input data and the classification of the input data; as well as According to the second task, the second sample is obtained based on the input data, the classification of the input data, and the reason.

13. The electronic device of claim 12, wherein obtaining the first sample comprises: Obtain a first template corresponding to the first task. The first template is represented in natural language format and includes a first position for inserting the input data and a second position for inserting the classification. as well as A first sample is obtained by updating the first template with the input data and the classification of the input data.

14. The electronic device of claim 13, wherein obtaining the first sample by updating the first template with the input data and the classification of the input data comprises: The first prompt word in the first sample is obtained by updating the prompt word portion in the first template with the input data; as well as The first response in the first sample is obtained by updating the response portion of the first template with the classification.

15. The electronic device of claim 12, wherein obtaining the second sample comprises: Obtain a second template corresponding to the second task. The second template is represented in natural language format and includes a third position for inserting the input data, a fourth position for inserting the category, and a fifth position for inserting the reason. as well as The second sample is obtained by updating the second template with the input data, the classification of the input data, and the reason.

16. The electronic device of claim 15, wherein obtaining the second sample by updating the second template with the input data, the classification of the input data, and the reason comprises: The second prompt word in the second sample is obtained by updating the prompt word portion in the second template with the input data and the classification; as well as The second response in the second sample is obtained by updating the response portion of the second template with the stated reason.

17. The electronic device of claim 11, wherein the method further comprises: Based on the purpose of the machine learning model, determine the ratio between the first number of the first plurality of first samples and the second number of the second plurality of second samples; as well as Based on the ratio, the first plurality of first samples and the second plurality of second samples are obtained.

18. The electronic device of claim 17, wherein updating the machine learning model based on the first sample and the second sample comprises: Based on a predetermined batch number, a batch of samples is selected from the first plurality of first samples and the second plurality of second samples; as well as The machine learning model is updated based on the batch of samples.

19. The electronic device of claim 18, wherein the method further comprises: In response to receiving a target prompt word including target input data, the machine learning model provides a target classification of the target input data and a reason why the target input data belongs to the target classification.

20. A computer program product comprising a computer-readable storage medium having program instructions embodied therein, the program instructions being executed by an electronic device to cause the electronic device to perform a method for data classification, the method comprising: Obtain samples for training a machine learning model, the samples including prompt words and responses to the prompt words, the prompt words including input data, and the responses including a classification of the input data and a reason why the input data belongs to the classification; A first sample is determined based on the input data and the classification of the input data, the first sample including a first prompt word and a first response; A second sample is determined based on the input data, the classification of the input data, and the reason, and the second sample includes a second prompt word and a second response; as well as The machine learning model is updated based on the first sample and the second sample.