Method and device for interpretable multi-label classification based on label grouping and logic chain
Through the method of tag grouping and logical chaining, the accuracy and efficiency of the generative large language model in multiple dialogue scenarios is solved, efficient multi-label classification is achieved, manual annotation costs are reduced, and the interpretability and controllability of the model are improved.
Patent Information
- Application Number
- CN202510297546.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-07-18
AI Technical Summary
The existing generative large language model is difficult to accurately classify multi-labels in multi-round dialogue scenarios, and it is expensive to rely on manual labeling. Existing methods such as CARP cannot effectively solve the multi-label classification problem of multi-round dialogue.
Using a method based on tag grouping and logical chain, by collecting dialogue data sets and grouping tags, building a first language model for few sample prompts and thinking chain guidance, generating a thinking process, building an enhanced dialogue data set to fine-tune the second language model, and realizing multi-label classification.
It reduces the cost of manual labeling, improves the accuracy and efficiency of multi-label classification, saves resources, and enhances the interpretability and controllability of the model.
Smart Images

Figure CN120336531A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of text classification and annotation, and particularly to an interpretable multi-label classification method and device based on label grouping and logical chains. Background Art
[0002] When exploring the applications of Generative Large Language Models (LLMs), one has to face their inherent challenges, especially in terms of the controllability of content generation. These models may absorb and replicate biases in the training data during the training process and may generate unsafe or non-compliant content when generating content, which not only involves legal risks but may also have negative impacts on society. Therefore, developing effective mechanisms for detecting inappropriate content to prevent the spread of illegal or harmful information has become an urgent task in the research field.
[0003] In multi-turn dialogue scenarios, LLMs need to understand the context of the dialogue and generate appropriate responses based on it. Although specific prompts can be used to guide the model to generate expected content, due to the inherent uncertainty of the generation process, the output of the model often fails to fully meet expectations. This has led to the need to invest a large amount of resources in manually annotating the types of problems existing in the dialogue content in practical applications to identify and correct the deficiencies in the model's performance in the dialogue.
[0004] To improve the performance of the model, it is usually necessary to fine-tune the LLMs using the annotated dialogue data. This process is not only time-consuming but also costly. In addition, to evaluate the effect of the fine-tuned model, it is necessary to manually annotate the dialogue content generated by the model again. Through this iterative process of annotation and fine-tuning, the performance of the model is gradually optimized to meet higher standards. However, manual annotation is not only costly but also inefficient. Therefore, it is necessary to seek automated solutions to reduce the dependence on manual annotation while improving the controllability and security of the model.
[0005] Existing work on improving the ability of models to handle complex problems mainly focuses on enhancing LLMs' performance in dealing with mathematical problems rather than complex language phenomena in text classification. Researchers proposed a method called CARP to address complex language phenomena in text classification through a step-by-step reasoning strategy and used KNN demonstration search to alleviate the problem of limited token numbers. CARP first prompts LLMs to find surface clues (such as keywords, tone, semantic relations, etc.), then conducts diagnostic reasoning based on these clues, and finally makes a decision. The design of CAPP aims to solve various complex language phenomena involved in text classification, such as emphasis, contrast, irony, etc. Through the step-by-step reasoning strategy, these complex language phenomena are processed. However, this method of guiding the model's reasoning through token words is not applicable to the analysis of multi-turn dialogue problems because when the model reasons about each dialogue sample, it needs to first reason out the content corresponding to each token word, and each label often involves multiple token words, and the selection of token words also affects the classification effect. Therefore, the CAPP method cannot well solve the multi-label classification problem of multi-turn dialogues. Summary of the Invention
[0006] The purpose of this application is to propose an interpretable multi-label classification method and device based on label grouping and logical chains for the above-mentioned technical problems.
[0007] In the first aspect, the present invention provides an interpretable multi-label classification method based on label grouping and logical chains, including the following steps:
[0008] Collect a dialogue dataset, where the samples in the dialogue dataset include multi-turn dialogues and the label array corresponding to the current response in the multi-turn dialogues, and the multi-turn dialogues include the dialogue history and its corresponding current response; group the labels in the label arrays corresponding to all multi-turn dialogues to obtain the labels corresponding to each label group.
[0009] Construct a first language model, and use few-shot prompting and chain-of-thought to guide the first language model to generate a thinking process for the label classification of the current response in the multi-turn dialogue based on the multi-turn dialogue and each label group and its corresponding labels; set a dynamic label array in the thinking process, the dynamic label array is initially empty, and first take each label group as the start node and perform the classification judgment of each label group, then further loop through each label in the corresponding label group according to the classification judgment results of each label group and perform the classification judgment of each label, and add the corresponding label to the dynamic label array according to the classification judgment results of each label, so that the labels in the dynamic label array obtained after the traversal are the same as the labels in the label array corresponding to the multi-turn dialogue; construct an enhanced dialogue dataset based on the samples in the dialogue dataset and their corresponding thinking processes.
[0010] Build a second language model, and fine-tune the second language model using an enhanced dialogue dataset to obtain a fine-tuned second language model;
[0011] Obtain a multi-turn dialogue to be classified, and input the multi-turn dialogue to be classified and a prompt for multi-label classification of the current response in the multi-turn dialogue to be classified according to the multi-turn dialogue to be classified into the fine-tuned second language model to obtain the corresponding thinking process and label array.
[0012] Preferably, during the fine-tuning process of the second language model, the input of the second language model is a multi-turn dialogue, a thinking process, and a prompt for multi-label classification of the current response in the multi-turn dialogue according to the multi-turn dialogue, and the output is a label array, and the label array is composed of at least one label.
[0013] Preferably, after all the labels in all label groups in the thinking process are traversed, enter the end node. In the end node, determine whether the dynamic label array obtained by traversing all the labels in all label groups is empty. If so, add "no label" to the dynamic label array obtained by traversing all the labels in all label groups to obtain the label array; otherwise, directly use the dynamic label array obtained by traversing all the labels in all label groups as the label array.
[0014] Preferably, group the labels in the label arrays corresponding to all multi-turn dialogues to obtain the labels corresponding to each label group, specifically including:
[0015] Set the admission conditions corresponding to each label group;
[0016] Traverse each multi-turn dialogue and its corresponding labels, and use the keyword-triggered method to classify the labels corresponding to the multi-turn dialogues whose content meets the admission conditions corresponding to the label group into the corresponding label group to obtain the labels corresponding to each label group.
[0017] Preferably, the first language model uses the Qwen2-72B-Instruct model, and the second language model uses the Qwen2-7B-Instruct model.
[0018] Preferably, the fine-tuning method of the second language model uses full-parameter fine-tuning.
[0019] In a second aspect, the present invention provides an interpretable multi-label classification device based on label grouping and logical chain, including:
[0020] The label grouping module is configured to collect a conversation dataset, where the samples in the conversation dataset include multi-turn conversations and the label arrays corresponding to the current responses in the multi-turn conversations, and the multi-turn conversations include the conversation history and its corresponding current response; group the labels in the label arrays corresponding to all the multi-turn conversations to obtain the labels corresponding to each label group.
[0021] The thinking process generation module is configured to construct a first language model and use few-shot prompting and chain of thought to guide the first language model to generate the thinking process of the label classification of the current response in the multi-turn conversation based on the multi-turn conversation and each label group and its corresponding labels; set a dynamic label array during the thinking process, where the dynamic label array is initially empty, and first use each label group as the starting node and perform the classification judgment of each label group, then further loop through each label in the corresponding label group according to the classification judgment result of each label group and perform the classification judgment of each label, and add the corresponding labels to the dynamic label array according to the classification judgment result of each label, so that the labels in the dynamic label array obtained after the traversal are the same as the labels in the label array corresponding to the multi-turn conversation; construct an enhanced conversation dataset based on the samples in the conversation dataset and their corresponding thinking processes.
[0022] The fine-tuning module is configured to construct a second language model and fine-tune the second language model using the enhanced conversation dataset to obtain the fine-tuned second language model.
[0023] The label generation module is configured to obtain the multi-turn conversation to be classified, and input the multi-turn conversation to be classified and the prompt words for multi-label classification of the current response in the multi-turn conversation to be classified according to the multi-turn conversation to be classified into the fine-tuned second language model to obtain the corresponding thinking process and label array.
[0024] In a third aspect, the present invention provides an electronic device, including one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner of the first aspect.
[0025] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in any implementation manner of the first aspect is implemented.
[0026] In a fifth aspect, the present invention provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method described in any implementation manner of the first aspect is implemented.
[0027] Compared with the prior art, the present invention has the following beneficial effects:
[0028] (1) The interpretable multi-label classification method based on label grouping and logical chain proposed by the present invention combines the interpretability of text classification with the chain of thought. Compared with previous research on the interpretability of text classification, this method does not require additional models and similarity judgment methods to extract prototypes. Instead, it can achieve multi-label classification of multi-turn conversations by using two pre-trained large language models. This method can accurately label the current response in multi-turn conversations, effectively reducing the manual labeling cost and improving the labeling efficiency of the generated current response.
[0029] (2) The interpretable multi-label classification method based on label grouping and logical chain proposed by the present invention avoids unnecessary inference processes of the model through the method of label grouping, reduces the output of tokens, and saves resources.
[0030] (3) The interpretable multi-label classification method based on label grouping and logical chain proposed by the present invention uses the method of few-shot prompting and chain of thought to guide the first language model to generate the thinking process of label classification for the current response in multi-turn conversations according to multi-turn conversations and each label group and its corresponding labels. It can achieve data augmentation for the dialogue dataset, and the constructed augmented dialogue dataset can complete the full-parameter fine-tuning of the second language model, effectively improving the accuracy of the second language model in multi-label classification of the current response. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for description in the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0032] Figure 1 It is a flowchart of the interpretable multi-label classification method based on label grouping and logical chain according to the embodiment of the present application;
[0033] Figure 2 It is a schematic diagram of the thinking process of the interpretable multi-label classification method based on label grouping and logical chain according to the embodiment of the present application;
[0034] Figure 3 It is a schematic diagram of the result of the interpretable multi-label classification method based on label grouping and logical chain according to the embodiment of the present application;
[0035] Figure 4 It is a schematic diagram of the multi-label classification result without label grouping in the multi-label classification method;
[0036] Figure 5Schematic diagram for comparing the training results with and without chain of thought in multi-label classification methods;
[0037] Figure 6 Schematic diagram of the interpretable multi-label classification device based on label grouping and logical chain according to the embodiment of the present application;
[0038] Figure 7 Schematic diagram of the hardware structure of the electronic device provided by the embodiment of the present invention. Detailed implementation manners
[0039] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0040] Figure 1 An interpretable multi-label classification method based on label grouping and logical chain provided by the embodiment of the present application is shown, including the following steps:
[0041] S1. Collect a dialogue data set. The samples in the dialogue data set include multi-turn dialogues and the label arrays corresponding to the current responses in the multi-turn dialogues, where the multi-turn dialogues include the dialogue history and its corresponding current response; group the labels in the label arrays corresponding to all the multi-turn dialogues to obtain the labels corresponding to each label group.
[0042] In a specific embodiment, grouping the labels in the label arrays corresponding to all the multi-turn dialogues to obtain the labels corresponding to each label group specifically includes:
[0043] Set the admission conditions corresponding to each label group;
[0044] Traverse each multi-turn dialogue and its corresponding label, and use the keyword trigger method to classify the labels corresponding to the multi-turn dialogues that meet the admission conditions corresponding to the label group in the content into the corresponding label group to obtain the labels corresponding to each label group.
[0045] Specifically, in the embodiments of the present application, the collected multi-turn conversations are first disassembled and divided into two parts: conversation history and current response. In order to add the interpretation of the labels to the fine-tuning dataset of the second language model and to simulate the thinking mode of humans for text classification, the labels are first grouped. The specific method is as follows: corresponding admission conditions are set for each label group, and the multi-turn conversations corresponding to each label are screened by keyword triggering. The labels of the multi-turn conversations that meet the same admission conditions are grouped under the same label group, so that the label group corresponding to each label is determined. Thus, the original multiple labels are divided into several label groups, avoiding unnecessary judgment branches and thinking logics. For example, only when the content of requesting to obtain contact information appears in the current response of the multi-turn conversation, all the labels under the contact information category will be judged. The grouping method can speed up the inference of the model and also make the entire multi-label classification process more structured and easy for the large model to understand.
[0046] In one example, the consultation category and the contact information category are used as two label groups. Among them, the consultation category includes two labels: logical error consultation and repeated consultation, and the contact information category includes two labels: forced call after refusal to connect and pure call for contact information. The admission condition corresponding to the consultation category is that the current response in the multi-turn conversation must contain consultation information expressed in an interrogative sentence; the admission condition corresponding to the contact information category is that the current response in the multi-turn conversation must contain the content of requesting to obtain the contact information of the visitor.
[0047] S2, construct a first language model, and use the few-shot prompting and chain-of-thought methods to guide the first language model to generate the thinking process of the label classification of the current response in the multi-turn conversation according to the multi-turn conversation and each label group and its corresponding labels; set a dynamic label array in the thinking process, the dynamic label array is initially empty, and first use each label group as the start node and execute the classification judgment of each label group, and then further loop through each label in the corresponding label group according to the classification judgment result of each label group and execute the classification judgment of each label, and add the corresponding label to the dynamic label array according to the classification judgment result of each label, so that the labels in the dynamic label array obtained after the traversal are the same as the labels in the label array corresponding to the multi-turn conversation; construct an enhanced conversation dataset based on the samples in the conversation dataset and their corresponding thinking processes.
[0048] In a specific embodiment, the first language model uses the Qwen2-72B-Instruct model, and the second language model uses the Qwen2-7B-Instruct model.
[0049] In a specific embodiment, after all the tags in all the tag groups during the thinking process are traversed, it enters the end node. In the end node, it is judged whether the dynamic tag array obtained by traversing all the tags in all the tag groups is empty. If so, "no tag" is added to the dynamic tag array obtained by traversing all the tags in all the tag groups to obtain the tag array. Otherwise, the dynamic tag array obtained by traversing all the tags in all the tag groups is directly used as the tag array.
[0050] Specifically, the embodiment of the present application uses the Qwen2-72B-Instruct model as the first language model. The Qwen2-72B-Instruct model, as a large language model pre-trained with a large amount of data, has strong learning ability and reasoning ability, and can provide better performance and more accurate predictions in more complex tasks. This first language model is used to generate the corresponding thinking process according to the multi-round dialogue and each tag group and its corresponding tags in the way of few-shot prompting and chain of thought (CoT). The enhanced dialogue dataset composed of this thinking process and the multi-round dialogue is used as the fine-tuning dataset of the second language model.
[0051] Construct the prompt used in the qwen2-72B-instruct model using each label in the label array of the dialogue dataset. The writing specification of this prompt combines the few-shot prompting and chain-of-thought methods to generate a thinking process that is more in line with human thinking and can combine the label groups generated after label grouping and their corresponding labels. Input each sample and the prompt in the dialogue dataset into the first language model to guide the first language model to generate a thinking process for the label classification of the current response in the multi-turn dialogue based on the multi-turn dialogue and each label group and its corresponding label. The template of the prompt can be set to be diverse, but it needs to satisfy setting a dynamic label array in this thinking process. The dynamic label array is initially empty. First, take one of the label groups as the starting node and perform the classification judgment of this label group, and then further perform the classification judgment on each label in this label group. Take the remaining label groups as the starting nodes respectively and repeat the above classification judgment process of the label group and the label until all labels in all label groups are traversed. Finally, enter the end node to judge whether the dynamic label array obtained after all labels in all label groups are traversed is empty. If so, add "no label" to the dynamic label array obtained after all labels in all label groups are traversed and output it as the label array. Otherwise, directly output the dynamic label array obtained after all labels in all label groups are traversed as the label array. And the finally generated label array is the label array corresponding to the multi-turn dialogue in each sample of the dialogue dataset. The classification judgment flowchart of the thinking process generated in one example is as Figure 2 shown. Construct an enhanced dialogue dataset from the multi-turn dialogue, label array, and their corresponding thinking processes of each sample in the dialogue dataset. In one embodiment, expand the dialogue dataset containing 300 multi-turn dialogues to an enhanced dialogue dataset containing 1000 multi-turn dialogues and their corresponding thinking processes. Among them, the test set is 100 high-quality data.
[0052] S3. Construct a second language model and fine-tune the second language model using the enhanced dialogue dataset to obtain the fine-tuned second language model.
[0053] In a specific embodiment, during the fine-tuning process of the second language model, the input of the second language model is the multi-turn dialogue, the thinking process, and the prompt for multi-label classification of the current response in the multi-turn dialogue based on the multi-turn dialogue, and the output is the label array, which is composed of at least one label.
[0054] In a specific embodiment, the fine-tuning method of the second language model adopts full-parameter fine-tuning.
[0055] Specifically, in the embodiments of the present application, the Qwen2-7B-Instruct model is used as the second language model. As a pre-trained large language model, the Qwen2-7B-Instruct model is relatively small in size and has low computing power requirements for devices, which facilitates installation and deployment. The enhanced dialogue dataset is used to fine-tune the second language model so that it can learn the thinking process in the enhanced dialogue dataset. During the fine-tuning process of the second language model, the input includes the dialogue history, the current response, the thinking process, and the prompt words for multi-label classification of the current response based on the dialogue history and the current response.
[0056] In one example, an example of the input during the fine-tuning process is provided, where the dialogue history and the current response are omitted, and only the prompt words and the thinking process are given, as follows:
[0057] Prompt words: You are good at multi-label classification of the current response according to multi-turn conversations. Your entire thinking process should follow the following process. The dynamic label array is initially empty, and finally, please output the label array.
[0058] Thinking process:
[0059] 1. Start node -> Inquiry type
[0060] Check whether the current response contains inquiry information expressed in an interrogative sentence. If not, directly go to 2. If so, go to 1.1.
[0061] 1.1. Inquiry type -> Logical error inquiry
[0062] All of the following 6 cases belong to the "logical error inquiry" label (the description of the 6 cases is omitted here). If one of the above 6 cases is met, add the "logical error inquiry" label to the dynamic label array. Otherwise, go to 1.2.
[0063] 1.2. Inquiry type -> Repeated inquiry
[0064] If the inquiry content in the current response can be answered from the dialogue history, then the inquiry content is redundant, and the "repeated inquiry" label is added to the dynamic label array. Otherwise, go to 2.
[0065] 2. Start node -> Contact information type
[0066] The current response must contain content requesting to obtain the contact information of the visitor. If not, go to 3. If so, go to 2.1.
[0067] 2.1. Contact information type -> Forcing to obtain contact information after refusal to connect
[0068] In the conversation history, the visitor clearly states that they do not want to add the contact information of the customer service or implicitly indicates that they do not want to leave their own contact information. However, in the current reply, the customer service still requests to add their own contact information or requests to obtain the contact information of the customer service. If the above situation exists, add the label "Forcing to obtain contact information after refusal to connect" to the dynamic label array; otherwise, proceed to 2.2.
[0069] 2. Contact Information Category -> Pure Contact Information Obtaining
[0070] If any of the following 3 situations (description of the 3 situations omitted here) exist in the current reply, it belongs to "pure contact information obtaining", and add "pure contact information obtaining" to the label array. Otherwise, proceed to 3.
[0071] 3. End Node
[0072] If the dynamic label array is empty at this time, add the label "No label" to the dynamic label array; otherwise, directly output the dynamic label array, and the entire traversal ends.
[0073] S4. Obtain the multi-turn conversation to be classified, and input the multi-turn conversation to be classified and the prompt words for multi-label classification of the current reply in the multi-turn conversation to be classified according to the multi-turn conversation to be classified into the fine-tuned second language model to obtain the corresponding thinking process and label array.
[0074] Specifically, deploy the fine-tuned second language model, input the current reply to be classified and its conversation history into the fine-tuned second language model, and set the prompt words for multi-label classification of the current reply to be classified according to the current reply to be classified and its conversation history to guide the fine-tuned second language model to generate the label array and thinking process corresponding to the current reply to be classified. In the inference stage, output the thinking process, which increases the interpretability of label generation.
[0075] To verify the necessity of label grouping, this application compares the method without label grouping and the method with the label grouping process as follows Figure 3 and 4 As shown, it can be seen that by grouping the labels, the repetitive judgment of the model can be reduced, the output of tokens can be reduced, resources can be saved, and the inference time of the model can be accelerated.
[0076] Furthermore, this application also verifies the necessity of the thought chain in processing complex text classification, and the results are shown in Table 1. The performance of the model on the test set with and without the thought chain training under different data volumes, as Figure 5As shown in the figure. It can be seen that by grouping tags, the output of the model tokens can be reduced, the inference of the model can be accelerated, and resources can be saved. Fine-tuning the full parameters of the training set containing the Chain of Thought (CoT) can improve the performance of the model by at least 15% compared to the model without CoT in the training set. Secondly, when the model is inferring the category label of the current response, the F1 value of the classification with CoT guidance is about 8% higher than that without CoT guidance.
[0077] Table 1
[0078]
[0079] Further reference Figure 6 , as an implementation of the methods shown in the above figures, an embodiment of an interpretable multi-label classification device based on tag grouping and logical chain is provided in the present application. This device embodiment corresponds to the method embodiment shown in Figure 1 and can be specifically applied to various electronic devices.
[0080] An embodiment of the present application provides an interpretable multi-label classification device based on tag grouping and logical chain, including:
[0081] A tag grouping module 1, configured to collect a dialogue data set, where the samples in the dialogue data set include multi-round dialogues and the tag arrays corresponding to the current responses in the multi-round dialogues, and the multi-round dialogues include the dialogue history and its corresponding current response; group the tags in the tag arrays corresponding to all the multi-round dialogues to obtain the tags corresponding to each tag group;
[0082] A thinking process generation module 2, configured to construct a first language model, and use few-shot prompting and the chain of thought method to guide the first language model to generate the thinking process for the tag classification of the current response in the multi-round dialogue according to the multi-round dialogue and each tag group and its corresponding tags; set a dynamic tag array in the thinking process, where the dynamic tag array is initially empty, and first use each tag group as the start node and execute the classification judgment of each tag group, then further loop through each tag in the corresponding tag group according to the classification judgment result of each tag group and execute the classification judgment of each tag, and add the corresponding tags to the dynamic tag array according to the classification judgment result of each tag, so that the tags in the obtained dynamic tag array after traversal are the same as the tags in the tag array corresponding to the multi-round dialogue; construct an enhanced dialogue data set based on the samples in the dialogue data set and their corresponding thinking processes;
[0083] A fine-tuning module 3, configured to construct a second language model, and use the enhanced dialogue data set to fine-tune the second language model to obtain a fine-tuned second language model;
[0084] The label generation module 4 is configured to obtain a multi-turn conversation to be classified, input the multi-turn conversation to be classified and a prompt for multi-label classification of the current response in the multi-turn conversation to be classified according to the multi-turn conversation to be classified into a fine-tuned second language model, and obtain a corresponding thinking process and label array.
[0085] Figure 7 The following is a schematic hardware structure diagram of the electronic device provided by the embodiment of the present invention. As Figure 7 shown, the electronic device of this embodiment includes: a processor 701 and a memory 702; wherein the memory 702 is used to store computer execution instructions; the processor 701 is used to execute the computer execution instructions stored in the memory to implement each step executed by the electronic device in the above embodiment. For details, reference can be made to the relevant descriptions in the foregoing method embodiments.
[0086] Optionally, the memory 702 can be either independent or integrated with the processor 701.
[0087] When the memory 702 is independently provided, the electronic device further includes a bus 703 for connecting the memory 702 and the processor 701.
[0088] The embodiment of the present invention also provides a computer storage medium, in which computer execution instructions are stored. When the processor 701 executes the computer execution instructions, the above method is implemented.
[0089] The embodiment of the present invention also provides a computer program product, including a computer program. When the computer program is executed by the processor 701, the above method is implemented.
[0090] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be indirect couplings or communication connections through some interfaces, devices or modules, and can be electrical, mechanical or other forms.
[0091] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to implement the solution of this embodiment.
[0092] In addition, in each embodiment of the present invention, each functional module can be integrated in a processing unit, or each module can exist physically alone, or two or more modules can be integrated in a unit. The unit formed by the above modules can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.
[0093] The integrated module implemented in the form of a software functional module can be stored in a computer-readable storage medium. The above software functional module is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor 701 to execute some steps of the methods in various embodiments of the present application.
[0094] It should be understood that the above processor 701 can be a central processing unit (Central Processing Unit, abbreviated as CPU), or other general-purpose processors, digital signal processors (Digital Signal Processor, abbreviated as DSP), application specific integrated circuits (Application Specific Integrated Circuit, abbreviated as ASIC), etc. The general-purpose processor can be a microprocessor or the processor 701 can also be any conventional processor 701, etc. The steps of the method disclosed in combination with the invention can be directly embodied as being executed and completed by the hardware processor 701, or by a combination of hardware and software modules in the processor 701.
[0095] The memory 702 may include a high-speed RAM memory, and may also include non-volatile storage NVM, such as at least one disk memory, and may also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk, or an optical disc, etc.
[0096] The bus 703 can be an Industry Standard Architecture (ISA), a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus 703 can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, the bus 703 in the drawings of the present application is not limited to only one bus 703 or one type of bus 703.
[0097] The above storage medium may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc. The storage medium may be any available medium accessible by a general-purpose or special-purpose computer.
[0098] An exemplary storage medium is coupled to the processor 701 such that the processor 701 can read information from the storage medium and write information to the storage medium. Of course, the storage medium may also be a component of the processor 701. The processor 701 and the storage medium may be located in an application specific integrated circuit (ASIC). Of course, the processor 701 and the storage medium may also exist as discrete components in an electronic device or a master device.
[0099] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program may be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes various media that can store program codes, such as ROM, RAM, magnetic disks, or optical discs.
[0100] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An interpretable multi-label classification method based on label grouping and logical chains, characterized in that Including the following steps: Collect a dialogue dataset, where the samples in the dialogue dataset include multi-turn dialogues and the label arrays corresponding to the current responses in the multi-turn dialogues, and the multi-turn dialogues include the dialogue history and its corresponding current response; group the labels in the label arrays corresponding to all the multi-turn dialogues to obtain the labels corresponding to each label group; Construct a first language model, and use few-shot prompting and chain of thought to guide the first language model to generate the thinking process of label classification for the current response in the multi-turn dialogues based on the multi-turn dialogues and each label group and its corresponding labels; Set a dynamic label array in the thinking process, where the dynamic label array is initially empty, and first use each label group as the starting node and perform the classification judgment of each label group, then further loop through each label in the corresponding label group according to the classification judgment result of each label group and perform the classification judgment of each label, and add the corresponding label to the dynamic label array according to the classification judgment result of each label, so that the labels in the dynamic label array obtained after the traversal are the same as the labels in the label array corresponding to the multi-turn dialogues; construct an enhanced dialogue dataset based on the samples in the dialogue dataset and their corresponding thinking processes; Construct a second language model, and use the enhanced dialogue dataset to fine-tune the second language model to obtain a fine-tuned second language model; Obtain multi-turn dialogues to be classified, and input the multi-turn dialogues to be classified and the prompt words for multi-label classification of the current response in the multi-turn dialogues to be classified into the fine-tuned second language model to obtain the corresponding thinking process and label array.
2. The interpretable multi-label classification method based on label grouping and logical chain according to claim 1, wherein During the fine-tuning process of the second language model, the input of the second language model is multi-turn dialogues, thinking processes, and prompt words for multi-label classification of the current response in the multi-turn dialogues, and the output is a label array, which consists of at least one label.
3. The interpretable multi-label classification method based on label grouping and logical chain according to claim 1, wherein After all the labels in all the label groups in the thinking process are traversed, enter the end node. In the end node, judge whether the dynamic label array obtained after all the labels in all the label groups are traversed is empty. If it is, add "no label" to the dynamic label array obtained after all the labels in all the label groups are traversed to obtain a label array, otherwise directly use the dynamic label array obtained after all the labels in all the label groups are traversed as the label array.
4. The interpretable multi-label classification method based on label grouping and logical chain according to claim 1, wherein Group the labels in the label arrays corresponding to all the multi-turn dialogues to obtain the labels corresponding to each label group, specifically including: Set the admission conditions corresponding to each label group; Traverse each multi-turn dialogue and its corresponding label, and use the keyword-triggered method to classify the labels corresponding to the multi-turn dialogues whose content meets the admission conditions corresponding to the label group into the corresponding label group to obtain the labels corresponding to each label group.
5. The interpretable multi-label classification method based on label grouping and logical chain according to claim 1, characterized in that, The first language model uses the Qwen2-72B-Instruct model, and the second language model uses the Qwen2-7B-Instruct model.
6. The interpretable multi-label classification method based on label grouping and logical chain according to claim 1, characterized in that The fine-tuning method of the second language model adopts full-parameter fine-tuning.
7. An interpretable multi-label classification device based on label grouping and logical chains, characterized in that, It includes: A label grouping module, configured to collect a dialogue dataset, where the samples in the dialogue dataset include multi-turn dialogues and the label array corresponding to the current response in the multi-turn dialogues, and the multi-turn dialogues include the dialogue history and its corresponding current response; group the labels in the label arrays corresponding to all multi-turn dialogues to obtain the labels corresponding to each label group; A thinking process generation module, configured to construct a first language model, and use few-shot prompting and chain of thought methods to guide the first language model to generate a thinking process for classifying the labels of the current response in the multi-turn dialogues according to the multi-turn dialogues and each label group and its corresponding labels; Set a dynamic label array in the thinking process, the dynamic label array is initially empty, and first use each label group as the start node and execute the classification judgment of each label group, then further loop through each label in the corresponding label group according to the classification judgment result of each label group and execute the classification judgment of each label, and add the corresponding label to the dynamic label array according to the classification judgment result of each label, so that the labels in the dynamic label array obtained after the traversal are the same as the labels in the label array corresponding to the multi-turn dialogues; construct an enhanced dialogue dataset based on the samples in the dialogue dataset and their corresponding thinking processes; A fine-tuning module, configured to construct a second language model, and use the enhanced dialogue dataset to fine-tune the second language model to obtain a fine-tuned second language model; A label generation module, configured to obtain multi-turn dialogues to be classified, and input the multi-turn dialogues to be classified and the prompt words for multi-label classification of the current response in the multi-turn dialogues to be classified into the fine-tuned second language model to obtain the corresponding thinking process and label array.
8. An electronic device, including: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method according to any one of claims 1-6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the method according to any one of claims 1-6.