Classification model training method and device, classification method and device, equipment and medium
By supervised training and unsupervised training of text data, a target classification model is formed, which solves the problems of cumbersome and inconsistency in the traditional classification process and improves classification efficiency and accuracy.
Patent Information
- Application Number
- CN202510243922.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-05-27
AI Technical Summary
The traditional classification process of people's livelihood demands is cumbersome, time-consuming and difficult to cope with the needs of large-scale data processing, and is prone to introduce subjectivity and inconsistency, affecting the accuracy and reliability of classification results.
By obtaining text data and its classification labels, supervised training is performed to obtain the initial classification model, using this model to predict and obtain negative class labels, forming a training data set based on semantic description, and unsupervised training is performed to obtain the target classification model.
It improves classification efficiency and accuracy, reduces manual intervention, and enhances the reliability and consistency of classification results.
Smart Images

Figure CN120045981A_ABST
Abstract
Description
Technical Field
[0001] This application is applicable to the classification technology field, and particularly relates to a training method, a classification method, a device, a device and a medium for a classification model. Background Art
[0002] In the business scenario of people's livelihood demands, processing the events reported by the public is a heavy and complex task. In the traditional work process, after receiving the reported events, staff need to manually select the most appropriate classification label from a large number of preset event categories. However, this process encounters multiple challenges in actual operation: First, the data volume of people's livelihood demands is huge, the classification system is intricate, and there are numerous classification criteria, making the screening process extremely difficult; Second, manual classification is not only time-consuming and laborious, with low efficiency, but also difficult to handle the large-scale data processing requirements; Third, staff must memorize and understand a large number of classification criteria to ensure that the demands are accurately classified into the corresponding categories, which poses a severe test to their memory and understanding ability, and also significantly increases the learning cost. In addition, manual classification may introduce subjectivity and inconsistency, affecting the accuracy and reliability of the classification results. Therefore, how to improve the classification efficiency has become an urgent problem to be solved. Summary of the Invention
[0003] In view of this, the embodiments of this application provide a training method, a classification method, a device, a device and a medium for a classification model to solve the problem of how to improve the classification efficiency.
[0004] In the first aspect, the embodiments of this application provide a training method for a classification model, including: Obtain first text data, a first positive classification label corresponding to the first text data, and second text data; According to the first text data and the first positive classification label corresponding to the first text data, perform supervised training on the initial classification model to obtain a trained initial classification model, and use the trained initial classification model to perform actual prediction on the second text data to obtain a second positive classification label corresponding to the second text data; Obtain a first negative classification label of the first text data during the supervised training process, and a second negative classification label of the second text data during the actual prediction process, and respectively match corresponding semantic descriptions for the first positive classification label, the first negative classification label, the second positive classification label, and the second negative classification label from the label expression mapping table; Form a first training data set by combining the first text data with the semantic descriptions of the corresponding first positive classification label and the first negative classification label, and form a second training data set by combining the second text data with the semantic descriptions of the corresponding second positive classification label and the second negative classification label; Use the first training data set and the second training data set to perform unsupervised training on the trained initial classification model to obtain a target classification model.
[0005] In a second aspect, an embodiment of the present application provides a classification method, including: Obtain the target classification model obtained by the training method of the text classification model in the first aspect; Obtain the text data to be classified, input the text data to be classified into the target classification model, and output the classification result of the text data to be classified.
[0006] In a third aspect, an embodiment of the present application provides a training device for a classification model, including: A first acquisition module for acquiring first text data, a first positive classification label corresponding to the first text data, and second text data; A supervised training module for performing supervised training on an initial classification model according to the first text data and the first positive classification label corresponding to the first text data to obtain a trained initial classification model, and using the trained initial classification model to perform actual prediction on the second text data to obtain a second positive classification label corresponding to the second text data; A matching module for acquiring a first negative classification label of the first text data during the supervised training process and a second negative classification label of the second text data during the actual prediction process, and respectively matching corresponding semantic descriptions for the first positive classification label, the first negative classification label, the second positive classification label, and the second negative classification label from a label expression mapping table; A first construction module for forming a first training data set by combining the first text data with the semantic descriptions of the corresponding first positive classification label and the first negative classification label, and forming a second training data set by combining the second text data with the semantic descriptions of the corresponding second positive classification label and the second negative classification label; An unsupervised training module for using the first training data set and the second training data set to perform unsupervised training on the trained initial classification model to obtain a target classification model.
[0007] In a fourth aspect, an embodiment of the present application provides a classification device, including: A second acquisition module for acquiring the target classification model obtained by the training method of the text classification model in the first aspect; A classification module, configured to obtain text data to be classified, input the text data to be classified into the target classification model, and output a classification result of the text data to be classified.
[0008] In a fifth aspect, an embodiment of the present application provides a computer device, which includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the training method of the classification model described in the first aspect or the classification method described in the second aspect is implemented.
[0009] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the training method of the classification model described in the first aspect or the classification method described in the second aspect is implemented.
[0010] The beneficial effects of the embodiments of the present application compared with the prior art are as follows: The present application performs supervised training on the initial classification model according to the first text data and the first positive classification label corresponding to the first text data to obtain a trained initial classification model. The trained initial classification model is used to actually predict the second text data to obtain a second positive classification label corresponding to the second text data. The first negative classification label of the first text data during the supervised training process and the second negative classification label of the second text data during the actual prediction process are obtained. From the label expression mapping table, semantic descriptions corresponding to the first positive classification label, the first negative classification label, the second positive classification label, and the second negative classification label are respectively matched. The first text data and the semantic descriptions of the corresponding first positive classification label and the first negative classification label form a first training data set. The second text data and the semantic descriptions of the corresponding second positive classification label and the second negative classification label form a second training data set. The first training data set and the second training data set are used to perform unsupervised training on the trained initial classification model to obtain a target classification model. The text data to be classified is input into the target classification model, and a classification result is output.
[0011] Among them, compared with the traditional semi-supervised learning method, in the unsupervised training process of this application, contrast training is carried out based on the semantic descriptions corresponding to the second positive classification labels and the semantic descriptions corresponding to the second negative classification labels of the second text data. So that during the training process, the model can more deeply learn the semantic connection between the text data and the semantic description of the corresponding positive classification label, as well as the semantic connection between the text data and the semantic description of the corresponding negative classification label, better narrowing the distance between the text data and the semantic description of the corresponding positive classification label, and widening the distance between the text data and the semantic description of the corresponding negative classification label. Therefore, when classifying the text data to be classified based on the target classification model, not only the classification efficiency is improved, but also the classification accuracy is improved. Brief Description of the Drawings
[0012] To more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0013] Figure 1 It is a schematic diagram of an application environment of a training method and a classification method of a classification model provided in Embodiment 1 of this application; Figure 2 It is a schematic flowchart of a training method of a classification model provided in Embodiment 2 of this application; Figure 3 It is a schematic flowchart of a training method of a classification model provided in Embodiment 3 of this application; Figure 4 It is a schematic flowchart of a training method of a classification model provided in Embodiment 4 of this application; Figure 5 It is a schematic flowchart of a classification method provided in Embodiment 5 of this application; Figure 6 It is a schematic structural diagram of a training device of a classification model provided in Embodiment 6 of this application; Figure 7 It is a schematic structural diagram of a classification device provided in Embodiment 7 of this application; Figure 8 It is a schematic structural diagram of a computer device provided in Embodiment 8 of this application. Detailed Description of the Invention
[0014] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from obscuring the description of the present application.
[0015] It should be understood that when used in the specification and claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0016] It should also be understood that the term "and / or" as used in the specification and claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0017] As used in the specification and claims of the present application, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" depending on the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" depending on the context.
[0018] In addition, in the description of the specification and claims of the present application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and should not be construed as indicating or implying relative importance.
[0019] The reference to "one embodiment" or "some embodiments" or the like described in the specification of the present application means that a specific feature, structure, or characteristic described in connection with the embodiment is included in one or more embodiments of the present application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way.
[0020] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0021] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0022] It should be understood that the magnitudes of the sequence numbers of the steps in the following embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0023] To illustrate the technical solution of the present application, specific embodiments will be used for illustration below.
[0024] A training method and a classification method for a classification model provided in the first embodiment of the present application can be applied in an application environment such as Figure 1 . Among them, the server communicates with the client, the server provides analysis services, and the client triggers an analysis task to the server. Among them, the client includes, but is not limited to, devices such as a palm computer, a desktop computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a cloud computer device, and a personal digital assistant (PDA). The computer device corresponding to the server can be implemented by an independent server or a server cluster composed of multiple servers.
[0025] See Figure 2 , which is a schematic flowchart of a training method for a classification model provided in the second embodiment of the present application. This method is applied to the server in Figure 1 . The server is connected to the client to obtain the first text data, the first positive classification label corresponding to the first text data, and the second text data sent by the client. As shown in Figure 2 , the following steps may be included: Step S201, obtain the first text data, the first positive classification label corresponding to the first text data, and the second text data.
[0026] Step S202: Supervise and train the initial classification model according to the first text data and the first positive classification label corresponding to the first text data to obtain a trained initial classification model, and use the trained initial classification model to perform actual prediction on the second text data to obtain a second positive classification label corresponding to the second text data.
[0027] In this embodiment, the first text data and the second text data may refer to text-type data. The first text data and the second text data are from different data sources. The positive classification label may refer to the positive class label corresponding to the text data, which is used to represent the classification category that the model needs to identify or predict. The negative classification label may refer to the negative class label corresponding to the text data, which is used to represent the classification category that the model does not need to identify or predict. The first positive classification label may refer to the positive class label corresponding to the first text data, and the second positive classification label may refer to the positive class label corresponding to the second text data. The initial classification model may refer to a preset initial text classification model, and this initial classification model may be a Bidirectional Encoder Representations from Transformers (BERT) model based on Transformer.
[0028] Specifically, input the first text data and the first positive classification label corresponding to the first text data into the initial classification model to supervise and train the initial classification model. During the supervision training process, obtain the training prediction result corresponding to the first text data, that is, the prediction result of the classification category of the first text data by the initial classification model, calculate the training error value between the training prediction result and the first positive classification label, and update the parameters of the initial classification model according to the training error value. When the termination condition of the supervision training is reached, obtain a trained initial classification model; Input the second text data into the trained initial classification model. The trained initial classification model is processed through its neural network structure and finally reaches the softmax layer. The softmax layer converts the output of the trained initial classification model into a probability distribution, where each classification category has a corresponding probability value. In this probability distribution, the classification category with a probability value reaching the threshold is output as the second positive classification label corresponding to the second text data.
[0029] Step S203: Obtain the first negative classification label of the first text data during the supervision training process and the second negative classification label of the second text data during the actual prediction process, and match the corresponding semantic descriptions for the first positive classification label, the first negative classification label, the second positive classification label, and the second negative classification label from the label expression mapping table.
[0030] In this embodiment, the first negative classification label may refer to the negative class label corresponding to the first text data, the second negative classification label may refer to the negative class label corresponding to the second text data, the semantic description may refer to the written description provided for the first positive classification label, the first negative classification label, the second positive classification label, and the second negative classification label, which can accurately express their class meanings and have clear semantic information, and the label expression mapping table is used to store the mapping relationships between the first positive classification label, the first negative classification label, the second positive classification label, the second negative classification label, and the corresponding semantic descriptions.
[0031] Specifically, in the process of supervised training according to the content in step S202, for the training prediction result of any first text data, the training prediction result with a training error value exceeding the threshold is used as the first negative classification label of the first text data; in the process of actual prediction according to the content in step S202, for any second text data, after inputting the second text data into the trained initial classification model and processing it through its neural network structure, when reaching the softmax layer, the classification category with a probability value not reaching the threshold is used as the second negative classification label of the second text data. After obtaining the first positive classification label and the first negative classification label of the first text data, the semantic description corresponding to the first positive classification label and the semantic description corresponding to the first negative classification label are matched from the label mapping table. After obtaining the second positive classification label and the second negative classification label of the second text data, the semantic description corresponding to the second positive classification label and the semantic description corresponding to the second negative classification label are matched from the label mapping table.
[0032] Step S204: Form a first training data set by using the first text data, the semantic description of the corresponding first positive classification label, and the semantic description of the first negative classification label, and form a second training data set by using the second text data, the semantic description of the corresponding second positive classification label, and the semantic description of the second negative classification label.
[0033] Step S205: Use the first training data set and the second training data set to perform unsupervised training on the trained initial classification model to obtain a target classification model.
[0034] In this embodiment, the first training data set may refer to the data set for training the trained initial classification model formed by using the first text data, the semantic description of the corresponding first positive classification label, and the semantic description of the first negative classification label. The second training data set may refer to the data set for training the trained initial classification model formed by using the second text data, the semantic description of the corresponding second positive classification label, and the semantic description of the second negative classification label. The target classification model may refer to the model obtained by performing unsupervised training on the trained initial classification model based on the first training data set and the second training data set.
[0035] Specifically, an initial classification model that has been trained is subjected to unsupervised training using a first training data set formed by first text data and semantic descriptions corresponding to a first positive classification label and a first negative classification label, and a second training data set formed by second text data and semantic descriptions corresponding to a second positive classification label and a second negative classification label. When the termination condition for unsupervised training is reached, a target classification model is obtained.
[0036] In the embodiments of the present application, compared with traditional semi-supervised learning methods, during the unsupervised training process of the present application, contrast training is performed based on the semantic descriptions corresponding to the second positive classification label and the second negative classification label of the second text data, so that during the training process, the model can more deeply learn the semantic connection between the text data and the semantic description corresponding to the positive classification label, as well as the semantic connection between the text data and the semantic description corresponding to the negative classification label, better narrowing the distance between the text data and the semantic description corresponding to the positive classification label, and widening the distance between the text data and the semantic description corresponding to the negative classification label. Therefore, when classifying the text data to be classified based on the target classification model, not only the classification efficiency is improved, but also the classification accuracy is improved.
[0037] See Figure 3 , which is a schematic flowchart of a method for training a classification model provided in Embodiment 3 of the present application. As Figure 3 shown, in step S204 above, forming a first training data set by the first text data and the semantic descriptions corresponding to the first positive classification label and the first negative classification label of the first text data, and forming a second training data set by the second text data and the semantic descriptions corresponding to the second positive classification label and the second negative classification label of the second text data may include the following steps: Step S301, for any first text data, construct a triple corresponding to the first text data according to the first text data, the semantic description of the first positive classification label of the first text data, and the semantic description of the first negative classification label of the first text data.
[0038] Step S302, for any second text data, construct a triple corresponding to the second text data according to the second text data, the semantic description of the second positive classification label of the second text data, and the semantic description of the second negative classification label of the second text data.
[0039] Step S303, form a first training data set from all the triples corresponding to the first text data, and form a second training data set from all the triples corresponding to the second text data.
[0040] Specifically, for any first text data, use the first text data as the anchor sample, the semantic description of the first positive classification label of the first text data as the positive sample, and the semantic description of the first negative classification label of the first text data as the negative sample to construct a triple corresponding to the first text data. Correspondingly, for any second text data, use the second text data as the anchor sample, the semantic description of the second positive classification label of the second text data as the positive sample, and the semantic description of the second negative classification label of the second text data as the negative sample to construct a triple corresponding to the second text data; form a first training data set with the triples corresponding to all the first text data, and form a second training data set with the triples corresponding to all the second text data.
[0041] In the embodiment of the present application, by constructing triples based on the first text data, the semantic description corresponding to the first positive classification label, the semantic description corresponding to the first negative classification label, the second text data, the semantic description corresponding to the second positive classification label, and the semantic description corresponding to the second negative classification label, and forming a training data set with the constructed triples to train the trained initial classification model. It helps the model to more accurately understand the semantic relationship between the text data and the semantic description of the classification label, better shorten the distance between the text data and the semantic description of the corresponding positive classification label, and widen the distance between the text data and the semantic description of the corresponding negative classification label. Therefore, when classifying the text data to be classified based on the target classification model, not only the classification efficiency is improved, but also the classification accuracy is improved.
[0042] See Figure 4 , which is a schematic flowchart of a method for training a classification model provided in Embodiment 4 of the present application. As Figure 4 shown, in the above step S205, using the first training data set and the second training data set to perform unsupervised training on the trained initial classification model to obtain the target classification model may include the following steps: Step S401, for any triple in the first training data set, calculate the positive similarity value between the first text data of the triple and the semantic description of the first positive classification label, and calculate the negative similarity value between the first text data of the triple and the semantic description of the first negative classification label.
[0043] Step S402, for any triple in the second training data set, calculate the positive similarity value between the second text data of the triple and the semantic description of the second positive classification label, and calculate the negative similarity value between the second text data of the triple and the semantic description of the second negative classification label.
[0044] Step S403, through a preset loss function, update the parameters of the trained initial classification model according to the positive similarity values and negative similarity values corresponding to all the triples to obtain the target classification model.
[0045] In this embodiment, the positive similarity value may refer to the similarity value between the first text data and the semantic description of the first positive classification label, or the similarity value between the second text data and the semantic description of the second positive classification label. The negative similarity value may refer to the similarity value between the first text data and the semantic description of the first negative classification label, or the similarity value between the second text data and the semantic description of the second negative classification label. The preset loss function may refer to a loss function set in advance. For example, the preset loss function may be a triplet loss function.
[0046] Specifically, for any triplet in the first training dataset, the positive similarity value between the first text data of the triplet and the semantic description of the first positive classification label can be calculated through the cosine similarity function, and the negative similarity value between the first text data of the triplet and the semantic description of the first negative classification label can be calculated. Correspondingly, for any triplet in the second training dataset, the positive similarity value between the second text data of the triplet and the semantic description of the second positive classification label can be calculated through the cosine similarity function, and the negative similarity value between the second text data of the triplet and the semantic description of the second negative classification label can be calculated. For any triplet, the training error value between the corresponding positive similarity value and negative similarity value of the triplet is calculated through the preset loss function. According to the training error value, the parameters of the trained initial classification model are updated. When the unsupervised training termination condition is reached, the target classification model is obtained.
[0047] In the embodiment of the present application, by calculating the positive similarity value between the text data and the semantic description of the positive classification label, and the negative similarity value between the text data and the semantic description of the negative classification label, based on the preset loss function, the training error value between the positive similarity value and the negative similarity value is calculated. Based on the training error value, the parameters of the trained initial classification model are updated to obtain the target classification model. This helps the model to more accurately understand the semantic relationship between the text data and the semantic description of the classification label, better narrow the distance between the text data and the semantic description of the corresponding positive classification label, and widen the distance between the text data and the semantic description of the corresponding negative classification label. Therefore, when classifying the text data to be classified based on the target classification model, not only the classification efficiency is improved, but also the classification accuracy is improved.
[0048] See Figure 5 , which is a schematic flowchart of a classification method provided in Embodiment 5 of the present application. This method is applied to Figure 1 the server in, and the server is connected to the client to obtain the text data to be classified sent by the client. As Figure 5 shown, the following steps may be included: Step S501, obtain the target classification model obtained by the training method of the text classification model; In step S502, obtain the text data to be classified, input the text data to be classified into the target classification model, and output the classification result of the text data to be classified.
[0049] In this embodiment, the text data to be classified may refer to the data of the text type to be classified, and the classification result may refer to the classification result of the target classification model for the text data to be classified.
[0050] Specifically, input the text data to be classified into the target classification model, and after the classification process of the target classification model, output the classification result; Optionally, after obtaining the text data to be classified in the above step S502, it is also possible to perform feature extraction on the text data to be classified to obtain the key data in the text data to be classified; according to the key data, query the classification mapping table to obtain the classification result of the text data to be classified.
[0051] The key data may refer to the data that can reflect the main content of the text data to be classified extracted from the text data to be classified, and the classification mapping table is used to store the mapping relationship between the key data and the classification result.
[0052] After obtaining the text data to be classified, it is also possible to extract the key data in the text data to be classified, perform a pre-query on the classification result of the text data to be classified according to the classification mapping table. If the corresponding classification result is not queried, then input the text data to be classified into the target classification model to output the classification result.
[0053] In the embodiment of the present application, compared with the traditional semi-supervised learning method, in the unsupervised training process of the present application, based on the semantic description corresponding to the second positive classification label and the semantic description corresponding to the second negative classification label of the second text data for comparative training, so that during the training process, the model can more deeply learn the semantic connection between the text data and the semantic description of the corresponding positive classification label, and the semantic connection between the text data and the semantic description of the corresponding negative classification label, better narrowing the distance between the text data and the semantic description of the corresponding positive classification label, and widening the distance between the text data and the semantic description of the corresponding negative classification label. Therefore, when classifying the text data to be classified based on the target classification model, not only the classification efficiency is improved, but also the classification accuracy is improved.
[0054] Corresponding to the training method of the classification model in the above embodiment, Figure 6 shows the structural block diagram of the training device of the classification model provided in the sixth embodiment of the present application. This device is applied to Figure 1 the server in. For the sake of convenience of description, only the parts related to the embodiment of the present application are shown.
[0055] See Figure 6, the training device of the classification model includes: The first acquisition module 61 is configured to acquire first text data, a first positive classification label corresponding to the first text data, and second text data; The supervised training module 62 is configured to perform supervised training on the initial classification model according to the first text data and the first positive classification label corresponding to the first text data, obtain a trained initial classification model, and use the trained initial classification model to perform actual prediction on the second text data to obtain a second positive classification label corresponding to the second text data; The matching module 63 is configured to acquire a first negative classification label of the first text data during the supervised training process, and a second negative classification label of the second text data during the actual prediction process, and match corresponding semantic descriptions for the first positive classification label, the first negative classification label, the second positive classification label, and the second negative classification label from the label expression mapping table; The first construction module 64 is configured to form a first training data set by using the first text data and the semantic descriptions of the corresponding first positive classification label and the first negative classification label of the first text data, and form a second training data set by using the second text data and the semantic descriptions of the corresponding second positive classification label and the second negative classification label of the second text data; The unsupervised training module 65 is configured to perform unsupervised training on the trained initial classification model by using the first training data set and the second training data set to obtain a target classification model.
[0056] Optionally, the first construction module 64 includes: The second construction unit is configured to construct a triple corresponding to any first text data according to the first text data, the semantic description of the first positive classification label of the first text data, and the semantic description of the first negative classification label of the first text data; The third construction unit is configured to construct a triple corresponding to any second text data according to the second text data, the semantic description of the second positive classification label of the second text data, and the semantic description of the second negative classification label of the second text data; The fourth construction unit is configured to form the first training data set by using all the triples corresponding to the first text data, and form the second training data set by using all the triples corresponding to the second text data.
[0057] Optionally, the unsupervised training module 65 includes: A first calculation unit, configured to calculate a positive similarity value between the first text data of any triple in the first training dataset and the semantic description of the first positive classification label, and calculate a negative similarity value between the first text data of the triple and the semantic description of the first negative classification label; A second calculation unit, configured to calculate a positive similarity value between the second text data of any triple in the second training dataset and the semantic description of the second positive classification label, and calculate a negative similarity value between the second text data of the triple and the semantic description of the second negative classification label; A model update unit, configured to update the parameters of the trained initial classification model according to the positive similarity values and negative similarity values corresponding to all triples through a preset loss function, so as to obtain the target classification model.
[0058] Optionally, the model update unit includes: An error calculation sub-unit, configured to calculate a training error value between the positive similarity value and the negative similarity value corresponding to any triple through the preset loss function; A parameter update sub-unit, configured to update the parameters of the trained initial classification model according to the training error value, so as to obtain the target classification model.
[0059] It should be noted that the information interaction, execution process, etc. between the above modules, due to being based on the same concept as the method embodiment of the present application, for their specific functions and the technical effects brought, please refer to the method embodiment part for details, and will not be elaborated here.
[0060] Corresponding to the classification method in the above embodiment, Figure 7 FIG. shows a structural block diagram of a classification device provided in Embodiment 7 of the present application. This device is applied to Figure 1 the server in. For the sake of convenience of description, only the parts related to the embodiments of the present application are shown.
[0061] See Figure 7 , this classification device includes: A second acquisition module 71, configured to acquire the target classification model obtained by the training method of the text classification model; A classification module 72, configured to acquire the text data to be classified, input the text data to be classified into the target classification model, and output the classification result of the text data to be classified.
[0062] Optionally, this classification device further includes: A feature extraction module, configured to extract features from the text data to be classified to obtain the key data in the text data to be classified; A query module, configured to query a classification mapping table according to the key data to obtain a classification result of the text data to be classified, where the classification mapping table is used to store the mapping relationship between the key data and the classification result.
[0063] It should be noted that for the information interaction, execution process, etc. between the above modules, since they are based on the same concept as the method embodiments of this application, their specific functions and the technical effects brought about can be specifically referred to in the method embodiment part, and will not be elaborated here.
[0064] Figure 8 This is a schematic structural diagram of a computer device provided in Embodiment 8 of this application. As Figure 8 shown, the computer device in this embodiment includes: at least one processor ( Figure 8 only one is shown in the figure), a memory, and a computer program stored in the memory and executable on at least one processor. When the processor executes the computer program, it implements the steps in the training method of any of the above classification models or the classification method embodiments.
[0065] The computer device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that Figure 8 merely an example of a computer device, which does not constitute a limitation on the computer device. The computer device may include more or fewer components than shown in the figure, or combine some components, or different components. For example, it may also include a network interface, a display screen, and an input device, etc.
[0066] The so-called processor may be a CPU, and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0067] The memory includes a readable storage medium, an internal memory, etc. Among them, the internal memory can be the memory of a computer device, and the internal memory provides an environment for the operation of the operating system and computer-readable instructions in the readable storage medium. The readable storage medium can be the hard disk of a computer device, and in some other embodiments, it can also be an external storage device of a computer device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device. Further, the memory can also include both the internal storage unit of the computer device and external storage devices. The memory is used to store an operating system, application programs, a boot loader (BootLoader), data, and other programs, such as the program code of a computer program. The memory can also be used to temporarily store data that has been output or is to be output.
[0068] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above device can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above method embodiments of this application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry computer program code, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.
[0069] To implement all or part of the processes in the above method embodiments of this application, it can also be completed by a computer program product. When the computer program product runs on a computer device, it enables the computer device to execute and implement the steps in the above method embodiments.
[0070] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0071] Those of ordinary skill in the art will realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.
[0072] In the embodiments provided in this application, it should be understood that the disclosed device / computer equipment and method can be implemented in other ways. For example, the device / computer equipment embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the device or unit can be in electrical, mechanical or other forms.
[0073] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0074] The above embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included in the protection scope of this application.
Claims
1. A classification model training method, characterized in that: include: Acquire first text data, a first positive classification label corresponding to the first text data, and second text data; According to the first text data and the first positive classification label corresponding to the first text data, supervised training is performed on the initial classification model to obtain a trained initial classification model, and the trained initial classification model is used to perform actual prediction on the second text data to obtain a second positive classification label corresponding to the second text data; Obtain a first negative classification label of the first text data in a supervised training process and a second negative classification label of the second text data in an actual prediction process, and match the first positive classification label, the first negative classification label, the second positive classification label, and the second negative classification label to corresponding semantic descriptions from a label representation mapping table; The first text data and the corresponding semantic description of the first positive classification label and the semantic description of the first negative classification label form a first training data set, and the second text data and the corresponding semantic description of the second positive classification label and the semantic description of the second negative classification label form a second training data set; The first training data set and the second training data set are used to perform unsupervised training on the trained initial classification model to obtain a target classification model.
2. The training method of the classification model according to claim 1, characterized in that: The first text data and the corresponding semantic description of the first positive classification label and the semantic description of the first negative classification label are used to form a first training data set, and the second text data and the corresponding semantic description of the second positive classification label and the semantic description of the second negative classification label are used to form a second training data set, including: For any first text data, construct a triple corresponding to the first text data according to the first text data, the semantic description of the first positive classification label of the first text data, and the semantic description of the first negative classification label of the first text data; For any second text data, construct a triple corresponding to the second text data according to the second text data, the semantic description of the second positive classification label of the second text data, and the semantic description of the second negative classification label of the second text data; The triplets corresponding to all the first text data form the first training data set, and the triplets corresponding to all the second text data form the second training data set.
3. The training method of the classification model according to claim 2, characterized in that: The step of using the first training data set and the second training data set to perform unsupervised training on the trained initial classification model to obtain a target classification model includes: For any triple in the first training data set, calculating a positive similarity value between the first text data of the triple and the semantic description of the first positive classification label, and calculating a negative similarity value between the first text data of the triple and the semantic description of the first negative classification label; For any triplet in the second training data set, calculating a positive similarity value between the second text data of the triplet and the semantic description of the second positive classification label, and calculating a negative similarity value between the second text data of the triplet and the semantic description of the second negative classification label; By means of a preset loss function, the parameters of the trained initial classification model are updated according to the positive similarity values and negative similarity values corresponding to all triples to obtain the target classification model.
4. The training method of the classification model according to claim 3, characterized in that: The method of updating the parameters of the trained initial classification model according to the positive similarity values and negative similarity values corresponding to all triples through a preset loss function to obtain the target classification model includes: For any triple, using the preset loss function, calculate the training error value between the positive similarity value and the negative similarity value corresponding to the triple; According to the training error value, the parameters of the trained initial classification model are updated to obtain the target classification model.
5. A classification method, characterized in that: include: Obtain a target classification model obtained by the training method of a text classification model according to any one of claims 1 to 4; Acquire text data to be classified, input the text data to be classified into the target classification model, and output the classification result of the text data to be classified.
6. The classification method according to claim 5, characterized in that: After obtaining the text data to be classified, the method further includes: Performing feature extraction on the text data to be classified to obtain key data in the text data to be classified; According to the key data, a classification mapping table is queried to obtain a classification result of the text data to be classified, and the classification mapping table is used to store a mapping relationship between the key data and the classification result.
7. A training device for a classification model, characterized in that: include: A first acquisition module, used to acquire first text data, a first positive classification label corresponding to the first text data, and second text data; A supervised training module is used to perform supervised training on the initial classification model according to the first text data and the first positive classification label corresponding to the first text data to obtain a trained initial classification model, and use the trained initial classification model to perform actual prediction on the second text data to obtain a second positive classification label corresponding to the second text data; A matching module, used to obtain a first negative classification label of the first text data in a supervised training process and a second negative classification label of the second text data in an actual prediction process, and to match the first positive classification label, the first negative classification label, the second positive classification label and the second negative classification label to corresponding semantic descriptions from a label representation mapping table; A first construction module is used to form a first training data set with the first text data and the corresponding semantic description of the first positive classification label and the semantic description of the first negative classification label, and to form a second training data set with the second text data and the corresponding semantic description of the second positive classification label and the semantic description of the second negative classification label; The unsupervised training module is used to use the first training data set and the second training data set to perform unsupervised training on the trained initial classification model to obtain a target classification model.
8. A classification device, characterized in that: include: A second acquisition module, used to acquire a target classification model obtained by the training method of a text classification model according to any one of claims 1 to 4; The classification module is used to obtain text data to be classified, input the text data to be classified into the target classification model, and output the classification result of the text data to be classified.
9. A computer device, characterized in that: The computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the training method for the classification model as described in any one of claims 1 to 4, or the classification method as described in any one of claims 5 to 6.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it implements the training method of the classification model as described in any one of claims 1 to 4, or the classification method as described in any one of claims 5 to 6.