A multi-label text classification method, apparatus and device

By training text samples and label sequences, a multi-label text classification method was developed. This method utilizes techniques such as the binary cross-entropy loss function and the positive point cross-PPMI correlation matrix to optimize the multi-label classification model, thereby solving the problem of low accuracy in existing technologies and achieving higher classification accuracy.

CN119829769BActive Publication Date: 2025-11-14CHINA UNIV OF GEOSCIENCES (WUHAN)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411912631.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-11-14
Estimated Expiration
2044-12-24

AI Technical Summary

Technical Problem

Existing fine-grained multi-label classification tasks generally have low accuracy, especially in multi-label text classification, where the problems of uneven data distribution and the complexity of label number prediction have not been effectively solved.

Method used

By acquiring text samples and label sequences, the initial prediction model is trained using the binary cross-entropy loss function, the positive point cross-PPMI correlation matrix, and the text label similarity matrix, combined with the boundary ranking loss function. The internal parameters are then adjusted to optimize the model's prediction results.

Benefits of technology

It improves the overall accuracy of multi-label classification, captures the semantic correlation between labels, and optimizes the model's prediction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119829769B_ABST
    Figure CN119829769B_ABST
Patent Text Reader

Abstract

This invention provides a multi-label text classification method, apparatus, and device, relating to the field of deep learning technology. The method includes: acquiring multiple text samples and corresponding label sequences; training an initial prediction model using the text samples and label sequences; determining the primary loss function value of the initial prediction model using a binary cross-entropy loss function; determining a first auxiliary loss function value of the initial prediction model by calculating the positive point cross-PPMI correlation matrix and the correlation difference between each label sequence; determining a second auxiliary loss function value of the initial prediction model based on a text label similarity matrix and a boundary ranking loss function; and adjusting the internal parameters of the initial prediction model based on the aforementioned loss function values ​​to obtain a multi-label text classification model. This invention effectively improves the overall accuracy of multi-label classification by capturing the semantic correlation between labels and comparing it with the results obtained from text feature training to optimize the model's prediction results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of deep learning technology, and in particular to a multi-label text classification method, apparatus, and device. Background Technology

[0002] Text classification is one of the most fundamental tasks in natural language processing (NLP). It involves determining the categories of a text sequence based on its features, requiring the assignment of category labels. Generally, it can be divided into supervised and unsupervised text classification. Supervised text classification uses labeled data during training, where each data point has a category label. This labeled data is fed into the model, which learns the relationship between text and categories through training, thus fitting a functional relationship between text features and categories. A trained model can then predict the category of unclassified text. Unsupervised text classification, on the other hand, does not require labeled data for model training. Instead, it uses the text's inherent features and structure to cluster or group the text, achieving classification. Generally, this type of clustering task can only perform coarse-grained hierarchical classification, and its effectiveness is often insufficient for certain domain requirements. Typical text classification tasks include sentiment analysis and topic classification. Less obvious tasks such as translation, named entity recognition, and summarization also require text classification techniques. One typical task in supervised text classification is multi-label text classification. In practical applications, texts for this task usually have multiple labels. With the development of the Internet and big data, the information contained in the data is often multifaceted. Therefore, research on multi-label text classification is more valuable, but also more complex. The classification process requires predicting the number of labels, and there may be uneven data distribution. Summary of the Invention

[0003] The purpose of this invention is to address the problem of generally low accuracy in current fine-grained multi-label classification tasks by providing a multi-label text classification method, apparatus, and device.

[0004] The technical solution of this application embodiment is implemented as follows:

[0005] The first aspect of this application provides a multi-label text classification method, including:

[0006] Obtain multiple text samples and corresponding tag sequences for the text samples; the tag sequence includes multiple tags related to the text samples, and the text samples include character sequences consisting of at least one character;

[0007] The initial prediction model is trained using the text samples and the label sequence, and the main loss function value of the initial prediction model is determined by the binary cross-entropy loss function.

[0008] A co-occurrence matrix is ​​constructed using the label co-occurrence information of the label sequences to determine the positive point mutual PPMI correlation matrix. The mean squared error loss function is used to calculate the correlation difference between the positive point mutual PPMI correlation matrix and each label sequence to determine the first auxiliary loss function value of the initial prediction model.

[0009] Based on the text samples and the label sequence, a text label similarity matrix is ​​determined, and the value of the second auxiliary loss function of the initial prediction model is determined by combining the boundary ranking loss function;

[0010] Based on the main loss function value, the first auxiliary loss function value, and the second auxiliary loss function value, the internal parameters of the initial prediction model are adjusted, and iterative training is performed until the training termination condition is met to obtain a multi-label text classification model.

[0011] Optionally, training the initial prediction model using the text samples and the label sequence, and determining the main loss function value of the initial prediction model using the binary cross-entropy loss function, includes:

[0012] The text sample is input into the initial prediction model to obtain the predicted label;

[0013] The difference between the predicted label and the label sequence is calculated using the binary cross-entropy loss function to determine the main loss function value of the initial prediction model.

[0014] Optionally, the difference between the predicted label and the label sequence is calculated using the binary cross-entropy loss function to determine the principal loss function value of the initial prediction model. The calculation formula is as follows:

[0015] ;

[0016] in, The true label represents the probability of a positive sample and a negative sample. For predicted values, , representing the probability of a positive sample. Indicates batch size, This represents the average loss for a batch of samples.

[0017] Optionally, the step of constructing a co-occurrence matrix using the tag co-occurrence information of the tag sequence to determine the positive point mutual PPMI correlation matrix includes:

[0018] Iterate through the label sequence of each text sample, count the frequency of each label sequence, and determine the co-occurrence matrix corresponding to the label sequence;

[0019] Calculate the positive point cross-PPMI correlation matrix between the tag sequences based on the co-occurrence matrix.

[0020] Optionally, calculating the positive point cross-PPMI correlation matrix between the tag sequences based on the co-occurrence matrix includes determining the positive point cross-PPMI correlation matrix using the following formula:

[0021] ;

[0022] ;

[0023] in, This represents the probability that labels i and j appear simultaneously. This represents the joint probability of label i and label j under independent events.

[0024] Optionally, the step of calculating the correlation difference between the positive point cross-PPMI correlation matrix and each of the label sequences using the mean squared error loss function to determine the first auxiliary loss function value of the initial prediction model includes determining the first auxiliary loss function value using the following formula:

[0025] ;

[0026] in, The true label represents the probability of a positive sample and a negative sample. For predicted values, , representing the probability of a positive sample. Indicates batch size, This represents the mean squared error loss function value.

[0027] Optionally, determining the text label similarity matrix based on the text samples and the label sequence, and determining the second auxiliary loss function value of the initial prediction model in conjunction with the boundary ranking loss function, includes:

[0028] The text samples and the label sequences are encoded using ALBERT, and the text samples and the label sequences are mapped into low-dimensional semantic vectors to obtain text embedding vectors and label embedding vectors;

[0029] The cosine similarity of the text embedding vector and the label embedding vector is calculated to obtain a similarity matrix; each element in the similarity matrix represents the semantic similarity between each sample and each label.

[0030] During the training iteration cycle, when the number of training rounds is less than the threshold, random negative sampling is performed; when the number of training rounds is greater than the threshold, hard sample negative sampling is performed. The second auxiliary loss function value of the initial prediction model is determined using the boundary ranking loss function.

[0031] A second aspect of this application provides a multi-label text classification device, comprising: an acquisition module, a first determination module, a second determination module, a third determination module, and a parameter adjustment module, wherein...

[0032] The acquisition module is configured to acquire multiple text samples and a tag sequence corresponding to the text samples; the tag sequence includes multiple tags related to the text samples, and the text samples include a character sequence consisting of at least one character;

[0033] The first determining module is configured to train an initial prediction model using the text samples and the label sequence, and determine the main loss function value of the initial prediction model using a binary cross-entropy loss function;

[0034] The second determining module is configured to construct a co-occurrence matrix through the label co-occurrence information of the label sequence, determine the positive point mutual PPMI correlation matrix, calculate the correlation difference between the positive point mutual PPMI correlation matrix and each label sequence using the mean squared error loss function, and determine the first auxiliary loss function value of the initial prediction model;

[0035] The third determining module is configured to determine a text label similarity matrix based on the text sample and the label sequence, and to determine the second auxiliary loss function value of the initial prediction model in combination with the boundary ranking loss function;

[0036] The parameter adjustment module is configured to adjust the internal parameters of the initial prediction model based on the main loss function value, the first auxiliary loss function value, and the second auxiliary loss function value, and perform iterative training until the training termination condition is met to obtain a multi-label text classification model.

[0037] A third aspect of this application provides an electronic device, including a processor and a memory; the memory stores a computer program, wherein the computer program, when executed by the processor, implements the multi-label text classification method described in the first aspect.

[0038] A fourth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in the first aspect.

[0039] Compared with the prior art, the beneficial effects of the technical solution provided in this application are:

[0040] This invention provides a multi-label text classification method, apparatus, and device. It acquires multiple text samples and their corresponding label sequences, trains an initial prediction model using the text samples and label sequences, determines the primary loss function value of the initial prediction model using a binary cross-entropy loss function, determines the first auxiliary loss function value of the initial prediction model by calculating the positive point cross-PPMI correlation matrix and the correlation difference between each label sequence, determines the second auxiliary loss function value of the initial prediction model based on the text label similarity matrix and the boundary ranking loss function, and adjusts the internal parameters of the initial prediction model based on the above loss function values ​​to obtain a multi-label text classification model. This invention effectively improves the overall accuracy of multi-label classification by capturing the semantic correlation between labels and comparing it with the results obtained from text feature training to optimize the model's prediction results. Attached Figure Description

[0041] Figure 1 A flowchart illustrating a multi-label text classification method provided in this application embodiment;

[0042] Figure 2 This is a schematic diagram of the structure of the network model to be trained provided in an embodiment of this application;

[0043] Figure 3 A schematic diagram of the structure of a multi-label text classification device provided in an embodiment of this application;

[0044] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0045] The embodiments of this application will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of this application. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of this application for ease of explanation. However, it will be apparent that one or more embodiments may be implemented without these specific details. Furthermore, descriptions of well-known structures and technologies are omitted in the following description to avoid unnecessarily obscuring the concepts of this application.

[0046] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0047] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0048] The accompanying drawings show some block diagrams and / or flowcharts. It should be understood that some blocks or combinations thereof in the block diagrams and / or flowcharts can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when executed by the processor, these instructions can create means for implementing the functions / operations described in these block diagrams and / or flowcharts.

[0049] In some embodiments, please refer to Figure 1 , Figure 1 A flowchart illustrating the multi-label text classification method provided in this application embodiment; the multi-label text classification method provided in this application embodiment includes:

[0050] S110, Obtain multiple text samples and the corresponding label sequences of the text samples; the label sequence includes multiple labels related to the text samples, and the text samples include character sequences consisting of at least one character.

[0051] Tagging text allows users to quickly understand its content, improves search efficiency, and enhances the accuracy of related text recommendations. Multi-tag text classification can be applied in various scenarios, such as text recommendation and text categorization. By identifying multiple tags associated with the text, more accurate text progression and categorization can be achieved. In this embodiment, a text sample can be represented as... The label sequence can be represented as .

[0052] S120 uses text samples and label sequences to train the initial prediction model, and determines the main loss function value of the initial prediction model through the binary cross-entropy loss function.

[0053] In some embodiments, an initial prediction model is trained using text samples and label sequences, and the main loss function value of the initial prediction model is determined using a binary cross-entropy loss function, including:

[0054] The text sample is input into the initial prediction model to obtain the predicted label;

[0055] The difference between the predicted label and the label sequence is calculated using the binary cross-entropy loss function to determine the main loss function value of the initial prediction model.

[0056] In some embodiments, the difference between the predicted label and the label sequence is calculated using the binary cross-entropy loss function to determine the main loss function value of the initial prediction model. The calculation formula is as follows:

[0057] ;

[0058] in, The true label represents the probability of a positive sample and a negative sample. For predicted values, , representing the probability of a positive sample. Indicates batch size, This represents the average loss for a batch of samples.

[0059] S130, construct a co-occurrence matrix using the label co-occurrence information of the label sequences, determine the positive point mutual PPMI correlation matrix, use the mean squared error loss function to calculate the correlation difference between the positive point mutual PPMI correlation matrix and each label sequence, and determine the first auxiliary loss function value of the initial prediction model.

[0060] Label co-occurrence information reflects the frequency of label occurrences. A label co-occurrence matrix is ​​a two-dimensional matrix used to represent the co-occurrence of different labels in a multi-label dataset. Point mutual information (PPMI) is a variant of point mutual information (PMI). This represents the probability that labels i and j appear simultaneously. The PPMI represents the joint probability of label i and label j as independent events. The PPMI uses the max function to avoid negative values ​​in the PMI. When the PMI is negative, it means that the probability of label i and label j occurring simultaneously is lower than the probability of them occurring independently, so it is set to 0.

[0061] In some embodiments, a co-occurrence matrix is ​​constructed using tag co-occurrence information of tag sequences to determine the positive point mutual PPMI association matrix, including:

[0062] Iterate through the label sequence of each text sample, count the frequency of each label sequence, and determine the co-occurrence matrix corresponding to the label sequences;

[0063] Calculate the positive point cross-PPMI correlation matrix between label sequences based on the co-occurrence matrix.

[0064] In some embodiments, calculating the positive point cross-PPMI correlation matrix between tag sequences based on the co-occurrence matrix includes determining the positive point cross-PPMI correlation matrix using the following formula:

[0065] ;

[0066] ;

[0067] in, This represents the probability that labels i and j appear simultaneously. This represents the joint probability of label i and label j under independent events.

[0068] In some embodiments, the mean squared error loss function is used to calculate the correlation difference between the positive point cross-PPMI correlation matrix and each label sequence to determine the first auxiliary loss function value of the initial prediction model, including determining the first auxiliary loss function value using the following formula:

[0069] ;

[0070] in, The true label represents the probability of a positive sample and a negative sample. For predicted values, , representing the probability of a positive sample. Indicates batch size, This represents the mean squared error loss function value.

[0071] In this embodiment, please refer to Figure 2 , Figure 2 This is a schematic diagram of the network model to be trained provided in this application embodiment. A fully connected layer can be defined, with an input dimension equal to the hidden layer size of a BiGRU and an output dimension equal to the total number of label categories in the multi-label text classification task. The output features from the previous step are used as input, and the final output is transformed using a non-linear function, the sigmoid function, to obtain the predicted label results. The label results are normalized by standardizing the vector of each sample to a unit vector to prevent scale differences between different samples from affecting the correlation calculation. The correlation between each label in the model output is calculated to obtain the correlation matrix. The mean squared error loss function is used to calculate the difference between the positive point cross-PPMI correlation matrix and the label co-occurrence matrix. This loss measures the difference between the label correlation of the model's predicted results and the true label PPMI matrix. The smaller the loss, the more consistent the correlation between the model's output labels is with the PPMI correlation of the true labels, and the better the model can capture the co-occurrence patterns of the true labels.

[0072] In one optional embodiment, an attention mechanism is used to process the output features to obtain the total number of corresponding label categories; the principle of the attention mechanism is as follows:

[0073] ;

[0074] ;

[0075] ;

[0076] ;

[0077] ;

[0078] ;

[0079] in, This indicates the number of heads with multi-head self-attention. Represents the dimension processed by each head. This represents the final result after all heads have been calculated and spliced ​​together.

[0080] S140, based on text samples and label sequences, determines the text label similarity matrix, and combines the boundary ranking loss function to determine the value of the second auxiliary loss function of the initial prediction model.

[0081] In some embodiments, a text label similarity matrix is ​​determined based on text samples and label sequences, and a second auxiliary loss function value for the initial prediction model is determined by combining the boundary ranking loss function, including:

[0082] ALBERT is used to encode text samples and label sequences, mapping them into low-dimensional semantic vectors to obtain text embedding vectors and label embedding vectors.

[0083] The cosine similarity of the text embedding vector and the label embedding vector is calculated to obtain the similarity matrix; each element in the similarity matrix represents the semantic similarity between each sample and each label.

[0084] During the training iteration cycle, when the number of training rounds is less than the threshold, random negative sampling is performed; when the number of training rounds is greater than the threshold, hard sample negative sampling is performed. The boundary ranking loss function is used to determine the value of the second auxiliary loss function of the initial prediction model.

[0085] Here, ALBERT can be used to encode text samples and label sequences, mapping them into low-dimensional semantic vectors to obtain text embedding vectors. Label embedding vector Based on this, the label sequence is transformed into a one-hot encoding format using MultiLabelBinarizer, where each label is represented by an independent binary feature, and the total number of label categories determines the dimension of the embedding vector.

[0086] S150: The internal parameters of the initial prediction model are adjusted based on the main loss function value, the first auxiliary loss function value, and the second auxiliary loss function value. Iterative training is performed until the training termination condition is met, resulting in a multi-label text classification model.

[0087] This application embodiment acquires multiple text samples and their corresponding label sequences. An initial prediction model is trained using these text samples and label sequences. The primary loss function value of the initial prediction model is determined using a binary cross-entropy loss function. A first auxiliary loss function value is determined by calculating the positive point cross-PPMI correlation matrix and the correlation difference between each label sequence. A second auxiliary loss function value is determined based on the text label similarity matrix and the boundary ranking loss function. The internal parameters of the initial prediction model are adjusted based on these loss function values ​​to obtain a multi-label text classification model. This invention effectively improves the overall accuracy of multi-label classification by capturing the semantic correlation between labels and comparing it with the results obtained from text feature training to optimize the model's prediction results.

[0088] In some embodiments, please refer to Figure 3 , Figure 3 This application provides a schematic diagram of the structure of a multi-label text classification device according to an embodiment of the present application. The embodiment of the present application provides a multi-label text classification device 300, including: an acquisition module 310, a first determination module 320, a second determination module 330, a third determination module 340, and a parameter adjustment module 350.

[0089] The acquisition module 310 is configured to acquire multiple text samples and the corresponding label sequences of the text samples; the label sequence includes multiple labels related to the text samples, and the text samples include a character sequence consisting of at least one character;

[0090] The first determining module 320 is configured to train the initial prediction model using text samples and label sequences, and determine the main loss function value of the initial prediction model through the binary cross-entropy loss function;

[0091] The second determining module 330 is configured to construct a co-occurrence matrix through the label co-occurrence information of the label sequence, determine the positive point mutual PPMI correlation matrix, use the mean squared error loss function to calculate the correlation difference between the positive point mutual PPMI correlation matrix and each label sequence, and determine the first auxiliary loss function value of the initial prediction model.

[0092] The third determining module 340 is configured to determine the text label similarity matrix based on text samples and label sequences, and combine the boundary ranking loss function to determine the value of the second auxiliary loss function of the initial prediction model;

[0093] The parameter adjustment module 350 is configured to adjust the internal parameters of the initial prediction model based on the main loss function value, the first auxiliary loss function value, and the second auxiliary loss function value, and to perform iterative training until the training termination condition is met, thereby obtaining a multi-label text classification model.

[0094] In some embodiments, the first determining module 320 is specifically configured as follows:

[0095] The text sample is input into the initial prediction model to obtain the predicted label;

[0096] The difference between the predicted label and the label sequence is calculated using the binary cross-entropy loss function to determine the main loss function value of the initial prediction model.

[0097] In some embodiments, the calculation formula for the specific configuration of the first determining module 320 is as follows:

[0098] ;

[0099] in, The true label represents the probability of a positive sample and a negative sample. For predicted values, , representing the probability of a positive sample. Indicates batch size, This represents the average loss for a batch of samples.

[0100] In some embodiments, the second determining module 330 is specifically configured as follows:

[0101] Iterate through the label sequence of each text sample, count the frequency of each label sequence, and determine the co-occurrence matrix corresponding to the label sequences;

[0102] Calculate the positive point cross-PPMI correlation matrix between label sequences based on the co-occurrence matrix.

[0103] In some embodiments, the second determining module 330 is specifically configured to determine the Positive Point Mutual PPMI correlation matrix using the following formula:

[0104] ;

[0105] ;

[0106] in, This represents the probability that labels i and j appear simultaneously. This represents the joint probability of label i and label j under independent events.

[0107] In some embodiments, the second determining module 330 is specifically configured to determine the value of the first auxiliary loss function using the following formula:

[0108] ;

[0109] in, The true label represents the probability of a positive sample and a negative sample. For predicted values, , representing the probability of a positive sample. Indicates batch size, This represents the average loss for a batch of samples.

[0110] In some embodiments, the third determining module 340 is specifically configured as follows:

[0111] ALBERT is used to encode text samples and label sequences, mapping them into low-dimensional semantic vectors to obtain text embedding vectors and label embedding vectors.

[0112] The cosine similarity of the text embedding vector and the label embedding vector is calculated to obtain the similarity matrix; each element in the similarity matrix represents the semantic similarity between each sample and each label.

[0113] During the training iteration cycle, when the number of training rounds is less than the threshold, random negative sampling is performed; when the number of training rounds is greater than the threshold, hard sample negative sampling is performed. The boundary ranking loss function is used to determine the value of the second auxiliary loss function of the initial prediction model.

[0114] The multi-label text classification device provided in this application embodiment can realize the various processes in the embodiments corresponding to the above-mentioned multi-label text classification method. To avoid repetition, it will not be described again here.

[0115] It should be noted that the multi-label text classification device and the multi-label text classification method provided in this application are based on the same application concept. Therefore, the specific implementation of this embodiment can refer to the implementation of the aforementioned multi-label text classification method, and the repeated parts will not be described again.

[0116] In some embodiments, please refer to Figure 4 , Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device 400 provided in this application includes a processor 410 and a memory 420; the memory 420 stores a computer program, wherein the computer program, when executed by the processor, implements the aforementioned multi-label text classification method.

[0117] Specifically, processor 410 may include, for example, a general-purpose microprocessor, an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. Processor 410 may also include onboard memory for caching purposes. Processor 410 may be a single processing unit or multiple processing units for performing different actions of the method flow according to embodiments of this application.

[0118] Memory 420 may be any medium capable of containing, storing, transmitting, propagating, or transmitting instructions. For example, memory 420 may include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, instruments, or propagation media. Specific examples of memory 420 include: magnetic storage devices such as magnetic tape or hard disk drives (HDDs); optical storage devices such as optical discs (CD-ROMs); and may also be random access memory (RAM) or flash memory; and / or wired / wireless communication links.

[0119] This application also provides a computer-readable medium storing a computer program thereon, which, when executed by a processor, implements the multi-label text classification method described above. This computer-readable medium may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into that device / apparatus / system. The aforementioned computer-readable medium carries one or more programs, which, when executed, implement the method according to the embodiments of this application.

[0120] According to embodiments of this application, a computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wired, optical fiber, radio frequency signals, etc., or any suitable combination thereof.

[0121] Those skilled in the art will understand that the features described in the various embodiments and / or claims of this application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this application. In particular, the features described in the various embodiments and / or claims of this application can be combined and / or combined in various ways without departing from the spirit and teachings of this application. All such combinations and / or combinations fall within the scope of this application. Therefore, the scope of this application should not be limited to the above embodiments, but should be defined not only by the appended claims, but also by the equivalents of the appended claims.

Claims

1. A multi-label text classification method, characterized in that, include: Obtain multiple text samples and corresponding tag sequences for the text samples; the tag sequence includes multiple tags related to the text samples, and the text samples include character sequences consisting of at least one character; The initial prediction model is trained using the text samples and the label sequence, and the main loss function value of the initial prediction model is determined by the binary cross-entropy loss function. A co-occurrence matrix is ​​constructed using the label co-occurrence information of the label sequences to determine the positive point mutual PPMI correlation matrix. The mean squared error loss function is used to calculate the correlation difference between the positive point mutual PPMI correlation matrix and each label sequence to determine the first auxiliary loss function value of the initial prediction model. Based on the text samples and the label sequence, a text label similarity matrix is ​​determined, and the value of the second auxiliary loss function of the initial prediction model is determined by combining the boundary ranking loss function; Based on the main loss function value, the first auxiliary loss function value, and the second auxiliary loss function value, the internal parameters of the initial prediction model are adjusted, and iterative training is performed until the training termination condition is met to obtain a multi-label text classification model.

2. The multi-label text classification method according to claim 1, characterized in that, The process of training an initial prediction model using the text samples and the label sequence, and determining the main loss function value of the initial prediction model using the binary cross-entropy loss function, includes: The text sample is input into the initial prediction model to obtain the predicted label; The difference between the predicted label and the label sequence is calculated using the binary cross-entropy loss function to determine the main loss function value of the initial prediction model.

3. The multi-label text classification method according to claim 2, characterized in that, The difference between the predicted label and the label sequence is calculated using the binary cross-entropy loss function to determine the principal loss function value of the initial prediction model. The calculation formula is as follows: ; in, The true label represents the probability of a positive sample and a negative sample. For predicted values, , representing the probability of a positive sample. Indicates batch size, This represents the average loss for a batch of samples.

4. The multi-label text classification method according to claim 1, characterized in that, The step of constructing a co-occurrence matrix using the tag co-occurrence information of the tag sequence to determine the positive point mutual PPMI correlation matrix includes: Iterate through the label sequence of each text sample, count the frequency of each label sequence, and determine the co-occurrence matrix corresponding to the label sequence; Calculate the positive point cross-PPMI correlation matrix between the tag sequences based on the co-occurrence matrix.

5. The multi-label text classification method according to claim 4, characterized in that, The step of calculating the positive point cross-PPMI correlation matrix between the tag sequences based on the co-occurrence matrix includes determining the positive point cross-PPMI correlation matrix using the following formula: ; ; in, This represents the probability that labels i and j appear simultaneously. This represents the joint probability of label i and label j under independent events.

6. The multi-label text classification method according to claim 1, characterized in that, The step of calculating the correlation difference between the positive point cross-PPMI correlation matrix and each of the label sequences using the mean squared error loss function, and determining the first auxiliary loss function value of the initial prediction model, includes determining the first auxiliary loss function value using the following formula: ; in, The true label represents the probability of a positive sample and a negative sample. For predicted values, , representing the probability of a positive sample. Indicates batch size, This represents the mean squared error loss function value.

7. The multi-label text classification method according to claim 1, characterized in that, The step of determining the text label similarity matrix based on the text samples and the label sequence, and determining the second auxiliary loss function value of the initial prediction model in conjunction with the boundary ranking loss function, includes: The text samples and the label sequences are encoded using ALBERT, and the text samples and the label sequences are mapped into low-dimensional semantic vectors to obtain text embedding vectors and label embedding vectors; The cosine similarity of the text embedding vector and the label embedding vector is calculated to obtain a similarity matrix; each element in the similarity matrix represents the semantic similarity between each sample and each label. During the training iteration cycle, when the number of training rounds is less than the threshold, random negative sampling is performed; when the number of training rounds is greater than the threshold, hard sample negative sampling is performed. The second auxiliary loss function value of the initial prediction model is determined using the boundary ranking loss function.

8. A multi-label text classification device, characterized in that, include: The module comprises an acquisition module, a first determination module, a second determination module, a third determination module, and a parameter adjustment module, wherein... The acquisition module is configured to acquire multiple text samples and a tag sequence corresponding to the text samples; the tag sequence includes multiple tags related to the text samples, and the text samples include a character sequence consisting of at least one character; The first determining module is configured to train an initial prediction model using the text samples and the label sequence, and determine the main loss function value of the initial prediction model using a binary cross-entropy loss function; The second determining module is configured to construct a co-occurrence matrix through the label co-occurrence information of the label sequence, determine the positive point mutual PPMI correlation matrix, calculate the correlation difference between the positive point mutual PPMI correlation matrix and each label sequence using the mean squared error loss function, and determine the first auxiliary loss function value of the initial prediction model; The third determining module is configured to determine a text label similarity matrix based on the text sample and the label sequence, and to determine the second auxiliary loss function value of the initial prediction model in combination with the boundary ranking loss function; The parameter adjustment module is configured to adjust the internal parameters of the initial prediction model based on the main loss function value, the first auxiliary loss function value, and the second auxiliary loss function value, and perform iterative training until the training termination condition is met to obtain a multi-label text classification model.

9. An electronic device comprising a processor and a memory; said memory having a storage for a computer program, wherein, When the computer program is executed by the processor, it implements the multi-label text classification method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps of the method of any one of claims 1 to 7.

Citation Information

Patent Citations

  • Multi-label text classification model and classification method based on improved GraphRNN

    CN113297385A

  • Semi-supervised detection method for self-adaptive routing inspection of overhead line based on unmanned aerial vehicle

    CN118506221A