Multi-label text interpretation method based on thinking chain prompt project

By integrating the thinking chain prompt engineering of pre-trained language models and large language models into multi-label text theme modeling, the problem of insufficient interpretation ability of multi-label text in the existing technology is solved, and efficient interpretation of multi-label text corpus and the accuracy of topic representation is achieved.

CN120124758AActive Publication Date: 2025-06-10NANJING UNIV OF POSTS & TELECOMM
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510595252.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-06-10
Estimated Expiration
2045-05-09

AI Technical Summary

Technical Problem

The existing multi-label text theme modeling methods cannot effectively incorporate document labels into the topic modeling process, resulting in insufficient interpretation ability in multi-label text corpus scenarios and inability to incorporate external semantic knowledge.

Method used

The multi-label text interpretation method based on thinking chain prompt engineering is adopted. By integrating pre-trained language models and large language models, the text fragments in the multi-label text corpus are aligned to semantic-related tags, and the pre-trained language model is used to obtain the topic vector representation matrix to ensure the accuracy of the topic vector representation.

Benefits of technology

It realizes effective interpretation of multi-label text corpus, improves sentence-level and word-level interpretation capabilities, ensures semantic consistency between labels and topics, and improves the quality and interpretability of topics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124758A_ABST
    Figure CN120124758A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of natural language processing, and discloses a multi-label text interpretation method based on thinking chain prompt engineering, which comprises the following steps: dividing documents in a multi-label text according to sentence-level fragments, constructing a sentence set corresponding to each label, averaging the expression of a pre-training model of the sentences in each set, and obtaining a multi-label text interpretation result; obtaining a vector representation of each label; according to the obtained topic vector, visualizing the word according to the cosine similarity between the vector representation of each word in the sentence and the topic vector representation; and the theme consistent with the label semantics is mined, so that fine-grained theme interpretation is provided. According to the method, the advantages of the pre-training language model and the large language model are utilized, and multi-granularity multi-label text interpretation is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of natural language processing, and specifically relates to a multi-label text interpretation method based on chain-of-thought prompting engineering. Background Art

[0002] The topic model is an effective text mining method in the field of natural language processing and is often used to mine potential semantic patterns in text corpora in an unsupervised manner. Although topic models have been widely explored, from traditional probabilistic topic models to neural topic models and existing methods based on pre-trained and large language models, unsupervised topic modeling methods cannot effectively incorporate document labels into the topic modeling process, resulting in their inability to effectively adapt to the scenario of multi-label text corpora. Existing topic modeling methods for multi-label texts only utilize word co-occurrence information, cannot incorporate external semantic knowledge such as pre-trained language models and large language models, and include irrelevant fragments in the text into topic inference, resulting in insufficient text interpretation ability.

[0003] Multi-label text interpretation aims to solve the credit attribution problem in multi-label texts, that is, different segments in the text correspond to different labels, different words in the text also correspond to different labels, and the text also contains some meaningless segments. Multi-label text interpretation needs to provide three levels of interpretation: (1) Map the segments in the text to their semantically related labels to provide sentence-level interpretation. (2) For the words contained in the sentences that have been assigned labels, map them to semantically related labels. (3) For each label in the text corpus, mine the topic with semantic consistency with it to provide label interpretation.

[0004] However, existing multi-label topic models that only rely on word co-occurrence information and cannot incorporate external semantic knowledge into the modeling method cannot accurately provide effective interpretations and cannot identify meaningless segments in the text, resulting in insufficient interpretation ability. Summary of the Invention

[0005] To solve the defects existing in the prior art, this application provides a multi-label text interpretation method based on chain-of-thought prompting engineering. By incorporating pre-trained language models and large language models, this application provides effective interpretations for multi-label text corpora.

[0006] To achieve the above object, this application is implemented through the following technical solutions:

[0007] This application is a multi-label text interpretation method based on chain-of-thought prompting engineering, including the following steps:

[0008] Step 1: Obtain the corpus on the multi-label text platform, clean the corpus, and obtain the multi-label text corpus set;

[0009] Step 2: For each label in the multi-label text corpus set, use the large language model to generate label description information, and use the chain of thought prompting engineering based on the large language model to align the text fragments in each document in the multi-label text corpus set to the labels semantically related to the text fragments, that is, assign labels to each text fragment. At the same time, use a label to align the text fragments that are not related to any label, and obtain the topic vector representation matrix based on the pre-trained language model ;

[0010] Step 3: According to the labels assigned to the text fragments by the large language model, visualize the words in the sentence by calculating the correlation between the word vectors of each word in the text fragment and the corresponding vectors of the labels assigned to the text fragment;

[0011] Step 4: Based on the labels corresponding to the text fragments feedback by the large language model and the topic vector representation matrix obtained in Step 2 , extract the topics that are semantically consistent with the labels in the multi-label text corpus set, and provide explanations for the labels. Different from the existing multi-label topic modeling methods, the method of this application uses the large language model for sentence-level topic alignment, and at the same time avoids irrelevant text fragments from participating in topic inference, ensuring the accuracy of the topic vector representation and facilitating subsequent word visualization and topic mining.

[0012] A further improvement of this application is that: in Step 1, obtaining the corpus on the multi-label text platform and cleaning the corpus specifically means: filtering the text in the corpus with less than three text fragments, and deleting the irrelevant content in the corpus text such as emojis and website links.

[0013] A further improvement of this application is that: Step 2 specifically includes the following steps:

[0014] Step 2.1: For each label in the multi-label text corpus set, use the large language model to generate a label description information, use the label description information to supplement the semantics of the label, and at the same time check the generated label description information to avoid the introduction of incorrect information affecting the model performance;

[0015] Step 2.2: The set of labels corresponding to each document is:

[0016] ,

[0017] wherein, represents the number of sentences in the document , Indicates the number of tags corresponding to the document, Indicates the document The corresponding tag set, Indicates the document Corresponding to the last tag in the tag set, Indicates the document The number of tags in the corresponding tag set.

[0018] For each tag corresponding to a document in the multi-tag text corpus set, first Add the tag to the multi-tag corpus set corresponding to the document to obtain the document The corresponding enhanced tag set ;

[0019] Step 2.3. For each text segment in the document , first let the large language model summarize the keywords in each text segment , extract the key information, combine the context corresponding to the keywords, purify the extracted keywords, and based on the purified keywords and the tag description information generated in Step 2.1, assign at least one semantically related tag to each text segment ;

[0020] Step 2.4. For all texts in the corpus, construct a sentence set of tags , use the pre-trained language model to obtain the word vectors of each word in the sentence set The th text segment , average the word vector representations of each word to obtain the vector representation of the text segment , and the above transformation is expressed as follows:

[0021]

[0022]

[0023] Among them, is the tag set; Indicates the number of words in the text segment , Indicates the context word vector of the th word in the text segment , is the sentence set in the number of sentences, Indicates the last word in the text segment ;

[0024] Based on the above transformation, the topic vector representation is obtained:

[0025]

[0026] Among them, represents a pre-trained language model, and a topic vector representation matrix is obtained based on the pre-trained language model ;

[0027]

[0028] Among them, represents the hidden dimension of the pre-trained language model, represents the number of labels in the label set corresponding to the multi-label corpus.

[0029] A further improvement of this application is that: Step 3 specifically includes the following steps:

[0030] Step 3.1, obtaining context word vector representation: Given the in the document th text fragment , the word sequence included in the text fragment is , and the label set assigned to the text fragment , and the context representation of each word, that is, the word vector, is obtained through the pre-trained language model :

[0031]

[0032] Among them, represents the th word in the word sequence , , represents the number of words included in the text fragment , is the word vector of represents the number of labels assigned by the large language model to the text fragment , represents the word in

[0033] Step 3.2, calculating the label-word correlation matrix: Based on the word vector calculated in Step 3.1, the word sequence and the semantic correlation between the th label in the label set ​ Calculated by the following formula:

[0034]

[0035] Wherein, represents the topic vector representation corresponding to the label ; is the topic vector representation of the label corresponding to .

[0036] Step 3.2: Based on the word vector and the vector representation of the label construct a label-word correlation matrix to capture the semantic correlation between words and labels, visualize the words in the sentence. The greater the correlation weight, the more it indicates that the current word can indicate the topic of the sentence, making the word-level interpretation meaningful.

[0037] A further improvement of this application is that the specific steps of step 4 are as follows:

[0038] Step 4.1: For the -th word in the vocabulary containing different words in the multi-label text corpus , , represents the number of different words in the multi-label corpus. Take each occurrence of the word in the multi-label corpus , and then for the -th occurrence of the word , denoted as , , obtain the context representation of the word through the pre-trained language model:

[0039]

[0040] Wherein, represents the context of the -th occurrence, represents the number of occurrences of

[0041] The local word-topic distribution of the word is calculated as follows:

[0042]

[0043] Wherein, represents the cosine similarity, is the normalization operation Indicates the topic vector representation corresponding to the th label, , the word corresponding corpus-level word-topic distribution is obtained through the following formula:

[0044]

[0045] Calculate the corpus-level word-topic distribution for each word in the vocabulary, and then obtain the global word-topic distribution matrix by taking the average ;

[0046] Step 4.2, based on the obtained global word-topic distribution matrix , through column normalization operation, obtain the final topic-word distribution matrix , specifically:

[0047]

[0048] Among them, represents the column normalization operation. For the topic-word distribution matrix corresponding to each label, obtain the top 10 words with the highest probability in the topic-word distribution matrix to represent the label and provide a fine-grained interpretation of the label.

[0049] The beneficial effect of this application is that by integrating the pre-trained language model and the large language model, this application provides an effective interpretation for the multi-label text corpus.

[0050] This application uses the thought chain prompting engineering based on the large language model to provide an effective sentence-level interpretation for multi-label text, which helps to quickly locate the text fragments related to the label.

[0051] This application uses the label-related word visualization mechanism based on the pre-trained language model to effectively extract the correlation between words and labels in the text fragment and provide a fine-grained interpretation of the words.

[0052] This application integrates the knowledge of the pre-trained language model in the topic extraction process to improve the quality and interpretability of the extracted topics. At the same time, the topic alignment module ensures the semantic consistency between the label and the topic and provides an interpretation of the label at the corpus level. Brief Description of the Drawings

[0053] Figure 1 is a schematic flowchart of the multi-label text interpretation method of this application.

[0054] Figure 2 is a schematic diagram of the main body alignment and representation of this application.

[0055] Figure 3It is a schematic diagram of the visualization of the related words of the basic application label.

[0056] Figure 4 It is a schematic diagram of the extraction of the related topics of the labels of this application. Detailed implementation manners

[0057] The following will disclose the implementation manners of the present invention in diagrams. For the sake of clear illustration, many practical details will be described together in the following description. However, it should be understood that these practical details are not used to limit the present invention. That is to say, in some implementation manners of this application, these practical details are not necessary.

[0058] To provide an effective tool for multi-label text interpretation, the present invention proposes a multi-label topic modeling method (CPTM) based on integrating multi-source semantic knowledge, which is based on the current research hotspot large language model and the effective text representation tool pre-trained model. This method uses the chain of thought prompting engineering based on the large model, integrates keyword information, context, and description information of the label, aligns text fragments to their semantically related labels to provide sentence-level explanations, and ensures the accuracy of subsequent word visualization and the semantic consistency between the mined topics and the labels. For word visualization and topic mining, this application integrates the context knowledge of the pre-trained language model to ensure the effect of word explanation and label explanation.

[0059] As Figure 1 shown, this application is a multi-label text interpretation method based on the chain of thought prompting engineering. The multi-label text interpretation method includes the following steps:

[0060] Step 1: Use web crawler technology to obtain the corpus on the multi-label text platform, clean the corpus, filter the texts with less than three text fragments, and delete the irrelevant content in the texts such as emojis, website links, etc., to obtain a multi-label text corpus set;

[0061] Step 2: For each label in the multi-label text corpus set, use the large language model to generate label description information, use the chain of thought prompting engineering based on the large language model to align the text fragments in each document in the multi-label text corpus set to the labels semantically related to the text fragments, that is, assign labels to each text fragment. At the same time, use a label to align the text fragments that are not related to any label, and obtain the topic vector representation matrix , as shown in the appendix Figure 2 shown, which specifically includes the following steps:

[0062] Step 2.1: For each label in the label set corresponding to the multi-label corpus, such as labels like'mysterious', 'romantic', 'war','sports', etc., first use the large language model to generate corresponding description information for each label. This step avoids the problem of insufficient semantic information provided solely by the labels in the theme alignment stage and ensures the accuracy of the subsequent theme representation matrix.

[0063] Step 2.2: Subsequently, for each text in the multi-label text corpus , first add the label ' ' to the corresponding label set of the text to obtain .

[0064] Step 2.3: Subsequently, use chain-of-thought prompting engineering to align the text fragments to the labels in . As shown in Figure 2 , for the document , the label set corresponding to its three text fragments is {mysterious, romantic}, and the final result of theme alignment is: fragment 1 corresponds to 'romantic' and'mysterious', fragment 2 corresponds to 'romantic', and fragment 3 is not related to any label, being ' '.

[0065] Step 2.4: Based on the above steps, for each label in the label set corresponding to the multi-label text corpus, this application can obtain the set of text fragments corresponding to each label. Subsequently, based on the pre-trained model, by taking the average of the representations of the pre-trained language model of the text fragments corresponding to each label, the vector representation of each label is obtained, and the theme representation matrix is obtained for subsequent visualization of label-related words and extraction of label-related themes.

[0066] The theme vector representation matrix is used for subsequent visualization of label-related words and mining of label-related themes. Existing methods usually assume that text fragments are a mixture of document labels, which will cause each word to be relevant to the label, resulting in inaccurate word visualization. However, this application ensures the accurate relevance of words and labels by assigning relevant labels to sentences.

[0067] Step 3: According to the labels assigned to the text fragments by the large language model, by calculating the correlation between the word vectors of each word in the text fragment and the corresponding vectors of the labels assigned to the text fragment, the words in the sentence are visualized; the correlation is calculated by the cosine similarity between the vector representation of the word and the vector representation of the label. As shown in the example in Figure 3 , step 3 specifically includes the following steps:

[0068] Step 3.1: The attached Figure 3 example is the text The visualization result of the first text fragment in , which is assigned two labels, 'romantic' and'mysterious', by the large language model through chain of thought prompting engineering. First, this application obtains the vector representation of each word in the fragment through a pre-trained language model.

[0069] Step 3.2: Subsequently, calculate the cosine similarity between the vector representation of each word and the vector representations of the assigned labels 'romantic' and'mysterious' to obtain a label-word correlation matrix. .

[0070] Step 3.3: Visualize the words according to the correlation scores between the labels and each word in the label-word correlation matrix. In the existing multi-label topic models, only word co-occurrence information is relied on during the process of mining topics, and external semantic knowledge is not incorporated, resulting in lower quality and interpretability of the mined topics. At the same time, the semantic consistency between the labels and the topics is also low. To address the above defects, this application incorporates the knowledge of the pre-trained language model for topic mining.

[0071] Existing multi-label topic models rely only on word co-occurrence information during the process of mining topics and do not incorporate external semantic knowledge, resulting in lower quality and interpretability of the mined topics. At the same time, the semantic consistency between the labels and the topics is also low. To address the above defects, this application incorporates the knowledge of the pre-trained language model for topic mining.

[0072] Step 4: Based on the labels corresponding to the text fragments feedback by the large language model and the topic vector representation matrix obtained in Step 2, extract the topics that are semantically consistent with the labels in the multi-label text corpus set to provide interpretations of the labels. Different from the existing multi-label topic modeling methods, the method of this application uses the large language model for sentence-level topic alignment, and at the same time avoids irrelevant text fragments from participating in topic inference, ensuring the accuracy of the topic vector representation, facilitating subsequent word visualization and topic mining, as shown in . Specifically, it includes the following steps: Figure 4 shown below:

[0073] Step 4.1: Word-topic distribution: To mine the topics that are semantically consistent with each label in the label set and provide fine-grained interpretations of the labels, this application designs a topic extraction mechanism based on the pre-trained language model . Specifically, for the th word in the vocabulary containing different words in the multi-label text corpus set, , , , represents the number of different words in the multi-label corpus. Obtain each occurrence of the word in the multi-label corpus, and then for the th occurrence of the word , denoted as , , , obtain the word Context representation :[[]]

[0074]

[0075] Among them, represents the th occurrence of the context, represents the number of occurrences in the multi-label corpus;

[0076] word 's local word-topic distribution The calculation method is as follows:

[0077]

[0078] Among them, represents the cosine similarity, is a normalization operation, represents the th topic vector representation corresponding to the label, , word 's corpus-level word-topic distribution is obtained through the following formula:

[0079]

[0080] Calculate the corpus-level word-topic distribution of each word in the vocabulary, and then obtain the global word-topic distribution matrix by taking the average ;

[0081] Step 4.2, Based on the obtained global word-topic distribution matrix , through column normalization operation, obtain the final topic-word distribution matrix , specifically:

[0082]

[0083] Among them, represents the column normalization operation. For each label's corresponding topic-word distribution matrix, obtain the top 10 words with the highest probability in the topic-word distribution matrix to represent the label and provide a fine-grained interpretation of the label. As shown in the appendix Figure 4 , for each label in the label set, such as'sports','mystery', 'romance', 'war', etc., extract the topics related to each label.

[0084] This application conducts experiments on the plot text data of IMDB movies for the effectiveness of multi-label text interpretation. For the interpretation of labels, the commonly used metrics for measuring topic coherence in topic models (CP, CA, NPMI, UCI) and the topic diversity metric (UTR) are used. At the same time, two metrics are designed to measure the semantic consistency between the extracted topics and labels, namely Label Discovered Rate (LDR) and Correlation Score (CS). For sentence-level interpretation, this application uses a manually annotated dataset and measures it using the MicroF1 and MacroF1 classification metrics. As a comparison, this application has unsupervised topic models for Comparative Example 1 and Comparative Example 2, and for multi-label topic models, we adopt Comparative Example 3 and Comparative Example 4. For word-level interpretation, through a manual evaluation method, the top 5 words most relevant to the label obtained by the model in each sentence are calculated, and then the 5 words corresponding to each scrambled sentence are manually reordered according to their relevance to the label, and then the difference between the two reordered sequences is calculated. Here, two evaluation metrics for measuring the difference between sequences are used, the Kendall correlation coefficient and the Spearman coefficient . Table 1 is a comparative statistical table of the numerical results of label interpretation, i.e., topic quality. Table 2 is a comparative statistical table of the numerical results of the semantic consistency between labels and topics. Table 3 is a comparative statistical table of the numerical results of sentence-level interpretation accuracy. Table 4 is a statistical table of the manual evaluation results of word-level interpretation.

[0085] Table 1

[0086]

[0087] Table 2

[0088]

[0089] Table 3

[0090]

[0091] Table 4

[0092]

[0093] The experimental results show that for label-level interpretation, the average results of the topic quality measurement metrics are shown in Table 1. The LDR and CS results are shown in Table 2. For sentence-level interpretation, the MicroF1 and MacroF1 are shown in Table 3. For word-level interpretation, the Kendall correlation coefficient and the Spearman coefficient are shown in Table 4. It can be seen that the evaluation results of this application are higher than those of the comparative methods.

[0094] In this experiment, the comparative models include unsupervised topic model Comparative Example 1, Comparative Example 2, and multi-label topic model Comparative Example 3 and Comparative Example 4.

[0095] The comparison results of this application include evaluations of labels, text fragments, and word visualization. First is the label-level explanation. As shown in Table 1, the comparison between Comparative Examples 1-4 and this application shows that the method of this application outperforms Comparative Examples 1-4 in terms of topic quality indicators (CP, CA, NPMI, UCI) and topic diversity indicator (UTR). At the same time, as shown in Table 2, the performance in terms of label-topic semantic consistency indicator is also better than that of Comparative Examples 1-4. For the explanation of text fragments, this application is compared with Comparative Examples 3 and 4, and the comparison results are shown in Table 3. It can be seen that the method of this application is better than Comparative Examples 3 and 4. Finally is the word-level explanation. This application uses the method of manual evaluation, and the results are shown in Table 4. It can be seen that the method of this application can provide effective word-level explanations. In summary, this application is superior to Comparative Examples 1-4 in all evaluation indicators.

[0096] The above are only the preferred embodiments of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A multi-label text interpretation method based on thought chain prompt engineering, characterized by: The multi-label text interpretation method comprises the following steps: Step 1: Obtain the corpus on the multi-label text platform, clean the corpus, and obtain a multi-label text corpus set; Step 2: For each label in the multi-label text corpus, use the large language model to generate label description information, and use the thought chain prompting project based on the large language model to align the text fragments in each document in the multi-label text corpus to the labels that are semantically related to the text fragments, that is, assign a label to each text fragment, and use a text fragment that is not related to any label to align the text fragments in each document in the multi-label text corpus to the labels that are semantically related to the text fragments. The labels are aligned and the topic vector representation matrix is ​​obtained based on the pre-trained language model. ; Step 3: Based on the labels assigned to the text fragments by the large language model, the words in the sentence are visualized by calculating the correlation between the word vector of each word in the text fragment and the vector corresponding to the label assigned to the text fragment; Step 4: Obtain the topic vector representation matrix based on the labels of the text segments fed back by the large language model and step 2 , extract topics that are semantically consistent with the tags in the multi-label text corpus and provide explanations for the tags.

2. According to claim 1, a multi-label text interpretation method based on thought chain prompt engineering is characterized by: The step 1 obtains the corpus on the multi-label text platform, and cleans the corpus by filtering texts with less than three sentences in the corpus, and deleting emoticons and website links in the corpus text.

3. The multi-label text interpretation method based on thought chain prompting engineering according to claim 1 is characterized by: The step 2 specifically includes the following steps: Step 2.1, for each label in the multi-label text corpus set, a label description information is generated using a large language model, the label description information is used to supplement the semantics of the label, and the generated label description information is checked; Step 2.2: Each document The corresponding tag set is: , in, Representation Document The number of sentences in Indicates that the document corresponds The number of labels, Representation Document The corresponding tag set, Representation Document Corresponding to the last tag in the tag set, Representation Document The number of tags in the corresponding tag set; For the label corresponding to each document in the multi-label text corpus, first The tag is added to the multi-tag corpus corresponding to the document to obtain the document Corresponding enhanced tag set ; Step 2.3: Documentation Each text segment in , first let the large language model summarize each text segment The keywords in the text are combined with the contextual context of the keywords to purify the extracted keywords. Based on the purified keywords and the label description information generated in step 2.1, a tag is generated for each text segment. Assign at least one semantically related tag; Step 2.4: For all the texts in the corpus, construct labels A collection of sentences , use the pre-trained language model to obtain a sentence set No. Text snippet The word vector of each word in , average the word vector representation of each word, and obtain the text fragment Vector representation of , the above transformation is expressed as follows: ; ; in, is a tag collection, Represents a text fragment The number of words in Represents a text fragment The last word in Represents a text fragment The The context word vector of the word; Based on the above transformation, we get the theme Vector representation of : ; in, represents the pre-trained language model, For a set of sentences The number of sentences in the topic vector representation matrix is ​​obtained based on the pre-trained language model ; ; in, represents the hidden dimension of the pre-trained language model, Indicates the number of tags in the tag set corresponding to the multi-label corpus.

4. The multi-label text interpretation method based on thought chain prompting engineering according to claim 3 is characterized by: Step 3 specifically includes the following steps: Step 3.1, context word vector representation acquisition: given a document The Text snippet , text snippet The word sequence contained is , and text snippets The set of assigned tags , through the pre-trained language model Get the context representation of each word, i.e. the word vector: ; in, Representing word sequence The words, , Represents a text fragment The number of words contained in for The word vector of Represents the assignment of a large language model to a text fragment The number of labels, Representing word sequence The words in; Step 3.2: Calculate the label-word correlation matrix: Based on the word vector calculated in step 3.1 , word sequence The Words With label collection Middle Tags Semantic relevance Calculated by the following formula: ; in, Representation and Labeling The corresponding topic vector representation; For The corresponding label's topic vector representation; Step 3.3: Based on word vector Vector representation with labels Constructing a label-word correlation matrix Capture semantic correlations between words and labels and visualize words in sentences.

5. The multi-label text interpretation method based on thought chain prompting engineering according to claim 1 is characterized by: The step 4 specifically includes the following steps: Step 4.1: For the first word in the vocabulary containing different words in the multi-label text corpus Words , , Indicates the number of different words in the multi-label corpus, and obtains the word Every occurrence in the multi-label corpus ,word No. Occurrences are represented by , , get words through pre-trained language model Contextual representation of : ; in, Indicates The context of the occurrence is represented by Expressive words The number of occurrences in the multi-label corpus; word Local word-topic distribution of The calculation method is: ; in, represents the cosine similarity, for Normalization operation, Indicates The topic vector corresponding to the label is represented by ,word Corresponding corpus-level word-topic distribution It is obtained by the following formula: ; Calculate the corpus-level word-topic distribution of each word in the vocabulary, and then get the global word-topic distribution matrix by taking the average ; Step 4.2: Based on the obtained global word-topic distribution matrix , through column normalization operation, we get the final topic-word distribution matrix , specifically: ; in, Represents a column normalization operation. For each label’s corresponding topic-word distribution matrix, obtain the top 10 words with the highest probability in the topic-word distribution matrix to represent the label, providing a fine-grained explanation of the label.

Citation Information

Patent Citations

  • Determination method and device for similarities of text semantics

    CN106776503A

  • Multi-label classification method based on topic attention mechanism for biomedical text

    CN112732872A

  • Multi-granularity text recommendation method based on context semantics

    CN112784013A

  • Material labeling method, model training method, multimedia content labeling method and related products

    CN119441516A

  • Pitman-Yor process topic modeling pre-seeded by keyword groupings

    US11429901B1