Online forum-oriented low-resource topic key topic extraction method

Through the LLM-enhanced low-resource theme modeling framework, the data sparsity, noise sensitivity and model complexity faced by theme modeling technology in online forums are solved, and higher quality theme discovery and computing efficiency are achieved.

CN120146046AActive Publication Date: 2025-06-13NANJING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
CN202510615488.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-06-13
Estimated Expiration
2045-05-14

AI Technical Summary

Technical Problem

Traditional theme modeling technology faces problems of data sparsity, noise sensitivity and model complexity in low-resource online forum environments, resulting in insufficient theme consistency and diversity and low computing efficiency.

Method used

The low-resource topic modeling framework enhanced by LLM is adopted to enhance semantically through large language models, generate enhanced document collections, extract document-level representations using pre-trained language models, build learnable topic embedding matrices, and design semantic-aware contrast learning frameworks, optimize topic embedding matrices to ensure topic consistency.

Benefits of technology

Improves theme consistency and diversity, improves computing efficiency, enables higher quality topics to be discovered from low-resource online forums, and is suitable for semantic analysis and public opinion insights of user-generated content on social media platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146046A_ABST
    Figure CN120146046A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of natural language processing and text mining, and discloses an online forum-oriented low-resource topic key topic extraction method, which comprises the following steps of: performing semantic-preserving data enhancement on an original text through a large language model to generate an enhanced document set; utilizing a pre-training language model to extract context-aware semantic representation of the document; constructing a learnable topic embedding matrix, and calculating and generating topic distribution; designing a semantic perception contrast learning framework, and optimizing theme diversity by adopting a dynamic negative sample screening strategy; and meanwhile, priori alignment loss is used for ensuring theme consistency. According to the invention, an LLM enhanced data expansion mechanism and a lightweight theme coding architecture are creatively fused, and through dual optimization of contrast learning regularization and prior distribution matching, three technical problems of data sparsity, model over-fitting and noise sensitivity in a low-resource scene are effectively solved; and an efficient and reliable theme modeling solution is provided for social media public opinion analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of natural language processing and text mining, and specifically relates to a method for extracting key topics of low-resource topics for online forums. Background Art

[0002] With the popularization of the Internet, online forums have become a popular platform for users to share their daily lives and exchange opinions. These user-generated texts provide valuable resources for analyzing the public's opinions on various social phenomena. However, most sub-forums on the forum show low-resource attributes, with limited active members and posts. Although the size of the corpus is small, the articles submitted by users contain important aspects, such as "mental health and family issues", such as "loneliness and romantic relationships", such as "health and medical struggles", all of which can lead to feelings of unhappiness and have potential analytical value.

[0003] Traditional topic modeling techniques face three core challenges: a) Data sparsity: The number of documents in low-activity communities is often less than 3000, resulting in insufficient neural network training. b) Noise sensitivity: There are spelling mistakes and grammar irregularities in user-generated texts, which affect the performance of the bag-of-words model. c) Model complexity: The neural topic model with the VAE architecture has too many parameters and is prone to overfitting on small datasets.

[0004] Existing solutions have significant defects. For example, traditional Bayesian methods, such as LDA and neural variants based on BOW, such as ECRTM, cannot solve the challenges of data scarcity and noisy information; similarly, recently developed context-aware topic models based on VAE, such as the context-aware topic model CTMNeg with negative sampling and the context-aware word-topic model CWTM, are less capable of dealing with the challenges of data scarcity and model complexity. Summary of the Invention

[0005] To solve the above technical defects, this application provides a method for extracting key topics of low-resource topics for online forums. This method uses an LLM-enhanced low-resource topic modeling framework to solve the three major technical problems of data sparsity, noise interference, and model complexity, achieving improved topic consistency, excellent topic diversity, and improved computational efficiency, and discovering higher-quality topics from low-resource online forums, which is applicable to semantic analysis and public opinion insight of user-generated content (UGC) in social media platforms.

[0006] To achieve the above object, this application is implemented through the following technical solutions:

[0007] This application is a method for extracting key topics of low-resource topics for online forums. The method for extracting key topics of low-resource topics specifically includes the following steps:

[0008] Step 1: Obtain low-resource documents from an online forum, and perform semantic-preserving data augmentation on the obtained low-resource documents through a large language model to generate an augmented document set;

[0009] Step 2: Use a pre-trained language model to extract document-level representations from the augmented document set;

[0010] Step 3: Construct a learnable topic embedding matrix, and calculate the document-topic distribution through document-topic similarity;

[0011] Step 4: Design a semantic-aware contrastive learning framework, screen dynamic negative samples in the augmented documents of the same batch in the contrastive learning framework, calculate the contrastive learning loss, optimize the topic embedding matrix, and ensure topic consistency with a prior alignment loss to obtain topic words;

[0012] Step 5: Use a large language model to expand and generalize the topic insights obtained in Step 4 to help better understand the low-resource corpus.

[0013] A further improvement of this application is that in Step 1, obtaining low-resource documents from an online forum and performing semantic-preserving data augmentation on the obtained low-resource documents through a large language model to generate an augmented document set specifically includes the following steps:

[0014] Step 1.1: Construct a document augmentation prompt template based on a large language model to generate large language model results, where the document augmentation prompt template contains semantic-preserving constraints:

[0015] a) The principle of minimum semantic variation: requires that each text in the generated augmented document set be as semantically similar as possible to the low-resource document;

[0016] b) Optimization of sentence fluency: Eliminate spelling mistakes and ungrammatical expressions;

[0017] Step 1.2: Iterative generation: Use a pre-trained language model to calculate the embedding similarity between the low-resource document and the large language model generation result corresponding to the low-resource document;

[0018] Step 1.3: When the embedding similarity of the generation result corresponding to the th low-resource document is lower than the threshold , trigger the screening mechanism, that is, repeat Steps 1.1 - 1.2 to regenerate the generation result corresponding to the th low-resource document , where is the low-resource corpus set, and retain the highest embedding similarity of the generation result in each iteration, and the low-resource corpus set , where is the low-resource corpus set, retain the highest embedding similarity of the generation result in each iteration, and the low-resource corpus set After two data augmentations, the first augmented document set is obtained and the second augmented document set , where is the number of documents in the low-resource corpus set

[0019] A further improvement of this application is that in step 2, the document-level representation of the documents in the augmented corpus is extracted, which specifically includes the following steps

[0020] Step 2.1: Using a pre-trained language model, encode each augmented document to obtain the word embedding set corresponding to this augmented document i.e., the set of words contained in this augmented document

[0021]

[0022] where , is the word embedding identifier is the augmented document the th word is the Transformer encoder , is the augmented document the number of words in is the dimension is the latent variable dimension size the th word's document-level representation

[0023] Step 2.2: Generate the document-level embedding representation corresponding to the augmented document :

[0024]

[0025] A further improvement of this application is that in step 3, a learnable topic embedding matrix is constructed, and the document topic distribution is calculated through document-topic similarity, which specifically includes the following steps

[0026] Step 3.1: Construct a learnable topic embedding matrix , where is the number of topics

[0027] Step 3.2: The topic distribution is calculated by the dot product of the document-level embedding representation and the topic embedding matrix. The topic distribution corresponding to each augmented document is calculated as follows

[0028] ​

[0029] Among them, is to enhance the document corresponding document-level representation, is to calculate the topic distribution, is the instance normalization operation to achieve independent normalization of each topic dimension.

[0030] A further improvement of this application lies in: in step 4, a semantic-aware contrastive learning framework is designed, dynamic negative samples are screened from the enhanced documents in the same batch of the contrastive learning framework, the contrastive learning loss is calculated, the topic embedding matrix is optimized, and the prior alignment loss is used to ensure topic consistency to obtain topic words, which specifically includes the following steps:

[0031] Step 4.1: Using the topic distribution of the enhanced document as a basis, calculate the th enhanced document embedding and the th enhanced document embedding the relative semantic correlation score between:

[0032]

[0033] Among them, represents calculating the cosine similarity of two vectors.

[0034] Step 4.2: Screen dynamic negative samples from the enhanced documents in the same batch of the contrastive learning framework through the following formula:

[0035]

[0036] Among them, is the hyperparameter of the threshold;

[0037] Step 4.3: For the low-resource document , there are the first enhanced document corresponding to the low-resource document and the second enhanced document . When the contrastive learning framework is trained, randomly sample from the low-resource corpus set , is the batch size, take the first batch of enhanced documents and the second batch of enhanced documents of the random sample , calculate the topic distribution of the first batch of enhanced documents and the topic distribution of the second batch of enhanced documents respectively, and construct positive sample pairs . The negative sample pairs within the same batch are passed through Perform screening and calculate the contrastive learning loss for the positive sample pairs and negative sample pairs in the first batch , and the contrastive learning loss for the positive sample pairs and negative sample pairs in the second batch :

[0038]

[0039]

[0040] Among them, is the hyperparameter of the trade-off factor between positive sample pairs and negative sample pairs, represents the temperature parameter of the document-level embedding representation, and the total contrastive learning loss for the same batch is:[[]]

[0041]

[0042] Step 4.3: Randomly sample two batch sizes, namely the prior topic distribution , and calculate the prior alignment loss :

[0043]

[0044] Among them, represents the topic distribution inferred from the two batches of augmented documents, is the moment order, is the number of topics calculated currently, is the th topic distribution, is the th prior distribution, represents the mean of the dth topic in the inferred topic distribution domain, represents the mean of the dth topic in the prior distribution domain, represents 's th topic, represents 's th topic;

[0045] Step 4.4: The total loss function :

[0046]

[0047] Among them, is the hyperparameter.[[]]

[0048] A further improvement of this application lies in: The training process of the semantic-aware contrastive learning framework designed in the said Step 4 includes:

[0049] Step T1: Initialize the topic embedding matrix and the pre-trained language model parameters, configure the non-zero hyperparameters including the size of the latent variable dimension, and input the low-resource documents into the large language model to generate an enhanced document set;

[0050] Step T2: Randomly sample the enhanced document set in batches and input it into the contrastive learning framework, perform prior sampling in the Dirichlet parameter space, and obtain the document topic distribution of the enhanced document set;

[0051] Step T3: Comprehensively evaluate the total contrastive learning loss and the prior alignment loss in the same batch, and repeatedly perform the operation of updating the parameters of the topic embedding matrix. Terminate the training when the semantic-aware contrastive learning framework converges.

[0052] A further improvement of this application is that the method for expanding and generalizing topic insights of the large language model in step 5 is specifically as follows: Give the name of the low-resource document and the 10 topic words corresponding to each of the corresponding topics in the prompt, and require it to generalize for each topic to generate human-readable topic insights.

[0053] The beneficial effects of this application are:

[0054] This application designs a semantic-aware contrastive learning framework, which is the first attempt to use a neural-based method to establish a low-resource topic model.

[0055] This application designs a semantic-aware text augmentation and corpus augmentation scheme based on the large language model, which can solve the problem of lack of data in the low-resource corpus.

[0056] This application proposes to explore incorporating semantic knowledge in the transformer language model into the modeling process to address the challenges of model complexity and noisy information.

[0057] In summary, this method has been successfully applied to the low-resource text analysis scenario of online forums, realizing a complete technical path from data augmentation to multi-level topic parsing. Compared with traditional topic models, this application shows significant advantages in topic semantic coherence and fine-grained knowledge discovery, providing a new solution for text analysis in low-resource environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 is the specific flowchart of this application.

[0059] Figure 2 is the implementation model architecture diagram of this application.

[0060] Figure 3 ​It is a comparison line chart of different comparison models in this application on four indicators. Detailed implementation manners

[0061] The embodiments of the present application will be disclosed below with reference to the drawings. For the sake of clarity, many practical details will be described together in the following description. However, it should be understood that these practical details are not used to limit the present application. That is to say, in some embodiments of the present application, these practical details are not necessary. In addition, for the purpose of simplifying the drawings, some conventional structures and components will be shown in a simple schematic manner in the drawings.

[0062] As Figure 1 shown, a low-resource topic key theme extraction method for an online forum in the present application specifically includes the following steps:

[0063] Step 1: Obtain low-resource documents of the online forum, and perform semantic-preserving data augmentation on the obtained low-resource documents through a large language model to generate an enhanced document set, as Figure 2 described in part (a) below, specifically:

[0064] Step 1.1: Construct a document augmentation prompt template based on the large language model to generate large language model results, where the document augmentation prompt template includes semantic-preserving constraint conditions:

[0065] a) Principle of minimum semantic variation: It is required that each text in the generated enhanced document set is as semantically similar as possible to the low-resource document;

[0066] b) Optimization of sentence fluency: Eliminate spelling mistakes and ungrammatical expressions;

[0067] The prompt is specifically:

[0068] .

[0069] Step 1.2: Iterative generation: Use a pre-trained language model, in this embodiment, Sentence-BERT is adopted, and calculate the embedding similarity between the low-resource document and the large language model generation result corresponding to the low-resource document, where specifically, cosine similarity is used to measure;

[0070] Step 1.3: When the th low-resource document corresponding generation result has an embedding similarity lower than the threshold , trigger the screening mechanism, that is, repeat steps 1.1 - 1.2 to regenerate the generation result corresponding to the th low-resource document corresponding generation result , where is a low-resource corpus set, is a hyperparameter that retains the embedding similarity of the highest generated result in each iteration. The low-resource corpus set obtains the first augmented document set after two data augmentations and the second augmented document set , where is the number of documents in the low-resource corpus set.

[0071] Step 2: Use the pre-trained language model to extract the document-level representations of the augmented document sets, which specifically includes the following steps:

[0072] Step 2.1: Adopt the pre-trained language model Sentence-BERT to encode each augmented document to obtain the corresponding word embedding set of this augmented document i.e., the set of words contained in this augmented document:

[0073]

[0074] Among them, , is the word embedding identifier, is the augmented document in the th word, is the Transformer encoder, , is the augmented document in the number of words, is the dimension, is the size of the latent variable dimension, which is set to 768 in Sentence-BERT, is the th word's document-level representation;

[0075] Step 2.2: Generate the document-level embedding representation corresponding to the augmented document :

[0076] .

[0077] Step 3: Construct a learnable topic embedding matrix, and calculate the document-topic distribution through the document-topic similarity, as described in part 2(b) of the figure, which specifically includes the following steps:

[0078] Step 3.1: Construct a learnable topic embedding matrix , this matrix is randomly initialized and parameterized throughout the training process, where is the number of topics;

[0079] Step 3.2. The topic distribution is calculated by the dot product of the document-level embedding representation and the topic embedding matrix. For each enhanced document the corresponding topic distribution is calculated as follows:

[0080]

[0081] where is the document-level representation corresponding to the enhanced document is used to calculate the topic distribution, and is the instance normalization operation to achieve independent normalization of each topic dimension.

[0082] Step 4. Design a semantic-aware contrastive learning framework. Implement a dynamic negative sample strategy in the same batch of enhanced documents in the contrastive learning framework, calculate the contrastive learning loss, optimize the topic embedding matrix, and ensure topic consistency with the prior alignment loss to obtain topic words;

[0083] Step 5. Use a large language model to expand and generalize the topic insights of the topic words obtained in Step 4 to better understand the low-resource corpus, as shown in (c) and (d) below. Specifically, it includes the following steps: Figure 2

[0084] Step 4.1. Using the topic distribution of the enhanced document as a basis, calculate the relative semantic correlation score between the th enhanced document embedding and the th enhanced document embedding :

[0085]

[0086] where represents calculating the cosine similarity of two vectors.

[0087] Step 4.2. Screen the dynamic negative samples in the same batch of enhanced documents in the contrastive learning framework through the following formula:

[0088]

[0089] where is the hyperparameter of the threshold; in this way, the negative samples with high semantic similarity to the original sentence will be regarded as false negative samples and masked.

[0090] Step 4.3. For the low-resource document , there is the first enhanced document corresponding to the low-resource document and the second enhanced document When training the contrastive learning framework, randomly sample from the low-resource corpus and , is the batch size, take the first batch of enhanced documents of the random sampling and the second batch of enhanced documents , calculate the topic distribution of the first batch of enhanced documents and the topic distribution of the second batch of enhanced documents respectively, and construct positive sample pairs . The negative sample pairs within the same batch are screened through , calculate the contrastive learning loss of the first batch of positive sample pairs and negative sample pairs , the contrastive learning loss of the second batch of positive sample pairs and negative sample pairs : :

[0091]

[0092]

[0093] Among them, is the hyperparameter of the trade-off factor between positive sample pairs and negative sample pairs, represents the temperature parameter of the document-level embedding representation. The total contrastive learning loss of the same batch is:

[0094]

[0095] Step 4.3. Randomly sample two batch sizes from the Dirichlet( ) distribution, that is, the prior topic distribution of dimensions , and calculate the prior alignment loss :

[0096]

[0097] Among them, represents the topic distribution inferred from the two batches of enhanced documents, is the moment order, is the number of topics calculated currently, is the th topic distribution, is the th prior distribution, represents the mean of the dth topic in the inferred topic distribution domain, represents the mean of the dth topic in the prior distribution domain, represents th a topic, indicating the th topic;

[0098] Step 4.4, the total loss function :

[0099]

[0100] wherein, is a hyperparameter.

[0101] The method for expanding the large language model and generalizing topic insights in step 5 is specifically as follows: Give the name of the low-resource document and the 10 topic words corresponding to each of the topics in the prompt, and require it to generalize for each topic to generate human-readable topic insights. The method for expanding the large language model and generalizing topic insights in step S6 is specifically as follows:

[0102] Give the name of the post and the 10 topic words corresponding to each of the topics in the prompt, and require it to generalize for each topic to generate human-readable topic insights. The prompt is specifically:

[0103] .

[0104] The training process of the semantic-aware contrastive learning framework includes:

[0105] Step T1, initialize the topic embedding matrix and the pre-trained language model parameters, configure the non-zero hyperparameters including the latent variable dimension size and input the low-resource document into the large language model to generate an enhanced document set;

[0106] Step T2, randomly sample the enhanced document set in batches and input it into the contrastive learning framework, perform prior sampling in the Dirichlet parameter space to obtain the document topic distribution of the enhanced document set;

[0107] Step T3, comprehensively evaluate the total contrastive learning loss and the prior alignment loss , and repeatedly execute the operation of updating the parameters of the topic embedding matrix, and terminate the training when the semantic-aware contrastive learning framework converges.

[0108] The training algorithm is as follows:

[0109]

[0110] To verify the effectiveness of this application, an experimental dataset was constructed by selecting a sub-forum section named "What troubles you", and a systematic evaluation was conducted through four topic coherence indicators. The average performance of the core evaluation indicators is shown in Table 1.

[0111] Table 1

[0112]

[0113] As Figure 3 shown, the indicators of this application are: CP: 0.2301, CA: 0.1651, NPMI: 0.02136, UCI: -0.2558, all of which are higher than those of the comparison models. The highest values of the comparison models are respectively: CP: 0.0163, CA: 0.1413, NPMI: 0.0046, UCI: -0.3225. Three types of representative methods were selected for the experiment: Comparison Model 1, Comparison Model 2, and Comparison Model 3 as the comparison benchmarks. The specific models are:

[0114] Comparison Model 1 is the LDA method. Comparison Model 2 is the vONT method. Comparison Model 3 is the LLMTopic method.

[0115] CP( ), CA( ), NPMI (Normalized Pointwise Mutual Information), and UCI( ) are four types of indicators that are common evaluation indicators for quantifying the quality of topic modeling from dimensions such as semantic relevance and vocabulary distribution.

[0116] This exemplary embodiment is only used to illustrate the feasible implementation path of the technical principle and does not constitute a substantial limitation on the patent claims. R & D personnel in related fields, on the premise of strictly following the innovation boundary defined by the claims of this invention, for deductive behaviors such as parameter adaptation, equivalent architecture conversion, or application scenario extension of the technical features disclosed in the specification, should all be regarded as within the radiation range of the original patent's creative contribution and legally enjoy the patent protection effect.

Claims

1. A low-resource topic key theme extraction method for online forums, characterized by: The low-resource topic key theme extraction method specifically includes the following steps: Step 1: Obtain low-resource documents from online forums, perform semantically preserved data enhancement on the obtained low-resource documents through a large language model, and generate an enhanced document set; Step 2: Use the pre-trained language model to extract document-level representations in the enhanced document collection; Step 3: Construct a learnable topic embedding matrix and calculate the document topic distribution through document-topic similarity; Step 4: Design a semantic-aware contrastive learning framework, implement a dynamic negative sample strategy in the same batch of enhanced documents in the contrastive learning framework, calculate the contrastive learning loss, optimize the topic embedding matrix, use a priori alignment loss to ensure topic consistency, and obtain topic words; Step 5: Use the large language model to expand the subject terms obtained in step 4 and summarize the subject insights to help better understand the low-resource corpus.

2. According to claim 1, a method for extracting key themes of low-resource topics for online forums is characterized by: In step 1, low-resource documents of online forums are obtained, and semantically preserved data enhancement is performed on the obtained low-resource documents through a large language model to generate an enhanced document set, which specifically includes the following steps: Step 1.1: construct a document enhancement prompt template based on a large language model to generate a large language model result, wherein the document enhancement prompt template contains semantic preservation constraints: a) Minimum semantic variation principle: It requires that every text in the generated enhanced document set is semantically similar to the low-resource document; b) Optimize sentence fluency: eliminate spelling errors and grammatical irregularities; Step 1.2, iterative generation: Use the pre-trained language model to calculate the embedding similarity between the low-resource document and the large language model generation result corresponding to the low-resource document; Step 1.3: Low-resource documents The corresponding generated results The embedding similarity of When the screening mechanism is triggered, steps 1.1 to 1.2 are repeated to regenerate the Low-resource documents The corresponding generated results ,in For low-resource corpus collections, retain the embedding similarity of the highest generated result in each iteration, low-resource corpus collection After two data enhancements, the first enhanced document set is obtained And the second enhanced document collection ,in is the number of documents in the low-resource corpus.

3. According to the method for extracting key themes of low-resource topics for online forums, it is characterized by: The step 2 extracts the document level representation of the document in the enhanced corpus, specifically comprising the following steps: Step 2.1: Use the pre-trained language model to enhance each document Encode to get the word embedding set corresponding to the enhanced document That is, the set of words contained in the enhanced document: ; in, , is the word embedding tag, To enhance the documentation Middle words, is the Transformer encoder, , To enhance the documentation The number of words in is the dimension, is the size of the latent variable dimension, For the Document-level representation of words; Step 2.2: Generate document-level embedding representation corresponding to the enhanced document : 。 4. The method for extracting key themes of low-resource topics for online forums according to claim 3, characterized in that: The step 3 constructs a learnable topic embedding matrix and calculates the document topic distribution through document-topic similarity, specifically including the following steps: Step 3.1: Construct a learnable topic embedding matrix ,in is the number of topics; Step 3.2: The topic distribution is calculated by multiplying the document-level embedding representation and the topic embedding matrix. Corresponding topic distribution The calculation method is: ; in, To enhance the documentation The corresponding document-level representation, is to calculate the topic distribution, This is the instance normalization operation.

5. The method for extracting key themes of low-resource topics for online forums according to claim 4, characterized in that: In step 4, a semantically-aware contrastive learning framework is designed, dynamic negative samples are screened in the same batch of enhanced documents in the contrastive learning framework, contrastive learning loss is calculated, topic embedding matrix is ​​optimized, prior alignment loss is used to ensure topic consistency, and topic words are obtained, which specifically includes the following steps: Step 4.1: Use the topic distribution of the enhanced document as a basis to calculate the The document-level embedding corresponding to the augmented document represents the embedding and The document-level embedding corresponding to the augmented document represents the embedding The relative semantic relevance score between : ; in, Indicates calculating the cosine similarity of two vectors; Step 4.2: Use the following formula to filter dynamic negative samples in the same batch of enhanced documents in the comparative learning framework: ; in, is the hyperparameter of the threshold; Step 4.3: For low-resource documents , both have low resource documents Corresponding first enhancement document And the second enhanced document , when training the contrastive learning framework, from a low-resource corpus Random sampling , is the batch size, take random samples The first batch of enhanced documents And the second batch of enhanced documents , calculate the first batch of enhanced document topic distributions And the second batch of enhanced document topic distribution , and construct positive sample pairs , the negative sample pairs in the same batch are Screen and calculate the contrastive learning loss of the first batch of positive and negative sample pairs , the contrastive learning loss of the second batch of positive and negative sample pairs : ; ; in, is a hyperparameter, represents the temperature parameter of the document-level embedding representation and the total contrastive learning loss of the same batch for: ; Step 4.3, randomly sample two batch sizes, that is Prior topic distribution , and calculate the prior alignment loss : ; in, represents the topic distribution inferred from two batches of augmented documents, is the moment order, is the number of topics currently being calculated, For the The distribution of topics, For the A prior topic distribution, Represents the inferred topic distribution domain The mean of the d-th topic, Represents the prior distribution domain The mean of the dth topic, express No. Themes, express No. Themes Step 4.4, total loss function : ; in, is a hyperparameter.

6. The method for extracting key themes of low-resource topics for online forums according to claim 5, characterized in that: The step 4 designs a semantically-aware contrastive learning framework training process including: Step T1: Initialize the topic embedding matrix The configuration includes the size of the latent variable dimension and the pre-trained language model parameters. Non-zero hyperparameters including , where low-resource documents are fed into a large language model to generate an enhanced document set; Step T2: Randomly extract the enhanced document set in batches and input it into the contrastive learning framework, perform prior sampling in the Dirichlet parameter space, and obtain the document topic distribution of the enhanced document set; Step T3: Comprehensively evaluate the total contrastive learning loss of the same batch Alignment loss with priors , the topic embedding matrix parameter update operation is performed cyclically, and the training is terminated when the semantic-aware contrastive learning framework converges.

7. The method for extracting key themes of low-resource topics for online forums according to claim 6, characterized in that: The method of expanding the large language model and summarizing topic insights in step 5 is specifically: giving the name of the low-resource document and its corresponding The 10 keywords corresponding to each topic are required to summarize each topic to generate human-readable topic insights.

Citation Information

Patent Citations

  • Text abstract automatic extraction method based on semantic matching

    CN115965027A

  • Cross-language abstract abstract method based on robust self-learning strategy in low-resource scene

    CN117271761A

  • Ancient book knowledge base intelligent question and answer method, device and equipment based on large model technology

    CN118332091A

  • Online human activity identification method based on robust reinforcement learning

    CN118897984A

  • Dangerous behavior identification and early warning method based on multi-modal analysis

    CN119360278A

Cited By

  • Multi-modal topic modeling method based on semantic consistency driving

    CN121859997A