Controllable theme system based on semantic potential space joint learning and optimization method

Through a manipulated theme system based on semantic latent spatial joint learning, combined with an autoencoder and visual editing system, the problems of unclear semantics and poor interpretability in neural theme models are solved, and efficient optimization of the theme and user-friendly editing experience are achieved.

CN120337871APending Publication Date: 2025-07-18TIANJIN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510416781.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing neural theme models are prone to produce unclear semantics or irrelevant words when generating topics, resulting in a decrease in the interpretability and practicality of topics. It is difficult for traditional visualization tools to effectively display the internal semantic relationships of neural theme models, limiting users' in-depth understanding and editing of topics.

Method used

The manipulated theme system based on semantic latent space joint learning is adopted, combined with a visual editing system and a guided neural theme model, the theme embedding representation is generated through the autoencoder and the clusterer, and the semantic relationship between objects is reflected using the RGB color space, supporting users to add or delete the association constraints between words, documents and topics during the training process, and optimize the theme quality.

Benefits of technology

It realizes refined editing and optimization of the topic, improves users' confidence in editing results, provides flexible and intuitive analysis tools, and improves the efficiency and quality of topic modeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337871A_ABST
    Figure CN120337871A_ABST
Patent Text Reader

Abstract

The invention discloses a maneuverable theme system based on semantic potential space joint learning. The system comprises a visual editing system and a guidable neural theme model based on an auto-encoder. A neural topic model can be guided, and semantic relations of words, documents and topics are mapped into a unified potential space in a joint embedding mode, so that semantically related objects are closer in the potential space; a constraint loss function is set to optimize a theme; and the visual editing system is provided with a man-machine interaction component and a quantitative index, performs quality scoring on a theme generated by the guidable neural theme model, and supports a user to add or delete association constraints between words, documents and themes based on the quality score of the theme in the process of training the guidable neural theme model, so that the user experience is improved. And enabling the guidable neural topic model to generate a topic meeting user requirements. According to the method, theme controllable optimization is realized through joint embedding and constraint loss, and a flexible and visual controllable analysis tool is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning, and particularly relates to a manipulable topic system and an optimization method based on joint learning of semantic latent space. Background Art

[0002] Currently, with the rapid development of information technology, the generation and accumulation of massive text data have become the norm in modern society. From news articles, social media posts to academic papers, the scale and complexity of text data are constantly increasing, and how to extract valuable semantic information from these data has become an important challenge. Topic modeling, as an unsupervised learning method, can automatically extract the latent topic structure from a large number of documents and is widely used in fields such as text mining, information retrieval, and recommendation systems. Traditional topic modeling methods, such as Latent Dirichlet Allocation (LDA) and Non-negative Matrix Factorization (NMF), capture the probability distribution between documents and words through statistical models and generate two distributions, namely "document-topic" and "topic-word", to represent topics. However, when dealing with large-scale and high-dimensional text data, these methods often face problems such as high computational complexity and unstable topic quality, and it is difficult to flexibly introduce user prior knowledge or constraints.

[0003] In recent years, with the rise of deep learning technology, Neural Topic Models (NTMs) have gradually become a research hotspot in the field of topic modeling. Different from traditional statistical models, neural topic models learn the distributed representations (i.e., embeddings) of words and documents through neural networks, can better capture semantic information, and have stronger scalability and flexibility. For example, neural topic models based on Variational Autoencoder (VAE) (such as ProdLDA) generate topic distributions by reconstructing the Bag of Words (BoW) model, while topic models based on pre-trained language models (such as BERT) (such as BERTopic) extract topics by clustering document embeddings. These methods perform well in topic modeling tasks, especially when dealing with large-scale text data, and can generate more semantically consistent topics.

[0004] However, existing neural topic models still have some limitations. First, although neural topic models can generate high-quality topics, in practical applications, topic models inevitably generate words with unclear semantics or unrelated to the topic, resulting in a decline in the interpretability and practicality of the topics. For example, some topics may contain words with vague semantics, or some high-ranked words are irrelevant to the core semantics of the topic. Therefore, some studies have proposed visualization-based topic editing systems, but these systems are mainly aimed at traditional statistical models and are difficult to be directly applied to neural topic models. Due to their complex neural network structure and embedded representation, traditional visualization tools are difficult to effectively display the internal semantic relationships of neural topic models, resulting in great difficulties for users in editing and optimizing topics. In addition, existing visualization tools usually cannot support the interactive exploration of the joint embedded representation of words, documents, and topics, restricting users' in-depth understanding of topic semantics and refined editing. Summary of the Invention

[0005] The present invention provides a manipulable topic system and optimization method based on joint learning of semantic latent space to solve the technical problems existing in the prior art.

[0006] The technical solution adopted by the present invention to solve the technical problems existing in the prior art is:

[0007] A manipulable topic system based on joint learning of semantic latent space, the system includes a visualization editing system and a steerable neural topic model based on an autoencoder;

[0008] The steerable neural topic model includes an encoder, a decoder, and a clustering module. The encoder is used to input the embeddings of words and documents and generate corresponding embedding representations. The clustering module is used to aggregate the embedding representations of words to generate topic embedding representations; the decoder is used to reconstruct the information of the generated word, document, and topic embedding representations, so that the latent space captures the structural features of the data distribution.

[0009] The visualization editing system sets up a human-computer interaction component and a quantization index, which scores the quality of the topics generated by the steerable neural topic model, and supports users to add or delete the association constraints between words, documents, and topics based on the quality score of the topics during the process of training the steerable neural topic model, so that the steerable neural topic model generates topics that meet the user's needs.

[0010] Furthermore, the visualization editing system includes: a topic quality module for evaluating the quality of topics and / or a semantic relevance module for evaluating the semantic relevance of objects; the topic quality module quantifies the quality of topics according to the compactness of clustering; the semantic relevance module measures the semantic relevance of objects based on the proximity between objects.

[0011] Furthermore, the visual editing system includes an embedded object color configuration module for assigning a unique color to each object and reflecting the semantic relationship between objects through colors; the embedded object color configuration module maps the embedded position of the object to the color space and assigns a unique color code to each object.

[0012] Furthermore, the color space is the RGB color space. The embedded object color configuration module takes the three-dimensional coordinate values of the embedded position of each object as the R, G, and B values of the color respectively, and assigns the corresponding color to it.

[0013] Furthermore, the visual editing system includes a theme editing module for assisting users in editing the theme content. The theme editing module includes a theme list sub-module, a word editing sub-module, and / or a document editing sub-module; the theme list sub-module is used to display the generated themes in a list form for users to select, manage, and edit the themes. It consists of multiple line charts, each line chart corresponding to a theme, and showing the change trend of the quality score of the theme in different editing rounds; the word editing sub-module is used to display the word similarity matrix of the selected theme to assist users in evaluating the theme quality and editing related words; the document editing sub-module is used to display the document list related to the selected theme, support document expansion and keyword annotation, and assist users in understanding the document semantics and performing related editing.

[0014] Furthermore, the theme editing module also includes a constraint operation icon sub-module. When the theme editing module performs a theme editing operation, the constraint operation icon sub-module uses corresponding icons to represent a constraint operation performed on the object being operated on.

[0015] Furthermore, the visual editing system includes an auxiliary function module for helping users complete the theme editing task more efficiently. The auxiliary function module includes a process management sub-module, an object recommendation sub-module, and / or a color application sub-module; the process management sub-module sets a rollback function, and quickly restores all themes to the state of a certain round through the rollback function; the object recommendation sub-module recommends other objects most relevant to the selected object to the user and highlights the recommended objects; the color application sub-module adopts a color mechanism in the theme list sub-module, the word editing sub-module, and the document editing sub-module, and intuitively reflects the ranking of objects by coloring related objects.

[0016] The present invention also provides an optimization method for using the above-mentioned manipulable theme system based on joint learning in the semantic latent space, which can guide the neural theme model. It maps the semantic relationships of words, documents, and themes to a unified latent space through a joint embedding method, so that semantically related objects are closer in the latent space; it sets constraints and constraint loss functions to optimize the theme;

[0017] Let: w iDenote the $i$-th word, where $i$ is the word sequence number; $i = 1, 2, \ldots, N$; $W = [w_1, w_2, \ldots, w N $ is a set containing $N$ words, and $d j Denote the $j$-th document; $j$ is the document sequence number; $j = 1, 2, \ldots, M$; $D = [d_1, d_2, \ldots, d M $ is a set containing $M$ documents; $N$ represents the number of words; $M$ represents the number of documents;

[0018] The encoder uses the BERT model to extract the embedding features of each word in the document. Let denote the embedding of the $i$-th word in the $j$-th document extracted; Let denote the $i$-th word embedding, which is characterized by the average of the embeddings of the $i$-th word in all documents containing the $i$-th word; Let $D i be the number of documents containing the $i$-th word; The set of word embeddings is $H w ,

[0019] The calculation formula of is as follows:

[0020]

[0021] In the formula:

[0022] is the word embedding in the document containing the $i$-th word;

[0023] $d$ is the document containing the $i$-th word;

[0024] Let be the $j$-th document embedding; The set of document embeddings is $H d , For any document, the attention mechanism is introduced to calculate The calculation formula of is as follows:

[0025]

[0026] Among them, $a i denotes the attention weight corresponding to $w i $, and $a i is calculated by the following formula:

[0027]

[0028] In the formula:

[0029] $l i is the new representation corresponding to ;

[0030] $n$ is the corresponding $H wThe number of new representations of word embeddings;

[0031] s corresponds to H w The serial number of the new representation of word embeddings;

[0032] l s corresponds to H w The s-th new representation in the set of new representations of word embeddings;

[0033] is l i transpose vector of;

[0034] is l s transpose vector of;

[0035] v is a learnable vector of the attention mechanism;

[0036] E is the weight of the neuron;

[0037] b is the bias of the neuron;

[0038] exp() represents the power function of the base e of the natural logarithm;

[0039] tanh() represents the hyperbolic tangent function;

[0040] E and b are learnable parameters of the linear transformation;

[0041] l i The dot product of l and v reflects w i and d j correlation;

[0042] Let: be the i-th word embedding representation; the set of word embedding representations is Z w ,

[0043] be the j-th document embedding representation; the set of document embedding representations is Z d ,

[0044] be the k-th topic embedding representation; k = 1, 2, …, K; the set of topic embedding representations is Z t , K is the number of topic embedding representations; k is the serial number of the topic embedding representation; t k be the k-th topic;

[0045] Among them, Z w and Z d are directly output by the encoder, while Z t is obtained by the clusterer for Zw Obtained by performing K-means clustering; each cluster corresponds to a topic, The calculation formula is as follows:

[0046]

[0047] where B k represents the k-th cluster; w represents the word included in B k ; z w represents the word embedding corresponding to w;

[0048] The guided neural topic model includes the following four losses: word reconstruction loss, document reconstruction loss, topic clustering loss, and constraint loss;

[0049] Let the word reconstruction loss be l w , the document reconstruction loss be l d , the topic clustering loss be l t , and the constraint loss be l c ;

[0050] Let be the set of reconstructed word embeddings, be the word reconstruction embedding corresponding to ;

[0051] l w The calculation formula is as follows:

[0052]

[0053] Let be the topic reconstruction embedding output by the decoder corresponding to ; the set of topic reconstruction embeddings is

[0054] Let be the document reconstruction embedding corresponding to , which is the weighted sum of , that is:

[0055]

[0056] In the formula:

[0057] is the and cosine similarity;

[0058] l d The calculation formula is as follows:

[0059]

[0060] By setting ld , increase Z t and Z d to enhance the relevance of Z, and better capture the semantic relationship between Z t and Z d ;

[0061] Introduce l t to enhance the importance of topic clustering in the latent space; Let be and 's cosine similarity, and define another similarity between and as:

[0062]

[0063] In the formula:

[0064] is and 's similarity; it is denser than clustering;

[0065] c k is the sum of the cosine similarities between and all words;

[0066] k′ is the topic embedding serial number in Z t ;

[0067] c k′ is the sum of the cosine similarities between the k′-th topic embedding and all words;

[0068] Define l t as and 's cross entropy, then there is:

[0069]

[0070] will make sharper, and if t k is more similar to w i , then its value will be larger; Therefore, l t will pull w i closer to its most similar topic in the latent space, thus forming a clear topic clustering;

[0071] Integrate all objects by constructing a matrix: words, documents, and topics; Assume there are M documents, N words, and K topics, and the matrix size is (M + N + K) × (M + N + K); Initially, the matrix is all 0; The user associates two objects by setting the value of the corresponding cell in the matrix to 1, or disassociates them by setting it to -1; Let aef is the value of the cell in the e-th row and f-th column of the matrix; e is the row index of the matrix; f is the column index of the matrix;

[0072] l c The calculation formula of

[0073]

[0074] In the formula:

[0075] z e is the embedding of the object in the e-th row;

[0076] z f is the embedding of the object in the f-th row;

[0077] u is the number of all embeddings.

[0078] Furthermore, the set constraints include one or several combinations of the following constraints:

[0079] The first constraint: Add a word to the topic; Let t k be a topic, w x be a word related to t k ; I(w x ) and I(t k ) are their indexes in the matrix; By setting add w x to t k ;

[0080] The second constraint: Remove a word from the topic; Let t k be a topic, w y be an irrelevant high-ranking word in t k ; I(w y ) and I(t k ) are their indexes in the matrix; By setting remove w y from t k ;

[0081] The third constraint: Add a document to the topic; Let t k be a topic, d m be a document related to t k ; I(d m ) and I(t k ) are their indexes in the matrix; By setting add d m to t k ;

[0082] The fourth constraint: Remove a document from the topic; Let t kFor a topic, d n For t k An irrelevant high-ranked document in it; I(d n ) and I(t k ) are their indices in the matrix; By setting Remove d n From t k In it;

[0083] Directly control Z w 、Z d And Z t To implement the following fourth to eighth constraints:

[0084] The fifth constraint: Delete a word; Let w u Be a meaningless word; By removing from Z w In it To delete the word from the corpus;

[0085] The sixth constraint: Delete a document; Let d u Be a relatively ambiguous document; By removing from Z d In it To delete the document from the corpus;

[0086] The seventh constraint: Delete a topic; Let t q Be a topic with relatively low quality; By deleting from Z t In it To delete the topic from the corpus;

[0087] The eighth constraint: Add a topic; Let t c Be a new topic, Be its embedding vector; By initializing with the embedding vector of any object related to t c Then Put Into Z t In it, thus adding t c To the corpus;

[0088] Use the above constraints to implement two other relatively more complex constraints, specifically as follows:

[0089] The ninth constraint: Merge topics; Let t p And t a Be two related topics, And Be their embedding vectors; By using And The average value of to initialize Let t p And t aMerge into a new topic t k ; then add to the corpus and remove t p and t a ;

[0090] Tenth constraint: split topic; let t d be a semantically rich topic; split t e into two topics by creating a new topic t d , where t e contains one semantic aspect of t d ; the embedding vector of the object most relevant to the semantics should be used to initialize and add the relevant objects in t d to t e .

[0091] Furthermore, the user views the initially generated topics and their quality scores through the visual editing system, selects the topics that need to be optimized; sets button components, and the user adds or deletes the association constraints between words, documents and topics by operating the corresponding button components, and cancels unreasonable editing operations; the visual editing system updates the quality scores of the topics in real time.

[0092] The advantages and positive effects of the present invention are: A manipulable topic system based on joint learning of semantic latent space of the present invention integrates a topic list, a word editor and a document editor, and supports the user to perform refined editing on topics in a unified context; through the coloring scheme of joint embedding and the object recommendation mechanism, visual clutter is reduced, user operations are simplified, and the user's confidence in the editing results is improved.

[0093] An optimization method for a manipulable topic system based on joint learning of semantic latent space of the present invention adopts a neural topic modeling method based on an autoencoder, and realizes controllable optimization of topics through joint embedding and a constraint loss function. It provides a flexible, intuitive and highly interpretable analysis tool for the controllable optimization of document topics, and is applicable to natural language processing tasks such as text understanding, topic modeling, and semantic analysis. Brief Description of the Drawings

[0094] Figure 1 is a flowchart of an optimization method for a manipulable topic system based on joint learning of semantic latent space of the present invention.

[0095] Figure 2 is a working principle diagram of an embedded object color configuration module in the visual editing system of the present invention.

[0096] Figure 3 is a visual interface diagram of the visual editing system of the present invention.

[0097] Figure 4 It is the operation icon of the constraint operation icon sub-module in the visual editing system of the present invention.

[0098] In the figure: f, encoder; g, decoder; h w , word embedding; h d , document embedding; Reconstructed word embedding; Reconstructed document embedding; Reconstructed topic embedding. l w is the word reconstruction loss, l d is the document reconstruction loss, l t is the topic clustering loss, l c is the constraint loss.

[0099] 1. Parameter setting window; 2. Topic list sub-module; 3. "+Word" button; 4. "-Word" button; 5. "×Word" button; 6. "+Doc" button; 7. "-Doc" button; 8. "×Doc" button; 9. "∨Topic" button; 10. "∧Topic" button; 11. "+Topic" button; 12. "×Topic" button; 13. Rollback function button; 14. Word editing sub-module; 15. Document editing sub-module; 16. Constraint operation icon; 17. Word and document switching button. Detailed implementation manners

[0100] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments. It should be understood that the preferred embodiments described herein are only for explaining and illustrating the present invention, and are not used to limit the present invention.

[0101] The Chinese interpretations of the following English words, abbreviations and phrases are as follows:

[0102] SNTM: The guided neural topic model in the present invention.

[0103] LDA: It is a probabilistic generative model that realizes text topic discovery by representing a document as a distribution of latent topics and representing a topic as a distribution of words.

[0104] Bert: It is a pre-trained language model based on bidirectional Transformer that understands sentence structure and semantics through deep contextualized representations and is used for various NLP tasks.

[0105] Attention: A mechanism that enables the model to focus on the most relevant information when processing sequential data by assigning different importance weights, improving performance and context understanding ability.

[0106] Weight: Trainable parameters in a neural network that determine the impact of input features on the final output and are continuously adjusted during training using optimization methods such as gradient descent.

[0107] Latent Space: The mapping of data in a low-dimensional or hidden feature representation that captures its core patterns and structures, commonly used in dimensionality reduction, generative models, and representation learning.

[0108] Topic: The distribution of a set of words with similar semantics, representing latent concepts in a document collection.

[0109] Word: The basic building block of text, regarded as an element generated from the probability distribution of a certain topic in a topic model.

[0110] Document: A text fragment composed of multiple words, represented as a mixed distribution of multiple topics in a topic model.

[0111] editer switch: The switch between the word editor and the document editor.

[0112] Add: Apply a constraint loss for linking to the selected element.

[0113] Remove: Apply a separate constraint loss for the selected element.

[0114] Delete: Delete the corresponding element from the database.

[0115] the current topic: The current topic being edited by the user.

[0116] other topic: Topics other than the current topic being edited by the user.

[0117] Data of New Topic: The data added for the newly created topic.

[0118] global: Global.

[0119] encoder: Encoder.

[0120] decoder: Decoder.

[0121] Please refer to Figures 1 to 4 , a manipulable topic system based on joint learning in semantic latent space, which includes a visual editing system and a guided neural topic model based on an autoencoder;

[0122] The guided neural topic model includes an encoder, a decoder, and a clustering module. The encoder is used to input the embeddings of words and documents and generate corresponding embedding representations. The clustering module is used to aggregate the embedding representations of words to generate topic embedding representations. The decoder is used to reconstruct information from the generated word, document, and topic embedding representations, enabling the latent space to capture the structural features of the data distribution.

[0123] A visualization editing system sets up a human-computer interaction component and a quantization metric, which scores the quality of the topics generated by the guided neural topic model, and supports the user in adding or deleting the association constraints between words, documents, and topics during the training of the guided neural topic model, so that the guided neural topic model generates topics that meet the user's needs.

[0124] Preferably, the visualization editing system may include: a topic quality module for evaluating the quality of topics and / or a semantic relevance module for evaluating the semantic relevance of objects; the topic quality module quantifies the quality of topics according to the compactness of clustering; the semantic relevance module measures the semantic relevance of objects based on the proximity between objects.

[0125] Preferably, the visualization editing system may include an embedded object color configuration module for assigning a unique color to each object and reflecting the semantic relationship between objects through colors; the embedded object color configuration module maps the embedding positions of objects to the color space and assigns a unique color code to each object.

[0126] Preferably, the color space may be the RGB color space, and the embedded object color configuration module takes the three-dimensional coordinate values of the embedding positions of each object as the R, G, and B values of the color respectively, and assigns the corresponding color to it.

[0127] Preferably, the visualization editing system may include a topic editing module for assisting the user in editing the topic content. The topic editing module includes a topic list sub-module 2, a word editing sub-module 14, and / or a document editing sub-module 15. The topic list sub-module 2 is used to display the generated topics in the form of a list for the user to select, manage, and edit topics. It may consist of multiple line charts, each line chart corresponding to a topic, showing the changing trend of the quality scores of the topic in different editing rounds. The word editing sub-module 14 is used to display the word similarity matrix of the selected topic, assisting the user in evaluating the topic quality and editing relevant words. The document editing sub-module 15 is used to display the list of documents related to the selected topic, support document expansion and keyword annotation, and assist the user in understanding the document semantics and performing relevant editing.

[0128] Preferably, the topic editing module may further include a constraint operation icon sub-module. When the topic editing module performs a topic editing operation, the constraint operation icon sub-module uses corresponding icons to represent a constraint operation performed on the object being operated on.

[0129] Preferably, the visual editing system may include an auxiliary function module for helping the user to more efficiently complete the topic editing task. The auxiliary function module includes a process management sub-module, an object recommendation sub-module, and / or a color application sub-module; the process management sub-module sets a rollback function to quickly restore all topics to the state of a certain round through the rollback function; the object recommendation sub-module recommends other objects most relevant to the selected object to the user and highlights the recommended objects; the color application sub-module adopts a color mechanism in the topic list sub-module 2, the word editing sub-module 14, and the document editing sub-module 15 to intuitively reflect the ranking of objects by coloring the relevant objects.

[0130] As Figure 3 shown, the auxiliary function module may further include various operation button components, such as: "+Word" button 3, "-Word" button 4, "×Word" button 5, "+Doc" button 6, "-Doc" button 7, "×Doc" button 8, "∨Topic" button 9, "∧Topic" button 10, "+Topic" button 11, "×Topic" button 12, rollback function button 13, word and document switching button 17.

[0131] The auxiliary function module may further include a parameter setting window 1 for setting various parameters.

[0132] The present invention also provides an optimization method for a manipulable topic system based on the above-mentioned joint learning of semantic latent space, which can guide a neural topic model. The neural topic model maps the semantic relationships of words, documents, and topics into a unified latent space through a joint embedding method, so that semantically related objects are closer in the latent space; it sets constraints and constraint loss functions to optimize the topics;

[0133] Let: w i represent the i-th word, where i is the word sequence number; i = 1, 2,..., N; W = [w1, w2,..., w N is a set containing N words, and d j represents the j-th document; j is the document sequence number; j = 1, 2,..., M; D = [d1, d2,..., d M is a set containing M documents; N represents the number of words; M represents the number of documents;

[0134] The encoder uses the BERT model to extract the embedding features of each word in the document. Let represent the embedding of the i-th word extracted in the j-th document; let represent the i-th word embedding, which is characterized by the average value of the embeddings of the i-th word in all documents containing the i-th word; let D i be the number of documents containing the i-th word; the word embedding set is Hw ,

[0135] The calculation formula of is as follows:

[0136]

[0137] Wherein:

[0138] is the word embedding in the document containing the i-th word;

[0139] d is the document containing the i-th word;

[0140] Let be the j-th document embedding; the set of document embeddings is H d , For any document, an attention mechanism is introduced to calculate The calculation formula of is as follows:

[0141]

[0142] where a i represents the attention weight corresponding to w i a i is calculated by the following formula:

[0143]

[0144] Wherein:

[0145] l i is the new representation corresponding to ;

[0146] n is the number of new representations of word embeddings in H w ;

[0147] s is the serial number of the new representation of word embeddings in H w ;

[0148] l s is the s-th new representation in the set of new representations of word embeddings in H w ;

[0149] is the transposed vector of l i ;

[0150] is the transposed vector of l s ;

[0151] v is the learnable vector of the attention mechanism;

[0152] E is the weight of the neuron;

[0153] b is the bias of the neuron;

[0154] exp() represents the power function of the base e of the natural logarithm;

[0155] tanh() represents the hyperbolic tangent function;

[0156] E and b are learnable parameters of the linear transformation;

[0157] l i The dot product of l and v reflects the i correlation with d j ;

[0158] Let: be the i-th word embedding representation; the set of word embedding representations is Z w ,

[0159] be the j-th document embedding representation; the set of document embedding representations is Z d ,

[0160] be the k-th topic embedding representation; k = 1, 2, …, K; the set of topic embedding representations is Z t , K is the number of topic embedding representations; k is the serial number of the topic embedding representation; t k is the k-th topic;

[0161] where, Z w and Z d are directly output by the encoder, while Z t is obtained by performing K-means clustering on Z w ; each cluster corresponds to a topic, The calculation formula of is as follows:

[0162]

[0163] where, B k represents the k-th cluster; w represents the word included in B k ; z w represents the word embedding corresponding to w;

[0164] The guided neural topic model includes the following four losses: word reconstruction loss, document reconstruction loss, topic clustering loss, and constraint loss;

[0165] Let the word reconstruction loss be l w , and the document reconstruction loss be l d, the topic clustering loss is \(l\). t , the constraint loss is \(l\). c ;

[0166] Let be the set of reconstructed word embeddings, and be the corresponding word reconstruction embedding;

[0167] The calculation formula of \(l\) w is as follows:

[0168]

[0169] Let be the corresponding topic reconstruction embedding output by the decoder for ; the set of topic reconstruction embeddings is

[0170] Let be the corresponding document reconstruction embedding for , which is the weighted sum of , that is:

[0171]

[0172] In the formula:

[0173] is and 's cosine similarity;

[0174] The calculation formula of \(l\) d is as follows:

[0175]

[0176] By setting \(l\) d , the relevance between \(Z\) t and \(Z\) d is increased, and the semantic relationship between \(Z\) t and \(Z\) d is better captured;

[0177] Introduce \(l\) t to enhance the importance of topic clustering in the latent space; Let be and 's cosine similarity, and define and 's another similarity as:

[0178]

[0179] In the formula:

[0180] For and similarity; it is denser than clustering;

[0181] c k is the sum of the cosine similarities between

[0182] k′ is the topic embedding serial number in Z t ;

[0183] c k′ is the sum of the cosine similarities between the k′-th topic embedding and all words;

[0184] Define l t as and cross entropy, then there is:

[0185]

[0186] will make sharper, and if t k is more similar to w i then its value will be larger; therefore, l t pulls w i closer to its most similar topic in the latent space, thus forming a clear topic clustering;

[0187] Integrate all objects by constructing a matrix: words, documents, and topics; assume there are M documents, N words, and K topics, and the matrix size is (M + N + K) × (M + N + K); initially, the matrix is all 0; the user associates two objects by setting the value of the corresponding cell in the matrix to 1, or disassociates them by setting it to -1; let a ef be the value of the cell in the e-th row and f-th column of the matrix; e is the matrix row serial number; f is the matrix column serial number;

[0188] l c is calculated as follows:

[0189]

[0190] In the formula:

[0191] z e is the embedding of the object in the e-th row;

[0192] z f is the embedding of the object in the f-th row;

[0193] u is the number of all embeddings.

[0194] Preferably, the set constraints may include one or several combinations of the following constraints:

[0195] The first constraint: adding a word to the topic; Let t k be a topic, and w x be a word related to t k ; I(w x ) and I(t k ) are their indexes in the matrix; By setting add w x to t k ;

[0196] The second constraint: removing a word from the topic; Let t k be a topic, and w y be an irrelevant high-ranking word in t k ; I(w y ) and I(t k ) are their indexes in the matrix; By setting remove w y from t k ;

[0197] The third constraint: adding a document to the topic; Let t k be a topic, and d m be a document related to t k ; I(d m ) and I(t k ) are their indexes in the matrix; By setting add d m to t k ;

[0198] The fourth constraint: removing a document from the topic; Let t k be a topic, and d n be an irrelevant high-ranking document in t k ; I(d n ) and I(t k ) are their indexes in the matrix; By setting remove d n from t k ;

[0199] Directly control Z w 、Z d and Z t to implement the following fourth to eighth constraints:

[0200] The fifth constraint: deleting a word; Let w u be a meaningless word; By removing from Z w ​ Delete the word from the corpus;

[0201] The sixth constraint: Delete a document; Let d u be a relatively ambiguous document; By removing it from Z d to delete the document from the corpus;

[0202] The seventh constraint: Delete a topic; Let t q be a topic with low quality; By deleting it from Z t to delete the topic from the corpus;

[0203] The eighth constraint: Add a topic; Let t c be a new topic, be its embedding vector; By initializing with the embedding vector of any object related to t c and then putting into Z to add t t to the corpus; c

[0204] Use the above constraints to implement two other relatively more complex constraints as follows:

[0205] The ninth constraint: Merge topics; Let t p and t a be two related topics, and be their embedding vectors; By using the average of and to initialize merge t p and t a into a new topic t k ; Then add to the corpus and remove t p and t a from the corpus;

[0206] The tenth constraint: Split a topic; Let t d be a semantically rich topic; By creating a new topic t e to split t d into two topics, where t e contains one semantic aspect of t d ; The embedding vector of the object most semantically related should be used to initialize and add the relevant objects in t d to t e . ​​​

[0207] Preferably, the user can view the initially generated topics and their quality scores through a visual editing system, and select the topics that need to be optimized; button components can be set, and the user can add or delete the association constraints between words, documents, and topics by operating the corresponding button components, and undo unreasonable editing operations; the visual editing system updates the quality scores of the topics in real time.

[0208] The following uses a preferred embodiment of the present invention to further illustrate the working process and working principle of the present invention:

[0209] A manipulable topic system based on joint learning of semantic latent space, the system includes a visual editing system and a guided neural topic model based on an autoencoder;

[0210] The guided neural topic model includes an encoder, a decoder, and a clustering algorithm. The encoder is used to input the embeddings of words and documents and generate corresponding embedding representations. The clustering algorithm is used to aggregate the embedding representations of words to generate topic embedding representations; the decoder is used to reconstruct the information of the generated word, document, and topic embedding representations, so that the latent space captures the structural features of the data distribution;

[0211] The visual editing system sets up human-computer interaction components and quantization metrics, which perform quality scoring on the topics generated by the guided neural topic model, and support the user to add or delete the association constraints between words, documents, and topics based on the quality scores of the topics during the process of training the guided neural topic model, so that the guided neural topic model generates topics that meet the user's needs.

[0212] (1) Guided Neural Topic Model (SNTM)

[0213] The core framework of the guided neural topic model is based on the autoencoder structure, aiming to model the topic distribution in a large-scale document corpus through embedding learning and unsupervised clustering methods. Traditional topic models (such as LDA) face challenges of insufficient computational efficiency and semantic consistency when dealing with high-dimensional sparse data, while the guided neural topic model of the present invention significantly improves the model's ability to capture complex semantic patterns through the introduction of a deep learning architecture.

[0214] Figure 1 As shown, the guided neural topic model of the present invention adopts an autoencoder structure. W = [w1, w2,..., w N is a set containing N words, and D = [d1, d2,..., d M is a set containing M documents.

[0215] The encoder and decoder respectively implement the low-dimensional representation and information reconstruction of high-dimensional sparse data. The encoder receives the embeddings of words and documents, generates their corresponding latent representations, and simultaneously performs clustering operations through a clustering algorithm to generate topic embeddings. The decoder then uses these latent representations for information reconstruction to ensure that the latent space can effectively capture the structural features of the data distribution.

[0216] The encoder f takes the word embedding and the document embedding as inputs and outputs the joint embedding representations of words, documents, and topics, denoted as and respectively, where Z w and Z d are directly output by the encoder, while Z t is obtained by performing K-means clustering on Z w . Each cluster corresponds to a topic, and its calculation formula is as follows:

[0217]

[0218] where B k represents the k-th cluster; w represents the word included in B k ; and z w represents the word embedding corresponding to w.

[0219] The decoder g takes Z w , Z d , and Z t as inputs and outputs the reconstructed representations of Z w , Z d , and Z t , denoted as and respectively. The decoder directly outputs and while is calculated based on by calculating the differences between H w and , and between H d and to define two basic loss functions of the model.

[0220] For the above structural technical solutions, the following two points need to be noted:

[0221] The calculation of Z t is based on Z w , denoted as Z t →Zw .

[0222] The calculation of is denoted as

[0223] These two technical solutions aim to tightly connect words, documents, and topics through embedded semantic associations, so that the embedded distances can effectively reflect the semantic similarities of the corresponding objects. The output of the encoder (Z w , Z d , Z t ) is the word, document, and topic embedding representations learned by the model.

[0224] Word embedding. Use the BERT model to extract the embedding representation of each word in the document, denoted as where w i and d j represent the word and the document respectively. For any word w i , it is calculated by the average value of its embeddings in all documents containing the word The formula is as follows:

[0225]

[0226] Document embedding. Let be the j-th document embedding; the set of document embeddings is H d , For any document d j , introduce the attention mechanism to calculate The formula is as follows:

[0227]

[0228] where, a i represents the attention weight corresponding to w i , and is calculated by the following formula:

[0229]

[0230] where, l i is 's new representation, and the dot product of it and the learnable vector v reflects the correlation between w i and d j . l i can be calculated by the following formula:

[0231]

[0232] In the formula:

[0233] is the word embedding in the document containing the i-th word;

[0234] d is the document containing the i-th word;

[0235] l i is the corresponding new representation;

[0236] n is the number of new representations of the word embeddings in H w ;

[0237] s is the serial number of the new representation of the word embeddings in H w ;

[0238] l s is the s-th new representation in the set of new representations of the word embeddings in H w ;

[0239] is the transposed vector of l i ;

[0240] is the transposed vector of l s ;

[0241] v is the learnable vector of the attention mechanism;

[0242] E is the weight of the neuron;

[0243] b is the bias of the neuron;

[0244] exp() represents the power function of the base e of the natural logarithm;

[0245] tanh() represents the hyperbolic tangent function;

[0246] E and b are the learnable parameters of the linear transformation;

[0247] l i The dot product of l and v reflects the correlation between w i and d j ;

[0248] This model integrates four loss functions: (1) word reconstruction loss l w , (2) document reconstruction loss l d , (3) topic clustering loss l t , and (4) constraint loss l c , as follows:

[0249] Word reconstruction loss. Let be the reconstructed word embeddings. Define l w as follows:

[0250]

[0251] Document reconstruction loss. Let For the reconstruction embedding of the document, which is calculated as the weighted sum of the reconstruction topic embeddings That is: Where

[0252]

[0253] is the weight, obtained by calculating the cosine similarity between and and That is:

[0254] Based on Define l d as:

[0255]

[0256] This technical solution aims to better capture their semantic relationship by increasing the relevance between Z t and Z d That is:

[0257] Topic clustering loss. Introduce l t to enhance the importance of topic clustering in the latent space. Let be the cosine similarity between and Define another similarity between them as:

[0258]

[0259] Where c k represents the sum of the cosine similarities between and all words:

[0260]

[0261] In the formula:

[0262] is the similarity between and ; it is denser than the cluster;

[0263] c k is the sum of the cosine similarities between and all words;

[0264] k′ is the serial number of the topic embedding in Z t ;

[0265] c k′ is the sum of the cosine similarities between the k′-th topic embedding and all words.

[0266] Define lt For and the cross - entropy between:

[0267]

[0268] As can be seen from formula (9), will make sharper, and if t k is particularly similar to w i then its value will be larger. Therefore, l t draws w i closer to its most similar topic in the latent space, thus forming a clear topic clustering.

[0269] Constraint loss. w i is achieved by constructing a matrix. This matrix integrates all objects (words, documents, and topics). Suppose there are M documents, N words, and K topics, and the matrix size is (M + N + K)×(M + N + K). Initially, the matrix is all 0. Analysts can associate two objects by setting the value of the corresponding cell in the matrix to 1, or disassociate them by setting it to - 1. Let a ef be the value of the cell in the e - th row and f - th column of the matrix; e is the row index of the matrix; f is the column index of the matrix; l c is calculated as follows:

[0270]

[0271] In the formula:

[0272] z e is the embedding of the object in the e - th row;

[0273] z f is the embedding of the object in the f - th row;

[0274] u is the number of all embeddings.

[0275] The bootstrapped neural topic model of the present invention supports multiple topic constraints TC1 to TC10, including the topic constraints recommended by existing research, specifically as follows:

[0276] (TC1) Add a word to a topic. Let t k be a topic and w x be a word related to t k . I(w x ) and I(t k ) are their indices in the matrix. By setting add w x to t k .

[0277] (TC2) Remove a word from a topic. Let t k be a topic and w y be an irrelevant high-ranking word in t k . I(w y ) and I(t k ) are their indices in the matrix. Remove w from t y by setting k .

[0278] (TC3) Add a document to a topic. Let t k be a topic and d m be a document related to t k ; I(d m ) and I(t k ) are their indices in the matrix; add d to t m by setting k .

[0279] (TC4) Remove a document from a topic. Let t k be a topic and d n be an irrelevant high-ranking document in t k ; I(d n ) and I(t k ) are their indices in the matrix; remove d from t n by setting k .

[0280] Four other topic constraints can be directly controlled by Z w , Z d and Z t as follows:

[0281] (TC5) Delete a word. Let w u be a meaningless word; delete the word from the corpus by removing w from Z .

[0282] (TC6) Delete a document. Let d u be a relatively ambiguous document; delete the document from the corpus by removing d from Z .

[0283] (TC7) Delete a topic. Let t q be a low-quality topic; delete the topic from the corpus by deleting t from Z .

[0284] (TC8) Add a topic. Let t c be a new topic, and c initialize its embedding vector; by using the embedding vector of any object related to t and then place into Z t so as to add t c to the corpus.

[0285] Use the above topic constraints to implement two other relatively more complex topic constraints, specifically as follows:

[0286] (TC9) Merge topics. Let t p and t a be two related topics, and be their embedding vectors; by using the average value of and to initialize merge t p and t a into a new topic t k ; then add to the corpus and remove t p and t a from the corpus.

[0287] (TC10) Split a topic. Split a topic; let t d be a semantically rich topic; split t e into two topics by creating a new topic t d where t e contains one semantic aspect of t d ; the embedding vector of the object (word or document) most semantically relevant should be used to initialize and add the relevant objects (words or documents) in t d to t e .

[0288] (II) Visualization Editing System

[0289] The visualization editing system proposes multiple functional modules based on the joint embedding of vocabulary, documents, and topics. These functional modules not only fully utilize the potential of the joint embedding but also provide clear guiding principles for the system technical solutions to better present the semantic relationships and structural characteristics of the data.

[0290] The visual editing system includes: a topic quality module for evaluating the quality of topics, a semantic relevance module for evaluating the semantic relevance of objects, an embedded object color configuration module for assigning a unique color to each object and reflecting the semantic relationship between objects through colors, a topic editing module for assisting users in editing topic content, an auxiliary function module for helping users complete topic editing tasks more efficiently, and a constraint operation icon sub-module.

[0291] The present invention can jointly learn the embedded representations of vocabulary, documents, and topics and support the following operations for facilitating topic editing:

[0292] A. A topic quality module for evaluating the quality of topics, which quantifies the topic quality based on the compactness of vocabulary embedding clustering.

[0293] SNTM generates topics by clustering vocabulary embeddings. Therefore, the quality of topics can be quantified according to the compactness of clustering. The more compact the clustering, the more relevant the involved vocabulary, and the higher the quality of the topic. This feature is of great significance for determining the target topics that need to be optimized and evaluating the editing effect. Analysts can select low-quality topics for editing or evaluate the editing effect by comparing the changes in topic quality before and after editing. For example, when the vocabulary distribution of a certain topic is relatively scattered, it may indicate that the topic lacks a clear semantic orientation and thus needs further optimization. In this way, the editing process becomes more data-driven and transparent. In addition, the compactness of this clustering also provides strong support for the automatic adjustment and optimization of the model, helping users efficiently screen out the content that needs to be focused on among a large number of topics and predicting the potential impact of editing operations on topic quality based on historical data, thus providing more decision-making basis for analysts.

[0294] B. A semantic relevance module for evaluating the semantic relevance of objects, which measures the semantic relevance of objects based on the proximity between embeddings.

[0295] The proximity between embeddings can reflect the semantic relevance between objects. Utilizing this feature, various topic constraints can be executed more efficiently. For example, by calculating the relevance between high-ranked vocabulary or documents in a certain topic and the topic, analysts can quickly identify irrelevant objects (TC2 and TC4). In addition, based on the proximity of embeddings, several objects most relevant to the selected object can be recommended to analysts (such as TC1 and TC3), thus avoiding the need for analysts to search among a large number of candidate objects. This mechanism not only improves the editing efficiency but also significantly reduces the complexity of manual operations, making the topic editing process more intelligent. This proximity measurement is not limited to vocabulary but can also be extended to articles, paragraphs, or even higher-level text structures, further enhancing the flexibility and diversity of topic editing.

[0296] C. Embedded object color configuration module, which assigns unique colors to each object by projecting the embedding into a color space.

[0297] The embeddings of all objects can be projected into an RGB color space, and the color corresponding to the embedding position of each object is assigned to it (the three-dimensional coordinates after projection are used as the R, G, and B values of the color), as Figure 2 shown. The similarity of object colors can reflect their correlation. This technical solution can assign a unique color to each object and, at the same time, reflect the semantic relationship between objects through colors. Therefore, a large number of colors can be used without visual chaos, which is particularly crucial for a topic editing system that needs to display a large number of objects and their relationships simultaneously. For example, analysts can intuitively distinguish the vocabulary, documents, and their semantic associations of different topics by color, so as to quickly locate the targets that need to be optimized. In addition, color mapping also provides intuitive visual feedback to help analysts identify potential problems in real time during the editing process, such as which vocabulary or documents have too large a semantic distance and need to be further adjusted or deleted. In this way, analysts can make more accurate decisions in a shorter time.

[0298] Generally speaking, these technical solutions provide intuitive and efficient support for the topic editing system, fully utilize the potential of the joint embedding model, and significantly improve the efficiency of topic analysis and optimization.

[0299] D. Topic editing module

[0300] The visualization editing system includes a topic editing module for assisting users in editing topic content. The topic editing module consists of three core visualization components, namely the topic list sub-module 2, the word editing sub-module 14, and the document editing sub-module 15, as Figure 3 shown. These components are closely combined to jointly support the efficient editing of topics. This section will detail the main technical solution concepts, functional characteristics of these components, and how to use them to implement all topic constraints, including simple operations on single objects (vocabulary, documents, or topics) and more complex constraints such as topic merging and splitting.

[0301] D1. The topic list sub-module 2 is located on the left side of the interface and consists of multiple line charts arranged vertically, as Figure 3 shown. The technical solutions of these line charts are intuitive and concise. Each line chart corresponds to a topic and shows the change trend of the quality score of the topic in different editing rounds. Through this technical solution, analysts can clearly observe how the quality of the topic changes over time and with editing operations at a glance, thus having a clearer understanding of the evolution of the topic quality. The horizontal axis of the line chart represents the editing round, while the vertical axis represents the quality score of the topic. The fluctuations of the curve reflect the impact of editing on the topic.

[0302] The theme list sub-module 2 provides a series of interactive functions to support analysts in performing various theme-related operations, such as theme constraint operations, specifically including the following four types:

[0303] Create a new theme (TC7): Analysts can click the "+Topic" button 11 and then select an object (such as a word or a document) from the interface to create a new theme based on that object. This function is applicable to scenarios where new themes need to be introduced quickly. For example, when analysts find that some words or documents do not match the semantics of existing themes, they can group them into an independent theme through this operation.

[0304] Delete a theme (TC8): Analysts can click the "×Topic" button 12 and then select a certain line chart to remove the corresponding theme from the corpus. This function is usually used to remove low-quality or redundant themes. For example, when the quality score of a certain theme remains low, or the theme content is irrelevant to the analysis goal.

[0305] Merge themes (TC9): Analysts can click the "∨Topic" button 9 and then select the themes corresponding to multiple line charts to merge these themes into a new theme. This function is suitable for handling themes with overlapping semantics or similar content. For example, when the core content of two themes is highly similar, the theme structure can be simplified through the merge operation, thereby improving the analysis efficiency.

[0306] Split a theme (TC10): Analysts can click the "∧Topic" button 10 and then select an object (such as a certain keyword or document) from a theme as a basis to create a new theme and split the original theme into two themes. This function is applicable to scenarios where semantic-rich and content-complex themes need to be further refined. For example, when a theme contains multiple different semantic levels, the split operation can be used to more accurately express each semantic level.

[0307] Through these interactive functions, the theme list sub-module 2 not only provides powerful editing capabilities, but also helps analysts manage themes efficiently. At the same time, the visual display method makes the analysis process more intuitive. In short, the theme list sub-module is an indispensable and important part of the system. It not only undertakes the role of information display, but also provides rich theme editing support, making the analysis work more convenient and efficient.

[0308] D2. The word editing sub-module 14 is located on the right side of the interface and is used to display the word similarity matrix of the selected theme, such as Figure 3As shown. This matrix intuitively reveals the impact of word pairs on the quality of the theme, helping analysts to more deeply understand the internal structure and relevance of the theme. The technical solution of the matrix is based on high-ranking words. The horizontal axis represents the top 10 high-ranking words of the theme, and the vertical axis represents the top 50 high-ranking words. These words are sorted in decreasing order of relevance to the theme. From left to right on the horizontal axis and from top to bottom on the vertical axis, the relevance gradually decreases.

[0309] Each cell in the matrix is represented by a circular marker indicating the cosine similarity between two words, and its size intuitively reflects the relevance: the larger the circle, the higher the relevance between the corresponding words. Therefore, the size distribution of these circles reveals the quality of the theme and which word pairs have a significant impact on the quality of the theme. A closely distributed large circle usually indicates a relatively high semantic consistency within the theme, while a loosely distributed small circle may imply a low theme quality or the existence of words that are semantically unrelated to the theme, thereby affecting the overall theme effect.

[0310] The word editing sub-module 14 provides three operation constraints related to words (i.e., word constraints):

[0311] Adding a word to a theme (TC1): Analysts can click the "+Word" button 3 and then drag a word to a certain theme to add the word to the theme. This operation is applicable to situations where it is necessary to expand the theme semantics or supplement new words to improve the theme quality.

[0312] Removing a word from a theme (TC2): Analysts can click the "-Word" button 4 or the "×Word" button 5 and then select a word to remove it from the theme. This operation is usually used to clean up high-ranking words that are semantically unrelated to the theme, thereby improving the semantic purity of the theme.

[0313] Deleting a word from the corpus (TC5): Analysts can also click the "×Word" button 5 to directly delete a certain word from the corpus. This function is applicable to removing meaningless or noise words, such as stop words that appear frequently but are worthless for analysis.

[0314] Through these functions, the word editing sub-module 14 not only provides analysts with more fine-grained word editing capabilities, but also helps them quickly identify and adjust the keywords of the theme through the visual similarity matrix, thereby significantly improving the quality and relevance of the theme.

[0315] D3. The document editing sub-module 15 shares the interface space with the word editing sub-module 14, displaying a column of vertically arranged documents, whose relevance gradually weakens from top to bottom, as Figure 3 shown. Analysts can switch to the word editor or the document editor through the interface control, as Figure 3As shown. The document editor defaults to showing the first sentence of each document. Analysts can click on a document to expand and view the full text, thus saving display space and effectively showing more documents. High-ranking words are also highlighted throughout the full text to help analysts quickly grasp the semantic structure of the documents.

[0316] The document editing sub-module 15 provides three types of operation constraints related to documents (i.e., document constraints):

[0317] Adding a document to a theme (TC3): Analysts can click the "+Doc" button 6 and then drag a document into a certain theme to add the document to the theme. This function helps introduce new information related to the theme and enhance the semantic depth and coverage of the theme.

[0318] Removing a document from a theme (TC4): Analysts can click the "-Doc" button 7 or the "×Doc" button 8 and then select a document to remove it from the theme. This operation is applicable for cleaning up documents that are semantically irrelevant to the theme and improving the semantic relevance of the theme.

[0319] Deleting a document from the corpus (TC6): Analysts can click the "×Doc" button 8 to directly delete a certain document from the corpus. This operation is usually used to remove irrelevant, low-quality or useless documents, thereby reducing noise and improving the accuracy of topic modeling.

[0320] Through these functions, the document editing sub-module 15 provides analysts with powerful document editing capabilities, helping them precisely control the document composition of the theme and improve the quality and semantic relevance of the theme.

[0321] E. Auxiliary function module

[0322] Auxiliary function modules are adopted to help analysts complete the theme editing tasks more efficiently. The technical solutions of these auxiliary function modules are designed to simplify the operation steps, enhance the user experience, and ensure the accuracy and scientific nature of the editing.

[0323] E1. Process management sub-module. The theme list sub-module 2 is one of the core components in the whole system responsible for managing the editing process. In the initial state, the theme list sub-module 2 is blank. Users need to first run the SNTM model to generate initial themes. At this time, the model runs without any constraint loss, which is used to provide a basic version for editing. The generated set of initial themes will be displayed in the theme list sub-module 2 for analysts to select and operate on.

[0324] During the actual editing process, analysts usually prefer to optimize themes with lower scores first, because these themes may contain more obvious redundant semantics or lower semantic consistency. The process management sub-module provides an important auxiliary function - the rollback function. When analysts find that a certain constraint operation has had an adverse impact on the quality scores of multiple themes, they can use the rollback function to quickly restore all themes to the state of a certain round, thereby reducing the impact of irreversible incorrect operations on the entire theme set.

[0325] To manage the editing process more scientifically, we propose a method for quantifying the quality scores of themes. Let t O represent a certain theme, and its quality score s O is calculated as follows:

[0326]

[0327] In the formula:

[0328] z wp represents the word embedding of the current cluster;

[0329] wp represents the word of the current cluster;

[0330] C O represents the word cluster associated with theme t O ;

[0331] represents the theme embedding of the current cluster.

[0332] cos(·) represents the cosine similarity between two embedding vectors, and the larger the value, the higher the semantic relevance. By calculating this score, analysts can intuitively judge the semantic tightness of each theme and select the target themes that need to be optimized accordingly.

[0333] The object recommendation sub-module recommends other objects that are most relevant to the selected object (such as words, documents, or themes) to analysts in an automated manner. The core principle of the function of the object recommendation sub-module is based on the semantic similarity of embedding vectors, that is, the cosine similarity between embedding vectors is calculated to measure the semantic relevance between two objects.

[0334] The object recommendation sub-module helps analysts complete operations more quickly and accurately by highlighting relevant objects, reducing the time cost of manually searching among a large number of candidate objects. Figure 3 The specific highlighting effects of the object recommendation sub-module in three visualization components are shown as follows:

[0335] (1) When an analyst selects a certain theme in the theme list sub-module 2, the system will highlight four other themes in the theme list sub-module 2 that are most relevant to this theme. At the same time, 50 high-ranked words related to this theme will be displayed in the word editing sub-module 14, and 50 high-ranked documents related to this theme will be displayed in the document editing sub-module 15.

[0336] (2) When an analyst selects a certain word in the word editing sub-module 14 or the document editing sub-module 15 (such as a colored word in the full text in the document editing sub-module 15), the theme list sub-module 2 will highlight four themes that are most relevant to this word. The word editing sub-module 14 will highlight ten other words that are most relevant to this word, and the document editing sub-module 15 will highlight ten documents that are most relevant to this word.

[0337] (3) When an analyst selects a certain document in the document editing sub-module 15, the theme list sub-module 2 will highlight four themes related to this document. The word editing sub-module 14 will highlight ten words related to this document, and the document editing sub-module 15 will highlight ten other documents related to this document.

[0338] The technical solution of the object recommendation sub-module significantly narrows the search scope, helps analysts quickly locate relevant objects, and thus more efficiently perform various constraint operations. For example, analysts can select a target theme to check whether the other themes recommended by the system contain redundant semantics, so as to decide whether to delete redundant themes (TC7). Similarly, they can also select a word or document related to the current theme, check whether the recommended objects need to be added to the theme (TC1, TC3), or judge whether the un-recommended objects need to be removed from the theme (TC2, TC4), or even deleted from the entire corpus (TC5, TC6). In addition, analysts can also select a target word or document and perform semantic expansion according to the most relevant themes recommended, and add it to the theme (TC1, TC3).

[0339] In summary, the auxiliary function significantly improves the efficiency and accuracy of theme editing by providing scientific calculation methods and efficient interaction technical solutions, enabling analysts to more easily complete complex theme adjustment tasks.

[0340] The color application sub-module. The color application sub-module utilizes color mechanisms in multiple technical solutions (see Figure 2 ) to reduce visual clutter and enhance the perception effect of analysts, as follows:

[0341] First, each line chart in the theme list sub-module 2 contains two ribbon areas, as Figure 3As shown in the figure. The upper strip area consists of 25 small rectangles, each corresponding to a high-ranking word of the topic and showing the color of that word. The lower strip area shows the color of the topic. These two strip areas can visually reflect the semantic differences between high-ranking words and the quality of the topic. For a high-quality topic, the colors of its rectangular bars should be consistent.

[0342] Secondly, each circle in the matrix of the word editing sub-module 14 contains two colors, as Figure 3 shown. The upper right part of the circle shows the color of the word represented by the column, while the lower left part shows the color of the word represented by the row. These two colors can reveal the semantic correlation between the two words. The higher the correlation, the more consistent the two colors are. This technical solution can help analysts quickly judge the semantic connection between word pairs.

[0343] Finally, in the document editing sub-module 15, the document number in front of each line shows the color of that document, and each high-ranking word in the expanded full text is also highlighted in the color of that word, as Figure 3 shown. This technical solution can visually show which words associate the document with the topic and their degree of relevance to the document, thus helping analysts better understand the connection between the document and the topic.

[0344] Constraint operation icons 16 (TC Logo), multiple graphic icons may be shown at the end of each line in the word editing sub-module 14 and the document editing sub-module 15, as Figure 3 shown. Each icon represents a constraint operation performed on the object of that line. Figure 4 A technical solution showing all six icons is presented. Specifically: The first two icons respectively indicate that the object is added to the current topic or removed from the current topic; the middle two icons indicate that the object is added to or removed from other topics; the fifth icon indicates that the object has been deleted from the corpus; the last icon indicates that the embedding of the object is used to initialize a new topic. Analysts can click on any icon to undo the corresponding constraint operation. This technical solution provides a convenient backtracking function for analysts, avoiding unnecessary error accumulation, while maintaining the transparency and controllability of the entire editing process.

[0345] The above-mentioned functional modules such as the encoder, decoder, clustering module, topic quality module, semantic correlation module, topic list sub-module, word editing sub-module, document editing sub-module, process management sub-module, object recommendation sub-module, color application sub-module, constraint operation icon sub-module, etc. can adopt applicable functional modules in the prior art, or adopt functional modules in the prior art and be constructed by conventional technical means.

[0346] The embodiments described above are only used to illustrate the technical idea and characteristics of the present invention. The purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. The patent scope of the present invention cannot be limited only by these embodiments. That is, any equivalent changes or modifications made in accordance with the spirit disclosed by the present invention still fall within the patent scope of the present invention.

Claims

1. A manipulable theme system based on joint learning of semantic latent space, characterized in that, The system includes a visual editing system and a bootable neural topic model based on an autoencoder; The bootable neural topic model includes an encoder, a decoder, and a clustering module. The encoder is used to input the embeddings of words and documents and generate corresponding embedding representations. The clustering module is used to aggregate the embedding representations of words to generate topic embedding representations; the decoder is used to reconstruct the information of the generated word, document, and topic embedding representations, so that the latent space captures the structural features of the data distribution; The visual editing system sets up human-computer interaction components and quantization metrics, which perform quality scoring on the topics generated by the bootable neural topic model, and support users to add or delete the association constraints between words, documents, and topics based on the quality scores of the topics during the process of training the bootable neural topic model, so that the bootable neural topic model generates topics that meet the user's needs.

2. The manipulable topic system based on joint learning of semantic latent space according to claim 1, characterized in that The visual editing system includes: a topic quality module for evaluating the quality of topics and / or a semantic relevance module for evaluating the semantic relevance of objects; the topic quality module quantifies the quality of topics according to the compactness of clustering; the semantic relevance module measures the semantic relevance of objects based on the proximity between objects.

3. The manipulable topic system based on joint learning of semantic latent space according to claim 1, characterized in that The visual editing system includes an embedded object color configuration module for assigning a unique color to each object and reflecting the semantic relationship between objects through colors; the embedded object color configuration module maps the embedding positions of objects to the color space and assigns a unique color code to each object.

4. The manipulable topic system based on semantic latent space joint learning according to claim 3, wherein The color space is the RGB color space. The embedded object color configuration module takes the three-dimensional coordinate values of the embedding positions of each object as the R, G, and B values of the color respectively, and assigns the corresponding color to it.

5. The manipulable topic system based on joint learning of semantic latent space according to claim 1, characterized in that, The visual editing system includes a topic editing module for assisting users in editing topic content. The topic editing module includes a topic list sub-module, a word editing sub-module, and / or a document editing sub-module; the topic list sub-module is used to display the generated topics in the form of a list for users to select, manage, and edit topics. It consists of multiple line charts, each line chart corresponding to a topic, showing the change trend of the quality scores of the topic in different editing rounds; the word editing sub-module is used to display the word similarity matrix of the selected topic to assist users in evaluating the topic quality and editing related words; the document editing sub-module is used to display the list of documents related to the selected topic, support document expansion and keyword annotation, and assist users in understanding the document semantics and performing related editing.

6. The manipulable topic system based on semantic latent space joint learning according to claim 5, wherein The topic editing module further includes a constraint operation icon sub-module. When the topic editing module performs a topic editing operation, the constraint operation icon sub-module uses corresponding icons to represent a constraint operation performed on the object being operated on.

7. The manipulable topic system based on semantic latent space joint learning according to claim 5, characterized in that, The visual editing system includes an auxiliary function module for helping users complete topic editing tasks more efficiently. The auxiliary function module includes a process management sub-module, an object recommendation sub-module, and / or a color application sub-module; The process management sub-module sets up a rollback function to quickly restore all topics to the state of a certain round through the rollback function; the object recommendation sub-module recommends other objects most relevant to the selected object to the user and highlights the recommended objects; the color application sub-module adopts a color mechanism in the topic list sub-module, the word editing sub-module, and the document editing sub-module to intuitively reflect the ranking of objects by coloring relevant objects.

8. An optimization method for a manipulable topic system using the semantic latent space joint learning-based system according to any one of claims 1 to 7, characterized in that The bootstrapped neural topic model maps the semantic relationships of words, documents, and topics into a unified latent space through joint embedding, making semantically related objects closer in the latent space; it sets constraints and constraint loss functions to optimize topics. Let: w i represent the i-th word, where i is the word sequence number; i = 1, 2, …, N; W = [w1, w2, ..., w N is a set containing N words, and d j represents the j-th document; j is the document sequence number; j = 1, 2, …, M; D = [d1, d2, ..., d M is a set containing M documents; N represents the number of words; M represents the number of documents; The encoder uses the BERT model to extract the embedding features of each word in the document. Let represent the embedding of the i-th word extracted in the j-th document; Let represent the i-th word embedding, which is characterized by the average value of the embeddings of the i-th word in all documents containing the i-th word; Let D i be the number of documents containing the i-th word; The set of word embeddings is H w , The calculation formula is as follows: In the formula: is the word embedding in the document containing the i-th word; d is the document containing the i-th word; Let be the j-th document embedding; the set of document embeddings is H d , For any document, an attention mechanism is introduced to calculate The calculation formula of is as follows: where a i represents the attention weight corresponding to w i and a i is calculated by the following formula: In the formula: l i For correspondence with the new representation n is the number of new representations of the word embeddings corresponding to H w in the middle; s corresponds to H w The serial number of the new representation of the word embedding; l s corresponding to H w the s-th new representation in the new representation set of word embeddings; is the transposed vector of l i ; is the transposed vector of l s ; v is the learnable vector of the attention mechanism; E is the weight of the neuron; b is the bias of the neuron; exp() represents the power function of the base e of the natural logarithm; tanh() represents the hyperbolic tangent function; E and b are the learnable parameters of the linear transformation; l i The dot product of l and v reflects w i and d j is relevant to; Let: be the i-th word embedding representation; the set of word embedding representations is Z w , is the embedding representation of the j-th document; the set of document embedding representations is Z d , is the k-th topic embedding representation; k = 1, 2, …, K; the set of topic embedding representations is Z t , K is the number of topic embedding representations; k is the serial number of the topic embedding representation; t k is the k-th topic; Among them, Z w and Z d are directly output by the encoder, while Z t is obtained by performing K-means clustering on Z w ; each cluster corresponds to a theme, The calculation formula of Among them, B k represents the k-th cluster; w represents the word included in B k ; z w represents the word embedding corresponding to w The bootstrapped neural topic model includes the following four losses: word reconstruction loss, document reconstruction loss, topic clustering loss, and constraint loss; Let the word reconstruction loss be \(l\). w Let the document reconstruction loss be \(l\). d Let the topic clustering loss be \(l\). t Let the constraint loss be \(l\). c ; Let be the set of reconstructed word embeddings, and be the word reconstruction embedding corresponding to l w The calculation formula is as follows: Let be the topic reconstruction embedding corresponding to the output of the decoder ; the set of topic reconstruction embeddings is Let be the document reconstruction embedding corresponding to , which is the weighted sum of , i.e.: In the formula: For and cosine similarity; l d The calculation formula is as follows: By setting l d to increase Z t and the relevance of Z d to better capture the semantic relationship between Z t and Z d ; Introduce l t to enhance the importance of topic clustering in the latent space; Let be and 's cosine similarity, define and Another similarity between them is: In the formula: For and similarity; compared with the clusters are denser; c k is the sum of the cosine similarities with all words; k′ is the topic embedding serial number in Z t ; c k′ is the sum of the cosine similarities between the k'-th topic embedding and all words; Define l t as the cross entropy between, then there is: will make sharper, and if t k is more similar to w i then its value will be larger; therefore, l t will bring w closer to its most similar topic in the latent space, thus forming a clear topic clustering; i ​ Integrate all objects: words, documents, and topics by constructing a matrix; assume there are M documents, N words, and K topics, and the size of the matrix is (M + N + K) × (M + N + K); initially, the matrix is all 0; the user associates two objects by setting the value of the corresponding cell in the matrix to 1, or disassociates them by setting it to -1; let a ef be the value of the cell in the e-th row and f-th column of the matrix; e is the row index of the matrix; f is the column index of the matrix; l c The calculation formula is as follows: In the formula: z e The embedding of the object in the e-th row; z f Embedding for the object of the f-th row; u is the number of all embeddings.

9. The optimization method of the manipulable topic system based on semantic latent space joint learning according to claim 8, characterized in that The set constraints include one or several combinations of the following constraints: The first constraint: Add a word to the topic; Let t k be a topic, and w x be a word related to t k ; I(w x ) and I(t k ) are their indices in the matrix; By setting , add w x to t k ; Second constraint: Remove a word from the topic; Let t k be a topic, w y be an irrelevant high-ranking word in t k ; I(w y ) and I(t k ) are their indices in the matrix; By setting , remove w y from t k ; The third constraint: adding a document to a topic; let t k be a topic and d m be a document related to t k ; I(d m ) and I(t k ) are their indices in the matrix; by setting add d m to t k ; Fourth constraint: Remove a document from the topic; Let t k be a topic and d n be an irrelevant high-ranked document in t k ; I(d n ) and I(t k ) are their indices in the matrix; Remove d from t n by setting k ; Direct control of Z w , Z d and Z t to implement the following fourth to eighth constraints: Fifth constraint: Delete a word; Let w u be a meaningless word; Delete the word from the corpus by removing it from Z w from ; The sixth constraint: deleting a document; let d u be a relatively ambiguous document; by removing it from Z d to delete the document from the corpus; ​ The seventh constraint: deleting a topic; let t q be a topic with low quality; by deleting t from Z to delete the topic from the corpus; Eighth constraint: Add a topic; Let t c be a new topic, and its embedding vector; Initialize c by using the embedding vector of any object related to t and then put into Z t to add t c to the corpus; The above constraints are used to implement another two relatively more complex constraints, specifically as follows: The ninth constraint: merging topics; let t p and t a be two related topics, and be their embedding vectors; by using and initialize the average value of Merge t p and t a into a new topic t k ; then add to the corpus and remove t p and t a from the corpus; Tenth constraint: split the theme; let t d be a semantically rich theme; split t e into two themes by creating a new theme t d where t e contains a semantic aspect of t d ; the embedding vectors of the objects most relevant semantically should be used to initialize and the relevant objects in t d should be added to t e .

10. The optimization method of the manipulable topic system based on semantic latent space joint learning according to claim 8, characterized in that, The user views the initially generated topics and their quality scores through the visual editing system and selects the topics that need to be optimized; button components are set up, and the user adds or deletes the association constraints between words, documents, and topics and cancels unreasonable editing operations by operating the corresponding button components; the visual editing system updates the quality scores of the topics in real time.