Content topic analysis method based on large language model

Through the content topic analysis method based on the large language model, the shortcomings of the existing technology in extracting fine-grained topics and combining network hot topics are solved, and efficient and accurate topic extraction and abstract generation of dialogue content are achieved.

CN120144753AActive Publication Date: 2025-06-13NAT COMP NETWORK & INFORMATION SECURITY MANAGEMENT CENT

Patent Information

Application Number
CN202411984631.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-06-13
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

The prior art is difficult to accurately extract fine-grained topics when dealing with dialogue content, and fails to effectively combine the latest online news or public opinion hotspots, resulting in insufficient clear topic extraction and lack of summary overview of conversation content topics.

Method used

A content topic analysis method based on a large language model is proposed. By pre-processing the conversation content, the discussion objects, abbreviations, key phrases, etc. are extracted, and topic classification and summary are combined with the BERT model to form topic names and abstracts, and topic analysis report on the discussion content is generated.

Benefits of technology

It realizes efficient and accurate topic extraction of dialogue content, forms clear topic reports, and accurately captures important information in rapidly changing discussion topics in the Internet era.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144753A_ABST
    Figure CN120144753A_ABST
Patent Text Reader

Abstract

The invention discloses a topic analysis method based on a large language model. The method comprises the following steps: performing data preprocessing on dialogue content; performing topic classification on the preprocessed dialogue contents, and collecting contents belonging to the same topic together; discussion objects, abbreviations and key phrases appearing in the content are extracted, and additional background knowledge is collected; summarizing contents under the topic by using a large language model to form a topic name and a topic abstract; and performing the same processing on a plurality of topics in the dialogue process to form a topic analysis report of the discussion content. Important information can be efficiently and accurately extracted, and a topic report is formed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to speech recognition, and in particular to a content topic analysis method based on a large language model. Background Art

[0002] Benefiting from the continuous innovation and development of information technology, online social network applications have rapidly spread around the world and become an important way for people to communicate in daily life. They have also generated a large amount of interactive data, which has the characteristics of many participating users, fast information update frequency, and a lot of useless information. In particular, there are multiple topics entangled, that is, in a continuous message, the content belonging to different topics will appear alternately. Therefore, timely and accurate acquisition of important topic information in interactive massive data has gradually become a hot topic of research at home and abroad. To achieve topic extraction, topic segmentation is usually required, and the key to segmentation is how to correctly judge the topic attribution of a conversation content, and then form a topic name and summary based on the content summary. When someone joins the conversation discussion, there is no need to check all the historical content one by one, only to check the topic analysis report, which can greatly improve the efficiency of information acquisition. In addition, it can also be used to assist in finding topics of interest.

[0003] At present, the commonly used methods for extracting conversation text topics are as follows:

[0004] 1. Topic detection technology based on multiple strategies

[0005] Wu Xu et al. proposed a multi-strategy topic detection technology for conversation content based on conversation content and auxiliary information such as user, time, and type. The topic sequence is constructed to solve the problem of topic intersection, and the auxiliary information is used to reduce the impact of sparse short text features on clustering effects, thus realizing topic detection for continuous conversation content mixed with multiple types of messages.

[0006] First, based on the context of the conversation, possible topics for discussion are identified and sorted through topic sequences, so that new conversation content can be matched to its topic with a higher probability, thereby solving the problem of overlapping and parallel topics.

[0007] There are two types of topics in the topic sequence: current topics and expired topics. Expired topics are topics that are eliminated from the current topic sequence and are no longer updated. New conversation content will not be added to these topics, and they will be stored as historical data; the current topic sequence is divided into a common topic sequence and a hot topic sequence. New conversation content will be added to the topics in these two sequences, or new topics will be opened in these two sequences.

[0008] Secondly, due to the short length of the dialogue content text, it has the disadvantage of sparse features. It is difficult to improve the topic detection performance solely relying on semantic information. Moreover, the continuity of the dialogue and the integrity of the topic participation group are also crucial for the analysis of the discussion topic. Therefore, a topic detection method assisted by artificial experience is proposed. There are mainly two types of auxiliary information. One is the type of the dialogue content. For text-type messages, combined with Chinese grammar habits, the following text features can be summarized to determine non-topic starting messages: the sentence starts with a conjunction or an adverb; the sentence ends with a specific modal particle such as "ba"; the sentence uses demonstrative pronouns containing characters such as "ni", "ta", "zhe", etc.; it does not contain pronouns and nouns. Therefore, in the clustering process, a message m with one of the above features in the text content should not be regarded as a topic starting sentence. The other is the time and user of the speech. According to the principle that "within a certain time, the content discussed by a user's speech is very likely to belong to the same topic", the topic heat detection time Ht is taken. When the new dialogue content is a meaningless message or the text content is not similar enough to any topic, if the dialogue content of the same user within the Ht time is found, the new dialogue content will be added to the topic to which the latest message that meets the above conditions belongs.

[0009] Finally, the proposed topic detection algorithm SPTSAI (Single-Pass Using Topic Sequence and Auxiliary Information) solves the topic intersection problem by constructing a topic sequence, and comprehensively uses text features and auxiliary information such as time, user, and message type to improve the clustering effect.

[0010] During the clustering process, the semantic similarity calculation mainly occurs between messages and topics, and between topics and topics. The latter is to address the problem that the Single-Pass clustering algorithm is prone to form small clusters, that is, to avoid the topic segmentation granularity being too small due to the sparse features of short texts. The two similarity calculation methods are as follows.

[0011]

[0012] 2. Topic Mining Model for Dialogue Content Based on BERT-LDA

[0013] The BERT-LDA model first merges the context of the historical dialogue content, retains the context relationship, then uses the BERT model to extract the semantic features of each merged dialogue content, and then uses the K-means model for text clustering. In this way, the dialogue content in each clustering cluster has a certain semantic similarity. Then, preprocessing operations such as word segmentation and stop word removal are performed on each dialogue content to obtain the bag-of-words representation of each dialogue content. Denote as the j-th dialogue content in the i-th clustering cluster, and merge the words included in each clustering cluster to form the bag-of-words representation of each clustering cluster Finally, when solving the topic distribution, different from the LDA topic model that samples , the BERT-LDA model samples W i , thus expanding the vocabulary contained in a single document and effectively solving the parameter estimation problem that occurs when the LDA model performs topic modeling on dialogue texts. On this basis, the semantic relationship between words is fused into the topic modeling. First, the external corpus is trained through the Word2Vec model to obtain the word vector representation of each word. After that, the cosine distance can be used to calculate the semantic similarity between any two words. Finally, when iteratively solving the document-topic distribution and the topic-word distribution, a double-layer semantic enhancement algorithm is used to increase the distribution probability of semantically similar words under the same topic. At the same time, the model also performs semantic enhancement on the document-topic distribution by calculating the semantic similarity between the word and the document as the semantic similarity between the corresponding topic and the document, and enhancing the distribution probability of the corresponding topic in the document. Its model structure is as Figure 3 shown:

[0014] 1. Disadvantages of the prior art:

[0015] Disadvantage 1: The existing BERT-LDA technical method is relatively simple when preprocessing historical dialogue content. It directly merges the context of historical dialogue content and then aggregates the dialogue content with semantic similarity through clustering. This method does not consider the situation where multiple topics may cross in a dialogue. When the clustering operation is completed, the topics in the clustering clusters will also be very messy, and this method requires a relatively large amount of dialogue data. Otherwise, it is impossible to extract representative topics, that is, this method cannot solve the fine-grained topic extraction in a single dialogue scenario. In today's Internet era, the topic changes very fast, and the discussion topic may change multiple times in a short period.

[0016] Disadvantage 2: The existing technical solutions do not consider combining the latest network news or public opinion hotspots when processing dialogue content messages. They may use some trendy words. Such dialogue content is difficult to extract clear topics without combining certain background information. In addition, the existing technical solutions do not give a summary of the topics of the dialogue content. Summary of the Invention

[0017] The object of the present invention is to propose a content topic analysis method based on a large language model to efficiently and accurately extract important information and form a topic report.

[0018] The technical solution for realizing the present invention is: A topic analysis method based on a large language model, including the following steps:

[0019] Perform data preprocessing on the dialogue content;

[0020] Classify the pre - processed conversation content into topics, and gather the content belonging to the same topic together;

[0021] Extract the discussion objects, abbreviations, and key phrases that appear in the content, and collect additional background knowledge;

[0022] Use a large - language model to summarize the content under the topic, forming a topic name and a topic summary;

[0023] Perform the same processing on multiple topics during the conversation to form a topic analysis report of the discussion content.

[0024] Furthermore, perform data pre - processing on the conversation content. The specific method is as follows:

[0025] Step 1.1, define the conversation content features:

[0026] Adopt a general speech - to - text engine to convert speech into text content, and extract the conversation content features, including the conversation content type, message reference, @ object, message interval time, conversation content index, and publisher activity. Among them, the conversation content type is divided into: text, picture, emoji, forward and share, and other types;

[0027] Step 1.2, feature processing of the conversation content:

[0028] Divide the 6 conversation content features into two categories. The conversation content type, interval time, conversation content index, and speaker activity need to be discretized; the message reference and @ object need to be texturized;

[0029] (1) Feature processing of the conversation content type

[0030] For the conversation content type, use a randomly initialized 50 - dimensional vector to represent each conversation content type;

[0031] (2) Feature processing of the interval time

[0032] For the interval time, first divide it into 7 levels: within half a minute, half a minute - 5 minutes, 5 minutes - 15 minutes, 15 minutes - 30 minutes, 30 minutes - 1 hour, 1 hour - 3 hours, and more than 3 hours. Then use a randomly initialized 50 - dimensional feature vector to represent the level of the message interval time;

[0033] (3) Feature processing of the conversation content index

[0034] For the conversation content index, adopt the initialization method of the position vector in the transformer method, and perform vector initialization according to the index of the conversation content in the entire conversation content list passed into the modeling window;

[0035] (4) Processing of speaker activity characteristics

[0036] Regarding the speaker activity, it is first divided into 4 levels: lurking, bubbling, active, and talkative. Then, a randomly initialized 50-dimensional feature vector is used to represent the levels of speaker activity;

[0037] (5) Processing of message reference characteristics

[0038] The referenced conversation content is also explicitly represented and processed into [Someone] posted [Something];

[0039] (6) Processing of @ object characteristics

[0040] The conversation content containing the @ object is processed into [Someone] @ [Someone], asking them to pay attention to the message [Something].

[0041] Furthermore, the preprocessed conversation content is classified by topic, and the content belonging to the same topic is gathered together. The specific method is as follows:

[0042] Step 2.1, adopt the strategy of a sliding window to complete the topic clustering of all conversation content;

[0043] Step 2.2, splice the consecutive conversation content. Each conversation content in the conversation content list is sequentially recorded as P 1 , P 2 , …, P 10 , …, where it is denoted that represents the k-th character of P i . Add [CLS] and [SEP] to the head and tail ends of each conversation content respectively to form the Input text string, which is expressed as follows:

[0044] Input = [CLS]P 1 [SEP][CLS]P 2 [SEP][CLS]P 3 [SEP][CLS]P 4 [SEP]……

[0045] Step 2.3, input the Input text string into the BERT model to obtain the deep semantic feature vector of each conversation content, and splice the deep semantic vector of each conversation content with the discrete feature vector to form the comprehensive feature vector of each conversation content;

[0046] Step 2.4, input the comprehensive feature vector of each conversation content into the bidirectional LSTM layer, and input the new vector output by the LSTM layer into the softmax layer for classification, so as to judge the topic attribution of each conversation content.

[0047] Furthermore, we extract the discussion objects, abbreviations, and key phrases that appear in the content and collect additional background knowledge. The specific method is as follows:

[0048] Step 3.1, obtain the background information of the elements:

[0049] First, for a topic that has been segmented, the large language model is used to extract the discussion objects, abbreviations, and key phrases that appear in the conversation content. The following prompt is used:

[0050] Suppose you are a language analysis expert and you need to deeply analyze the given conversation content data. The data discusses a topic event and may use some Internet hot words, abbreviations, homophones, etc. However, these words have a great influence on the semantic analysis. Your goal is to find these words in a reasonable and well-founded way.

[0051] Require:

[0052] 1. Before extracting a word, you need to think about whether you know the specific meaning of the word. If you know the exact meaning of the word, please ignore it;

[0053] 2. When you feel that you cannot accurately know the exact meaning of a word, you need to use an Internet search engine to find relevant information. This is the target word you need to extract;

[0054] 3. The words you extract need to occupy a relatively important semantic position in the conversation content. For some irrelevant words, just ignore them;

[0055] 4. In addition, you need to have an overall understanding of the core discussion topic, which will help you to determine the exact meaning of the words when searching on the Internet in the future;

[0056] The given conversation data is:

[0057] {chat}

[0058] Output Description

[0059] The output result must be in JSON format, with the following format: {"topic event":"","words to be confirmed":[{"words":"","explanation":""},...]}. Each word to be confirmed needs to explain why it needs to be confirmed. The topic event result must be brief and retain the core semantics.

[0060] let's do it step by step.

[0061] Notice

[0062] Answer all questions in Chinese, and be concise and to the point. Do not give meaningless or useless explanations. Directly ignore [emoji], [picture], [forward], etc.

[0063] Then, based on the large language model, for each of the core topic events obtained by combining the word extraction results and the conversation data, call an online search engine to obtain background knowledge documents related to the words.

[0064] Next, perform filtering based on the large language model, only retaining the relevant background knowledge text, specifically using the following prompt:

[0065]

【Please analyze the web page text returned by the online search engine and find the text fragments that are helpful for understanding the query keywords and the conversation content:

[0066] The web page text returned by the online search engine is: {text};

[0067] The query keywords input to the search engine are: {KeyWord};

[0068] The goal is to better understand the following conversation content: {chat};

[0069] Please directly provide the text fragments in the web page text returned by the search engine that can meet the conditions. If not found, return "none".

[0070] Finally, summarize and splice the extracted relevant background knowledge fragments to form external background information for the entire conversation content.

[0071] Step 3.2, Extract the topics of the conversation content by the large language model:

[0072] After obtaining the external background information of the conversation content, generate the final topic event name based on the large language model, and use the chain of thought reasoning method to discover the points of argument, similarities, and key points that appear in the entire conversation content text. Finally, give a summary of the topic, specifically using the following prompt:

[0073]

【As a language analysis expert, you are very good at summarizing and generalizing, and can clearly analyze the points of argument, similarities, and key points that appear in different conversation contents during the conversation process; since there may be some internet buzzwords, abbreviations, homophonic rewritten words, etc. in the conversation content, in order to help you better understand the conversation content, I will provide you with some reference information:

[0074] The given reference information is: {text}

[0075] Ultimate goal

[0076] Your goal is to use a short phrase or sentence to give the core discussion topic of the conversation. Secondly, you need to give a logical and easy-to-understand summary of the topic, and bring the speaker information into the topic summary so that any other speaker can quickly understand the full picture of the conversation content.

[0077] Require:

[0078] 1. When summarizing the core topic, you need to think and analyze, and use the content in the given reference information to complete the semantics of the conversation content. After summarizing the core topic, you need to perform backtesting to determine whether the summarized topic represents the main semantics of the conversation content.

[0079] 2. If you find that the topic you summarized is unacceptable after going back to verify, please rethink it and produce new topic sentences until you are satisfied;

[0080] 3. After summarizing the topic of the conversation, create a topic summary around the topic. Since it is a conversation scene, the readers of the topic summary need to understand the whole discussion content, so you need to include the name of the speaker in the summary and give their opinions, points of contention with others, etc.;

[0081] The given conversation content data is:

[0082] {chat}

[0083] Output Description

[0084] The output result must be in JSON format, with the following format: {"topic event":"","topic summary":""}. The topic summary cannot be a running account of the conversation content, but needs to summarize the core arguments.

[0085] let's do it step by step;

[0086] Notice

[0087] All answers should be in Chinese and be concise and clear. Pay attention to using the given reference information and do not give meaningless or useless thinking steps].

[0088] Furthermore, multiple topics in the conversation process are processed in the same way to form a topic analysis report of the discussion content. The specific method is as follows:

[0089] Step 4.1: Summarize all the conversation contents of a conversation, use the large language model to extract the speaker's portrait characteristics, set the topic or event name of interest, and use the following prompt:

[0090]

【Suppose you are an expert in summarizing the characteristics of a person's portrait, and you are very good at summarizing the topics or events that a person prefers based on the conversation content in the person's history, thereby forming the person's portrait;

[0091] Task description

[0092] I will provide you with the historical conversation content of a person in a conversation. You need to summarize this historical information in two sentences. At the same time, the user has also given some of their own preference information. You need to summarize the characteristics that are different from the user's self-given preferences to improve the user's portrait of the character;

[0093] Historical conversation content of the person

[0094] List of historical conversation content: {text}

[0095] Preference information given by the user himself

[0096] Preference information: {preference}

[0097] Requirements:

[0098] 1. When summarizing the characteristics of the character portrait, you need to think and analyze, and be faithful to the content in the given information. You can make some inferences based on your own experience and knowledge appropriately;

[0099] 2. The content of the summarized character characteristics must be very concise and general, and the length of two sentences is sufficient;

[0100] Output description

[0101] The output result must be in json format, as shown below: {"Content analysis and reasoning": "", "Character portrait characteristics": ""}

[0102] Let’s do it step by step;

[0103] Note

[0104] Answer all in Chinese, and be concise and clear. Pay attention to using the given reference information and do not give meaningless and useless thinking steps];

[0105] Step 4.2, based on the characteristics of the role portrait, use the BERT semantic similarity calculation model to judge the relevance of each topic and summary of the conversation content to the characteristics of the role portrait;

[0106] Step 4.3, sort according to the relevance to obtain the output of a personalized analysis report.

[0107] Furthermore, the BERT semantic similarity calculation model uses an embedding model.

[0108] A topic analysis system based on a large language model implements the above-mentioned topic analysis method based on a large language model to achieve topic analysis based on a large language model, which is executed in five modules respectively:

[0109] Perform data preprocessing on the conversation content;

[0110] Classify the preprocessed conversation content into topics and gather the content belonging to the same topic together;

[0111] Extract the discussion objects, abbreviations, and key phrases appearing in the content, and collect additional background knowledge;

[0112] Use a large language model to summarize the content under the topic to form a topic name and a topic summary;

[0113] Perform the same processing on multiple topics during the conversation to form a topic analysis report on the discussion content.

[0114] A computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the above-mentioned topic analysis method based on a large language model to achieve topic analysis based on a large language model.

[0115] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned topic analysis method based on a large language model to achieve topic analysis based on a large language model.

[0116] Compared with the prior art, the significant advantages of the present invention are as follows: 1) A topic segmentation method for conversation content combining multiple features is proposed. Through fine-grained mining and analysis of the conversation content, 6 data features are summarized, and different strategies are used to process the conversation content, laying a good data foundation for the next topic segmentation based on the BERT model. The topic segmentation method proposed by the present invention cleverly utilizes the model structure of BERT itself, combines the semantics of each conversation content with multiple feature vectors, and further blends the semantic information through a bidirectional LSTM neural network to achieve accurate topic segmentation. 2) A topic extraction method based on a large model is proposed. In order to extract topics and topic summaries more accurately, first, key elements such as discussion objects, abbreviations, and key phrases in the conversation content are identified through a large model, then background knowledge of these key elements is obtained with the help of an online plugin, and then the specific topic name and topic summary of each discussion topic are obtained based on the thinking chain reasoning ability of the large model. Finally, a personalized topic analysis report is generated by summarizing the portrait information of the speakers. Description of the Drawings

[0117] Figure 1 It is a topic sequence structure diagram.

[0118] Figure 2 It is the basic process of the SPTSAI algorithm.

[0119] Figure 3 It is a topic mining model based on BERT-LDA.

[0120] Figure 4 It is a schematic diagram of the dialogue content.

[0121] Figure 5 It is a schematic diagram of the dialogue content topic segmentation model based on BERT.

[0122] Figure 6 It is the output flow chart of the personalized topic analysis report. Specific implementation manners

[0123] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0124] A topic analysis method based on a large language model. First, it is necessary to classify the topics of the dialogue content and gather the content belonging to the same topic together; then extract the discussion objects, abbreviations, key phrases, etc. that appear in the content, and collect the additional background knowledge of this content, so as to help the model better understand the semantics of the corresponding content and some online hot memes; finally, use the large model to summarize the content under the topic to form a topic name and a topic summary. By performing similar processing on multiple topics in the entire dialogue process, a topic analysis report of the discussion content can be generated. Specifically as follows:

[0125] Step 1. Data preprocessing of the dialogue content

[0126] (1) Definition of dialogue content features

[0127] When speaking using an online social software, the types of dialogue content are rich and diverse, among which text, pictures, voice and expressions account for the vast majority. For voice-type dialogue content, the present invention adopts a general speech-to-text engine to convert the voice into text content. Since the present invention mainly solves the problem of topic extraction at the text level, for pictures, emoticons and other uncommon message types, only simple processing is performed, and only the message types are retained, such as [Picture], [Expression], [Forward and share], etc. In addition, through sufficient analysis of the dialogue content, the present invention has summarized the following features:

[0128] ① Type of conversation content: As described above, in the present invention, the type of conversation content is taken as one of the features, mainly divided into: text, picture, expression, forward sharing, and other types;

[0129] ② Message reference: In current online social communication software, there is a function to reference other messages. If you want to reply to a certain conversation content before multiple messages, by referencing that content, it is very convenient to make people understand the communication object and the meaning expressed in the conversation content. That is to say, in the scenario of the present invention, the conversation content with a reference relationship is very likely to belong to the same topic conversation, which is a relatively direct data feature.

[0130] ③ @ object: In daily online communication, people often @ other people when speaking to indicate that the corresponding person should pay attention to the message. Subsequently, the conversation content of the corresponding person being @ is very likely to be a response to the conversation content at the time of being @.

[0131] ④ Message interval time: An article has expressed such a view: "Conversation content within a certain time is very likely to belong to the same topic, and conversation content with a long interval may have started a new topic." In the present invention, the time interval between the sending times of the upper and lower two conversation contents is taken as a feature.

[0132] ⑤ Conversation content index: The present invention also takes the index position of the conversation content in the entire conversation process as a feature. The premise requirement of the model established subsequently in the present invention is to start from a message known to be the beginning of a topic. When it is necessary to model a certain conversation process from the first conversation content, the first conversation content must be the start of the topic.

[0133] ⑥ Speaker activity: In a conversation process, there will be people with high activity and those who are lurking. People with higher activity indicate a higher degree of participation in the topic. In the present invention, the activity of the speaker is also taken as one of the modeling features.

[0134] For the 6 features designed in the present invention, they can be generally divided into two categories. One category needs to be discretized, and the other category needs to be verbalized. The following is a separate description:

[0135] (2) Feature processing of conversation content

[0136] Discrete feature processing

[0137] In the present invention, the conversation content type, interval time, index, and speaker activity are data features that need to be discretized. Among them, the message type is already in a discrete form, that is, it includes text, picture, expression, forward sharing, and other types of messages. In order to be able to participate in the subsequent modeling process, the present invention represents the conversation content type in a vectorized manner:

[0138] Text = [0.992734, -0.476647, ……, 0.217249]

[0139] Image = [-0.135216, 0.156160, ……, 0.001139]

[0140] Emoji = [0.088582, 0.240145, ……, -0.006931]

[0141] Forward = [0.213415, -0.352543, ……, 0.234367]

[0142] ……

[0143] Each type of conversation content is represented by a randomly initialized 50 - dimensional vector, and these feature vectors of conversation content types will participate in the subsequent training and update of the model.

[0144] Regarding the processing of the interval time, it is divided into 7 levels: within half a minute, half a minute - 5 minutes, 5 minutes - 15 minutes, 15 minutes - 30 minutes, 30 minutes - 1 hour, 1 hour - 3 hours, and more than 3 hours. Generally, the more timely the reply, the higher the probability of continuing the same topic. For the messages sent after a period of time, there is a greater possibility that a new topic has been restarted. Similar to the processing of the conversation content type features, 50 - dimensional feature vectors randomly initialized are still used to represent the several levels of message interval time.

[0145] Within half a minute = [0.992734, -0.476647, ……, 0.217249]

[0146] Half a minute to 5 minutes = [-0.135216, 0.156160, ……, 0.001139]

[0147] ……

[0148] Regarding the processing of the conversation content index features, the vector initialization is performed according to the index of the conversation content in the entire list of conversation content passed into the modeling window. As described in the definition of this feature, the conversation content passed into the modeling window needs to be consecutive speaking messages at the beginning of a new topic. Since the number of conversation contents input into the modeling window can reach several hundred, the present invention refers to the initialization method of the position vector in the transformer method, which is a simple and effective position encoding method, and outputs a d - dimensional vector for each index position. Its specific definition is: Let t be the index position to be modeled, For its corresponding position encoding, the position encoding function f is defined as follows:

[0149]

[0150] Among them, From the definition, we can see that the frequency ω of the trigonometric function k decreases continuously along the vector dimension, so its wavelength forms a geometric sequence from 2π to 10000·2π. For the positional encoding of the t-th word can be regarded as a vector composed of sine-cosine pairs with different frequencies:

[0151]

[0152] Although this index feature will participate in the subsequent modeling process, it will not be updated during the training process.

[0153] For the processing of the speaker activity feature, referring to the definition in the Internet, it can be divided into several levels such as lurking, bubbling, active, and talkative. Similar to the processing means of the message interval and message type features, a randomly initialized 50-dimensional vector is also used for representation.

[0154] Literal processing

[0155] For the processing of the citation feature, the present invention chooses to explicitly represent the cited conversation content together. For example, Zhang San posts a message "A certain star...", and Li Si quotes this message and posts "Really? Is there a video?". For this kind of conversation message containing a citation relationship, the present invention directly performs plain text display processing, becoming: Cite the message [A certain star...] posted by [Zhang San], and [Li Si] replies and posts the message [Really? Is there a video?].

[0156] For the processing of the @ object feature, the present invention performs certain processing on the conversation content containing the @ object so that the structure is consistent with the citation feature. For example: Zhang San posts a message "@Li Si@Wang Wu There will be a meeting in the conference room at 8 o'clock tomorrow. Please reply if you receive it!". The present invention will process it as: [Zhang San] @ [Li Si] [Wang Wu], asking them to pay attention to the message [There will be a meeting in the conference room at 8 o'clock tomorrow. Please reply if you receive it!]

[0157] For ordinary conversation content, in order to maintain the form unity in the modeling process, in the present invention, each ordinary conversation content is processed as [Someone] posted [Something].

[0158] Step 2: Topic segmentation of conversation content

[0159] (1) Topic clustering

[0160] In the present invention, since the number of conversation contents may be very large, the present invention adopts a sliding window strategy for separate processing, that is, only a certain number of conversation contents are modeled each time, which is called a modeling window. And within each modeling window, the present invention assumes that the number of topic changes within a modeling window does not exceed 10. Coupled with some meaningless conversation contents, such as following the trend to reply "Received", emoticons, etc., a meaningless topic category is added. The 10 topics mentioned here are not topics with clear meanings and can be considered as 10 topic clusters. When some conversation contents are talking about the same topic content, they should all be classified into the same topic cluster. However, the modeling method of the present invention is completely inconsistent with the clustering method used in the existing technical solutions. The present invention uses a pre-trained language model as the base, combines the context semantics of the conversation content and combines multiple features for topic segmentation of the conversation content, while the traditional technical solutions only use the clustering scheme to roughly complete the topic clustering of all conversation contents. (Obtain different clustering clusters).

[0161] (2) BERT-based Topic Segmentation Model for Conversation Content

[0162] For each clustering cluster, the BERT-based topic segmentation model for conversation content first needs to splice the continuous conversation contents to form the input text that can be normally input into the BERT model. For example, there are the following conversation messages:

[0163] Using the text processing method described in the previous step, process each message content. For example, for conversation content 1, it needs to be transformed into: [Zhang San] posted [A major accident occurred due to the collapse of a certain highway. Have you all seen this news?]; for conversation content 4, it needs to be transformed into: Quote the message posted by [Zhang San] [A major accident occurred due to the collapse of a certain highway. Have you all seen this news?], [Li Si] replied and posted the message [Is it true? Do you have a video?]; for conversation content 9, it needs to be transformed into [Xiaohong] posted [Received! [Emoticon]]. The remaining conversation contents are also transformed in the manner described in the previous step.

[0164] Assume that each conversation content in the conversation content list is sequentially recorded as P 1 , P 2 , …, P 10 , …, where it is recorded that contains k characters, and k can be different for different conversation contents. Splice these conversation contents into the input data form of the BERT model, and add [CLS] and [SEP] to the head and tail ends of each conversation content respectively, as shown below:

[0165] Input = [CLS]P 1 [SEP][CLS]P2 [SEP][CLS]P 3 [SEP][CLS]P 4 [SEP]……

[0166] After the data splicing process is completed, the discrete features of each conversation content start to be processed. The discrete feature vectors of each conversation content are obtained according to the method described in the previous step: conversation content type, interval time, index, and publisher activity. The specific generation method of the discrete feature vectors will not be elaborated here.

[0167] The model framework of the present invention for modeling is as Figure 5 shown. First, for the list of conversation contents to be modeled, after concatenating [CLS] and [SEP] to form the Input text string, it is input into the BERT model to obtain the deep semantic feature vectors of each conversation content. Considering that different conversation contents are independent of each other, the present invention improves the data input of the BERT model by setting the segment_ids intervals of different conversation contents to 0 and 1 respectively, so as to introduce the overall information and interval information of the conversation contents. After passing through the BERT model, the semantic vector representation of each character is obtained, and the semantic vector of the [CLS] character concatenated for each conversation content is taken as the overall semantic representation of the entire conversation content. In addition, discrete feature vectors such as the conversation content type, interval time, index, and speaker activity of each conversation content are obtained, and these discrete feature vectors are concatenated with the deep semantic vectors of each conversation content to form the comprehensive feature vectors of each conversation content. Since the conversation content list has context continuity, a bidirectional LSTM layer is added in the present invention for semantic interaction of the comprehensive feature vectors of the conversation contents. The new vectors output by the LSTM layer are input into the softmax layer for classification, so as to judge the topic attribution of each conversation content.

[0168] Step 3: Topic extraction module based on large models

[0169] In the previous step, a conversation content is accurately segmented into multiple topics. The original form of the conversation content is still retained within each topic, and only the conversation contents of different topics are assigned to their respective topics. However, the names of each topic have not been obtained yet, and the analysis is still very laborious. The present invention introduces a topic extraction scheme based on large models in the following text.

[0170] (1) Obtaining element background information

[0171] In view of the fact that the traditional solution does not consider the latest network hot spots and hot topics when extracting topics, the present invention obtains the latest background knowledge information related to the topic of the conversation content through the network plug-in. First, the big model is prompted by prompts to understand the semantics of each conversation content, and the discussion objects, abbreviations, key phrases, etc. that appear in the conversation content are extracted. The main goal is to extract phrases that the big model itself cannot accurately judge. For example, the word "YYDS" has appeared for a long time, and the existing big model can already understand the meaning of the word, so there is no need to conduct an online query. For some newly emerging "lantern damage assessment", an online query is required to collect additional background knowledge of these contents, so as to help the model better understand the semantics of the conversation content and some network hot topics. Next, the information queried online is merged and filtered using the big model to obtain background knowledge related to the topic. With this background knowledge, the big model can be summarized to obtain the accurate topic name.

[0172] Specifically, the present invention designs a set of background information acquisition process of key elements of conversation content based on a large model. The algorithm flow is as follows:

[0173] (1) First, for a topic that has been segmented, extract words and phrases that cannot be well understood by the large model. Different from the traditional element extraction task, the key elements of the conversation content here are not named entities in the traditional sense (names of people, places, institutions or proper nouns), but some phrases and phrases that occupy a relatively important semantic role in the conversation content but cannot be understood by the model. For this purpose, the present invention adopts the following prompt:

[0174] [Suppose you are a language analysis expert, and you need to deeply analyze the given conversation content data. The data discusses a topic event and may use some Internet hot words, abbreviations, homophones, etc., but these words have a great impact on the analysis of semantics. Your goal is to find these words in a reasonable and well-founded way.

[0175] ##Require:

[0176] 1. Before extracting a word, you need to think about whether you know the specific meaning of the word. If you know the exact meaning of the word, please ignore it;

[0177] 2. When you feel that you cannot accurately know the exact meaning of a word, you need to use an Internet search engine to find relevant information. This is the target word you need to extract;

[0178] 3. The words you extract need to occupy a relatively important semantic position in the conversation content. For some irrelevant words, just ignore them;

[0179] 4. Additionally, you need to comprehensively understand the topic event under core discussion, which will help you be more specific when determining the exact meaning of words with the aid of Internet search in the future;

[0180] ## The given conversation data is as follows:

[0181] {chat}

[0182] ## Output instructions

[0183] The output result must be in JSON format, as shown below: {"topic event":"","words to be confirmed":[{"word":"","explanation":""},...]} Each word to be confirmed needs to be explained as to why it needs to be confirmed, and the topic event result must be brief, retaining the core semantics.

[0184] let’s do it step by step。

[0185] ## Note

[0186] Answer all in Chinese, and be concise and clear. Do not give meaningless or useless explanations. Directly ignore [emoji], [picture], [forward], etc.

[0187] For Figure 4 the example in, after inputting it to the large model, the following output can be obtained:

[0188]

【Step 1: Extract the topic event

[0189] Based on the content of the conversation data, the core topic event can be summarized as:

[0190] **Topic event**: A major accident occurred due to the collapse of a highway.

[0191] Step 2: Extract the words to be confirmed

[0192] Based on the conversation data, the following words need to be confirmed:

[0193] 1. **XX Expressway**: An abbreviation, need to determine which specific expressway it is.

[0194] 2. **M State**: May be a homophone rewrite or abbreviation referring to a certain region.

[0195] The following is the complete result:

[0196]

[0197] (2) Combine the word extraction results obtained from the previous prompt with the core topic event of the conversation data and call the online search engine one by one to obtain background knowledge documents related to the words;

[0198] (3) Since the search engine will return the top-N search results, these searched web pages may contain a large amount of text content that is irrelevant to the query terms and topic events. It is necessary to use a large model to filter and only retain relevant background knowledge text. The prompt designed by the present invention is as follows:

[0199] Please analyze the webpage text returned by online search engines and find the text fragments that help understand the query keywords and conversation content:

[0200] The webpage text returned by the online search engine is: {text};

[0201] The query keyword entered into the search engine is: {KeyWord};

[0202] The goal is to better understand the following conversation: {chat};

[0203] Please directly provide the text fragment that meets the conditions in the webpage text returned by the search engine. If it cannot be found, please return "None". 】

[0204] (4) Summarize and splice the relevant background knowledge fragments extracted in step 3 to form external background information for the entire conversation content;

[0205] (2) Large model extracts conversation topics

[0206] After obtaining the external background information of the conversation content, the next step is to extract the specific topic name with the help of the semantic understanding and summarization ability of the big model. Although the big model also generated the topic event content in the prompt of the previous step, because the previous step did not obtain enough background information, the topic event generated in the previous stage has no practical value. In this step, sufficient information support has been obtained through the Internet search engine, so the final topic event name can be generated based on the big model, and the controversial points, similarities, and key points that appear in the entire conversation content text can be found by thinking chain reasoning, and finally a summary of the topic is given. Through this strategy, a logically clear and easy-to-understand topic summary can be generated, rather than just giving a simple two-sentence description of the topic content. The prompt designed in this process is as follows:

[0207]

【As a language analysis expert, you are very good at summarizing and generalizing, and can clearly analyze the controversial points, similarities, and key points in different conversations. Since the conversations may use some Internet hot words, abbreviations, homophones, etc., in order to help you better understand the conversations, I will provide you with some reference information:

[0208] The given reference information is: {text}

[0209] Ultimate Goal

[0210] Your goal is to use a short phrase or sentence to give the core discussion topic event of the conversation content. Secondly, you need to give a logical and easy-to-understand topic summary, and bring the speaker information into the topic summary so that any other speaker can quickly understand the full picture of the conversation content.

[0211] ##Require:

[0212] 1. When summarizing the core topic, you need to think and analyze, and use the content in the given reference information to complete the semantics of the conversation content. After summarizing the core topic, you need to perform backtesting to determine whether the summarized topic represents the main semantics of the conversation content.

[0213] 2. If you find that the topic you summarized is unacceptable after going back to verify, please rethink it and produce new topic sentences until you are satisfied;

[0214] 3. After summarizing the topic of the conversation, create a topic summary around the topic. Since it is a conversation scene, the readers of the topic summary need to understand the whole discussion content, so you need to include the name of the speaker in the summary and give their opinions, points of contention with others, etc.;

[0215] ##The given conversation content data is:

[0216] {chat}

[0217] ##Output Description

[0218] The output result must be in JSON format, and the format is as follows: {"topic event":"","topic summary":""}. The topic summary cannot be a running account of the conversation content, but needs to summarize the core arguments.

[0219] let's do it step by step.

[0220] ##Notice

[0221] All answers should be in Chinese and should be concise and clear. Pay attention to using the given reference information and do not give meaningless and useless thinking steps. 】

[0222] After the prompt is input to the large model, the following sample output can be obtained:

[0223]

[0224]

[0225] Step 4: Generation of Topic Analysis Report

[0226] In a dialogue scenario, multiple topics may emerge. Based on the topic segmentation technology and the topic and abstract extraction technology based on large models mentioned above, the dialogue content can be basically classified and highly summarized. If you want to output an analysis report of the dialogue content, you can directly combine and splice each summarized topic and its corresponding abstract to form a report. However, this approach does not consider the situation that each person's interested topics may be different. Therefore, the present invention proposes a strategy for personalized output of topic analysis reports according to the speaker's role portrait. First, summarize all the dialogue content of a dialogue, and use a large model to summarize the portrait characteristics of the speaker. The speaker can also independently set the topics or event names of interest; secondly, based on the summarized role portrait characteristics, use the BERT semantic similarity calculation model to judge the relevance of each topic and abstract in the dialogue content to the role portrait characteristics; finally, sort according to the relevance to obtain a personalized analysis report output. The specific implementation process is as Figure 6 shown, where the role portrait feature summary Prompt is:

[0227] 【Suppose you are an expert in summarizing portrait characteristics, and you are very good at summarizing the topics or events preferred by a person based on the historical dialogue content of the person, so as to form the portrait of this person.

[0228] ## Task Description

[0229] I will provide you with the historical dialogue content of a person in a dialogue. You need to summarize this historical information in two sentences. At the same time, the user has also given some of their own preference information. You need to summarize the characteristics different from the user's self-given preferences to improve the user's character portrait.

[0230] ## Historical Dialogue Content of the Person

[0231] List of historical dialogue content: {text}

[0232] ## Preference Information Given by the Person Independently

[0233] Preference information: {preference}

[0234] ## Requirements:

[0235] 1. When summarizing the character portrait features of the person, you need to think and analyze, and be faithful to the content in the given information. You can make some inferences based on your own experience and knowledge appropriately;

[0236] 2. The summarized character features must be very concise and general, and the length of two sentences is sufficient;

[0237] ## Output Instructions

[0238] The output result must be in JSON format, as shown below: {"Content Analysis and Reasoning": "", "Portrait Features of Character Roles": ""}.

[0239] let’s do it step by step。

[0240] ## Note

[0241] Answer all in Chinese, and be concise and clear. Pay attention to using the given reference information and do not give meaningless or useless thinking steps.

[0242] In the present invention, the BERT semantic similarity calculation model adopts an open-source embedding model.

[0243] The present invention also proposes a topic analysis system based on a large language model, implements the above-mentioned topic analysis method based on a large language model, and realizes topic analysis based on a large language model.

[0244] A computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the above-mentioned topic analysis method based on a large language model and realizes topic analysis based on a large language model.

[0245] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned topic analysis method based on a large language model and realizes topic analysis based on a large language model.

[0246] In summary, in view of the problems that the traditional solution does not consider combining the latest network hotspots and memes when extracting discussion topics and has poor deep semantic modeling ability for dialogue content, the present invention proposes a method for topic analysis of dialogue content based on a large language model. Through an online plug-in, the large model can obtain the latest background knowledge information related to the discussion topic. Since the large model itself already has a vast amount of knowledge, in order to improve the efficiency of the large model in extracting discussion topic content, the present invention also designs a set of high-efficiency strategies for obtaining relevant background information. First, prompt the large model to understand the semantics of each piece of dialogue content and extract phrases that cannot be accurately judged in the dialogue content. For example, the term "YYDS" has been around for a long time, and existing large models can already understand its meaning, so there is no need to query the network. However, for some newly emerged terms such as "damage assessment by holding a lantern", it is necessary to query the network. Next, use the large model to merge and filter the information obtained from the network query to obtain background knowledge related to the discussion topic. With this background knowledge, then let the large model summarize to obtain the accurate topic name.

[0247] In view of the fact that the existing solutions do not use techniques based on pre-trained models to refine the processing of cross-topics that occur during conversations, the present invention proposes a method for segmenting conversation content topics based on the fusion of multiple feature vectors of BERT, realizing the refined processing of conversation content data, splitting and retaining the original conversation content data belonging to each topic during a conversation process. Specifically, the present invention first analyzes the conversation data and, in combination with the experience of predecessors, summarizes some common data features, such as: the message types of the conversation content [emoji, pictures, website links, etc.], the time interval between adjacent messages, the sending user, the @ relationship in the conversation content text, the citation relationship of the conversation content, the id index of the conversation content in the entire communication and discussion process, etc. Each feature is represented as a vector. In addition, semantic features of each message are extracted based on the BERT model under the context of the conversation content, and a bidirectional LSTM layer is additionally established to semantically fuse all the feature vectors of each message, and then each message is classified, thus realizing the processing of cross-topics.

[0248] In response to the problem that the existing solutions only extract topics without giving a summary overview of the topics, the present invention proposes a topic overview generation strategy based on the chain of thought summary of large models. First, a large language model is used to analyze the semantic information in a specific conversation process, enabling the large model to re-interpret the semantics behind each conversation content. Then, in the way of chain of thought reasoning, the points of argument, similarities, and key points that appear during the entire conversation process are discovered. Finally, a summary overview of the topic is given. Through this strategy, a topic summary with clear logic and easy to understand can be generated, rather than just giving a simple two-sentence description of the topic content.

[0249] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as within the scope described in this specification.

[0250] The above-described embodiments merely represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A topic analysis method based on a large language model, characterized in that: The steps include: Perform data preprocessing on the conversation content; Classify the preprocessed conversation content into topics and group the content belonging to the same topic together; Extract discussion objects, abbreviations, key phrases that appear in the content and collect additional background knowledge; Use a large language model to summarize the content under the topic to form a topic name and topic summary; The same processing is performed on multiple topics during the conversation to form a topic analysis report of the discussion content.

2. The topic analysis method based on a large language model according to claim 1, characterized in that: The data preprocessing of the conversation content is as follows: Step 1.1, define the characteristics of the conversation content: A general speech-to-text engine is used to convert speech into text content and extract conversation content features, including conversation content type, message reference, @ object, message interval, conversation content index, and publisher activity. Conversation content types are divided into: text, picture, emoticon, forwarding and sharing, and other types. Step 1.2, feature processing of conversation content: The six conversation content features are divided into two categories. Conversation content type, interval time, conversation content index, and speaker activity need to be discretized; message references and @ objects need to be textualized; (1) Conversation content type feature processing For the conversation content type, a randomly initialized 50-dimensional vector is used to represent each conversation content type; (2) Interval time feature processing The interval time is first divided into seven levels: less than half a minute, half a minute to 5 minutes, 5 minutes to 15 minutes, 15 minutes to 30 minutes, 30 minutes to 1 hour, 1 hour to 3 hours, and more than 3 hours. Then, a randomly initialized 50-dimensional feature vector is used to represent the level of message interval time. (3) Speech Content Index Feature Processing For the conversation content index, the position vector initialization method in the transformer method is used to initialize the vector according to the index of the conversation content in the entire conversation content list passed into the modeling window; (4) Processing of speaker activity characteristics The speaker activity is first divided into four levels: diving, bubbling, active, and chattering. Then a randomly initialized 50-dimensional feature vector is used to represent the speaker activity level. (5) Message reference feature processing The quoted conversation content is also explicitly indicated, and processed as [someone] published [some content]; (6)@Object feature processing Process the conversation content containing the @ object as [someone] @ [someone], and let them pay attention to the message [certain content].

3. The topic analysis method based on a large language model according to claim 1, characterized in that: The preprocessed conversation content is classified into topics, and the content belonging to the same topic is gathered together. The specific method is as follows: Step 2.1: Use the sliding window strategy to complete the topic clustering of all conversation contents; Step 2.2, concatenate the continuous conversation contents, and record each conversation content in the conversation content list in order as P1, P2, ..., P 10 ,…, where Indicates P i The kth character of the conversation is added with [CLS] and [SEP] at the beginning and end of each conversation content to form the Input text string, which is expressed as follows: Input=[CLS]P1[SEP][CLS]P2[SEP][CLS]P3[SEP][CLS]P4[SEP]…… Step 2.3: Input the input text string into the BERT model to obtain the deep semantic feature vector of each conversation content, and concatenate the deep semantic vector of each conversation content with the discrete feature vector to form a comprehensive feature vector of each conversation content. In step 2.4, the comprehensive feature vector of each conversation content is input into the bidirectional LSTM layer, and the new vector output by the LSTM layer is input into the softmax layer for classification, so as to determine the topic of each conversation content.

4. The topic analysis method based on a large language model according to claim 1, characterized in that: Extract discussion objects, abbreviations, key phrases that appear in the content and collect additional background knowledge by: Step 3.1, obtain the background information of the elements: First, for a topic that has been segmented, the large language model is used to extract the discussion objects, abbreviations, and key phrases that appear in the conversation content. The following prompt is used: Suppose you are a language analysis expert and you need to deeply analyze the given conversation content data. The data discusses a topic event and may use some Internet hot words, abbreviations, homophones, etc. However, these words have a great influence on the semantic analysis. Your goal is to find these words in a reasonable and well-founded way. Require:

1. Before extracting a word, you need to think about whether you know the specific meaning of the word. If you know the exact meaning of the word, please ignore it; 2. When you feel that you cannot accurately know the exact meaning of a word, you need to use an Internet search engine to find relevant information. This is the target word you need to extract; 3. The words you extract need to occupy a relatively important semantic position in the conversation content. For some irrelevant words, just ignore them; 4. In addition, you need to have an overall understanding of the core discussion topic, which will help you to determine the exact meaning of the words when searching on the Internet in the future; The given conversation data is: {chat} Output Description The output result must be in JSON format, with the following format: {"topic event":"","words to be confirmed":[{"words":"","explanation":""},...]}. Each word to be confirmed needs to explain why it needs to be confirmed. The topic event result must be brief and retain the core semantics. let's do it step by step. Notice All answers should be in Chinese and should be concise and clear. No meaningless or useless explanations should be given. [Emojis], [Pictures], [Reposts], etc. should be ignored.] Then, based on the word extraction results obtained from the large language model and combined with the core topic events of the conversation data, the online search engine is called one by one to obtain the background knowledge documents related to the words; Next, filter based on the large language model to retain only relevant background knowledge text, using the following prompt: Please analyze the webpage text returned by online search engines and find the text fragments that help understand the query keywords and conversation content: The webpage text returned by the online search engine is: {text}; The query keyword entered into the search engine is: {KeyWord}; The goal is to better understand the following conversation: {chat}; Please directly provide the text fragment that meets the conditions in the webpage text returned by the search engine. If it cannot be found, please return "None"] Finally, the extracted relevant background knowledge fragments are summarized and spliced ​​to form external background information for the entire conversation content; Step 3.2: The large language model extracts the conversation content topic: After obtaining the external background information of the conversation content, the final topic event name is generated based on the large language model, and the controversial points, similarities, and key points that appear in the entire conversation content text are discovered by thinking chain reasoning. Finally, a summary of the topic is given. The following prompt is used: [As a language analysis expert, you are very good at summarizing and generalizing, and can clearly analyze the controversial points, similarities, and key points in different conversations during the conversation. Since the conversation may use some Internet hot words, abbreviations, homophones, etc., in order to help you better understand the conversation content, I will provide you with some reference information: The given reference information is: {text} Ultimate Goal Your goal is to use a short phrase or sentence to give the core discussion topic of the conversation. Secondly, you need to give a logical and easy-to-understand summary of the topic, and bring the speaker information into the topic summary so that any other speaker can quickly understand the full picture of the conversation content. Require:

1. When summarizing the core topic, you need to think and analyze, and use the content in the given reference information to complete the semantics of the conversation content. After summarizing the core topic, you need to perform backtesting to determine whether the summarized topic represents the main semantics of the conversation content.

2. If you find that the topic you summarized is unacceptable after going back to verify, please rethink it and produce new topic sentences until you are satisfied; 3. After summarizing the topic of the conversation, create a topic summary around the topic. Since it is a conversation scene, the readers of the topic summary need to understand the whole discussion content, so you need to include the name of the speaker in the summary and give their opinions, points of contention with others, etc.; The given conversation content data is: {chat} Output Description The output result must be in JSON format, with the following format: {"topic event":"","topic summary":""}. The topic summary cannot be a running account of the conversation content, but needs to summarize the core arguments. let's do it step by step; Notice All answers should be in Chinese and be concise and clear. Pay attention to using the given reference information and do not give meaningless or useless thinking steps].

5. The topic analysis method based on a large language model according to claim 1, characterized in that: The same processing is performed on multiple topics in the conversation process to form a topic analysis report of the discussion content. The specific method is as follows: Step 4.1: Summarize all the conversation contents of a conversation, use the large language model to extract the speaker's portrait characteristics, set the topic or event name of interest, and use the following prompt: [Suppose you are an expert in character profile summary, and you are very good at summarizing the character's favorite topics or events based on the character's historical conversations, and thus forming a character profile of the character; Mission Statement I will provide you with a character's historical conversation content in a conversation. You need to summarize this historical information in two sentences. At the same time, the user himself also provides some of his own preference information. You need to summarize the characteristics that are different from the preferences given by the user, so as to improve the user's character portrait; Character history dialogue content List of historical conversation contents: {text} Preference information given by the character Preference information: {preference} Require:

1. When summarizing the characteristics of a character portrait, you need to think and analyze, and be faithful to the content of the given information. You can make some inferences based on your own experience and knowledge; 2. The character characteristics of the summary must be very concise and only two sentences long; Output Description The output result must be in JSON format, and the format is as follows: {"Content Analysis Reasoning":"","Character Portrait Features":""} let's do it step by step; Notice All answers should be in Chinese and should be concise and clear. Pay attention to using the given reference information and do not give meaningless and useless thinking steps]; Step 4.2: Based on the characteristics of the role portrait, use the BERT semantic similarity calculation model to determine the relevance of each topic and summary of the conversation content with the characteristics of the role portrait; Step 4.3, sort by relevance and obtain personalized analysis report output.

6. The topic analysis method based on a large language model according to claim 1, characterized in that: The BERT semantic similarity calculation model uses the embedding model.

7. A topic analysis system based on a large language model, implementing the topic analysis method based on a large language model according to any one of claims 1 to 6, realizing topic analysis based on a large language model, and executing in five modules respectively: Perform data preprocessing on the conversation content; Classify the preprocessed conversation content into topics and group the content belonging to the same topic together; Extract discussion objects, abbreviations, key phrases that appear in the content and collect additional background knowledge; Use a large language model to summarize the content under the topic to form a topic name and topic summary; The same processing is performed on multiple topics during the conversation to form a topic analysis report of the discussion content.

8. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the topic analysis method based on a large language model according to any one of claims 1 to 6 is implemented to realize topic analysis based on a large language model.

9. A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the topic analysis method based on a large language model according to any one of claims 1 to 6 is implemented to realize topic analysis based on a large language model.

Citation Information

Patent Citations

  • Technology for realizing search engine optimization by optimizing new keywords

    CN106776678A

  • Voice detection processing system

    CN113129895A

  • Theme recognition method for dialogue text

    CN113641778A

  • Text abstract method and device fused with knowledge extraction

    CN118312609A

  • Product recommendation method, device, equipment, system and program product

    CN119168729A

Cited By

  • Dialogue interaction method and system based on long historical dialogue semantic understanding

    CN120804270A

  • Generative financial report writing method based on role division

    CN121997900A

  • An agent topic attribution and tracing method based on multi-time point streaming decision

    CN122654320A