A context-aware based group chat conversation structure understanding optimization method

CN122476084BActive Publication Date: 2026-09-04SHANDONG SHENGDE INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610930283.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-26
Publication Date
2026-09-04
Estimated Expiration
2046-06-26

AI Technical Summary

Benefits of technology

[0009] Beneficial Effects: This invention calculates the comprehensive association strength between adjacent messages in the preceding message sequence based on the semantic vector similarity, sender continuity, and time interval between adjacent messages in the preceding message sequence of the currently analyzed message. This comprehensive association strength is then used to divide the preceding message sequence into session fragments. Each session fragment is defined as a context source, and the message sequence formed by the chronological order of all speeches by each participant in the preceding message sequence is defined as a context source, forming the current context source set for the currently analyzed message. Then, based on the semantic affinity, temporal proximity, and interaction tightness between the currently analyzed message and each context source in the current context source set, a relevance weight is calculated. The semantic vectors of the current context source set are weighted and fused based on the relevance weights to obtain the context semantic vector of the currently analyzed message. The comprehensive semantic vector obtained by fusing the context semantic vector of the currently analyzed message with the semantic vector of the corresponding message itself is input into the downstream task model, outputting the session structure understanding result of the currently analyzed message. Furthermore, by introducing a dynamic and hierarchical context-aware mechanism, this invention achieves a transformation from a one-size-fits-all approach or a static and flat perception of context to a precise and focused one. This significantly improves the accuracy and robustness of understanding the structure of group chat sessions and provides a high-quality, structured, and interpretable data foundation for subsequent advanced applications such as intelligent summarization, intent analysis, and collaborative efficiency evaluation. It fundamentally solves the problem of excessive contextual noise and large comprehension biases when dealing with real and complex group chat scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_40
    Figure SMS_40
  • Figure QLYQS_1
    Figure QLYQS_1
  • Figure QLYQS_16
    Figure QLYQS_16
Patent Text Reader

Abstract

The present application relates to the technical field of data processing, in particular to a group chat conversation structure understanding optimization method based on context perception. The method comprises the following steps: obtaining a current context source set, calculating a correlation weight based on the semantic affinity, time sequence proximity and interaction closeness between the message being analyzed and each context source in the current context source set, weighting and fusing the semantic vectors of the current context source set based on the correlation weight to obtain a context semantic vector of the message being analyzed, inputting the comprehensive semantic vector obtained by fusing the context semantic vector of the message being analyzed and the semantic vector of the corresponding message into a downstream task model, and outputting the conversation structure understanding result of the message being analyzed. The present application can improve the accuracy of group chat conversation structure understanding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and more specifically to a method for optimizing the structure of group chat sessions based on context awareness. Background Technology

[0002] With the widespread adoption of instant messaging applications, group chats have become a core scenario for people's daily collaboration, social interaction, and information exchange. Group chat sessions typically contain rich information about multi-turn dialogues, topic shifts, and multi-person interactions. Therefore, automated structural understanding of these sessions (such as topic segmentation, dialogue behavior recognition, and intent tracking) is fundamental to supporting advanced applications such as intelligent conversation summarization, key information extraction, and group intent analysis.

[0003] Existing technologies for understanding the structure of group chat sessions typically employ sequence modeling (such as RNNs and Transformers) or graph neural networks (GNNs), using sequential messages or pairwise relationships between messages as direct model input to predict the topic segment or dialogue behavior to which a message belongs. However, these methods are static and flat in their perception of context. That is, when understanding the structure of group chat sessions, existing methods usually treat historical messages within a fixed time window as context, or consider the impact of all historical messages on the current message equally. They fail to dynamically and finely distinguish the different contributions of different parts of the context (such as different speakers or different topic initiators) to the understanding of the current message, which can easily introduce irrelevant semantic noise, leading to problems such as misjudgment of topic boundaries or deviation in intent understanding, thus affecting the accuracy of the final understanding of the group chat session structure. Therefore, how to optimize the way the context is constructed by introducing a dynamic and hierarchical context-aware mechanism to improve the accuracy of understanding the structure of group chat sessions has become an urgent problem to be solved. Summary of the Invention

[0004] To address the aforementioned problems, this invention provides a context-aware group chat session structure understanding and optimization method, the specific technical solution of which is as follows:

[0005] One embodiment of the present invention provides a method for optimizing the structure of group chat sessions based on context awareness, comprising the following steps:

[0006] Obtain the semantic vector of each message in the processed group chat message sequence;

[0007] Based on the semantic vector similarity, sender continuity, and time interval between adjacent messages in the preceding message sequence of the currently analyzed message in the processed group chat session message sequence, the comprehensive association strength between adjacent messages in the preceding message sequence is calculated. The comprehensive association strength is used to divide the preceding message sequence into session fragments. Each session fragment is defined as a context source, and the message sequence formed by all the speeches of each participant in the preceding message sequence in chronological order is defined as a context source, forming the current context source set of the currently analyzed message.

[0008] Based on the semantic affinity, temporal proximity, and interaction tightness between the currently analyzed message and each context source in the current context source set, a relevance weight is calculated. The semantic vectors of the current context source set are then weighted and fused based on the relevance weights to obtain the context semantic vector of the currently analyzed message. The comprehensive semantic vector obtained by fusing the context semantic vector of the currently analyzed message with the semantic vector of the corresponding message itself is input into the downstream task model, and the conversation structure understanding result of the currently analyzed message is output.

[0009] Beneficial Effects: This invention calculates the comprehensive association strength between adjacent messages in the preceding message sequence based on the semantic vector similarity, sender continuity, and time interval between adjacent messages in the preceding message sequence of the currently analyzed message. This comprehensive association strength is then used to divide the preceding message sequence into session fragments. Each session fragment is defined as a context source, and the message sequence formed by the chronological order of all speeches by each participant in the preceding message sequence is defined as a context source, forming the current context source set for the currently analyzed message. Then, based on the semantic affinity, temporal proximity, and interaction tightness between the currently analyzed message and each context source in the current context source set, a relevance weight is calculated. The semantic vectors of the current context source set are weighted and fused based on the relevance weights to obtain the context semantic vector of the currently analyzed message. The comprehensive semantic vector obtained by fusing the context semantic vector of the currently analyzed message with the semantic vector of the corresponding message itself is input into the downstream task model, outputting the session structure understanding result of the currently analyzed message. Furthermore, by introducing a dynamic and hierarchical context-aware mechanism, this invention achieves a transformation from a one-size-fits-all approach or a static and flat perception of context to a precise and focused one. This significantly improves the accuracy and robustness of understanding the structure of group chat sessions and provides a high-quality, structured, and interpretable data foundation for subsequent advanced applications such as intelligent summarization, intent analysis, and collaborative efficiency evaluation. It fundamentally solves the problem of excessive contextual noise and large comprehension biases when dealing with real and complex group chat scenarios. Detailed Implementation

[0010] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention are within the protection scope of the embodiments of the present invention.

[0011] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art.

[0012] This embodiment provides a context-aware group chat session structure understanding and optimization method, which is described in detail below:

[0013] Step S001: Obtain the semantic vector of each message in the processed group chat message sequence.

[0014] Current methods for understanding group chat conversation structures typically treat historical messages within a fixed time window as context, or consider the impact of all historical messages on the current message equally. This fails to dynamically and meticulously differentiate the varying contributions of different parts of the context to the understanding of the current message, thus affecting the accuracy of the final understanding of the group chat conversation structure. For example, in a work group, inquiries about the progress of "Project A" may introduce irrelevant noise if messages related to Project A are treated equally with those unrelated to Project A, leading to misjudgment of topic boundaries or deviations in intent understanding, thereby affecting the accuracy of the final structure understanding. To improve the accuracy of understanding group chat conversation structures, this embodiment will introduce a dynamic and hierarchical context-aware mechanism to optimize the way the context is constructed. The optimized vectors will be used for conversation structure understanding, thereby improving the accuracy of understanding group chat conversation structures.

[0015] This embodiment first acquires the original chat data for understanding the group chat session structure, and then processes the original chat data to obtain processed data for subsequent context optimization; the specific data acquisition and processing process is as follows:

[0016] The raw group chat session data is obtained through open APIs provided by instant messaging platforms (such as WeChat Work and DingTalk robot APIs) or by parsing and exporting chat log files. The raw group chat session consists of multiple messages, and each message in the raw group chat session includes: a unique message ID, a sender ID, a sending timestamp, and the message text content. If the communication platform supports this, it may also include the ID of a referenced or duplicated message. Then, the raw group chat session is cleaned, basic features are extracted, and preliminary dialogue behavior is labeled to form a message sequence that includes at least the sender, sending timestamp, semantic vector, and preliminary dialogue behavior labels. This is recorded as the processed group chat session message sequence. That is, each message in the processed group chat session message sequence includes at least the sender, sending timestamp, semantic vector, and preliminary dialogue behavior labels. The semantic vector extracted at this time is the semantic vector of the message itself. The messages in the processed group chat session message sequence are arranged according to the sending timestamp or the order in which they were sent.

[0017] Data cleaning in this embodiment includes at least invalid message filtering, text normalization, and segmentation. Basic feature extraction includes at least word segmentation and part-of-speech tagging, named entity recognition, semantic vector representation, temporal feature extraction, and participant interaction matrix construction. Invalid message filtering refers to removing system notifications (e.g., someone joining a group chat), pure emoticons (without text), URL links, and other messages that contribute little to semantic structure analysis. Text normalization includes converting traditional Chinese to simplified Chinese, converting full-width characters to half-width characters, correcting common spelling errors, and standardizing date, time, or number formats. Segmentation refers to reasonably segmenting excessively long single messages (e.g., exceeding 200 characters) according to punctuation marks (period, question mark, exclamation mark), treating them as logically continuous sub-messages to refine the granularity of analysis. Word segmentation and part-of-speech tagging refer to using open-source tools (such as Jieba, etc.) to... HanLP performs word segmentation and part-of-speech tagging on message text. Named entity recognition refers to identifying entities such as person names, place names, organization names, time, and product names in messages. Semantic vector representation refers to encoding each message text into a fixed-dimensional semantic vector using a pre-trained language model (such as Sentence-BERT). Temporal feature extraction refers to calculating the time interval between adjacent messages in a conversation. The participant interaction matrix is ​​constructed by statistically analyzing the frequency of consecutive speeches between any two participants in a group chat conversation. Preliminary dialogue behavior labels refer to assigning a preliminary dialogue behavior label (such as "ask a question", "answer", "suggest", "agree", "disagree", "notify") to each message using rule-based or lightweight model methods.

[0018] Step S002: Based on the semantic vector similarity, sender continuity, and time interval between adjacent messages in the preceding message sequence of the currently analyzed message in the processed group chat session message sequence, calculate the comprehensive association strength between adjacent messages in the preceding message sequence. Use the comprehensive association strength to divide the preceding message sequence into session fragments, define each session fragment as a context source, and define the message sequence formed by all the speeches of each participant in the preceding message sequence in chronological order as a context source, thus forming the current context source set of the currently analyzed message.

[0019] The main purpose of this embodiment is to optimize the way context is constructed by introducing a dynamic and hierarchical context-aware mechanism, thereby improving the accuracy of understanding the structure of group chat sessions. Typically, the truly effective context of a message should consist of multiple dispersed context sources of varying importance. Therefore, after processing the group chat message sequence and the semantic vectors of each message, this embodiment uses the basic features of the historical session before the currently analyzed message (such as semantic similarity, time interval, and participant continuity) to subdivide the historical session and identify multiple potential context sources (e.g., different topic clusters, different question-and-answer threads). The purpose of this process is to identify possible and relatively independent dialogue substructures (i.e., context sources). For example, a continuous question-and-answer pair, a discussion block around a specific subject, or a speaking sequence dominated by a core participant may all be identified as independent context sources at this stage. Furthermore, this embodiment subsequently employs a lightweight, rule-based, and clustering-based hybrid method to identify context sources in the historical session before the currently analyzed message. The specific identification and acquisition process is as follows:

[0020] First, in the processed group chat message sequence, obtain the sequence of all messages preceding the currently analyzed message, denoted as the preceding message sequence. Then, based on the semantic vector similarity, sender continuity, and time interval between adjacent messages in the preceding message sequence, calculate the comprehensive association strength between adjacent messages in the preceding message sequence. Taking the calculation of the comprehensive association strength between the (i-1)th message and the ith message in the preceding message sequence as an example (where i is not equal to 1), the formula for calculating the comprehensive association strength between the (i-1)th message and the ith message in the processed group chat message sequence is as follows:

[0021]

[0022] in, Let exp() be the overall correlation strength between the (i-1)th message and the ith message in the preceding message sequence of the message being analyzed. exp() is an exponential function with base e. Let i be the semantic vector of the (i-1)th message itself. Let i be the semantic vector of the i-th message itself. , and All are balanced weights. Let be the cosine similarity between the semantic vector of message i and the semantic vector of message (i-1). Ti represents the time interval between the sending time of the (i-1)th message and the sending time of the ith message, in seconds. T0 is a preset reference time constant. For indicator functions, Represents the sender of the i-th message. Represents the sender of the (i-1)th message. Indicator Function Output 1, The indicator function outputs 0.

[0023] It can indicate whether the (i-1)th and i-th messages belong to the same topic segment or dialogue substructure. The larger the value, the stronger the correlation between the (i-1)th and the ith message, and the greater the probability that the (i-1)th and ith messages belong to the same topic segment or dialogue substructure; conversely, the smaller the value, the lower the probability. The smaller the value, the weaker the correlation between the (i-1)th and the ith message, and the less likely that the (i-1)th and ith messages belong to the same topic segment or dialogue substructure. In the above formula, it is required that... , , The sum of the three is 1, and the implementer can adaptively set the values ​​according to the actual scenario, such as the importance of different items. For example, in this embodiment, it can be set as follows: . Reflecting semantic continuity and relevance, cosine similarity in natural language processing typically has a range of [0,1]. This is because the semantic vectors of a message are usually trained on large-scale corpora (e.g., Word2Vec, BERT), and these models tend to map semantically similar words or sentences to similar directions in space. Therefore, values ​​less than 0 are not considered. The value ranges from 0 to 1; The closer a value is to 1, the higher the semantic overlap or semantic continuity between the two messages. The more likely the (i-1)th and ithth messages belong to the same topic segment or dialogue substructure, the more likely they are to belong to the same topic segment or dialogue substructure. The closer a value is to 0, the less semantically related the two messages are, and the less likely they are to belong to the same topic segment or dialogue substructure. The time decay term measures how closely two messages occur in time. In a conversation, messages that are closer in time are considered to have lower decay. The smaller the value, the higher the likelihood that they belong to the same discussion flow. The smaller the value, the closer the time decay term is to 1, and the greater the probability that the two messages belong to the same topic segment or dialogue substructure; the implementer can set T0 according to the actual situation, such as in this embodiment, T0 can be set to 300 seconds. For indicator functions, when When, that is, when the sender of the (i-1)th message is the same as the sender of the ith message, the indicator function F... Output 1, indicating that when the sender of the (i-1)th message is different from the sender of the ith message, the indicator function F is executed. Output 0, that is Represents the sender of the i-th message. The pointer function F represents the sender of the (i-1)th message. The output can reflect whether the same person speaks continuously. Since the same person sends multiple messages in a row, it is often to express the same point of view or to complete a complete expression. Therefore, when the output is 1, it indicates that the two messages are more likely to belong to the same topic segment or dialogue substructure.

[0024] After obtaining the comprehensive correlation strength between adjacent messages in the preceding message sequence of the currently being analyzed message, the preceding message sequence is divided based on the comparison between the comprehensive correlation strength and a preset segmentation threshold to generate the corresponding session fragment for the currently being analyzed message. The specific process for obtaining the session fragment for the currently being analyzed message is as follows: It is determined whether the comprehensive correlation strength between adjacent messages in the preceding message sequence is less than the preset segmentation threshold. If it is less, a potential topic breakpoint is marked between these two messages, indicating a possible topic switch. If it is greater, these two messages are considered to belong to the same topic fragment, and no topic breakpoint marking is required between them. That is, for the (i-1)th message and the ith message in the preceding message sequence of the currently being analyzed message, if... If the potential topic breakpoint is less than the preset segmentation threshold, a potential topic breakpoint is marked between the (i-1)th message and the ith message. Otherwise, no potential topic breakpoint is marked between the (i-1)th message and the ith message. This yields all potential topic breakpoints in the preceding message sequence of the currently being analyzed message. Then, the preceding message sequence of the currently being analyzed message is segmented using the potential topic breakpoints in the preceding message sequence, generating continuous message subsequences. Each generated message subsequence is recorded as a session fragment corresponding to the currently being analyzed message. If there are no potential topic breakpoints in the preceding message sequence, then the preceding message sequence is a session fragment. In specific applications, the implementer can set the preset segmentation threshold according to the actual situation. For example, in this embodiment, it can be set to 0.4.

[0025] Then, each session fragment corresponding to the currently analyzed message is defined as the first type of context source of the currently analyzed message, that is, each session fragment is a context source. Since all of a person's previous statements are highly relevant contexts when understanding someone's current statement, it is necessary to obtain the individual statement lines of the participants in the preceding message sequence and define them as the second type of context source. That is, in the preceding message sequence of the currently analyzed message, the message sequence formed by all the statements of each participant in chronological order is defined as the second type of context source of the currently analyzed message. In other words, each message sequence formed by all the messages belonging to the same sender in the preceding message sequence of the currently analyzed message in chronological order is a second type of context source. And each second type of context source may be a topic fragment or a sequence of all the statements of a certain participant. Then, the set formed by all the first type of context sources and all the second type of context sources of the currently analyzed message is denoted as the current context source set of the currently analyzed message.

[0026] Therefore, this embodiment obtains the current context source set of the message currently being analyzed through the above process. In addition, this embodiment processes the messages in the session one by one or according to the position order in the processed group chat session message sequence. That is, it starts from the first message in the processed group chat session message sequence to understand the session structure. Therefore, when a message is traversed in sequence and the session structure of the message is understood in sequence, the message that is focused on is defined as the message currently being analyzed. That is, the message currently being analyzed is not a specific message, but a message that points to any message being analyzed in the group chat session during the analysis process.

[0027] Step S003: Calculate the relevance weights based on the semantic affinity, temporal proximity, and interaction tightness between the currently analyzed message and each context source in the current context source set; perform weighted fusion of the semantic vectors of the current context source set based on the relevance weights to obtain the context semantic vector of the currently analyzed message; input the comprehensive semantic vector obtained by fusing the context semantic vector of the currently analyzed message with the semantic vector of the corresponding message itself into the downstream task model, and output the conversation structure understanding result of the currently analyzed message.

[0028] After obtaining the current context source set, this embodiment uses a dynamic, relevance-based optimization strategy to improve the reliability of subsequent context semantic vector extraction. The main content of the strategy is as follows: First, based on the current context source set of the message being analyzed obtained in the above steps, the relevance between the message being analyzed and each context source in the current context source set is evaluated. When evaluating the relevance, not only semantic similarity is considered, but also the unique interaction patterns in group chat are captured. The relevance evaluation result can reflect the extent to which the message being analyzed depends on or belongs to a certain context source. Therefore, the contribution of the context source to the message being analyzed or the importance of the context source to the message being analyzed can be obtained based on the relevance evaluation result. The importance is used for the weighted fusion of the semantic vectors of the context sources.

[0029] Based on the above description, this embodiment next needs to determine the context source relevance of the currently analyzed message based on the current context source set of the currently analyzed message, and then obtain the relevance weight. The relevance weight is the contribution or importance of the semantic vector of the context source when fusing. This embodiment mainly measures the relevance from three dimensions: semantic affinity, temporal proximity, and interaction tightness. In this embodiment, the specific process of obtaining the relevance weight between the currently analyzed message and each context source in the current context source set based on the semantic affinity, temporal proximity, and interaction tightness between the currently analyzed message and each context source in the current context source set is as follows:

[0030] Since the evaluation and calculation method for the relevance weight between the currently analyzed message and each context source in the current context source set of the currently analyzed message is the same, for ease of understanding, this embodiment will next describe the calculation and evaluation process of the relevance weight between the currently analyzed message and the k-th context source in the current context source set as an example. The specific process for obtaining the relevance weight between the currently analyzed message and the k-th context source in the current context source set is as follows:

[0031] Calculate the mean cosine similarity between the semantic vector of the currently analyzed message and the semantic vectors of all messages in the k-th context source. This is also the mean cosine similarity between the semantic vector of the currently analyzed message and the semantic vectors of all messages in the k-th context source, and is denoted as the semantic affinity between the currently analyzed message and the k-th context source. Obtain the latest message in the k-th context source, which is the message in the k-th context source whose sending time is closest to the sending time of the currently analyzed message. The result of applying a negative exponential decay mapping to the sending time interval between the latest message in the k-th context source and the currently analyzed message is denoted as the temporal proximity between the currently analyzed message and the k-th context source, expressed as: t represents the sending time of the message currently being analyzed. T1 represents the sending time of the latest message in the k-th context source, and T1 is the preset decay constant. In specific applications, implementers can set the preset decay constant according to actual conditions such as the bias of the business scenario. If it is necessary to capture instantaneous topic jumps, or if messages in the group are refreshed very quickly and old messages instantly lose their reference value, then the preset decay constant can be set to tens of seconds to minutes. However, if the group chat discussion is in-depth and the topic may not return for several hours, or if the interval between participants' speeches is long and it is necessary to cross a long silence period to maintain the continuity of the topic, then it can be set to tens of minutes or even several hours.

[0032] Obtain the number of messages in the k-th context source that have direct interaction with the sender of the message currently being analyzed. The ratio of this number to the total number of messages in the k-th context source is denoted as the interaction density between the currently analyzed message and the k-th context source. In other words, the proportion of messages in the k-th context source that have direct interaction with the sender of the currently analyzed message represents the interaction density between the currently analyzed message and the k-th context source. Furthermore, direct interaction with the sender of the currently analyzed message means that for any message in the k-th context source, if the message explicitly mentions the sender of the currently analyzed message, or if the sender of the currently analyzed message is replying to the message, then the message has direct interaction with the sender of the currently analyzed message.

[0033] Then, the semantic affinity, temporal proximity, and interaction density between the currently analyzed message and the k-th context source are weighted, fused, and normalized to obtain the relevance weight between the currently analyzed message and the k-th context source in the current context source set. The expression for the relevance weight between the currently analyzed message and the k-th context source in the current context source set is as follows:

[0034]

[0035] in, Let exp() be the relevance weight between the currently analyzed message and the k-th context source in the current context source set, and let e be the exponential function with base e. , and All are contribution factors. Currently, we are analyzing the semantic affinity between the message and the k-th context source. The current analysis focuses on the temporal proximity between the message being analyzed and the k-th context source. K0 represents the interaction density between the currently analyzed message and the k-th context source, where K0 is the number of context sources in the current context source set corresponding to the currently analyzed message.

[0036] The above formula is a variation of the Softmax function, where the input is a weighted sum of the semantic affinity, temporal proximity, and interaction tightness between the currently analyzed message and the k-th context source. Through the Softmax operation, it transforms the semantic affinity, temporal proximity, and interaction tightness between the currently analyzed message and the context sources in the current context source set into a probability distribution. Therefore... The value range is from 0 to 1. The larger the value, the more important and relevant the k-th context source is to understanding the message being analyzed, and the greater its weight will be in the subsequent construction of the dynamic context vector. Conversely, the smaller the value, the less important the context source is to understand the message being analyzed. The smaller the value, the less important and less relevant the k-th context source is to understanding the message being analyzed, and the smaller its weight will be when constructing the dynamic context vector later. Measure the overall semantic similarity or continuity between the currently analyzed message and the k-th context source. The larger the value, the more semantically similar and relevant the currently analyzed message is to the k-th context source; conversely, the smaller the value, the more relevant the message is to the context source. The smaller the value, the lower or less semantically related the currently analyzed message is to the k-th context source. This technique is used to capture the effect of temporal proximity, meaning that in a conversation, recently concluded sessions are generally more relevant to the message being analyzed than much older sessions. The larger the value, the closer the currently analyzed message is to the k-th context source in time, and thus the higher the correlation or relevance between the currently analyzed message and the k-th context source. The smaller the value, the further away the currently analyzed message is from the k-th context source in time, which in turn indicates a lower degree of correlation or relevance between the currently analyzed message and the k-th context source. This quantifies the strength of the social interaction between the currently analyzed message sender and the k-th context source. The larger the value, the deeper the discussion between the message sender and the k-th context source being analyzed, indicating a higher degree of correlation or relevance between the message and the k-th context source being analyzed, and vice versa. The smaller the value, the less discussion the message sender is involved in regarding the k-th context source, indicating a lower degree of correlation or relevance between the message and the k-th context source. , and For adjustment , , The relative contributions of the three feature dimensions to the relevance assessment can be used by implementers. , and Set as an experience value, such as setting , , .

[0037] After obtaining the relevance weights between the currently analyzed message and each context source in the current context source set, the semantic vectors of the current context source set are weighted and fused based on the relevance weights to obtain the context semantic vector of the currently analyzed message. Specifically: first, the semantic vectors of all messages in each context source in the current context source set are averaged and pooled, and this result is denoted as the representation vector of the corresponding context source. Average pooling refers to adding the values ​​of the corresponding positions of the semantic vectors of all messages in a context source and then dividing by the number of semantic vectors in that context source. Then, the relevance weights between the currently analyzed message and each context source in the current context source set are weighted and accumulated to obtain the vector of the corresponding context source representation vector, which is denoted as the context semantic vector of the currently analyzed message. The expression for the context semantic vector of the currently analyzed message is as follows: , Let be the representation vector of the k-th context source in the current context source set of the message being analyzed. The relevance weight between the currently analyzed message and the kth context source in the current context source set is used. The essence of this operation is the application of the attention mechanism at the context source granularity. The model does not treat all historical messages equally, but rather pays attention to those truly relevant dialogue substructures or individual speaking history based on the relevance weight, making the context representation more discriminative and information-dense.

[0038] In this embodiment, after obtaining the context semantic vector of the currently analyzed message, the context semantic vector of the currently analyzed message is fused with the semantic vector of the corresponding message itself to obtain the comprehensive semantic vector of the currently analyzed message. This comprehensive semantic vector is then input into the downstream task model, and the conversation structure understanding result of the currently analyzed message is output. The comprehensive semantic vector of the currently analyzed message is: , This is the semantic vector of the message itself that is currently being analyzed. This is the context semantic vector of the message currently being analyzed. Additionally, when the currently analyzed message is the first M messages in the processed group chat message sequence, its synthesized semantic vector is not constructed; instead, its self-generated semantic vector is directly input into the downstream task model. Furthermore, as another implementation, it is also possible to choose not to perform session structure understanding on the first M messages of the processed group chat message sequence. In specific applications, the implementer can set M according to the actual situation, but M must be no less than 2; for example, M can be set to 2.

[0039] The downstream task model is also a classifier. In this embodiment, the downstream task model is a fully connected neural network, which mainly receives the comprehensive semantic vector as input and is used to perform topic boundary binary classification tasks or dialogue behavior recognition tasks. That is, for topic segmentation tasks, the downstream task model determines whether the message being analyzed is the starting point of a topic. For dialogue behavior recognition tasks, the downstream task model can predict the final dialogue behavior label of the message being analyzed. The training process of the downstream task model is a well-known technique.

[0040] After obtaining the conversation structure understanding results of each message in the processed group chat message sequence, a structured conversation analysis report is generated based on the conversation structure understanding results of all messages in the processed group chat message sequence. The structured conversation analysis report includes at least topic thread reconstruction, dialogue behavior sequence analysis, participant contribution analysis, and key information summary generation. Topic thread reconstruction refers to organizing all messages marked with the same topic in chronological order to form clear topic discussion threads. Keywords, core participants, and start and end times can be extracted from each thread; dialogue behavior sequence analysis refers to analyzing the dialogue behavior sequence (such as "question-answer-agree", "proposal-discussion-disagree-correction-pass") within each topic thread, which can identify the decision-making process, consensus formation process, or conflict points of the thread; participant role and contribution analysis refers to combining the message sender and dialogue behavior to statistically analyze the speaking percentage, number of questions, number of answers, and proposal adoption of each participant in different topics, generating a participant profile; key information summary generation: key information summary generation refers to automatically generating a key information summary of the group chat session using the identified "notification" type messages and the "answer" content in "question-answer pairs".

[0041] Thus, this embodiment completes the optimization of understanding the group chat session structure.

[0042] In summary, this embodiment calculates the comprehensive association strength between adjacent messages in the preceding message sequence based on the semantic vector similarity, sender continuity, and time interval between adjacent messages in the preceding message sequence of the currently analyzed message. The comprehensive association strength is then used to divide the preceding message sequence into session fragments. Each session fragment is defined as a context source, and the message sequence formed by the chronological order of all speeches by each participant in the preceding message sequence is defined as a context source, forming the current context source set for the currently analyzed message. Then, based on the semantic affinity, temporal proximity, and interaction tightness between the currently analyzed message and each context source in the current context source set, a relevance weight is calculated. The semantic vectors of the current context source set are weighted and fused based on the relevance weights to obtain the context semantic vector of the currently analyzed message. The comprehensive semantic vector obtained by fusing the context semantic vector of the currently analyzed message with the semantic vector of the corresponding message itself is input into the downstream task model, and the session structure understanding result of the currently analyzed message is output. Furthermore, this embodiment introduces a dynamic and hierarchical context-aware mechanism, realizing a shift from a one-size-fits-all approach or a static and flat perception of context to a precise and focused one. This significantly improves the accuracy and robustness of understanding the structure of group chat sessions and provides a high-quality, structured, and interpretable data foundation for subsequent advanced applications such as intelligent summarization, intent analysis, and collaboration efficiency evaluation. It fundamentally solves the problem of excessive contextual noise and large comprehension biases when dealing with real and complex group chat scenarios.

[0043] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for optimizing the structure of group chat sessions based on context awareness, characterized in that, The method includes the following steps: Obtain the semantic vector of each message in the processed group chat message sequence; Based on the semantic vector similarity, sender continuity, and time interval between adjacent messages in the preceding message sequence of the currently analyzed message in the processed group chat session message sequence, the comprehensive association strength between adjacent messages in the preceding message sequence is calculated. The sender continuity is used to characterize whether adjacent messages in the preceding message sequence are sent by the same sender. The comprehensive association strength is used to divide the preceding message sequence into session fragments. Each session fragment is defined as a context source, and the message sequence formed by all the speeches of each participant in the preceding message sequence in chronological order is defined as a context source, forming the current context source set of the currently analyzed message. Based on the semantic affinity, temporal proximity, and interaction density between the currently analyzed message and each context source in the current context source set, a relevance weight is calculated. The semantic affinity is the average cosine similarity between the semantic vector of the currently analyzed message and the semantic vectors of all messages in the context source. The temporal proximity is the result of a negative exponential decay mapping between the sending time interval of the latest message in the context source and the currently analyzed message. The interaction density is the proportion of messages in the context source that have direct interaction with the sender of the currently analyzed message. The semantic vectors of the current context source set are weighted and fused based on the relevance weights to obtain the context semantic vector of the currently analyzed message. The comprehensive semantic vector obtained by fusing the context semantic vector of the currently analyzed message with the semantic vector of the corresponding message is input into the downstream task model, and the conversation structure understanding result of the currently analyzed message is output.

2. The method for optimizing the structure of a group chat session based on context awareness as described in claim 1, characterized in that, The overall correlation strength between adjacent messages in the preceding message sequence of the message currently being analyzed is calculated using the following formula: ; in, Let exp() be the overall correlation strength between the (i-1)th message and the ith message in the preceding message sequence of the message being analyzed. exp() is an exponential function with base e. Let i be the semantic vector of the (i-1)th message itself. Let i be the semantic vector of the i-th message itself. , and All are balanced weights. Let be the cosine similarity between the semantic vector of message i and the semantic vector of message (i-1). Ti represents the time interval between the sending time of the (i-1)th message and the sending time of the ith message, in seconds. T0 is a preset reference time constant. For indicator functions, Represents the sender of the i-th message. Represents the sender of the (i-1)th message. Indicator Function Output 1, The indicator function outputs 0.

3. The method for optimizing the structure of a group chat session based on context awareness as described in claim 1, characterized in that, The relevance weight is calculated using the following formula: ; in, Let exp() be the relevance weight between the currently analyzed message and the k-th context source in the current context source set, and let e be the exponential function with a constant e as the base. , and All are contribution factors. Currently, we are analyzing the semantic affinity between the message and the k-th context source. The current analysis focuses on the temporal proximity between the message being analyzed and the k-th context source. K0 represents the interaction density between the currently analyzed message and the k-th context source, where K0 is the number of context sources in the current context source set corresponding to the currently analyzed message.

4. The method for optimizing the structure of a group chat session based on context awareness as described in claim 1, characterized in that, The methods currently being analyzed for obtaining the context semantic vector of a message include: The average pooling result of the semantic vectors of all messages in each context source in the current context source set is denoted as the representation vector of the corresponding context source. The weighted sum of the representation vectors of the corresponding context sources using the correlation weights between the currently analyzed message and each context source in the current context source set is denoted as the context semantic vector of the currently analyzed message.

5. The method for optimizing the structure of a group chat session based on context awareness as described in claim 1, characterized in that, Semantic vectors are obtained through a pre-trained language model.

6. The method for optimizing the structure of a group chat session based on context awareness as described in claim 1, characterized in that, The downstream task model is a fully connected neural network that receives the comprehensive semantic vector as input and is used to perform topic boundary binary classification tasks or dialogue behavior recognition tasks.

Citation Information

Patent Citations

  • WeChat group chat record recognition method and system fusing session scene information

    CN113326373A

  • Intelligent interaction system and method based on dynamic intention recognition

    CN121144495A