A method for extracting key points from conversation content based on reinforcement learning
Through a reinforcement learning-based method, dialogue text preprocessing, Transformer multi-head attention layer and graph convolution network, focus state vectors and graphs are constructed to generate guided topics, which solves the shortcomings of intelligent dialogue systems in topic guidance and adaptability, and improves the accuracy and mediation efficiency of dialogue content extraction.
Patent Information
- Application Number
- CN202510850664.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-06-24
AI Technical Summary
The existing intelligent dialogue system based on online mediation corpus has insufficient topic guidance capabilities and adaptability, and cannot provide targeted topic guidance and adaptive answers.
Using a reinforcement learning method, through dialogue text preprocessing, Transformer multi-headed attention layer, graph convolutional network and reinforcement learning algorithm, focus state vectors and graphs are constructed, reward and punishment rules are designed, and guiding topics are generated.
It improves the richness and accuracy of dialogue content extraction, enhances the timeliness of identification of key information, supports interpretability analysis of the mediation process, and improves mediation efficiency.
Smart Images

Figure CN120353923B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method for extracting key points from conversation content based on reinforcement learning. Background Art
[0002] The current technical solutions for intelligent dialogue research based on online mediation corpus still have shortcomings in topic guidance ability and adaptability.
[0003] The technical solution for conducting intelligent dialogue research based on online mediation corpus still has shortcomings in topic guidance ability and adaptability. First, it is unable to conduct targeted topic guidance based on the mediation goals; second, it is unable to provide targeted and adaptive answers and dialogues based on the feedback content of the parties. Summary of the Invention
[0004] The present invention addresses the technical problems existing in the prior art and provides a method for extracting key points from conversation content based on reinforcement learning.
[0005] The present invention solves the above-mentioned technical problem with the following technical solution: A method for extracting key points from conversation content based on reinforcement learning, comprising the following steps:
[0006] S101. Preprocess the input text sentences and identify the behavioral intentions of verb phrases. Embed the semantic vectors, entity labels, sentiment scores, and action phrases into multi-dimensional features and splice them into a composite feature vector.
[0007] S102: Input the preprocessed dialogue composite feature vector sequence into the Transformer multi-head attention layer. After applying exponential decay to the attention weights of historical sentences, sentiment polarity and entity density are introduced as adjustment factors to optimize the attention weight calculation. Key statements are filtered based on the threshold and stored in the cache.
[0008] S103: Construct the focus state vector and graph evolution. Set the named entities and action phrases of the key statements as graph nodes and assign attributes. Then, establish semantic edges and temporal edges according to the rules to form a graph. Enhance the nodes through the graph convolutional network. Finally, pool the active nodes to generate a vector as the input of reinforcement learning.
[0009] S104. When using reinforcement learning to generate guiding topics, first construct a state space by concatenating the focus state vector and sentiment statistics, design an action space containing the guiding topic and a corresponding template, then set reward and punishment rules, train the policy network using the reinforcement learning algorithm, and select an action to call the template based on probability;
[0010] S105. Based on the current state vector, a greedy strategy is used to select the action with the largest Q value from the action space. The candidate topics are repeatedly verified and logically verified. A predefined template is called based on the action type. The pre-trained Seq2Seq model optimizes the fluency of the sentence and outputs a guiding topic that meets the mediation scenario.
[0011] In a preferred embodiment, in S101, a regular expression is first used in combination with a syntactic analysis tool to segment the input text T into independent sentences S=RegexSplit(T,Ω) according to sentence boundaries, where S represents a set of segmented sentences, and RegexSplit(T,Ω) represents a function that uses a regular expression to segment the text according to Ω. After denoising, the specific calculation formula for denoising is as follows:
[0012] s i =FilterNoise(s i )
[0013] Among them, s i Indicates the i-th sentence, FilterNoise(s i ) represents a text cleaning function that removes blank lines and noise, and uses the pre-trained BERT model to generate a semantic vector t for each sentence. i =Tokenize(s i ), where t i Indicates the word segmentation result of the i-th sentence, Tokenize(s i ) represents the word segmentation function to capture the contextual semantic association, and at the same time, the key entities E of name, time, place and amount are annotated by named entity technology. i =NER(s i ,τ), where τ represents the entity type set and NER represents the named entity function. The sentiment polarity score of the sentence is calculated by training the sentiment classification model. The specific calculation formula is as follows:
[0014] Emotion i =SentimentModel(Q i )
[0015] Among them, SentimentModel(s i ) represents the emotion classification model function, Emotion i represents the sentiment polarity score of the i-th sentence, where (Q i ∈[-1,1]) and then extract verb phrases through part-of-speech tagging to identify behavioral intentions. Finally, the semantic vectors, entity labels, sentiment scores, and action phrases are embedded in multi-dimensional features and spliced into a composite feature vector. After standardization, it is stored as structured data containing sentence text, feature vectors, entity sentiment, and action phrases.
[0016] In a preferred embodiment, in S102, the composite feature vector sequence obtained by preprocessing the dialogue text is input into the Transformer multi-head attention layer, and the query matrix is obtained by linear transformation. Bond Matrix Sum Matrix in, Indicates the i-th attention head, used to convert the composite feature vector f i The weight matrix mapped to the query matrix, Indicates the i-th attention head, used to convert the composite feature vector f i The weight matrix mapped to the key matrix, Indicates the i-th attention head, used to convert the composite feature vector f i Map the weight matrix to the value matrix and calculate the original attention score matrix of the current sentence and all historical sentences. The specific calculation formula of the single-head attention score is as follows:
[0017]
[0018] in, represents the raw attention score of the current round t and the historical round i in the hth attention head, t represents the current dialogue round, that is, the current sentence index when calculating the attention, i represents the historical dialogue round, that is, the historical sentence index when calculating the attention, h represents the attention head number, corresponding to the hth subspace in the multi-head attention mechanism, d k represents the feature dimension of each attention head, represents the query vector of the current round t in the h-th attention head, represents the key vector of historical round i in the h-th attention head, (·) T Represents matrix transposition, transposing the key vector into a column vector for matrix multiplication with the query vector, and applying exponential decay to the attention weight of the historical sentences. The specific calculation formula is as follows:
[0019] γ t,i =λ t-i
[0020] Among them, γ t,i represents the exponential decay coefficient, λ represents the time decay rate, and the specific calculation formula of the attenuated attention score is as follows:
[0021]
[0022] in, Represents the attenuated attention score, Score t,iRepresents the attention score after the fusion of the current round t and the historical round i, highlighting the impact of recent dialogues and introducing emotional polarity Emotion i and entity density Among them, n i Indicates the number of entities, n max Represents the maximum number of entities, which serves as a regulating factor to optimize the attention weight calculation. The specific calculation formula is as follows:
[0023]
[0024] Among them, α t,i It represents the adjusted normalized attention weight, which is used to measure the importance of historical round i to the current round t. exp represents the exponential function. μ and v represent hyperparameters, which adjust the influence weights of sentiment polarity and entity density. Key statements are filtered according to the set attention threshold and stored in the dynamic cache.
[0025] In a preferred embodiment, in S103, the named entities and action phrases in the key statements are used as graph nodes, and the nodes are given attributes such as frequency of occurrence, emotional tendency, and first appearance timestamp. Then, semantic edges and temporal edges are established according to the edge connection rules. If two nodes appear in the same statement or in adjacent rounds and the semantic similarity is higher than the threshold, an undirected semantic edge with a weight of semantic similarity is established. At the same time, a directed temporal edge reflecting the logical sequence relationship is established according to the order of the first appearance of the nodes to form a focus evolution graph. Then, the node neighborhood information is aggregated through a graph convolutional network to enhance the node representation. Finally, the enhanced representation of all currently active nodes is subjected to mean pooling operation to generate a focus state vector that integrates the semantics, emotion, and temporal relationship of the dispute focus as one of the input features of reinforcement learning.
[0026] In a preferred embodiment, in S104, when using the reinforcement learning strategy to generate the guiding topic, the focus state vector is first concatenated with the emotional statistics of the current conversation to construct a state space that fully reflects the semantic focus and emotional state of the conversation, providing basic input for subsequent strategic decision-making. Next, the action space is designed to divide the guiding topics into three categories: fact clarification, responsibility confirmation, and solution negotiation. Preset natural language templates are prepared for each type of topic. These templates can generate specific guiding content by filling in relevant information of the current focus node. In the reward function design link, effective guiding topics are generated for the incentive strategy network, and positive and negative reward rules are set: if the generated topic successfully guides to a new focus or promotes consensus, a reward of +R1 is given. The specific calculation formula is as follows:
[0027] R focus =R1·(I new_focus +I consensus )
[0028] Among them, R focus represents the new focus guidance reward value, R1 represents the reward coefficient, I new_focus Indicator function indicating the emergence of a new focus, I consensus This is an indicator function indicating consensus. If the topic increases the emotional polarity of the conversation and reduces negative emotions, the reward is +R2. The specific calculation formula is as follows:
[0029] R sentiment =R2·(Δs+Δp-)
[0030] Among them, R sentiment represents the reward value of the emotion reward optimization, R2 represents the reward function, Δs=s t -s t-1 represents the change in sentiment polarity between the current round and the previous round s∈[-1,1], represents the decrease in the proportion of negative emotions p - ∈[0,1], when the topic is repeated and fails to elicit an effective response from the user, causing the user to repeat an existing statement, a negative reward of -R3 is given. The specific calculation formula is as follows:
[0031] R penalty =-R3·(I duplicate +I invalid )
[0032] Among them, R penalty represents the penalty value, R3 represents the penalty coefficient, I duplicate Indicator function for topic repetition, I invalid An indicator function indicates that the user has not responded effectively. R1, R2, and R3 are dynamically adjustable core parameters used to balance the weights and influence of different reward conditions. The policy network is trained using a deep reinforcement learning algorithm. During the training process, the policy network takes the state vector as input and outputs the probability distribution of each action. It continuously interacts with the environment and optimizes the parameters based on reward feedback. In the inference stage, the policy network selects the action with the highest probability based on the action probability distribution output by the current state, then calls the corresponding natural language template, fills in the focus node parameters, and finally generates a guiding topic in natural language form.
[0033] In a preferred embodiment, in S105, in the reasoning stage, first, based on the current state vector, the action with the largest Q value is selected from the action space through a greedy strategy. After the action is selected, the candidate topics are subjected to duplication detection and logical consistency verification. By calculating the semantic similarity between the candidate topics and the current focus state vector and setting a threshold, topics that are irrelevant to the current focus or repeated are excluded. After completion of the verification, the natural language generation link is entered, and the predefined template is called according to the topic type to which the action belongs. The parameters such as entities and keywords in the focus state are filled into the corresponding positions of the template to generate a preliminary guiding sentence. Finally, the pre-trained Seq2Seq model is used to optimize the fluency of the sentences generated by the template, and the sentence structure is adjusted through the encoding-decoding process to ensure that the language is natural and meets the professional standards of the mediation scenario, and finally the guiding topic is output.
[0034] The beneficial effects of the present invention are: the present invention combines multimodal information of semantics, entities, emotions, and action phrases to improve the richness and accuracy of content representation, dynamically captures the temporal correlation and semantic focus in the conversation through time decay and multi-head attention, enhances the timeliness of key information identification, uses GCN to construct a focus evolution map, intuitively presents the logical evolution and interaction relationship of the dispute points, supports the interpretable analysis of the mediation process, designs the reward function driven by the mediation goal, realizes the adaptive optimization of the topic generation strategy, and improves the mediation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 Flow chart of the method of the present invention. DETAILED DESCRIPTION
[0036] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.
[0037] like Figure 1 This embodiment provides: a method for extracting key points from conversation content based on reinforcement learning, comprising the following steps:
[0038] S101. Preprocess the input text sentences and identify the behavioral intentions of verb phrases. Embed the semantic vectors, entity labels, sentiment scores, and action phrases into multi-dimensional features and splice them into a composite feature vector.
[0039] Furthermore, we first use regular expressions combined with syntactic analysis tools to split the input text T into independent sentences S = RegexSplit(T,Ω) according to sentence boundaries, where S represents the set of sentences after segmentation, and RegexSplit(T,Ω) represents the function that uses regular expressions to split the text according to Ω. After denoising, the specific calculation formula for denoising is as follows:
[0040] s i =FilterNoise(s i )
[0041] Among them, s i Indicates the i-th sentence, FilterNoise(s i ) represents a text cleaning function that removes blank lines and noise, and uses the pre-trained BERT model to generate a semantic vector t for each sentence. i =Tokenize(s i ), where t i Indicates the word segmentation result of the i-th sentence, Tokenize(s i ) represents the word segmentation function to capture the contextual semantic association, and at the same time, the key entities E of name, time, place and amount are annotated by named entity technology. i =NER(s i ,τ), where τ represents the entity type set and NER represents the named entity function. The sentiment polarity score of the sentence is calculated by training the sentiment classification model. The specific calculation formula is as follows:
[0042] Emotion i =SentimentModel(Q i )
[0043] Among them, SentimentModel(s i ) represents the emotion classification model function, Emotion i represents the sentiment polarity score of the i-th sentence, where (Q i ∈[-1,1]) and then extract verb phrases through part-of-speech tagging to identify behavioral intentions. Finally, the semantic vectors, entity labels, sentiment scores, and action phrases are embedded in multi-dimensional features and spliced into a composite feature vector. After standardization, it is stored as structured data containing sentence text, feature vectors, entity sentiment, and action phrases.
[0044] S102: Input the preprocessed dialogue composite feature vector sequence into the Transformer multi-head attention layer. After applying exponential decay to the attention weights of historical sentences, sentiment polarity and entity density are introduced as adjustment factors to optimize the attention weight calculation. Key statements are filtered based on the threshold and stored in the cache.
[0045] Furthermore, the composite feature vector sequence obtained by preprocessing the conversation text is input into the Transformer multi-head attention layer, and the query matrix is obtained through linear transformation Bond Matrix Sum Matrix in, Indicates the i-th attention head, used to convert the composite feature vector f i The weight matrix mapped to the query matrix, Indicates the i-th attention head, used to convert the composite feature vector f i The weight matrix mapped to the key matrix, Indicates the i-th attention head, used to convert the composite feature vector f i Map the weight matrix to the value matrix and calculate the original attention score matrix of the current sentence and all historical sentences. The specific calculation formula of the single-head attention score is as follows:
[0046]
[0047] in, represents the raw attention score of the current round t and the historical round i in the hth attention head, t represents the current dialogue round, that is, the current sentence index when calculating the attention, i represents the historical dialogue round, that is, the historical sentence index when calculating the attention, h represents the attention head number, corresponding to the hth subspace in the multi-head attention mechanism, d k represents the feature dimension of each attention head, represents the query vector of the current round t in the h-th attention head, represents the key vector of historical round i in the h-th attention head, (·) T Represents matrix transposition, transposing the key vector into a column vector for matrix multiplication with the query vector, and applying exponential decay to the attention weight of the historical sentences. The specific calculation formula is as follows:
[0048] γ t,i =λ t-i
[0049] Among them, γ t,i represents the exponential decay coefficient, λ represents the time decay rate, and the specific calculation formula of the attenuated attention score is as follows:
[0050]
[0051] in, Represents the attenuated attention score, Score t,i Represents the attention score after the fusion of the current round t and the historical round i, highlighting the impact of recent dialogues and introducing emotional polarity Emotion iand entity density Among them, n i Indicates the number of entities, n max Represents the maximum number of entities, which serves as a regulating factor to optimize the attention weight calculation. The specific calculation formula is as follows:
[0052]
[0053] Among them, α t,i It represents the adjusted normalized attention weight, which is used to measure the importance of historical round i to the current round t. exp represents the exponential function. μ and v represent hyperparameters, which adjust the influence weights of sentiment polarity and entity density. Key statements are filtered according to the set attention threshold and stored in the dynamic cache.
[0054] S103: Construct the focus state vector and graph evolution. Set the named entities and action phrases of the key statements as graph nodes and assign attributes. Then, establish semantic edges and temporal edges according to the rules to form a graph. Enhance the nodes through the graph convolutional network. Finally, pool the active nodes to generate a vector as the input of reinforcement learning.
[0055] Furthermore, the named entities and action phrases in the key statements are used as graph nodes, and the nodes are assigned attributes such as frequency of occurrence, emotional tendency, and first appearance timestamp. Then, semantic edges and temporal edges are established according to the edge connection rules. If two nodes appear in the same statement or in adjacent rounds and the semantic similarity is higher than the threshold (such as cosine similarity>0.7), an undirected semantic edge with the weight of semantic similarity is established. At the same time, a directed temporal edge reflecting the logical sequence relationship is established according to the order of the first appearance of the nodes to form a focus evolution graph. Then, the node neighborhood information is aggregated through the graph convolutional network to enhance the node representation. Finally, the enhanced representation of all currently active nodes is average pooled to generate a focus state vector that integrates the semantics, emotion, and temporal relationship of the dispute focus, which is used as one of the input features of reinforcement learning.
[0056] S104. When using reinforcement learning to generate guiding topics, first construct a state space by concatenating the focus state vector and sentiment statistics, design an action space containing the guiding topic and a corresponding template, then set reward and punishment rules, train the policy network using the reinforcement learning algorithm, and select an action to call the template based on probability;
[0057] Furthermore, when using reinforcement learning strategies to generate guiding topics, the focus state vector is first concatenated with the emotional statistics of the current conversation (including average emotional polarity, proportion of negative emotions, etc.) to construct a state space that fully reflects the semantic focus and emotional state of the conversation, providing basic input for subsequent strategic decision-making. Next, the action space is designed to divide the guiding topics into three categories: fact clarification, responsibility confirmation, and solution negotiation. Preset natural language templates are prepared for each type of topic. These templates can generate specific guiding content by filling in relevant information of the current focus node (such as entity name, dispute type). In the reward function design phase, effective guiding topics are generated for the incentive strategy network, and positive and negative reward rules are set: if the generated topic successfully guides to a new focus or promotes consensus, a reward of +R1 is given. The specific calculation formula is as follows:
[0058] R focus =R1·(I new_focus +I consensus )
[0059] Among them, R focus represents the new focus guidance reward value, R1 represents the reward coefficient, I new_focus Indicator function indicating the emergence of a new focus, I consensus This is an indicator function indicating consensus. If the topic increases the emotional polarity of the conversation and reduces negative emotions, the reward is +R2. The specific calculation formula is as follows:
[0060] R sentiment =R2·(Δs+Δp-)
[0061] Among them, R sentiment represents the reward value of the emotion reward optimization, R2 represents the reward function, Δs=s t -s t-1 represents the change in sentiment polarity between the current round and the previous round s∈[-1,1], represents the decrease in the proportion of negative emotions p - ∈[0,1], when the topic is repeated and fails to elicit an effective response from the user, causing the user to repeat an existing statement, a negative reward of -R3 is given. The specific calculation formula is as follows:
[0062] R penalty =-R3·(I duplicate +I invalid )
[0063] Among them, R penalty represents the penalty value, R3 represents the penalty coefficient, I duplicate Indicator function for topic repetition, I invalidAn indicator function indicates that the user has not responded effectively. R1, R2, and R3 are dynamically adjustable core parameters used to balance the weights and influence of different reward conditions. The policy network is trained using a deep reinforcement learning algorithm. During the training process, the policy network takes the state vector as input and outputs the probability distribution of each action. It continuously interacts with the environment and optimizes the parameters based on reward feedback. In the inference stage, the policy network selects the action with the highest probability based on the action probability distribution output by the current state, then calls the corresponding natural language template, fills in the focus node parameters, and finally generates a guiding topic in natural language form.
[0064] S105. Based on the current state vector, a greedy strategy is used to select the action with the largest Q value from the action space. The candidate topics are repeatedly verified and logically verified. A predefined template is called based on the action type. The pre-trained Seq2Seq model optimizes the fluency of the sentence and outputs a guiding topic that meets the mediation scenario.
[0065] Furthermore, in the reasoning stage, first, based on the current state vector, the action with the largest Q value is selected from the action space through a greedy strategy. After the action is selected, the candidate topics are subjected to duplication detection and logical consistency verification. By calculating the semantic similarity between the candidate topic and the current focus state vector and setting a threshold, topics that are irrelevant to the current focus or repeated are excluded. After completion of the verification, the natural language generation link is entered. According to the topic type to which the action belongs (such as fact clarification, responsibility determination, compensation negotiation), the predefined template is called, and the entities, keywords and other parameters in the focus state are filled into the corresponding positions of the template to generate a preliminary guiding sentence. Finally, the pre-trained Seq2Seq model is used to optimize the fluency of the sentences generated by the template, and the sentence structure is adjusted through the encoding-decoding process to ensure that the language is natural and meets the professional standards of the mediation scenario, and finally the guiding topic is output.
[0066] It should be noted that the method for extracting key points of the present invention is to divide the dialogue text into sentences, generate semantic vectors through BERT, simultaneously extract features such as entities, emotions, action phrases, etc., splice them into composite vectors, build a multi-dimensional data foundation, calculate the semantic weights between sentences based on Transformer multi-head attention, superimpose the time decay factor to highlight recent content, optimize the weights by combining emotional polarity and entity density, screen high-weight sentences as key statements, abstract the entities and action phrases in the key statements into graph nodes, construct a graph through semantic edges (similarity) and temporal edges (logical order), enhance node representation by aggregating neighborhood information using GCN, pool active nodes to generate focus state vectors, fuse focus states and emotion statistics to define state space, design three types of topic-guiding action sets, use focus promotion and emotion improvement as positive rewards, and repeat invalid actions as negative penalties, train the policy network through the PPO algorithm, select actions according to probability and call templates to generate natural language-guided topics.
[0067] It should be noted that, in the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0068] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
Claims
1. A method for extracting key points from conversation content based on reinforcement learning, characterized in that: The following steps are involved: S101. Preprocess the input text sentence and identify the behavioral intention of the verb phrase, embed the semantic vector, entity label, sentiment score, and action phrase into multi-dimensional features and splice them into a composite feature vector; S102: Input the preprocessed dialogue composite feature vector sequence into the Transformer multi-head attention layer. After applying exponential decay to the attention weights of historical sentences, sentiment polarity and entity density are introduced as adjustment factors to optimize the attention weight calculation. Key statements are filtered based on the threshold and stored in the cache. S103: Construct the focus state vector and graph evolution. Set the named entities and action phrases of the key statements as graph nodes and assign attributes. Then, establish semantic edges and temporal edges according to the rules to form a graph. Enhance the nodes through the graph convolutional network. Finally, pool the active nodes to generate a vector as the input of reinforcement learning. S104. When using reinforcement learning to generate guiding topics, first construct a state space by concatenating the focus state vector and sentiment statistics, design an action space containing the guiding topic and a corresponding template, then set reward and punishment rules, train the policy network using the reinforcement learning algorithm, and select an action to call the template based on probability; S105. Based on the current state vector, a greedy strategy is used to select the action with the largest Q value from the action space, and candidate topics are repeatedly checked and logically checked. Predefined templates are called according to the action type, and the pre-trained Seq2Seq model optimizes the fluency of sentences and outputs a guiding topic that meets the mediation scenario.
2. The method for extracting key points from conversation content based on reinforcement learning according to claim 1, characterized in that: In S101, regular expressions are first used in combination with syntactic analysis tools to split the input text into independent sentences according to sentence boundaries. After removing blank lines and noise, the pre-trained BERT model is used to generate a semantic vector for each sentence to capture the contextual semantic association. At the same time, the key entities of name, time, place and amount are annotated through named entity technology. The sentiment polarity score of the sentence is calculated with the help of a trained sentiment classification model. Then, verb phrases are extracted through part-of-speech tagging to identify behavioral intentions. Finally, the semantic vector, entity label, sentiment score and action phrase are embedded in multi-dimensional features and spliced into a composite feature vector. After standardization, it is stored as structured data containing sentence text, feature vectors, entity sentiment and action phrases.
3. The method for extracting key points from conversation content based on reinforcement learning according to claim 1, characterized in that: In S102, the composite feature vector sequence obtained by preprocessing the dialogue text is input into the Transformer multi-head attention layer, and the query matrix is obtained through linear transformation Bond Matrix Sum Matrix in, Indicates the i-th attention head, used to convert the composite feature vector f i The weight matrix mapped to the query matrix, Indicates the i-th attention head, used to convert the composite feature vector f i The weight matrix mapped to the key matrix, Indicates the i-th attention head, used to convert the composite feature vector f i Map the weight matrix to the value matrix and calculate the original attention score matrix of the current sentence and all historical sentences. The specific calculation formula of the single-head attention score is as follows: in, represents the raw attention score of the current round t and the historical round i in the hth attention head, t represents the current dialogue round, that is, the current sentence index when calculating the attention, i represents the historical dialogue round, that is, the historical sentence index when calculating the attention, h represents the attention head number, corresponding to the hth subspace in the multi-head attention mechanism, d k represents the feature dimension of each attention head, represents the query vector of the current round t in the h-th attention head, represents the key vector of historical round i in the h-th attention head, (·) T Represents matrix transposition, transposes the key vector into a column vector, and applies exponential decay to the attention weight of historical sentences. The specific calculation formula is as follows: c t,i =λ t-i Among them, γ t,i represents the exponential decay coefficient, λ represents the time decay rate, and the specific calculation formula of the attenuated attention score is as follows: in, Represents the attenuated attention score, Score t,i Represents the attention score after the fusion of the current round t and the historical round i, highlighting the impact of recent dialogues and introducing emotional polarity Emotion i and entity density Among them, n i Indicates the number of entities, n max Represents the maximum number of entities, which serves as a regulating factor to optimize the attention weight calculation. The specific calculation formula is as follows: Among them, α t,i It represents the adjusted normalized attention weight, which is used to measure the importance of historical round i to the current round t. exp represents the exponential function. μ and v represent hyperparameters, which adjust the influence weights of sentiment polarity and entity density. Key statements are filtered according to the set attention threshold and stored in the dynamic cache.
4. The method for extracting key points from conversation content based on reinforcement learning according to claim 1, characterized in that: In S103, named entities and action phrases in key statements are used as graph nodes, and the nodes are assigned attributes such as frequency of occurrence, emotional tendency, and first appearance timestamp. Then, semantic edges and temporal edges are established according to edge connection rules. If two nodes appear in adjacent rounds and the semantic similarity is higher than the threshold, an undirected semantic edge with a weight equal to the semantic similarity is established. At the same time, a directed temporal edge reflecting the logical sequence relationship is established according to the order of the first appearance of the nodes to form a focus evolution graph. Then, the node neighborhood information is aggregated through a graph convolutional network to enhance the node representation. Finally, the enhanced representation of all currently active nodes is average pooled to generate a focus state vector that integrates the semantics, emotion, and temporal relationship of the dispute focus, which is used as one of the input features of reinforcement learning.
5. The method for extracting key points from conversation content based on reinforcement learning according to claim 1, characterized in that: When using reinforcement learning strategies to generate guiding topics in S104, the focus state vector is first concatenated with the sentiment statistics of the current conversation to construct a state space that comprehensively reflects the semantic focus and emotional state of the conversation, providing basic input for subsequent strategic decision-making. Next, the action space is designed to divide guiding topics into three categories: fact clarification, responsibility confirmation, and solution negotiation. Preset natural language templates are prepared for each type of topic. In the reward function design phase, effective guiding topics are generated for the incentive strategy network, and positive and negative reward rules are set.
6. The method for extracting key points from conversation content based on reinforcement learning according to claim 5, characterized in that: Positive and negative reward rules: If the generated topic is successfully guided to a new focus, a +R1 reward will be given. The specific calculation formula is as follows: R focus =R1·(I new_focus +I consensus ) Among them, R focus represents the new focus guidance reward value, R1 represents the reward coefficient, I new_focus Indicator function indicating the emergence of a new focus, I consensus This is an indicator function indicating consensus. If the topic increases the emotional polarity of the conversation and reduces negative emotions, the reward is +R2. The specific calculation formula is as follows: R sentiment =R2·(Δs+Δp-) Among them, R sentiment represents the reward value of the emotion reward optimization, R2 represents the reward function, Δs=s t -s t-1 represents the change in sentiment polarity between the current round and the previous round s∈[-1,1], represents the decrease in the proportion of negative emotions p - ∈[0,1], when the topic is repeated and fails to elicit an effective response from the user, causing the user to repeat an existing statement, a negative reward of -R3 is given. The specific calculation formula is as follows: R penalty =-R3·(I duplicate +I invalid ) Among them, R penalty represents the penalty value, R3 represents the penalty coefficient, I duplicate Indicator function for topic repetition, I invalid An indicator function indicates that the user has not responded effectively. R1, R2, and R3 are dynamically adjustable core parameters used to balance the weights and influence of different reward conditions. The policy network is trained using a deep reinforcement learning algorithm. During the training process, the policy network takes the state vector as input and outputs the probability distribution of each action. It continuously interacts with the environment and optimizes the parameters based on reward feedback. In the inference stage, the policy network selects the action with the highest probability based on the action probability distribution output by the current state, then calls the corresponding natural language template, fills in the focus node parameters, and finally generates a guiding topic in natural language form.
7. The method for extracting key points from conversation content based on reinforcement learning according to claim 1, characterized in that: In S105, in the reasoning stage, first, based on the current state vector, the action with the largest Q value is selected from the action space through a greedy strategy. After the action is selected, the candidate topics are subjected to duplication detection and logical consistency verification. By calculating the semantic similarity between the candidate topics and the current focus state vector and setting a threshold, topics that are irrelevant to the current focus or repeated are excluded. After completing the verification, the natural language generation link is entered. The predefined template is called according to the topic type to which the action belongs. The entity and keyword parameters in the focus state are filled into the corresponding positions of the template to generate a preliminary guiding sentence. Finally, the pre-trained Seq2Seq model is used to optimize the fluency of the sentence generated by the template. The sentence structure is adjusted through the encoding-decoding process to finally output the guiding topic.
Citation Information
Patent Citations
Information diffusion prediction system based on space-time attention and heterogeneous graph convolutional network
CN113807616A
Method and system for context semantic extraction and intention recognition in voice dialogue
CN119862891A