Method and system for understanding long text user opinions based on memory enhancement
Through a memory-enhanced neural network model, combined with a hierarchical encoder, key-value memory network and pointer generation network, the complexity problem of long text user opinions is solved, high-quality structured abstracts and keyword generation are achieved, and the accuracy and robustness of complex opinions are improved.
Patent Information
- Application Number
- CN202411414525.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-11
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-10-11
AI Technical Summary
The prior art is difficult to effectively understand and analyze long-text user opinions, especially when dealing with scenarios with linguistic diversity, high information density and strong context dependence, and lacks robustness to low-frequency, novel and complex opinions expression methods.
The long text user opinion understanding method based on memory enhancement is adopted, and preprocessed and understood through the memory-enhanced neural network model, including a combination of hierarchical encoder, key-value memory network and pointer generation network, semantic features are extracted using multi-head attention mechanism and gated recurrent units, and long-term context information is stored and updated through the key-value memory network, and structured opinion summary and keywords are finally generated.
It realizes a deep understanding and information extraction of long-text user opinions, can effectively capture multi-grained semantic information in the text, generate high-quality structured abstracts and keywords, and improves the accuracy and robustness of complex opinions.
Smart Images

Figure CN119293148B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to artificial intelligence technology, and in particular to a method and system for understanding long-text user opinions based on memory enhancement. Background Art
[0002] User opinion texts usually have the characteristics of language diversity, high information density, and strong context dependence. Especially in the scenario of long texts, it is difficult to achieve ideal results by directly applying traditional natural language processing and text mining technologies. Long text user opinions often contain multiple topics or opinions. There are complex semantic associations and context dependencies between different topics. The expression of opinions is flexible and varied, and there are a lot of colloquial and non-standard language phenomena. These characteristics bring great challenges to the accurate understanding of user opinions and information extraction.
[0003] Most existing long text opinion understanding methods are based on local features, ignoring the global semantics and contextual connections of long texts, making it difficult to accurately grasp the core content and topic structure of opinions. In addition, existing methods lack robustness to low-frequency, novel, and complex opinion expressions, and have limited effectiveness in dealing with open domains and dynamically changing user opinions. Summary of the invention
[0004] The embodiments of the present invention provide a method and system for understanding long-text user opinions based on memory enhancement, which can solve the problems in the prior art.
[0005] According to a first aspect of the embodiments of the present invention,
[0006] Provides a long text user opinion understanding method based on memory enhancement, including:
[0007] Obtain the long text user opinions to be understood, and preprocess the long text user opinions, including word segmentation, part-of-speech tagging, named entity recognition and dependency syntactic analysis, to obtain a text representation rich in linguistic features; use the word vector model to map the text representation into a low-dimensional dense vector as the input of the memory-enhanced neural network model, the word vector model uses a pre-trained language model and performs incremental training on the domain corpus; the memory-enhanced neural network model is constructed using a multi-head attention mechanism and a gated recurrent unit, including three modules: a hierarchical encoder, a key-value memory network, and a pointer generation network; the hierarchical encoder is used to extract the semantic features of the text at the word, phrase, sentence and paragraph levels; the key-value memory network uses a multi-layer sparse memory matrix, realizes fast addressing through local sensitive hashing, and introduces a forgetting mechanism to dynamically update the memory, which is used to store long-term context information and enhance the semantic features extracted by the hierarchical encoder; the pointer generation network is used to decode the semantic features enhanced by memory using a beam search algorithm, combined with a vocabulary masking mechanism and a length penalty, to generate structured opinion summaries and keywords, and can selectively copy fragments in the input text;
[0008] Post-process the opinion summaries and keywords output by the memory-enhanced neural network model to obtain standardized structured opinion information; based on the structured opinion information, use the pre-built business domain knowledge base to understand and analyze user opinions and generate an opinion understanding report. The business domain knowledge base is built in a combination of ontology and rules, including domain concepts, entities, relationships, and constraint rules, and achieves deep opinion understanding through knowledge reasoning;
[0009] Based on the opinion understanding report, combined with the domain dialogue knowledge base and user portrait, responses are automatically generated in the form of text, voice, and charts. The domain dialogue knowledge base is learned using multi-round dialogue data and includes intent recognition, slot filling, and dialogue strategy learning functions to achieve context-related, logically self-consistent multi-round responses. Through the reinforcement learning algorithm, the response strategy is dynamically adjusted using user feedback to achieve human-computer interaction optimization.
[0010] In an optional embodiment,
[0011] The memory-enhanced neural network model is constructed using a multi-head attention mechanism and a gated recurrent unit, and includes three modules: a hierarchical encoder, a key-value memory network, and a pointer generation network. The steps include:
[0012] For each layer of the hierarchical encoder, a gated recurrent unit is used to model the input sequence. The reset gate and update gate are introduced to dynamically control the flow and update of information, alleviate the gradient vanishing problem, and obtain the hidden state sequence.
[0013] Use a multi-head attention mechanism to calculate the matching degree between the hidden state sequence and the memory matrix, obtain the query matrix, key matrix and value matrix through linear transformation, multiply the query matrix by the transpose of the key matrix and divide it by the scaling factor to obtain the attention score matrix, normalize the attention score matrix using a normalized exponential function to obtain the attention weight matrix, use the attention weight matrix to perform weighted summation on the value matrix to obtain the attention output matrix, and concatenate to obtain the multi-head attention output, enhance the ability of the memory-enhanced neural network model to capture the diverse interactions between text and memory, and obtain the memory readout vector;
[0014] The memory readout vector is concatenated with the hidden state sequence as the input of the next encoder layer to achieve information transfer and feature fusion between the layers of the hierarchical encoder;
[0015] In the decoder, a gated recurrent unit is used to model the sequence of memory readout vectors, and a multi-head attention mechanism is used to calculate the match between the decoder hidden state and the encoder output to obtain the context vector.
[0016] The context vector is concatenated with the decoder hidden state, and after linear transformation and normalized exponential function, the probability distribution of the generated words is obtained, and the structured opinion summary and keywords are generated by decoding.
[0017] In an optional embodiment,
[0018] The hierarchical encoder is used to extract the semantic features of the text at the word, phrase, sentence and paragraph levels; the key-value memory network adopts a multi-layer sparse memory matrix, realizes fast addressing through local sensitive hashing, and introduces a forgetting mechanism to dynamically update the memory, so as to store long-term context information and enhance the semantic features extracted by the hierarchical encoder; the pointer generation network is used to decode the semantic features enhanced by the memory using a beam search algorithm, and generate structured opinion summaries and keywords in combination with a vocabulary masking mechanism and a length penalty, and can selectively copy fragments in the input text, the steps include:
[0019] The hierarchical encoder is used to receive input long text user opinions, extract character-level local features through character-level convolutional neural networks, capture long-distance dependencies between words through word-level self-attention mechanisms, then use sentence-level convolutional neural networks to extract sentence-level local features, and finally use paragraph-level gated recurrent units to extract paragraph-level global features, thereby extracting multi-granularity hierarchical semantic representations of user opinions from the bottom up;
[0020] The key-value memory network includes a key matrix, a value matrix and a read-write controller, wherein the key matrix and the value matrix are both multi-layer sparse matrices for storing long-term context information; the read-write controller is used to receive the multi-granularity hierarchical semantic representation extracted by the hierarchical encoder, and through the attention mechanism and the local sensitive hashing algorithm, quickly retrieve the memory fragment most relevant to the current input in the key matrix, and read the corresponding memory content from the corresponding value matrix, so as to enhance the multi-granularity hierarchical semantic representation extracted by the hierarchical encoder; the read-write controller dynamically updates the key matrix and the value matrix through the gating mechanism, and introduces a forgetting mechanism to discard expired memory that is no longer needed, so that the key-value memory network can adaptively store and update long-term context information;
[0021] The pointer generation network takes the multi-granularity hierarchical semantic representation enhanced by the key-value memory network as input, and adopts a decoder with an attention mechanism to decode and generate respectively the summary generation task and the keyword extraction task. During the decoding process, the pointer network mechanism is used to allow key fragments and words to be copied from the input user opinions, and a beam search algorithm is introduced for decoding, a vocabulary mask mechanism is used to constrain the decoding space, and a length-based reward and punishment mechanism is used to control the length of the generated sequence, so as to generate structured summary text and keyword sequence as output.
[0022] In an optional embodiment,
[0023] The steps of introducing a beam search algorithm for decoding, a word list mask mechanism to constrain the decoding space, and a length-based reward and punishment mechanism to control the length of the generated sequence and generating a structured summary text and keyword sequence as output include:
[0024] Encode the input text into a hidden state sequence, initialize the decoder's hidden state, and the candidate set containing only the sequence start symbol;
[0025] In each decoding step, for each sequence in the candidate set, the decoder hidden state is updated, the attention distribution and context vector are calculated, and the final word probability distribution is obtained by combining the vocabulary probability distribution and the replication probability distribution;
[0026] Based on the beam search strategy, multiple extensions with the highest probability are selected from the word probability distribution and added to the candidate set. At the same time, a word list mask is introduced to remove the generated words from the word list to avoid repeated generation;
[0027] Calculate the score of each candidate sequence, including the log-likelihood score and the length-based reward and penalty items, and control the preference for shorter or longer sequences by setting the reward and penalty factors;
[0028] Repeat the decoding steps until the maximum length is reached or all sequences in the candidate set are terminated by a sequence terminator, and select the sequence with the highest score as the final generated result;
[0029] During the generation process, the decoding strategy is controlled by adjusting the bundle size, vocabulary mask switch, and length reward and penalty factors to meet different task requirements and performance optimization goals.
[0030] In an optional embodiment,
[0031] The steps of post-processing the opinion summary and keywords output by the memory-enhanced neural network model to obtain standardized structured opinion information include:
[0032] According to the number of occurrences of the keyword in the opinion text set, the document frequency of each keyword is calculated to obtain the document frequency result reflecting the distribution breadth of the keyword in the opinion text;
[0033] Based on the document frequency of the keyword and the total number of opinion texts, the inverse document frequency of each keyword is calculated by taking the logarithm operation to obtain the inverse document frequency result that measures the discrimination of the keyword in the entire text collection;
[0034] Count the number of times each keyword appears in the corresponding opinion summary to obtain the word frequency result that reflects the contribution of the keyword to the summary content;
[0035] Multiply the keyword's word frequency and inverse document frequency to get the word frequency-inverse document frequency (TF-IDF) value that comprehensively considers the keyword's importance in the abstract and its overall discrimination;
[0036] According to the TF-IDF value of the keyword, all the keywords in the opinion summary are sorted in descending order, and several keywords with the highest TF-IDF value are selected as the final keyword screening results to highlight the core content of the opinion summary.
[0037] In an optional embodiment,
[0038] The business domain knowledge base is constructed by combining ontology and rules, including domain concepts, entities, relationships and constraint rules. The steps of achieving deep opinion understanding through knowledge reasoning include:
[0039] According to the entities and relations identified in the opinion information, Simple Protocol and Resource Description Framework Query Language (SPARQL) query statements are constructed to retrieve triples related to the entities and relations from the knowledge base stored in the Resource Description Framework format;
[0040] Match the retrieved triple knowledge with the predefined IF-THEN reasoning rules in the knowledge base, trigger the rules through forward reasoning or backward reasoning algorithms, and generate new triple knowledge according to the consequences of the rules;
[0041] By using the category hierarchical relationship in the ontology, based on the category attributes of the entity, combined with the sub-category and parent-category relationships, the knowledge scope and granularity of the opinion information can be expanded through ontology reasoning;
[0042] Check the consistency of knowledge generated during the reasoning process, automatically identify attribute value conflicts, category conflicts, and relationship conflicts through rule-based confidence comparison, and resolve conflicts based on domain knowledge and business rules to ensure the consistency and reliability of reasoning results;
[0043] The new knowledge obtained by reasoning is integrated with the original opinion entities and relations. The confidence of knowledge and the credibility of the source are considered. The probabilistic graphical model is used to model and infer knowledge from different sources to obtain a unified opinion knowledge representation and form an opinion understanding report.
[0044] In an optional embodiment,
[0045] Through the reinforcement learning algorithm, the response strategy is dynamically adjusted using user feedback. The steps to achieve human-computer interaction optimization include:
[0046] The opinion interaction process is modeled as a Markov decision process, and the state space is defined to represent the interaction state based on the user's opinion query, system response history, user feedback, and dialogue turn information;
[0047] Define the system action space, including response action types such as opinion summary, opinion comparison, detail enumeration, multiple rounds of clarification, and greeting replies;
[0048] Combining the user's feedback rating on the system's response and the pre-trained dialogue quality evaluation model, we designed a reward function for the state-action pair, taking into account both user satisfaction and response quality.
[0049] The state-action value function is represented by a function approximation method, and a deep neural network is used to fit the value function. The input of the deep neural network is the interaction state feature, and the output is the value estimate of the response action.
[0050] Use the experience replay mechanism for offline strategy learning, store the state transition samples generated during the interaction process into the experience replay pool, and update the value network parameters by randomly sampling the replay samples;
[0051] When generating actual response actions, the ε-greedy strategy is used to balance exploration and exploitation, select the optimal action based on the value estimate of the current state, and gradually reduce the exploration probability as training progresses;
[0052] The trained value network is converted into a lightweight online decision-making model through model compression and knowledge distillation technology to improve decision-making efficiency and response speed.
[0053] According to a second aspect of the embodiments of the present invention,
[0054] Provides a long text user opinion understanding system based on memory enhancement, including:
[0055] The first unit is used to obtain the long text user opinions to be understood, and pre-process the long text user opinions, including word segmentation, part-of-speech tagging, named entity recognition and dependency syntactic analysis, to obtain a text representation rich in linguistic features; the word vector model is used to map the text representation into a low-dimensional dense vector as the input of the memory-enhanced neural network model, and the word vector model uses a pre-trained language model and performs incremental training on the domain corpus; the memory-enhanced neural network model is constructed using a multi-head attention mechanism and a gated recurrent unit, including three modules: a hierarchical encoder, a key-value memory network, and a pointer generation network; the hierarchical encoder is used to extract the semantic features of the text at the word, word, sentence and paragraph levels; the key-value memory network uses a multi-layer sparse memory matrix, realizes fast addressing through local sensitive hashing, and introduces a forgetting mechanism to dynamically update the memory, which is used to store long-term context information and enhance the semantic features extracted by the hierarchical encoder; the pointer generation network is used to decode the semantic features enhanced by memory using a beam search algorithm, combined with a vocabulary masking mechanism and a length penalty, to generate structured opinion summaries and keywords, and can selectively copy fragments in the input text;
[0056] The second unit is used to post-process the opinion summary and keywords output by the memory-enhanced neural network model to obtain standardized structured opinion information; based on the structured opinion information, the user opinions are understood and analyzed using a pre-built business domain knowledge base to generate an opinion understanding report. The business domain knowledge base is constructed in a combination of ontology and rules, including domain concepts, entities, relationships and constraint rules, and achieves deep opinion understanding through knowledge reasoning;
[0057] The third unit is used to automatically generate responses based on the opinion understanding report, combined with the domain dialogue knowledge base and user portrait. The response forms include text, voice, and charts. The domain dialogue knowledge base is learned using multi-round dialogue data and includes intent recognition, slot filling, and dialogue strategy learning functions to achieve context-related, logically self-consistent multi-round responses. Through the reinforcement learning algorithm, the response strategy is dynamically adjusted using user feedback to achieve human-computer interaction optimization.
[0058] According to a third aspect of the embodiments of the present invention,
[0059] An electronic device is provided, comprising:
[0060] processor;
[0061] a memory for storing processor-executable instructions;
[0062] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0063] A fourth aspect of the embodiments of the present invention is:
[0064] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.
[0065] This paper demonstrates the end-to-end operation of the memory-enhanced pointer generation network, demonstrating the effectiveness and interpretability of key technologies such as hierarchical encoding, memory enhancement, and pointer generation. The network can fully utilize the multi-granularity semantic information of long texts, store and update long-term context information through a memory mechanism, and capture and replicate key information through a pointer generation mechanism, generating high-quality structured summaries and keywords.
[0066] This paper uses deep opinion understanding based on knowledge reasoning to make full use of the ontology and rules in the business domain knowledge base, mine the implicit knowledge in opinion information, expand the knowledge scope and granularity of opinion information, and improve the comprehensiveness and accuracy of opinion understanding. At the same time, through consistency checking and knowledge fusion, the reliability and consistency of reasoning results are ensured, and high-quality opinion understanding reports are generated to provide valuable reference for subsequent business decisions. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 A flowchart of a method for understanding long text user opinions based on memory enhancement according to an embodiment of the present invention;
[0068] Figure 2 It is a structural diagram of a long text user opinion understanding system based on memory enhancement according to an embodiment of the present invention. DETAILED DESCRIPTION
[0069] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0070] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0071] Figure 1FIG. 1 is a flow chart of a method for understanding long text user opinions based on memory enhancement according to an embodiment of the present invention. Figure 1 As shown, the method includes:
[0072] S1. Obtain the long text user opinions to be understood, and preprocess the long text user opinions, including word segmentation, part-of-speech tagging, named entity recognition and dependency syntactic analysis, to obtain a text representation rich in linguistic features; use the word vector model to map the text representation into a low-dimensional dense vector as the input of the memory-enhanced neural network model, the word vector model uses a pre-trained language model and performs incremental training on the domain corpus; the memory-enhanced neural network model is constructed using a multi-head attention mechanism and a gated recurrent unit, including three modules: a hierarchical encoder, a key-value memory network, and a pointer generation network; the hierarchical encoder is used to extract the semantic features of the text at the word, word, sentence and paragraph levels; the key-value memory network uses a multi-layer sparse memory matrix, realizes fast addressing through local sensitive hashing, and introduces a forgetting mechanism to dynamically update the memory, which is used to store long-term context information and enhance the semantic features extracted by the hierarchical encoder; the pointer generation network is used to decode the semantic features enhanced by memory using a beam search algorithm, combined with a vocabulary masking mechanism and a length penalty, to generate structured opinion summaries and keywords, and can selectively copy fragments in the input text;
[0073] S2. Post-process the opinion summary and keywords output by the memory-enhanced neural network model to obtain standardized structured opinion information; based on the structured opinion information, use the pre-built business domain knowledge base to understand and analyze the user opinions and generate an opinion understanding report. The business domain knowledge base is built in a way that combines ontology and rules, including domain concepts, entities, relationships and constraint rules, and achieves deep opinion understanding through knowledge reasoning;
[0074] S3. Based on the opinion understanding report, combined with the domain dialogue knowledge base and user portrait, the response is automatically generated. The response form includes text, voice, and chart. The domain dialogue knowledge base is learned using multi-round dialogue data, and includes intent recognition, slot filling, and dialogue strategy learning functions to achieve context-related and logically self-consistent multi-round responses; through the reinforcement learning algorithm, the response strategy is dynamically adjusted using user feedback to achieve human-computer interaction optimization.
[0075] In an optional embodiment,
[0076] The memory-enhanced neural network model is constructed using a multi-head attention mechanism and a gated recurrent unit, and includes three modules: a hierarchical encoder, a key-value memory network, and a pointer generation network. The steps include:
[0077] For each layer of the hierarchical encoder, a gated recurrent unit is used to model the input sequence. The reset gate and update gate are introduced to dynamically control the flow and update of information, alleviate the gradient vanishing problem, and obtain the hidden state sequence.
[0078] Use a multi-head attention mechanism to calculate the matching degree between the hidden state sequence and the memory matrix, obtain the query matrix, key matrix and value matrix through linear transformation, multiply the query matrix by the transpose of the key matrix and divide it by the scaling factor to obtain the attention score matrix, normalize the attention score matrix using a normalized exponential function to obtain the attention weight matrix, use the attention weight matrix to perform weighted summation on the value matrix to obtain the attention output matrix, and concatenate to obtain the multi-head attention output, enhance the ability of the memory-enhanced neural network model to capture the diverse interactions between text and memory, and obtain the memory readout vector;
[0079] The memory readout vector is concatenated with the hidden state sequence as the input of the next encoder layer to achieve information transfer and feature fusion between the layers of the hierarchical encoder;
[0080] In the decoder, a gated recurrent unit is used to model the sequence of memory readout vectors, and a multi-head attention mechanism is used to calculate the match between the decoder hidden state and the encoder output to obtain the context vector.
[0081] The context vector is concatenated with the decoder hidden state, and after linear transformation and normalized exponential function, the probability distribution of the generated words is obtained, and the structured opinion summary and keywords are generated by decoding.
[0082] For example, the memory-enhanced neural network model proposed in this paper is constructed using a multi-head attention mechanism and a gated recurrent unit, including three modules: a hierarchical encoder, a key-value memory network, and a pointer generation network. The following will introduce the model construction process and key technical details in detail.
[0083] The hierarchical encoder uses multi-layer stacked gated recurrent units (GRU) to model the sequence of input text, extracting and abstracting the semantic features of the text layer by layer. For each layer of the encoder, the input sequence is first encoded using GRU, and the flow and update of information are dynamically controlled by introducing reset gates and update gates to alleviate the gradient vanishing problem.
[0084] Specifically, the calculation process of GRU's reset gate, update gate, and candidate hidden state is as follows:
[0085] Reset gate: The value of the reset gate is obtained by concatenating the current input and the hidden state of the previous moment, and then performing a linear transformation and a sigmoid activation function. This is used to control the extent to which information from the previous moment flows into the current moment.
[0086] Update gate: Similar to the reset gate, the update gate value is obtained by concatenating the current input and the hidden state of the previous moment, and then performing a linear transformation and a sigmoid activation function to control the degree of information retention at the previous moment.
[0087] Candidate hidden state: According to the value of the reset gate, the hidden state of the previous moment is selectively combined with the current input, and the candidate hidden state is obtained after linear transformation and tanh activation function.
[0088] By controlling the reset gate and the update gate, GRU can adaptively decide to forget or update the information of the previous moment and combine it with the candidate hidden state to generate the hidden state of the current moment.
[0089] After GRU sequence modeling, the hidden state sequence of each layer of the encoder can be obtained, denoted as H l , where l represents the number of encoder layers.
[0090] Next, use the multi-head attention mechanism to calculate the hidden state sequence H l The semantic association and matching degree between the text and the external memory matrix M. The multi-head attention mechanism captures the diverse interaction patterns between text and memory through multiple different attention heads, enhancing the representation ability of the model.
[0091] For each attention head, first the hidden state sequence H l , memory matrix M is converted into query matrix Q, key matrix K and value matrix V respectively, and then the similarity between query matrix Q and key matrix K is calculated to obtain the attention score matrix. Next, the normalized exponential function (softmax function) is applied to the attention score matrix to normalize it and obtain the attention weight matrix, which represents the matching degree between the query and the key. Finally, the attention weight matrix is used to perform weighted summation on the value matrix V to obtain the output of the attention head.
[0092] The outputs of all attention heads are concatenated to obtain the output matrix of the multi-head attention, denoted as O. Through the multi-head attention mechanism, the model can capture the interaction pattern between text and memory from different perspectives and enhance the richness and diversity of semantic representation.
[0093] The output matrix O of the multi-head attention is combined with the hidden state sequence H l The concatenation is used as the input of the next layer of encoder to realize information transfer and feature fusion between encoder layers.
[0094] Through the hierarchical encoder design, the model can extract and abstract the semantic features of the text layer by layer, while integrating external memory information to enhance the understanding and representation capabilities of long texts.
[0095] Key-value memory networks are used to store and retrieve external knowledge, providing relevant background information and domain knowledge for the model. Memory networks organize knowledge in the form of key-value pairs, where the key is used to match queries and the value is used to store the corresponding knowledge content.
[0096] Let the memory matrix of the key-value memory network be M, where each row represents a key-value pair, the number of rows of M is the number of memories, and the number of columns is the dimension of the key-value pair. The key and value can be represented by word vectors, knowledge graph entity vectors, etc., which are used to encode the semantic information of knowledge.
[0097] During model training, the memory matrix M can be initialized and updated in the following ways:
[0098] Random initialization: The keys and values of the memory matrix M are randomly initialized as vectors of fixed dimensions, which are automatically optimized and adjusted through backpropagation during training.
[0099] Pre-trained word vectors: Use pre-trained word vectors to initialize the keys and values of the memory matrix M, representing knowledge as distributed word embedding vectors.
[0100] Knowledge graph embedding: Using the structured information of the knowledge graph, entities and relationships are embedded into a low-dimensional vector space through knowledge representation learning algorithms (such as TransE, ComplEx, etc.), and the keys and values of the memory matrix M are initialized.
[0101] Dynamic update: During the model training process, the memory matrix M is dynamically updated according to the text content and task feedback, and the text information is written into the memory through the attention mechanism, continuously optimizing and enriching the memory knowledge.
[0102] Through the key-value memory network, the model can flexibly store and retrieve external knowledge, providing relevant background information and domain knowledge support for text understanding and generation.
[0103] The pointer generation network is used to decode and generate structured opinion summaries and keywords. By combining the pointer mechanism and the generation mechanism, the replication and generation of the input text are achieved.
[0104] In the decoder, GRU is first used to sequence model the output of the encoder to obtain the hidden state sequence S of the decoder.
[0105] Then, a multi-head attention mechanism is used to calculate the attention distribution a between the decoder hidden state and the encoder output H t ,Through attention distribution, the decoder can focus on the relevant information in the input text when generating each target word, ,and achieve the replication of keywords and fragments.
[0106] Next, the attention distribution a tThe weighted summation with the encoder output H gives the context vector c t , the context vector combines the current state of the decoder with relevant information about the input text.
[0107] The context vector c t It is concatenated with the decoder hidden state, and after linear transformation and softmax function, the generation probability distribution p of the target word is obtained. gen , represents the probability of generating a new word from the vocabulary.
[0108] In order to realize the pointer mechanism, the pointer probability p is introduced ptr Represents the probability of copying a word from the input text, and the context vector c is transformed into t and the decoder hidden states are mapped to pointer probabilities.
[0109] Finally, the decoder generates each target word y t When the probability of passing through the pointer is p ptr To decide is to generate the probability distribution p gen Sampling to generate new words, or sampling from the attention distribution a t Select Copy words from input text.
[0110] Through the pointer generation network, the model can flexibly generate new words and copy keywords, improving the accuracy and interpretability of opinion summarization and keyword generation.
[0111] The above is the construction process and key technical details of the memory-enhanced neural network model. Through the organic combination of hierarchical encoders, key-value memory networks, and pointer generation networks, the model can make full use of external knowledge to enhance the understanding and representation of long texts. At the same time, through the combination of pointer mechanisms and generation mechanisms, accurate and interpretable opinion summaries and keyword generation are achieved.
[0112] The following is a specific data case to illustrate the operation process of the model.
[0113] Suppose the input text is: "The hotel room was spacious and clean, with comfortable beds and good facilities. The staff was friendly and helpful and provided excellent service during our stay. The hotel is conveniently located with easy access to public transportation and nearby attractions. However, the room was a bit noisy due to street traffic outside. Overall, we enjoyed our stay and would recommend this hotel."
[0114] The text is encoded through a hierarchical encoder to obtain a hidden state sequence H. Assuming that the encoder has two layers and the hidden state dimension of each layer is 128, the shape of H is (number of text words, 2, 128).
[0115] The key-value memory network stores relevant background knowledge, such as knowledge vectors corresponding to keywords such as "hotel", "room", "staff", "service", and "location". Assume that the shape of the memory matrix M is (number of memories, knowledge vector dimension), where the knowledge vector dimension is 128.
[0116] Through the multi-head attention mechanism, the attention weights between the hidden state sequence H and the memory matrix M are calculated to obtain the representation O that integrates external knowledge. Assuming that 4 attention heads are used, the shape of O is (number of text words, 4*128).
[0117] Concatenate O and H as the input of the next layer of encoder to continue encoding the text and fusing knowledge.
[0118] In the decoder, a multi-head attention mechanism is used to calculate the attention distribution a between the decoder hidden state and the encoder output H t Assuming the decoder hidden state dimension is 128 and the target sequence length is 20, then a t The shape is (20, number of words in the text).
[0119] Through the pointer generation network, when generating each target word, according to the pointer probability p ptr The decision is to generate the probability distribution p gen Sampling to generate new words, or sampling from the attention distribution a t Select the words in the copied input text. Suppose when generating the tth target word, p ptr =0.7, then with a probability of 0.7 t Select the copied input word from p with a probability of 0.3 gen Generate new words in .
[0120] Ultimately, the resulting opinion summary might be: "Spacious, clean room with comfortable bed. Friendly staff, great service. Convenient location, close to public transportation. But a bit noisy due to street traffic. Overall, a pleasant stay, would recommend."
[0121] Keywords might be: "hotel room", "staff", "service", "location", "noisy".
[0122] Through the above data cases, the operation process and generation results of the memory-enhanced neural network model are demonstrated. Through the collaborative work of hierarchical encoders, key-value memory networks, and pointer generation networks, the model achieves in-depth understanding and knowledge fusion of long texts, and generates accurate and interpretable opinion summaries and keywords.
[0123] In an optional embodiment,
[0124] The hierarchical encoder is used to extract the semantic features of the text at the word, phrase, sentence and paragraph levels; the key-value memory network adopts a multi-layer sparse memory matrix, realizes fast addressing through local sensitive hashing, and introduces a forgetting mechanism to dynamically update the memory, so as to store long-term context information and enhance the semantic features extracted by the hierarchical encoder; the pointer generation network is used to decode the semantic features enhanced by the memory using a beam search algorithm, and generate structured opinion summaries and keywords in combination with a vocabulary masking mechanism and a length penalty, and can selectively copy fragments in the input text, the steps include:
[0125] The hierarchical encoder is used to receive input long text user opinions, extract character-level local features through character-level convolutional neural networks, capture long-distance dependencies between words through word-level self-attention mechanisms, then use sentence-level convolutional neural networks to extract sentence-level local features, and finally use paragraph-level gated recurrent units to extract paragraph-level global features, thereby extracting multi-granularity hierarchical semantic representations of user opinions from the bottom up;
[0126] The key-value memory network includes a key matrix, a value matrix and a read-write controller, wherein the key matrix and the value matrix are both multi-layer sparse matrices for storing long-term context information; the read-write controller is used to receive the multi-granularity hierarchical semantic representation extracted by the hierarchical encoder, and through the attention mechanism and the local sensitive hashing algorithm, quickly retrieve the memory fragment most relevant to the current input in the key matrix, and read the corresponding memory content from the corresponding value matrix, so as to enhance the multi-granularity hierarchical semantic representation extracted by the hierarchical encoder; the read-write controller dynamically updates the key matrix and the value matrix through the gating mechanism, and introduces a forgetting mechanism to discard expired memory that is no longer needed, so that the key-value memory network can adaptively store and update long-term context information;
[0127] The pointer generation network takes the multi-granularity hierarchical semantic representation enhanced by the key-value memory network as input, and adopts a decoder with an attention mechanism to decode and generate respectively the summary generation task and the keyword extraction task. During the decoding process, the pointer network mechanism is used to allow key fragments and words to be copied from the input user opinions, and a beam search algorithm is introduced for decoding, a vocabulary mask mechanism is used to constrain the decoding space, and a length-based reward and punishment mechanism is used to control the length of the generated sequence, so as to generate structured summary text and keyword sequence as output.
[0128] Exemplarily, this paper proposes a memory-enhanced pointer generation network for generating structured opinion summaries and keywords. The network consists of three modules: a hierarchical encoder, a key-value memory network, and a pointer generation network. Through bottom-up hierarchical encoding, storage and updating of long-term context information, and decoding generation combining pointer mechanism with beam search, it achieves in-depth understanding of long-text user opinions and interpretable summary generation. The following will introduce the construction steps and key technical details of the memory-enhanced pointer generation network in detail.
[0129] The hierarchical encoder aims to extract multi-granularity hierarchical semantic representations of user opinions from the bottom up, including character-level, word-level, sentence-level, and paragraph-level features.
[0130] First, the character-level local features of the input text are extracted through a character-level convolutional neural network. Specifically, multiple convolution kernels of different sizes are used to slide on the character embedding to extract character combination features of different sizes. Assume that the input text contains n characters and the character embedding dimension is d char , the convolution kernel size is [3,4,5], the number of convolution kernels of each size is c, then the output feature dimension of the character-level convolution is n*(c*3).
[0131] Based on the character-level features, the word-level self-attention mechanism is used to capture the long-distance dependencies between words. First, the character-level features are aggregated into word-level features through the maximum pooling operation to obtain a word-level feature matrix with a dimension of m*(c*3), where m is the number of words in the text. Then, the attention weights between words are calculated through the self-attention mechanism to capture the long-distance dependencies at the word level. Assume that the query, key, and value matrices of the self-attention are d w , the number of heads is h, then the output feature dimension of word-level self-attention is m*d w .
[0132] Based on the word-level features, sentence-level convolutional neural networks are used to extract sentence-level local features. Similar to character-level convolution, multiple convolution kernels of different sizes are used to slide on the word-level features to extract word combination features of different sizes and obtain sentence-level features. Assume that the text contains p sentences, the convolution kernel size is [3, 4, 5], and the number of convolution kernels of each size is c. s , then the output feature dimension of sentence-level convolution is p*(c s *3).
[0133] Finally, the paragraph-level global features are extracted through the paragraph-level gated recurrent unit (GRU). The sentence-level features are used as the input of the GRU, and the hidden state is gradually updated to capture the global information at the paragraph level. Assuming that the text contains q paragraphs, the hidden state dimension of the GRU is d p , then the output feature dimension of paragraph-level GRU is q*dp .
[0134] Through the above-mentioned hierarchical encoder, multi-granular semantic representations of user opinions are extracted from the bottom up to obtain character-level, word-level, sentence-level, and paragraph-level features, providing rich semantic information for the subsequent memory network and pointer generation network.
[0135] The purpose of the key-value memory network is to store long-term contextual information and enhance the semantic features extracted by the hierarchical encoder. The key-value memory network consists of three parts: the key matrix, the value matrix, and the read-write controller.
[0136] Both the key matrix and the value matrix are multi-layer sparse matrices used to store long-term context information. The key matrix is used to store the index information of the memory, and the value matrix is used to store the content information of the memory. In order to improve the storage efficiency and retrieval speed, the locality sensitive hashing (LSH) algorithm is used to compress and index the key matrix.
[0137] The read / write controller is used to receive the multi-granularity semantic representation extracted by the hierarchical encoder and interact with the key-value memory network to realize the reading and writing of the memory.
[0138] In the reading phase, the read / write controller first matches the semantic features extracted by the hierarchical encoder with the key matrix through the attention mechanism, and calculates the memory fragment most relevant to the current input. Specifically, the attention score is calculated using the query vector (the semantic features of the current input) and each key vector in the key matrix, and the attention weight is normalized by the softmax function. Then, the corresponding memory content is retrieved from the value matrix according to the attention weight, and fused with the semantic features of the current input to obtain the semantic representation after memory enhancement.
[0139] In the write phase, the read-write controller dynamically updates the key matrix and value matrix through a gating mechanism, and introduces a forgetting mechanism to discard expired memories that are no longer needed. Specifically, the weights of the write gate and the forget gate are calculated based on the semantic features of the current input and the state of the memory network. The write gate controls the extent to which new memory content is written into the key-value matrix, and the forget gate controls the extent to which expired memories are removed from the key-value matrix. Through the control of the gating mechanism, the key-value memory network can adaptively store and update long-term context information.
[0140] In summary, the key-value memory network stores long-term context information through multi-layer sparse matrices and uses the locality-sensitive hashing algorithm for fast retrieval. The read-write controller uses the attention mechanism and the gating mechanism to read and write the memory, so that the memory network can dynamically store and update the memory according to the semantic features of the current input, and enhance the semantic features extracted by the hierarchical encoder.
[0141] The purpose of the pointer generation network is to generate structured opinion summaries and keywords based on the semantic features after memory enhancement. The pointer generation network uses a beam search algorithm for decoding, and introduces a word list masking mechanism and length penalty to constrain the generation results.
[0142] The pointer generation network uses an attention mechanism decoder to decode the summary generation task and keyword extraction task respectively. The decoder takes the memory-enhanced semantic features as input and adaptively focuses on the output of the encoder through the attention mechanism to generate summaries and keywords.
[0143] The decoder uses GRU as the basic unit. At each time step, the attention weight between the current decoding state and the encoder output is calculated through the attention mechanism. The context vector is obtained by weighted summing the attention weight and the encoder output. Then, the context vector is concatenated with the current decoding state, and the generation probability distribution is obtained through linear transformation and softmax function to generate the next word.
[0144] In order to be able to copy key fragments and vocabulary from the input user opinions, the pointer generation network introduces a pointer network mechanism. During the decoding process, in addition to generating vocabulary by generating probability distribution, vocabulary can also be directly copied from the input text through the pointer network.
[0145] Specifically, at each decoding time step, the generation probability and copy probability are calculated. The generation probability represents the probability of generating a word from a fixed vocabulary, and the copy probability represents the probability of copying a word from the input text. The generation probability and copy probability are weighted and combined through a gating mechanism to obtain the final word generation probability distribution.
[0146] In order to generate higher quality summaries and keywords, the pointer generation network uses a beam search algorithm for decoding. The beam search algorithm retains the beam size optimal candidate sequences at each time step, and gradually generates the optimal solution by expanding the candidate sequences and selecting the top beam size sequences with the highest scores.
[0147] In the process of beam search, a length penalty mechanism is introduced to control the length of the generated sequence. The length penalty encourages the generation of sequences of moderate length and avoids the generation of overly long or short summaries and keywords by adding a penalty term related to the sequence length to the score of the candidate sequence.
[0148] The pointer generation network also introduces a word list mask mechanism to constrain the decoding space. When generating summaries, the word list mask is set to limit the decoder to only generate relevant words in the summary field, avoiding the generation of irrelevant or wrong words. When generating keywords, the word list mask is set to limit the decoder to only generate keyword-type words such as nouns and verbs, thereby improving the accuracy of keyword extraction.
[0149] In summary, the pointer generation network generates summaries and keywords through the decoder of the attention mechanism, uses the pointer network mechanism to achieve the replication of key fragments and vocabulary, introduces the beam search algorithm for decoding optimization, and uses the word list mask mechanism to constrain the decoding space. Finally, the pointer generation network can generate structured and interpretable opinion summaries and keywords based on the semantic features after memory enhancement.
[0150] This paper demonstrates the end-to-end operation of the memory-enhanced pointer generation network, demonstrating the effectiveness and interpretability of key technologies such as hierarchical encoding, memory enhancement, and pointer generation. The network can fully utilize the multi-granularity semantic information of long texts, store and update long-term context information through a memory mechanism, and capture and replicate key information through a pointer generation mechanism, generating high-quality structured summaries and keywords.
[0151] In an optional embodiment,
[0152] The steps of introducing a beam search algorithm for decoding, a word list mask mechanism to constrain the decoding space, and a length-based reward and punishment mechanism to control the length of the generated sequence and generating a structured summary text and keyword sequence as output include:
[0153] Encode the input text into a hidden state sequence, initialize the decoder's hidden state, and the candidate set containing only the sequence start symbol;
[0154] In each decoding step, for each sequence in the candidate set, the decoder hidden state is updated, the attention distribution and context vector are calculated, and the final word probability distribution is obtained by combining the vocabulary probability distribution and the replication probability distribution;
[0155] Based on the beam search strategy, multiple extensions with the highest probability are selected from the word probability distribution and added to the candidate set. At the same time, a word list mask is introduced to remove the generated words from the word list to avoid repeated generation;
[0156] Calculate the score of each candidate sequence, including the log-likelihood score and the length-based reward and penalty items, and control the preference for shorter or longer sequences by setting the reward and penalty factors;
[0157] Repeat the decoding steps until the maximum length is reached or all sequences in the candidate set are terminated by a sequence terminator, and select the sequence with the highest score as the final generated result;
[0158] During the generation process, the decoding strategy is controlled by adjusting the bundle size, vocabulary mask switch, and length reward and penalty factors to meet different task requirements and performance optimization goals.
[0159] For example, in order to generate high-quality summary text and keyword sequences, this paper introduces a beam search algorithm for decoding, and combines the word list mask mechanism to constrain the decoding space, and the length-based reward and punishment mechanism to control the length of the generated sequence. The following will introduce the technical details and implementation steps of the decoding strategy based on beam search and constraint mechanism.
[0160] First, the input text is mapped into a hidden state sequence through the encoder. Specifically, the hierarchical encoder is used to extract the multi-granular semantic features of the input text, and the memory is enhanced through the key-value memory network to obtain the final hidden state sequence. Then, the hidden state of the decoder is initialized, usually set to the last hidden state of the encoder. At the same time, the candidate set is initialized and set to a set containing only the sequence start symbol.
[0161] In each decoding step, the following operations are performed on each sequence in the candidate set:
[0162] Update decoder hidden state: The word embedding generated in the previous step of the current sequence and the previous decoder hidden state are input into the decoder, such as GRU, to obtain the updated decoder hidden state.
[0163] Calculate attention distribution: Use the attention mechanism to calculate the attention distribution based on the current decoder hidden state and the encoder hidden state sequence. Common attention mechanisms include dot product attention, additive attention, etc.
[0164] Calculate the context vector: According to the attention distribution, the encoder hidden state sequence is weighted summed to obtain the context vector as the context information of the current decoding step.
[0165] Calculate word probability distribution: Combine the decoder hidden state, context vector, and vocabulary embedding matrix, calculate the generation probability of each word in the vocabulary through linear transformation and softmax function, and get the vocabulary probability distribution.
[0166] Calculate the copy probability distribution: Use the pointer network mechanism to calculate the copy probability of each word in the input text based on the decoder hidden state and encoder hidden state sequence to obtain the copy probability distribution.
[0167] Combine word probability distribution and copy probability distribution: Through the gating mechanism, dynamically combine the vocabulary probability distribution and the copy probability distribution to obtain the final word probability distribution.
[0168] Based on the beam search strategy, several extensions with the highest probability are selected from the word probability distribution and added to the candidate set. Specifically, the beam size is set to k, which means that the k candidate sequences with the highest probability are retained in each decoding step. Each sequence in the candidate set is expanded separately to obtain k*v new candidate sequences, where v is the word list size. Then, based on the score of each new sequence, the k sequences with the highest score are selected as the new candidate set and enter the next decoding step.
[0169] The word list mask mechanism is introduced to dynamically remove the generated words from the word list to avoid repeated generation of the same words in the subsequent decoding steps. Specifically, for each sequence in the candidate set, a mask vector with the same size as the word list is generated, and the positions corresponding to the generated words are set to 0, and the remaining positions are set to 1. When calculating the word probability distribution, the mask vector is multiplied element by element with the word list probability distribution, so as to set the probability of the generated words to 0 to avoid repeated generation.
[0170] In order to control the length of the generated sequence, a length-based reward and penalty mechanism is introduced. Specifically, the score of each candidate sequence is adjusted, and the score consists of two parts: the log-likelihood score and the length reward and penalty item. The log-likelihood score represents the logarithmic value of the probability of sequence generation, and the length reward and penalty item rewards or penalizes based on the length of the sequence. By setting the reward and penalty factors, the preference for shorter or longer sequences can be controlled. For example, setting a positive reward factor will encourage the generation of longer sequences, while setting a negative penalty factor will inhibit the generation of overly long sequences.
[0171] Repeat the decoding steps until the preset maximum length is reached or all sequences in the candidate set end with a sequence terminator (such as <eos>) terminates. The sequence with the highest score is selected from the candidate set as the final generated result, i.e., the summary text and keyword sequence.
[0172] By adjusting hyperparameters such as beam size, wordlist mask switch, and length penalty factor, the decoding strategy can be controlled to meet different task requirements and performance optimization goals. For example, increasing the beam width can expand the search space and improve the generation quality, but at the same time increase the computational overhead. Enabling the wordlist mask mechanism can avoid generating repeated words, but may affect the diversity of generation. Adjusting the length penalty factor can control the length distribution of the generated sequence to adapt to different length preferences.
[0173] In an optional embodiment,
[0174] The steps of post-processing the opinion summary and keywords output by the memory-enhanced neural network model to obtain standardized structured opinion information include:
[0175] According to the number of occurrences of the keyword in the opinion text set, the document frequency of each keyword is calculated to obtain the document frequency result reflecting the distribution breadth of the keyword in the opinion text;
[0176] Based on the document frequency of the keyword and the total number of opinion texts, the inverse document frequency of each keyword is calculated by taking the logarithm operation to obtain the inverse document frequency result that measures the discrimination of the keyword in the entire text collection;
[0177] Count the number of times each keyword appears in the corresponding opinion summary to obtain the word frequency result that reflects the contribution of the keyword to the summary content;
[0178] Multiply the keyword's word frequency and inverse document frequency to get the word frequency-inverse document frequency (TF-IDF) value that comprehensively considers the keyword's importance in the abstract and its overall discrimination;
[0179] According to the TF-IDF value of the keyword, all the keywords in the opinion summary are sorted in descending order, and several keywords with the highest TF-IDF value are selected as the final keyword screening results to highlight the core content of the opinion summary.
[0180] For example, in order to improve the quality of opinion summaries and the representativeness of keywords, this study proposes a post-processing method based on TF-IDF to optimize and filter the opinion summaries and keywords output by the memory-enhanced neural network model to obtain standardized structured opinion information. The following will introduce the technical details and implementation steps of keyword filtering and summary optimization based on TF-IDF.
[0181] First, count the number of documents in which each keyword appears in the entire opinion text collection, that is, the document frequency (DF). Document frequency reflects the distribution breadth of keywords in opinion texts. Keywords that appear in more documents are usually more representative. Specifically, for each keyword, traverse the opinion text collection, count the number of documents containing the keyword, and obtain the document frequency of the keyword.
[0182] Based on the document frequency of the keyword and the total number of opinion texts, the inverse document frequency (IDF) of each keyword is calculated. The inverse document frequency measures the discrimination of the keyword in the entire text collection. Keywords that appear in fewer documents usually have higher discrimination. By taking the logarithm operation, the inverse document frequency of the keyword can be obtained. Specifically, the total number of documents is divided by the document frequency of the keyword, and then the logarithm is taken to obtain the inverse document frequency of the keyword.
[0183] For each opinion summary, count the number of times each keyword appears, i.e., term frequency (TF). Term frequency reflects the contribution of a keyword to the content of the summary. The more keywords appear, the more important they are to the expression of the summary. Specifically, for each keyword, count the number of times it appears in the corresponding opinion summary to obtain the keyword's term frequency.
[0184] Multiply the keyword's word frequency by its inverse document frequency to get the word frequency-inverse document frequency (TF-IDF) value, which comprehensively considers the keyword's importance in the abstract and its overall discrimination. The higher the TF-IDF value, the more important the keyword is in the abstract, and the higher its discrimination in the entire text collection. By calculating the TF-IDF value, you can make a comprehensive evaluation of the keywords and highlight those keywords that appear frequently in the abstract and have high overall discrimination.
[0185] According to the TF-IDF value of the keyword, all the keywords in the opinion summary are sorted in descending order. Several keywords with the highest TF-IDF value are selected as the final keyword screening results. Usually, Top-K keywords are selected, where K can be adjusted according to actual needs. The selected keywords can highlight the core content of the opinion summary and have high discrimination and representativeness.
[0186] According to the selected keywords, the summary of opinions is optimized and adjusted. Specifically, the selected keywords are matched with the content of the summary to ensure that the keywords are included in the summary, and the expression of the summary is appropriately adjusted to make it more concise, coherent and highlight key information. The optimized summary can better reflect the core content represented by the keywords and improve the quality and readability of the summary.
[0187] The following is a specific data case to illustrate the application of keyword screening and abstract optimization based on TF-IDF.
[0188] Assume there are the following 5 opinion texts and corresponding summaries:
[0189] Text 1: The hotel room was spacious and clean, the bed was comfortable and the facilities were complete. The staff was friendly and helpful. The location is convenient and well connected to public transportation. However, the room was a bit noisy due to street traffic noise. Overall, we had a pleasant stay and would recommend this hotel.
[0190] Summary 1: The room was spacious and clean, the bed was comfortable and the facilities were complete. The staff was friendly and helpful. The location was convenient and easy to get to. The room was a bit noisy due to street traffic noise. Overall a pleasant stay and highly recommended.
[0191] Keyword 1: room, staff, location, noise, pleasant;
[0192] Text 2: I stayed at this hotel recently on a business trip. The room was well appointed and equipped with all the necessary amenities. The hotel staff was professional and attentive. The food in the hotel restaurant was delicious and the breakfast buffet had a good variety. The only downside was the slow Wi-Fi connection in the room. Despite this, I had an efficient and comfortable stay.
[0193] Summary 2: The rooms are well furnished and well equipped. The staff is professional and attentive. The hotel restaurant has delicious meals and a wide variety of breakfast buffet options. The wireless internet connection in the room is slow. Overall, the accommodation is efficient and comfortable.
[0194] Keyword 2: room, staff, restaurant, wireless network, comfort;
[0195] Text 3: We had a wonderful family vacation at this resort. The resort offers a wide variety of activities and entertainment options for all ages. The kids loved the water park and mini golf course. The accommodation is spacious and well-maintained, surrounded by beautiful natural scenery. The resort staff is friendly and always ready to help. The only problem we encountered was the long waiting time at some of the popular restaurants. Despite this, it was still a memorable and enjoyable vacation.
[0196] Summary 3: Excellent family vacation with a wide variety of activities and entertainment for all ages. The kids loved the water park and mini golf. The accommodation was spacious, well maintained and beautifully landscaped. The staff was friendly and helpful. The waiting times at popular restaurants were long. A memorable and enjoyable vacation.
[0197] Keyword 3: family, vacation, activities, accommodation, staff;
[0198] Text 4: I had a very bad dining experience at this restaurant. The service was slow and inattentive, with the waiter often disappearing for long periods of time. The food was mediocre at best, with some dishes overcooked and others lacking flavor. The prices were high considering the quality of the food and service. The restaurant was also very noisy, making it difficult to hold a conversation. I would not recommend this restaurant to anyone.
[0199] Summary 4: Poor dining experience, slow and inattentive service, waiters disappeared for a long time. Food was average, dishes were overcooked and lacked flavor. Prices were high for the quality of food and service. The restaurant was noisy and conversation was difficult. Would not recommend this restaurant.
[0200] Keyword 4: bad, service, food, price, noisy;
[0201] Text 5: This product exceeded my expectations. It was very easy to install and use right out of the box. The product quality is excellent, the materials are sturdy, and the design is stylish. The performance is top notch, the processing speed is fast, and it runs smoothly. The battery life is also impressive, and a single charge lasts for a long time. The only minor issue I noticed is that the device can get a little warm during extended use, but never to the point of being uncomfortable. Overall, I would highly recommend this product to anyone looking for a reliable, high-performance device.
[0202] Summary 5: The product exceeded expectations, it worked right out of the box, and was easy to install and use. The product quality is excellent, the materials are solid, and the design is stylish. The performance is first-rate, the processing speed is fast, and it runs smoothly. The battery life is long, and a single charge can be used for a long time. The device will heat up slightly during long-term use, but it does not affect comfort. Highly recommended for those who are looking for a reliable and high-performance device.
[0203] Keyword 5: product, quality, performance, battery, recommendation;
[0204] Based on the above opinion text and summary, the document frequency, inverse document frequency, word frequency and TF-IDF value of the keywords are calculated.
[0205] Assuming that the top 3 keywords are selected as the final screening results, the filtered keywords are:
[0206] Summary 1: Room, Staff, Location;
[0207] Summary 2: Rooms, staff, restaurants;
[0208] Summary 3: Activities, accommodation, staff;
[0209] Summary 4: Service, food, price;
[0210] Summary 5: Quality, performance, battery;
[0211] According to the selected keywords, the opinion summary is optimized and adjusted to obtain the following optimized summary:
[0212] Optimized summary 1: The room is spacious and clean, the bed is comfortable, and the facilities are complete. The staff is friendly and helpful. The location is convenient and the transportation is convenient.
[0213] Optimized summary 2: The rooms are well furnished and well equipped. The staff is professional and attentive. The hotel restaurant has delicious meals and a wide variety of breakfast buffet options.
[0214] Improved summary 3: Provides a wide variety of activities and entertainment for all ages. Accommodation is spacious, well-maintained, and has beautiful views. Staff are friendly and helpful.
[0215] Optimized summary 4: Service was slow and inattentive. Food was average, dishes were overcooked and lacked flavor. Prices were high relative to the quality of food and service.
[0216] Optimized summary 5: Excellent product quality and first-rate performance. Long battery life and long use on a single charge. Highly recommended for those who seek reliable, high-performance devices.
[0217] Through keyword screening and summary optimization based on TF-IDF, more concise, prominent and representative opinion summaries and keywords can be obtained, thereby improving the structured degree and usability of opinion information.
[0218] In an optional embodiment,
[0219] The business domain knowledge base is constructed by combining ontology and rules, including domain concepts, entities, relationships and constraint rules. The steps of achieving deep opinion understanding through knowledge reasoning include:
[0220] According to the entities and relations identified in the opinion information, Simple Protocol and Resource Description Framework Query Language (SPARQL) query statements are constructed to retrieve triples related to the entities and relations from the knowledge base stored in the Resource Description Framework format;
[0221] Match the retrieved triple knowledge with the predefined IF-THEN reasoning rules in the knowledge base, trigger the rules through forward reasoning or backward reasoning algorithms, and generate new triple knowledge according to the consequences of the rules;
[0222] By using the category hierarchical relationship in the ontology, based on the category attributes of the entity, combined with the sub-category and parent-category relationships, the knowledge scope and granularity of the opinion information can be expanded through ontology reasoning;
[0223] Check the consistency of knowledge generated during the reasoning process, automatically identify attribute value conflicts, category conflicts, and relationship conflicts through rule-based confidence comparison, and resolve conflicts based on domain knowledge and business rules to ensure the consistency and reliability of reasoning results;
[0224] The new knowledge obtained by reasoning is integrated with the original opinion entities and relations. The confidence of knowledge and the credibility of the source are considered. The probabilistic graphical model is used to model and infer knowledge from different sources to obtain a unified opinion knowledge representation and form an opinion understanding report.
[0225] For example, in order to achieve a deep understanding of opinion information, this study proposes a method based on knowledge reasoning, which uses the ontology and rules in the business domain knowledge base, expands the knowledge scope and granularity of opinion information through knowledge retrieval, rule reasoning and ontology reasoning, and performs knowledge fusion and consistency check to form a unified opinion understanding report. The technical details and implementation steps of deep opinion understanding based on knowledge reasoning will be introduced in detail below.
[0226] First, based on the entities and relations identified in the opinion information, Simple Protocol and Resource Description Framework Query Language (SPARQL) query statements are constructed to retrieve triple knowledge related to entities and relations from the knowledge base stored in the Resource Description Framework (RDF) format. Specifically, the entity is used as the subject or object, and the relation is used as the predicate to construct a triple query pattern in the form of "subject-predicate-object". SPARQL statements are used to match and retrieve in the RDF knowledge base to obtain triple knowledge related to the opinion information.
[0227] The retrieved triple knowledge is matched with the predefined IF-THEN reasoning rules in the knowledge base, and the rules are triggered through forward reasoning or backward reasoning algorithms, and new triple knowledge is generated according to the consequent of the rules. Specifically, the reasoning rules in the knowledge base are traversed, and the triple knowledge is matched with the antecedent of the rule. If the match is successful, the rule is triggered, and new triple knowledge is generated according to the consequent of the rule.
[0228] For example, the following inference rules exist in the knowledge base:
[0229] IF(room,attributes,spacious) AND(room,has,bed) AND(bed,attributes,comfortable);
[0230] THEN(room,attributes,comfort);
[0231] This rule means that if the room is spacious and has a bed, and the bed is comfortable, then it can be inferred that the room is also comfortable. The retrieved triple knowledge is matched with the rule, the rule is triggered, and a new triple knowledge (room, attribute, comfort) is generated according to the consequence of the rule.
[0232] By using the category hierarchical relationship in the ontology, according to the category attributes of the entity, combined with the subclass and parent class relationship, the knowledge scope and granularity of the opinion information are expanded through ontology reasoning. Specifically, the category hierarchical relationship in the ontology is traversed to find the subclasses and parent classes related to the entity category in the opinion information, and new triple knowledge is generated based on the relationship between the categories.
[0233] For example, there are the following category hierarchical relationships in the ontology:
[0234] :hotelrdfs:subClassOf:hotel;
[0235] :Five-star hotelrdfs:subClassOf:hotel;
[0236] :Four-star hotelrdfs:subClassOf:hotel;
[0237] According to the category attribute of the entity "hotel", ontology reasoning can be used to determine that the hotel belongs to the "guesthouse" category, and further infer that the hotel may be a "five-star hotel" or a "four-star hotel". Based on these category hierarchical relationships, new triple knowledge can be generated, such as (hotel, belongs to, guesthouse), (hotel, may be, a five-star hotel), (hotel, may be, a four-star hotel), etc., which expands the knowledge scope and granularity of opinion information.
[0238] The knowledge generated in the reasoning process is checked for consistency. Through rule-based confidence comparison, attribute value conflicts, category conflicts, and relationship conflicts are automatically identified, and conflicts are resolved based on domain knowledge and business rules to ensure the consistency and reliability of the reasoning results. Specifically, for each newly generated triple knowledge, check whether it conflicts with existing knowledge. If there is a conflict, it is handled according to the predefined confidence comparison rules and conflict resolution strategies.
[0239] For example, the following triples of knowledge are generated during reasoning:
[0240] (room, area, 20 square meters);
[0241] (room, area, 30 square meters);
[0242] There is a conflict in the attribute values of these two triples of knowledge, that is, the area of the same room cannot be 20 square meters and 30 square meters at the same time. According to the confidence comparison rule, the triple with higher confidence can be selected as the final result. Assuming that (room, area, 30 square meters) has a higher confidence, the triple is retained and (room, area, 20 square meters) is discarded.
[0243] For example, the following triples of knowledge are generated during the reasoning process:
[0244] (hotel, belonging to, guesthouse);
[0245] (hotel,belongs to,restaurant);
[0246] There is a category affiliation conflict between these two triples, that is, the hotel cannot belong to two different categories, hotel and restaurant. According to domain knowledge and business rules, the hotel should belong to the hotel category, not the restaurant category. Therefore, keep (hotel, belongs to, hotel) and discard (hotel, belongs to, restaurant).
[0247] The new knowledge obtained by reasoning is integrated with the original opinion entities and relations. Considering the confidence of knowledge and the credibility of sources, the probabilistic graph model is used to model and infer knowledge from different sources to obtain a unified opinion knowledge representation and form an opinion understanding report. Specifically, the new knowledge obtained by reasoning is combined with the original opinion entities and relations to construct a directed acyclic graph. The nodes represent entities and attributes, and the edges represent the relationships between entities. The probabilistic graph model is used to model and infer the probabilities of nodes and edges, and the marginal probability distribution of each node and edge is obtained as a measure of its confidence. Finally, according to the confidence of nodes and edges, an opinion understanding report is generated to comprehensively express the core content and implicit knowledge of opinion information.
[0248] For example, for the opinion information "The rooms in this hotel are spacious and the beds are comfortable", after knowledge reasoning and fusion, the following opinion knowledge representation can be obtained:
[0249] (hotel, has, rooms) Confidence: 0.95;
[0250] (room, attributes, spaciousness) Confidence: 0.90;
[0251] (room, has, bed) confidence: 0.92;
[0252] (bed, attributes, comfort) confidence: 0.88;
[0253] (Room, Attributes, Comfort) Confidence: 0.85;
[0254] (hotel, belongs to, guesthouse) Confidence: 0.93;
[0255] (Hotel, possibly, a five-star hotel) Confidence: 0.78;
[0256] (Hotel, possibly, a four-star hotel) Confidence: 0.82;
[0257] Based on the above opinion knowledge representation, the following opinion understanding report can be generated:
[0258] This opinion reflects the user's evaluation of a hotel room. The user thinks that the hotel room is spacious, the bed is comfortable, and the overall accommodation experience is satisfactory. According to the knowledge reasoning results, the hotel belongs to the hotel category and may be a four-star or five-star high-end hotel.
[0259] This paper uses deep opinion understanding based on knowledge reasoning to make full use of the ontology and rules in the business domain knowledge base, mine the implicit knowledge in opinion information, expand the knowledge scope and granularity of opinion information, and improve the comprehensiveness and accuracy of opinion understanding. At the same time, through consistency checking and knowledge fusion, the reliability and consistency of reasoning results are ensured, and high-quality opinion understanding reports are generated to provide valuable reference for subsequent business decisions.
[0260] In an optional embodiment,
[0261] Through the reinforcement learning algorithm, the response strategy is dynamically adjusted using user feedback. The steps to achieve human-computer interaction optimization include:
[0262] The opinion interaction process is modeled as a Markov decision process, and the state space is defined to represent the interaction state based on the user's opinion query, system response history, user feedback, and dialogue turn information;
[0263] Define the system action space, including response action types such as opinion summary, opinion comparison, detail enumeration, multiple rounds of clarification, and greeting replies;
[0264] Combining the user's feedback rating on the system's response and the pre-trained dialogue quality evaluation model, we designed a reward function for the state-action pair, taking into account both user satisfaction and response quality.
[0265] The state-action value function is represented by a function approximation method, and a deep neural network is used to fit the value function. The input of the deep neural network is the interaction state feature, and the output is the value estimate of the response action.
[0266] Use the experience replay mechanism for offline strategy learning, store the state transition samples generated during the interaction process into the experience replay pool, and update the value network parameters by randomly sampling the replay samples;
[0267] When generating actual response actions, the ε-greedy strategy is used to balance exploration and exploitation, select the optimal action based on the value estimate of the current state, and gradually reduce the exploration probability as training progresses;
[0268] The trained value network is converted into a lightweight online decision-making model through model compression and knowledge distillation technology to improve decision-making efficiency and response speed.
[0269] For example, in order to improve the quality of human-computer interaction and user satisfaction, this study proposes a human-computer interaction optimization method based on reinforcement learning, which realizes adaptive learning and optimization of user feedback by dynamically adjusting the system's response strategy. The following will introduce the technical details and implementation steps of human-computer interaction optimization based on reinforcement learning in detail.
[0270] First, the opinion interaction process is modeled as a Markov decision process (MDP), and the state space is defined. The interaction state is represented according to the user opinion query, system response history, user feedback, and dialogue turn information. Specifically, a multi-dimensional state representation vector is designed, including the following information:
[0271] User opinion inquiry: represents the user's current opinion inquiry content, which can be represented using bag-of-words model, topic model or semantic vector.
[0272] System response history: represents the content and type of the system’s responses in previous rounds of interactions. A recurrent neural network can be used to encode the response sequence.
[0273] User feedback: represents the user's evaluation and feedback on the system's response, which can be an explicit rating or an implicit emotional tendency.
[0274] Dialogue turn: Indicates the current number of dialogue turns, reflecting the progress and context information of the interaction.
[0275] This information is combined into a fixed-dimensional state representation vector as the state in the MDP.
[0276] For example, in an interactive scenario of a hotel review, the user asks "How is the room in this hotel?", the system replies "The room is spacious and the bed is comfortable", and the user gives a 4-star rating. The current interactive state can be expressed as:
[0277] state = [0.2, 0.1, 0.3, ..., 0.8, 0.5, 4, 2];
[0278] Among them, the first several dimensions represent the topic vectors of the user's inquiry, the middle several dimensions represent the semantic vectors of the system's response, the second to last dimension represents the user's rating, and the last dimension represents the number of conversation rounds.
[0279] Next, define the system action space, including response action types such as opinion summary, opinion comparison, detail enumeration, multiple rounds of clarification, and greeting replies. Design corresponding response action templates for different interaction scenarios and opinion inquiry types. Each action type can have multiple specific response template instances. The system action space consists of all response action types and instances.
[0280] For example, in the hotel review interaction scenario, the following response action types and instances can be defined:
[0281] Summary of comments:
[0282] Action 1: "According to the reviews of most users, the rooms in this hotel are generally very good, spacious and comfortable, and the bedding is of good quality."
[0283] Action 2: "Users generally believe that the rooms in this hotel are clean and tidy, the layout is warm, and the accommodation experience is good."
[0284] Comparison of views:
[0285] Action 3: "Some users think the room is spacious, but a few users think the room is small, which may be related to the room type and floor."
[0286] Action 4: "Most users are satisfied with the quality of the bedding, but some users complain that the mattress is too soft and not supportive enough."
[0287] Details:
[0288] Action 5: "The room is equipped with basic furniture such as a large bed, desk, wardrobe, free Wi-Fi and electric kettle."
[0289] Action 6: "The floor-to-ceiling windows in the room provide good lighting, and you can see the beautiful city view from the window, especially the night view is very charming."
[0290] Multiple rounds of clarification:
[0291] Action 7: "What aspects of the room are you dissatisfied with? Can you explain it in detail?"
[0292] Action 8: "You think the mattress is too soft. Which type of mattress do you mean? Standard room or deluxe room?"
[0293] Greetings reply:
[0294] Action 9: "Thank you for your valuable comments. We will continue to work hard to provide a better accommodation experience."
[0295] Action 10: "Wish you a pleasant stay. If you need anything, please contact our staff."
[0296] Combining the user's feedback rating on the system's response and the pre-trained dialogue quality assessment model, a reward function on the state-action pair is designed, taking into account user satisfaction and response quality. Specifically, the user feedback rating and dialogue quality assessment results are weighted to obtain a scalar reward value as the optimization target of the reinforcement learning algorithm.
[0297] For example, for user feedback rating, the following reward function can be designed:
[0298] If the user's rating is 5 stars, the reward value is 2;
[0299] If the user's rating is 4 stars, the reward value is 1;
[0300] If the user's rating is 3 stars, the reward value is 0;
[0301] If the user's rating is 2 stars, the reward value is -1;
[0302] If the user's rating is 1 star, the reward value is -2.
[0303] For the dialogue quality assessment model, the quality score output by the model can be normalized to the interval [0,1], then multiplied by a weight coefficient, and weighted summed with the reward value of the user rating to obtain the final reward value.
[0304] For example, assuming that the score given by the dialogue quality assessment model is 0.8, the weight coefficient is 0.5, and the user rating is 4 stars, the reward value is calculated as follows:
[0305] Reward value = 0.5*0.8+1 = 1.4;
[0306] This shows that the system response takes into account user satisfaction and objective quality and achieves good results.
[0307] The state-action value function is represented by the function approximation method, and the value function is fitted by a deep neural network. The input of the deep neural network is the interaction state feature, and the output is the value estimation of different response actions. The network structure can use multi-layer perceptron, convolutional neural network or recurrent neural network, etc., and the appropriate network architecture is selected according to the characteristics of the state representation.
[0308] For example, for the state representation vector defined above, a three-layer multilayer perceptron can be designed, with the input layer dimension being the state vector dimension, the hidden layer dimension being 64, the output layer dimension being the action space size, and the activation function being ReLU. The network parameters are optimized by minimizing the temporal difference error, that is, minimizing the mean square error between the estimated value of the current state-action pair and the true value of the next state.
[0309] The network parameters are updated through the gradient descent algorithm, gradually bringing the value estimate closer to the true value, thus obtaining an accurate approximation of the value function.
[0310] The experience replay mechanism is used for offline strategy learning. The state transition samples (s, a, r, s') generated during the interaction are stored in the experience replay pool, where s is the current state, a is the action selected by the system, r is the reward given by the environment, and s' is the transition to the next state. The value network parameters are updated by randomly sampling and replaying samples. Experience replay can break the temporal correlation between samples, improve data utilization efficiency and learning stability.
[0311] Specifically, a fixed-size experience replay pool is set up, and some samples generated by random exploration are initially filled in. After each interaction with the user, the newly generated samples are added to the replay pool, and a small batch of samples are randomly sampled from it, the temporal difference error is calculated, and the parameters of the value network are updated through the back-propagation algorithm. This process is repeated until the value network converges or the preset number of training rounds is reached.
[0312] For example, set the experience replay pool size to 10000 and the batch size to 64. After each interaction, add the (s, a, r, s') sample to the replay pool. If the replay pool is full, randomly delete the earliest added sample. Then randomly sample 64 samples from the replay pool, calculate the time difference error, and update the value network parameters through the Adam optimizer, with the learning rate set to 0.001. Repeat this process for 1000 rounds to obtain a relatively stable value function estimate.
[0313] When generating actual response actions, the ε-greedy strategy is used to balance exploration and exploitation. The action with the highest value is selected based on the value estimate of the current state, and other actions are randomly selected for exploration with a probability of ε. As training progresses, the exploration probability ε is gradually reduced, so that the strategy gradually converges to the optimal action.
[0314] Specifically, at each interaction, a random number p between 0 and 1 is generated. If p is less than ε, an action is randomly selected; otherwise, the action with the highest estimated value in the current state is selected. The initial value of ε can be set to 0.1, and multiplied by a decay factor every certain number of rounds, such as 0.99, so that the exploration probability gradually decreases.
[0315] For example, the current state is s, and the value network estimates the value of each action as follows:
[0316] Action 1: 0.8;
[0317] Action 2: 0.6;
[0318] Action 3: 0.7;
[0319] Action 4: 0.5;
[0320] If p = 0.05, which is less than ε = 0.1, then randomly select an action, assuming action 2 is selected. If p = 0.15, which is greater than ε = 0.1, then select action 1 with the highest value.
[0321] After multiple rounds of interaction and strategy improvement, the system's response strategy will be continuously optimized, selecting more reasonable and effective response actions to improve user satisfaction and interaction quality.
[0322] The trained value network is converted into a lightweight online decision model through model compression, knowledge distillation and other technologies to improve decision efficiency and response speed. Specifically, the deep neural network model can be optimized by pruning, quantization, low-rank decomposition, etc. to reduce the model size and computational complexity.
[0323] In addition, the knowledge of the value network can be extracted and distilled into a smaller student model, so that the student model can achieve decision-making performance similar to that of the teacher model at a lower computational cost.
[0324] For example, the original value network model is compressed into a three-layer multilayer perceptron, the hidden layer dimension is reduced from 64 to 32, and the weight matrix is decomposed into a low-rank matrix to obtain the product of several low-rank matrices, thereby greatly reducing the model size.
[0325] Then, using knowledge distillation technology, the output of the original value network is used as a soft label to calculate the cross entropy loss with the output of the student model, and L2 regularization is applied to the output of the student model to make it as close as possible to the decision result of the teacher model. After distillation training, a streamlined and efficient online decision model is obtained.
[0326] Through the above-mentioned strategy learning, compression optimization and other technologies, the system can achieve real-time response strategy optimization with lower time and space overhead, improve the efficiency and quality of human-computer interaction, and better meet user needs.
[0327] Figure 2 FIG. 1 is a schematic diagram of the structure of a long text user opinion understanding system based on memory enhancement according to an embodiment of the present invention. Figure 2 As shown, the system comprises:
[0328] The first unit is used to obtain the long text user opinions to be understood, and pre-process the long text user opinions, including word segmentation, part-of-speech tagging, named entity recognition and dependency syntactic analysis, to obtain a text representation rich in linguistic features; the word vector model is used to map the text representation into a low-dimensional dense vector as the input of the memory-enhanced neural network model, and the word vector model uses a pre-trained language model and performs incremental training on the domain corpus; the memory-enhanced neural network model is constructed using a multi-head attention mechanism and a gated recurrent unit, including three modules: a hierarchical encoder, a key-value memory network, and a pointer generation network; the hierarchical encoder is used to extract the semantic features of the text at the word, word, sentence and paragraph levels; the key-value memory network uses a multi-layer sparse memory matrix, realizes fast addressing through local sensitive hashing, and introduces a forgetting mechanism to dynamically update the memory, which is used to store long-term context information and enhance the semantic features extracted by the hierarchical encoder; the pointer generation network is used to decode the semantic features enhanced by memory using a beam search algorithm, combined with a vocabulary masking mechanism and a length penalty, to generate structured opinion summaries and keywords, and can selectively copy fragments in the input text;
[0329] The second unit is used to post-process the opinion summary and keywords output by the memory-enhanced neural network model to obtain standardized structured opinion information; based on the structured opinion information, the user opinions are understood and analyzed using a pre-built business domain knowledge base to generate an opinion understanding report. The business domain knowledge base is constructed in a combination of ontology and rules, including domain concepts, entities, relationships and constraint rules, and achieves deep opinion understanding through knowledge reasoning;
[0330] The third unit is used to automatically generate responses based on the opinion understanding report, combined with the domain dialogue knowledge base and user portrait. The response forms include text, voice, and charts. The domain dialogue knowledge base is learned using multi-round dialogue data and includes intent recognition, slot filling, and dialogue strategy learning functions to achieve context-related, logically self-consistent multi-round responses. Through the reinforcement learning algorithm, the response strategy is dynamically adjusted using user feedback to achieve human-computer interaction optimization.
[0331] According to a third aspect of the embodiments of the present invention,
[0332] An electronic device is provided, comprising:
[0333] processor;
[0334] a memory for storing processor-executable instructions;
[0335] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0336] A fourth aspect of the embodiments of the present invention is:
[0337] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.
[0338] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.
[0339] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.< / eos>
Claims
1. A long text user opinion understanding method based on memory enhancement, characterized in that: include: Obtain long text user opinions to be understood, preprocess the long text user opinions, including word segmentation, part-of-speech tagging, named entity recognition and dependency syntactic analysis, to obtain a text representation rich in linguistic features; use a word vector model to map the text representation into a low-dimensional dense vector as the input of a memory-enhanced neural network model, the word vector model uses a pre-trained language model and performs incremental training on a domain corpus; the memory-enhanced neural network model is constructed using a multi-head attention mechanism and a gated recurrent unit, including three modules: a hierarchical encoder, a key-value memory network, and a pointer generation network; the hierarchical encoder is used to extract semantic features of the text at the word, word, sentence and paragraph levels; the key-value memory network uses a multi-layer sparse memory matrix, achieves fast addressing through local sensitive hashing, and introduces a forgetting mechanism to dynamically update the memory, which is used to store long-term context information and enhance the semantic features extracted by the hierarchical encoder; The pointer generation network is used to generate structured opinion summaries and keywords based on memory-enhanced semantic features, using a beam search algorithm for decoding, combined with a vocabulary masking mechanism and length penalty, and can selectively copy fragments from the input text; Post-process the opinion summary and keywords output by the memory-enhanced neural network model to obtain standardized structured opinion information; Based on the structured opinion information, the user opinions are understood and analyzed using the pre-built business domain knowledge base to generate an opinion understanding report. The business domain knowledge base is built in a combination of ontology and rules, including domain concepts, entities, relationships and constraint rules, and achieves deep opinion understanding through knowledge reasoning; Based on the opinion understanding report, combined with the domain dialogue knowledge base and user portrait, responses are automatically generated in the form of text, voice, and charts. The domain dialogue knowledge base is learned using multi-round dialogue data and includes intent recognition, slot filling, and dialogue strategy learning functions to achieve context-related, logically self-consistent multi-round responses. Through the reinforcement learning algorithm, the response strategy is dynamically adjusted using user feedback to achieve human-computer interaction optimization.
2. The method according to claim 1, characterized in that The memory-enhanced neural network model is constructed using a multi-head attention mechanism and a gated recurrent unit, and includes three modules: a hierarchical encoder, a key-value memory network, and a pointer generation network. The steps include: For each layer of the hierarchical encoder, a gated recurrent unit is used to model the input sequence. The reset gate and update gate are introduced to dynamically control the flow and update of information, alleviate the gradient vanishing problem, and obtain the hidden state sequence. Use a multi-head attention mechanism to calculate the matching degree between the hidden state sequence and the memory matrix, obtain the query matrix, key matrix and value matrix through linear transformation, multiply the query matrix by the transpose of the key matrix and divide it by the scaling factor to obtain the attention score matrix, normalize the attention score matrix using a normalized exponential function to obtain the attention weight matrix, use the attention weight matrix to perform weighted summation on the value matrix to obtain the attention output matrix, and concatenate to obtain the multi-head attention output, enhance the ability of the memory-enhanced neural network model to capture the diverse interactions between text and memory, and obtain the memory readout vector; The memory readout vector is concatenated with the hidden state sequence as the input of the next encoder layer to achieve information transfer and feature fusion between the layers of the hierarchical encoder; In the decoder, a gated recurrent unit is used to model the sequence of memory readout vectors, and a multi-head attention mechanism is used to calculate the match between the decoder hidden state and the encoder output to obtain the context vector. The context vector is concatenated with the decoder hidden state, and after linear transformation and normalized exponential function, the probability distribution of the generated words is obtained, and the structured opinion summary and keywords are generated by decoding.
3. The method according to claim 1, characterized in that The hierarchical encoder is used to extract semantic features of text at the word, phrase, sentence and paragraph levels; the key-value memory network adopts a multi-layer sparse memory matrix, realizes fast addressing through local sensitive hashing, and introduces a forgetting mechanism to dynamically update memory, which is used to store long-term context information and enhance the semantic features extracted by the hierarchical encoder; The pointer generation network is used to generate structured opinion summaries and keywords based on memory-enhanced semantic features, using a beam search algorithm for decoding, combined with a vocabulary masking mechanism and a length penalty, and can selectively copy fragments in the input text, including the following steps: The hierarchical encoder is used to receive input long text user opinions, extract character-level local features through character-level convolutional neural networks, capture long-distance dependencies between words through word-level self-attention mechanisms, then use sentence-level convolutional neural networks to extract sentence-level local features, and finally use paragraph-level gated recurrent units to extract paragraph-level global features, thereby extracting multi-granularity hierarchical semantic representations of user opinions from the bottom up; The key-value memory network includes a key matrix, a value matrix and a read-write controller, wherein the key matrix and the value matrix are both multi-layer sparse matrices for storing long-term context information; the read-write controller is used to receive the multi-granularity hierarchical semantic representation extracted by the hierarchical encoder, and through the attention mechanism and the local sensitive hashing algorithm, quickly retrieve the memory fragment most relevant to the current input in the key matrix, and read the corresponding memory content from the corresponding value matrix, so as to enhance the multi-granularity hierarchical semantic representation extracted by the hierarchical encoder; the read-write controller dynamically updates the key matrix and the value matrix through the gating mechanism, and introduces a forgetting mechanism to discard expired memory that is no longer needed, so that the key-value memory network can adaptively store and update long-term context information; The pointer generation network takes the multi-granularity hierarchical semantic representation enhanced by the key-value memory network as input, and adopts a decoder with an attention mechanism to decode and generate respectively the summary generation task and the keyword extraction task. During the decoding process, the pointer network mechanism is used to allow key fragments and words to be copied from the input user opinions, and a beam search algorithm is introduced for decoding, a word list mask mechanism is used to constrain the decoding space, and a length-based reward and punishment mechanism is used to control the length of the generated sequence, so as to generate structured summary text and keyword sequence as output.
4. The method according to claim 3, characterized in that The steps of introducing a beam search algorithm for decoding, a word list mask mechanism to constrain the decoding space, and a length-based reward and punishment mechanism to control the length of the generated sequence and generating a structured summary text and keyword sequence as output include: Encode the input text into a hidden state sequence, initialize the decoder's hidden state, and the candidate set containing only the sequence start symbol; In each decoding step, for each sequence in the candidate set, the decoder hidden state is updated, the attention distribution and context vector are calculated, and the final word probability distribution is obtained by combining the vocabulary probability distribution and the replication probability distribution; Based on the beam search strategy, multiple extensions with the highest probability are selected from the word probability distribution and added to the candidate set. At the same time, a word list mask is introduced to remove the generated words from the word list to avoid repeated generation; Calculate the score of each candidate sequence, including the log-likelihood score and the length-based reward and penalty items, and control the preference for shorter or longer sequences by setting the reward and penalty factors; Repeat the decoding steps until the maximum length is reached or all sequences in the candidate set are terminated by a sequence terminator, and select the sequence with the highest score as the final generated result; During the generation process, the decoding strategy is controlled by adjusting the bundle size, vocabulary mask switch, and length reward and penalty factors to meet different task requirements and performance optimization goals.
5. The method according to claim 1, characterized in that The steps of post-processing the opinion summary and keywords output by the memory-enhanced neural network model to obtain standardized structured opinion information include: According to the number of occurrences of the keyword in the opinion text set, the document frequency of each keyword is calculated to obtain the document frequency result reflecting the distribution breadth of the keyword in the opinion text; Based on the document frequency of the keyword and the total number of opinion texts, the inverse document frequency of each keyword is calculated by taking the logarithm operation to obtain the inverse document frequency result that measures the discrimination of the keyword in the entire text collection; Count the number of times each keyword appears in the corresponding opinion summary to obtain the word frequency result that reflects the contribution of the keyword to the summary content; Multiply the keyword's word frequency and inverse document frequency to get the word frequency-inverse document frequency (TF-IDF) value that comprehensively considers the keyword's importance in the abstract and its overall discrimination; According to the TF-IDF value of the keyword, all the keywords in the opinion summary are sorted in descending order, and several keywords with the highest TF-IDF value are selected as the final keyword screening results to highlight the core content of the opinion summary.
6. The method according to claim 1, characterized in that The business domain knowledge base is constructed by combining ontology and rules, including domain concepts, entities, relationships and constraint rules. The steps of achieving deep opinion understanding through knowledge reasoning include: According to the entities and relations identified in the opinion information, Simple Protocol and Resource Description Framework Query Language (SPARQL) query statements are constructed to retrieve triples related to the entities and relations from the knowledge base stored in the Resource Description Framework format; Match the retrieved triple knowledge with the predefined IF-THEN reasoning rules in the knowledge base, trigger the rules through forward reasoning or backward reasoning algorithms, and generate new triple knowledge according to the consequences of the rules; By using the category hierarchical relationship in the ontology, based on the category attributes of the entity, combined with the sub-category and parent-category relationships, the knowledge scope and granularity of the opinion information can be expanded through ontology reasoning; Check the consistency of knowledge generated during the reasoning process, automatically identify attribute value conflicts, category conflicts, and relationship conflicts through rule-based confidence comparison, and resolve conflicts based on domain knowledge and business rules to ensure the consistency and reliability of reasoning results; The new knowledge obtained by reasoning is integrated with the original opinion entities and relations. The confidence of knowledge and the credibility of the source are considered. The probabilistic graphical model is used to model and infer knowledge from different sources to obtain a unified opinion knowledge representation and form an opinion understanding report.
7. The method according to claim 1, characterized in that Through the reinforcement learning algorithm, the response strategy is dynamically adjusted using user feedback. The steps to achieve human-computer interaction optimization include: The opinion interaction process is modeled as a Markov decision process, and the state space is defined to represent the interaction state based on the user's opinion query, system response history, user feedback, and dialogue turn information; Define the system action space, including response action types such as opinion summary, opinion comparison, detail enumeration, multiple rounds of clarification, and greeting replies; Combining the user's feedback rating on the system's response and the pre-trained dialogue quality evaluation model, we designed a reward function for the state-action pair, taking into account both user satisfaction and response quality. The state-action value function is represented by a function approximation method, and a deep neural network is used to fit the value function. The input of the deep neural network is the interaction state feature, and the output is the value estimate of the response action. Use the experience replay mechanism for offline strategy learning, store the state transition samples generated during the interaction process into the experience replay pool, and update the value network parameters by randomly sampling the replay samples; When generating actual response actions, the ε-greedy strategy is used to balance exploration and exploitation, select the optimal action based on the value estimate of the current state, and gradually reduce the exploration probability as training progresses; The trained value network is converted into a lightweight online decision-making model through model compression and knowledge distillation technology to improve decision-making efficiency and response speed.
8. A long text user opinion understanding system based on memory enhancement, used to implement the method described in any one of claims 1 to 7, characterized in that: include: The first unit is used to obtain long text user opinions to be understood, and preprocess the long text user opinions, including word segmentation, part-of-speech tagging, named entity recognition and dependency syntactic analysis, to obtain a text representation rich in linguistic features; the word vector model is used to map the text representation into a low-dimensional dense vector as the input of the memory-enhanced neural network model, and the word vector model uses a pre-trained language model and performs incremental training on the domain corpus; the memory-enhanced neural network model is constructed using a multi-head attention mechanism and a gated recurrent unit, including three modules: a hierarchical encoder, a key-value memory network, and a pointer generation network; the hierarchical encoder is used to extract the semantic features of the text at the word, word, sentence and paragraph levels; the key-value memory network uses a multi-layer sparse memory matrix, realizes fast addressing through local sensitive hashing, and introduces a forgetting mechanism to dynamically update the memory, which is used to store long-term context information and enhance the semantic features extracted by the hierarchical encoder; The pointer generation network is used to generate structured opinion summaries and keywords based on memory-enhanced semantic features, using a beam search algorithm for decoding, combined with a vocabulary masking mechanism and length penalty, and can selectively copy fragments from the input text; The second unit is used to post-process the opinion summary and keywords output by the memory-enhanced neural network model to obtain standardized structured opinion information; Based on the structured opinion information, the user opinions are understood and analyzed using the pre-built business domain knowledge base to generate an opinion understanding report. The business domain knowledge base is built in a combination of ontology and rules, including domain concepts, entities, relationships and constraint rules, and achieves deep opinion understanding through knowledge reasoning; The third unit is used to automatically generate responses based on the opinion understanding report, combined with the domain dialogue knowledge base and user portrait. The response forms include text, voice, and charts. The domain dialogue knowledge base is learned using multi-round dialogue data and includes intent recognition, slot filling, and dialogue strategy learning functions to achieve context-related, logically self-consistent multi-round responses. Through the reinforcement learning algorithm, the response strategy is dynamically adjusted using user feedback to achieve human-computer interaction optimization.
9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Semantic matching method and system for knowledge retrieval and question answering of power transformer
CN113962219A
Session recommendation system fusing sparse graph and multi-hop attention
CN114817508A