Intelligent voice interaction system based on text travel big model

By integrating adjacent dialogue texts and combining sentiment analysis models and large language models, the problem of insufficient sentiment recognition in existing interaction models when the topic changes is solved, realizing accurate recognition of user emotions and dynamic adjustment of response content, thereby improving user experience and demand satisfaction.

CN120930808AActive Publication Date: 2025-11-11SICHUAN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511473186.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-15
Publication Date
2025-11-11
Estimated Expiration
2045-10-15

AI Technical Summary

Technical Problem

Existing interaction models are inadequate in terms of emotion recognition and response optimization when faced with topic changes. They are unable to dynamically adjust response strategies based on changes in user emotions, which affects user experience and the satisfaction of user needs.

Method used

By acquiring the correlation between adjacent dialogue texts, integrating them into a comprehensive text, and combining it with a sentiment analysis model and a quantitative scaling table, the system identifies user sentiment categories and score differences. It then uses sentiment keywords and semantic feature vectors to generate correlation weights, dynamically adjusts response content, and combines a large language model to generate personalized interactive conversations.

Benefits of technology

It achieves accurate emotion recognition and response strategy adjustment in scenarios with changing themes, improving user comfort and interaction satisfaction. It can dynamically adjust the response content according to changes in user emotions to meet user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120930808A_ABST
    Figure CN120930808A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent voice interaction system based on a text travel big model, and belongs to the technical field of man-machine interaction, and the system specifically comprises the steps: obtaining two adjacent dialogue texts A and B of a user, extracting semantic feature vectors, calculating the correlation degree, and if the correlation degree is greater than or equal to a preset threshold value, integrating the text A and the text B to obtain a comprehensive text; calculating an emotion difference value, and if the difference value does not exceed a threshold value, rewriting the dialogue text A to obtain a corrected text; performing demand keyword identification on the corrected text A and the reply text, comparing identification results to obtain a missing demand keyword set, and determining a reply text b corresponding to the text B according to the demand keyword set; and taking the comprehensive text emotion category as the current emotion of the user, inputting the current emotion of the user and the reply text b into the pre-trained large language model, and generating a personalized interaction session, so that the intelligent voice interaction efficiency is improved, and emotional interaction is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of human-computer interaction technology, specifically to an intelligent voice interaction system based on a large-scale cultural tourism model. Background Technology

[0002] In the current environment of digital transformation and intelligent upgrading of the cultural and tourism industry, the cultural and tourism interaction model creates an immersive experience for tourists through technology integration and scenario adaptation, providing personalized and convenient services throughout the entire process from pre-trip planning to post-trip sharing, greatly improving the user experience. However, in actual human-computer interaction scenarios, the existing interaction model has certain limitations in emotion recognition.

[0003] Existing interaction models often treat user dialogue text as a whole, relying on the contextual relationships between user statements to identify the emotional tone of the statements. This method is applicable to dialogues with relatively fixed topics, but in real-world human-model interactions, the dialogue topic is usually in a dynamic process of change. As the topic of the dialogue content changes, the emotions of the speakers also change accordingly. If the dialogue text is continued to be treated as a whole, it will seriously affect the model's ability to accurately identify the emotions in the dialogue, resulting in an inability to accurately grasp the user's true emotional state.

[0004] When the topic of adjacent conversations remains unchanged, changes in user emotions contain important interaction information. If a user's emotion worsens, it likely indicates that the model's initial response failed to effectively address the user's needs; conversely, if the user's emotion improves, it suggests that the user's needs may have been addressed or partially addressed in the initial response. However, current technologies lack an effective mechanism for analyzing the relationship between these emotional changes and need resolution, making it difficult to dynamically adjust response strategies based on emotional shifts. This results in a failure to optimize response content in a timely manner during the interaction, impacting user experience and need fulfillment. Summary of the Invention

[0005] The purpose of this invention is to provide an intelligent voice interaction system based on a large-scale cultural tourism model, which solves the following technical problems: Existing dialogue sentiment analysis technologies have significant shortcomings in recognizing emotions when faced with topic changes and optimizing responses based on emotion changes. There is an urgent need to propose an intelligent voice interaction system that can more accurately handle the correlation between dialogue emotions and topics and dynamically optimize responses.

[0006] The objective of this invention can be achieved through the following technical solutions: A smart voice interaction system based on a large-scale cultural tourism model includes the following modules: Data acquisition module: Acquires the dialogue text of two consecutive inputs from the user, denoted as text A and text B, and obtains the semantic feature vectors corresponding to text A and text B respectively; Text generation module: Calculates the correlation between text A and text B based on semantic feature vectors. If the correlation is greater than or equal to a preset threshold, text A and text B are integrated to obtain a composite text. Text analysis module: Input text A and the comprehensive text into the preset sentiment analysis model respectively to obtain the sentiment category and the corresponding sentiment score and calculate the sentiment score difference. If the sentiment score difference is less than the preset score threshold, obtain the sentiment keywords corresponding to text B and fuse them with the semantic feature vector of text A to obtain the association weight. Based on the association weight, rewrite text A to obtain the corrected text A. Result generation module: Based on the cultural tourism big data model, obtain the response text a corresponding to text A, identify the demand keywords for the corrected text A and the response text a respectively, compare the identification results to obtain the set of missing demand keywords, and determine the response text b corresponding to text B based on the set of missing demand keywords. Emotion Injection Module: Obtain the emotion category corresponding to the comprehensive text and label it as the current user's emotion category. Input the emotion category and the reply text b into the pre-trained large language model to generate a personalized interactive conversation for the corresponding emotion category.

[0007] As a further aspect of the present invention, it also includes preprocessing the dialogue text, the specific process of which is as follows: Remove special characters from the dialogue text using regular expressions; remove stop words from the dialogue text using a predefined stop word list and perform sentence segmentation on the dialogue text.

[0008] As a further aspect of the present invention, it also includes pre-constructing a quantitative scaling table for emotion categories, mapping a unique positive score to each emotion category, with the higher the degree of negativity, the smaller the positive score, and the higher the degree of positivity, the larger the positive score.

[0009] As a further aspect of the present invention: the specific process for obtaining the association weights in the text analysis module is as follows: Obtain the semantic feature vector of text A and the sentiment keywords of text B. Concatenate the semantic feature vector and the sentiment keywords to form a joint vector, which is then input into a preset multilayer perceptron model. The multilayer perceptron model includes a first fully connected layer, a first ReLU activation layer, a second fully connected layer, a second ReLU activation layer, and a third fully connected layer connected in sequence. Perform a linear transformation on the joint vector through the first fully connected layer and a non-linear transformation through the first ReLU activation layer to obtain the output of the first hidden layer. Perform a linear transformation on the output of the first hidden layer through the second fully connected layer and a non-linear transformation through the second ReLU activation layer to obtain the output of the second hidden layer. Perform a linear transformation on the output of the second hidden layer through the third fully connected layer to obtain an unnormalized weight vector. Input the unnormalized weight vector into a Softmax function for normalization to obtain the association weights.

[0010] As a further aspect of the present invention: the specific process for obtaining the corrected text A in the text analysis module is as follows: The semantic feature vector of text A is decomposed into multiple semantic components, and the original weights of each semantic component are obtained. By calculating the cosine similarity between the sentiment keywords corresponding to text B and each semantic component, a similarity matrix is ​​constructed to establish a mapping relationship between sentiment keywords and semantic components. Based on the mapping relationship and the association weights, each semantic component is weighted and adjusted to obtain the adjusted semantic components. According to the original weights and the association weights, the new weights of each semantic component are calculated. The adjusted semantic components are linearly combined according to the new weights to obtain the corrected semantic feature vector. All keywords and their corresponding weights are extracted from the corrected semantic feature vector, sorted in descending order of weight, and a word vector sequence is generated. The word vector sequence is integrated using natural language processing techniques to obtain the corrected text A.

[0011] As a further aspect of the present invention: the specific process of replying to text B in the result generation module is as follows: S1. Based on the set of missing requirement keywords, a graph traversal algorithm is used to retrieve first-level nodes and second-level nodes that are directly related to the missing requirement keywords from the domain knowledge graph. The retrieved knowledge node content is then processed in a structured manner to generate a set of candidate response content. S2, concatenate the set of missing demand keywords with the semantic feature vector of text B to obtain the target joint vector, calculate the correlation between the target joint vector and each candidate response content, and select the candidate response content with the highest correlation as the response text of text B.

[0012] As a further aspect of the present invention: the text analysis module further includes, if the difference in sentiment scores is greater than or equal to a preset threshold, identifying demand keywords in the response text a to obtain a set of demand keywords, using a graph traversal algorithm to retrieve first-level nodes and second-level nodes connected by relational paths from the domain knowledge graph that are directly related to the demand keywords, performing structured processing on the retrieved knowledge node content to generate a set of candidate response content, concatenating the set of demand keywords with the semantic feature vector of text B to obtain a target joint vector, calculating the correlation between the target joint vector and each candidate response content, and selecting the candidate response content with the highest correlation as the response text of text B.

[0013] As a further aspect of the present invention, the text generation module further includes a function that, if the relevance is less than a preset threshold, directly responds to text B.

[0014] The beneficial effects of this invention are: 1) This invention integrates topic-related dialogues into a comprehensive text by acquiring adjacent dialogue texts and calculating their relevance. Combining a sentiment analysis model with a quantitative scaling table, it achieves accurate identification of user emotion categories and calculation of score differences. Compared to existing technologies that process the entire dialogue, this mechanism effectively addresses scenarios involving topic changes, avoiding sentiment recognition bias caused by topic switching. This allows for accurate understanding of user emotional fluctuations during interaction, providing a reliable basis for adjusting subsequent response strategies.

[0015] 2) When the detected difference in emotion scores is small, the original text is rewritten to correct semantic comprehension biases by fusing emotion keywords and semantic feature vectors to generate association weights. If the difference is large, value-added information is supplemented based on demand coverage assessment and knowledge graph retrieval. This scenario-based response generation logic breaks through the limitations of existing technologies that lack emotion-demand association analysis. It can dynamically adjust the response content according to changes in user emotions, effectively solving the problem of initial responses failing to meet needs and improving the efficiency of demand resolution during the interaction process.

[0016] 3) By labeling the emotional categories of the comprehensive text and injecting them into a large language model, the generated response text can match the user's current emotional state, forming a personalized conversation with emotional resonance. Compared with the shortcomings of existing technologies that ignore the influence of emotions on responses, this mechanism can achieve more natural human-computer interaction in cultural and tourism scenarios. For example, it can provide enthusiastic recommendations based on the user's pleasant emotions and quickly respond to and improve responses to dissatisfaction, thereby enhancing the comfort and satisfaction of the user experience. Attached Figure Description

[0017] The invention will now be further described with reference to the accompanying drawings.

[0018] Figure 1This is a schematic diagram of the process of an intelligent voice interaction system based on a large-scale cultural tourism model according to the present invention. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] Please see Figure 1 As shown, this invention is an intelligent voice interaction system based on a large-scale cultural tourism model, comprising the following modules: Data acquisition module: Acquires the dialogue text of two consecutive inputs from the user, denoted as text A and text B, and obtains the semantic feature vectors corresponding to text A and text B respectively; Text generation module: Calculates the correlation between text A and text B based on semantic feature vectors. If the correlation is greater than or equal to a preset threshold, text A and text B are integrated to obtain a composite text. Text analysis module: Input text A and the comprehensive text into the preset sentiment analysis model respectively to obtain the sentiment category and the corresponding sentiment score and calculate the sentiment score difference. If the sentiment score difference is less than the preset score threshold, obtain the sentiment keywords corresponding to text B and fuse them with the semantic feature vector of text A to obtain the association weight. Based on the association weight, rewrite text A to obtain the corrected text A. Result generation module: Based on the cultural tourism big data model, obtain the response text corresponding to text A, identify the demand keywords for the corrected text A and the response text respectively, compare the identification results to obtain the set of missing demand keywords, and determine the response text b corresponding to text B based on the set of demand keywords; Emotion Injection Module: Obtain the emotion category corresponding to the comprehensive text and label it as the current user's emotion category. Input the emotion category and the reply text b into the pre-trained large language model to generate a personalized interactive conversation for the corresponding emotion category.

[0021] 1) This invention integrates topic-related dialogues into a comprehensive text by acquiring adjacent dialogue texts and calculating their relevance. Combining a sentiment analysis model with a quantitative scaling table, it achieves accurate identification of user emotion categories and calculation of score differences. Compared to existing technologies that process the entire dialogue, this mechanism effectively addresses scenarios involving topic changes, avoiding sentiment recognition bias caused by topic switching. This allows for accurate understanding of user emotional fluctuations during interaction, providing a reliable basis for adjusting subsequent response strategies.

[0022] 2) When the detected difference in emotion scores is small, the original text is rewritten to correct semantic comprehension biases by fusing emotion keywords and semantic feature vectors to generate association weights. If the difference is large, value-added information is supplemented based on demand coverage assessment and knowledge graph retrieval. This scenario-based response generation logic breaks through the limitations of existing technologies that lack emotion-demand association analysis. It can dynamically adjust the response content according to changes in user emotions, effectively solving the problem of initial responses failing to meet needs and improving the efficiency of demand resolution during the interaction process.

[0023] 3) By labeling the emotional categories of the comprehensive text and injecting them into a large language model, the generated response text can match the user's current emotional state, forming a personalized conversation with emotional resonance. Compared with the shortcomings of existing technologies that ignore the influence of emotions on responses, this mechanism can achieve more natural human-computer interaction in cultural and tourism scenarios. For example, it can provide enthusiastic recommendations based on the user's pleasant emotions and quickly respond to and improve responses to dissatisfaction, thereby enhancing the comfort and satisfaction of the user experience.

[0024] In a preferred embodiment of the present invention, the data analysis module further includes preprocessing the dialogue text, the specific process of which is as follows: Special characters in the dialogue text are removed using regular expressions; stop words in the case text are removed using a predefined stop word list; and the dialogue text is segmented into sentences.

[0025] By using regular expressions to remove special characters, interference from emojis and punctuation marks can be avoided in semantic feature extraction, allowing the subsequently obtained semantic feature vectors to focus more on core content. Using a stop word list to filter meaningless words reduces redundant information in the text, lowers data dimensionality, and improves model processing efficiency, while highlighting key semantic information (such as attraction names and consultation needs). Employing a sentence segmentation model combining bidirectional LSTM and CRF accurately divides the semantic units of the dialogue text, ensuring that the sentiment analysis model can accurately calculate the sentiment category and score for each independent sentence (e.g., distinguishing between the consultation sentiment of interrogative sentences and the exclamatory sentiment of exclamatory sentences). This provides a more refined semantic foundation for subsequent correlation calculations, text correction, and response generation, enabling the entire system to more accurately understand user needs and emotional states, laying the data processing foundation for generating personalized interactive conversations.

[0026] In another preferred embodiment of the present invention, the text analysis module further includes a pre-constructed emotion category quantification scale, which maps a unique positive score to each emotion category, with the higher the degree of negativity, the smaller the positive score, and the higher the degree of positivity, the larger the positive score.

[0027] When pre-constructing a quantitative scale for emotion categories, we first identify common user emotion categories in cultural and tourism interaction scenarios, such as classifying emotions into types like "pleasure," "expectation," "calm," "confusion," "dissatisfaction," and "anger." Then, we map a unique positive score to each category based on the degree of positive or negative emotion. For example, "pleasure," as a strong positive emotion, is assigned a relatively high positive score; "expectation," a moderately positive emotion, has a score lower than "pleasure"; "calm," as a neutral emotion, has a score in the middle; "confusion," a mildly negative emotion, has a score lower than the neutral value; "dissatisfaction," being more negative than "confusion," has a lower score; and "anger," as a strong negative emotion, is assigned the lowest positive score. This forms a quantitative system where "the higher the degree of positivity, the higher the score; the higher the degree of negativity, the lower the score." In actual construction, we can set "pleasure" to correspond to the high score range and "anger" to the low score range, with the scores of each emotion category arranged in order of positive or negative degree.

[0028] The purpose of constructing a quantitative scale table for emotion categories is to transform users' qualitative emotions into calculable quantitative values, facilitating subsequent emotion analysis models to calculate the difference between the emotion scores of text A and the composite text. By assigning a unique positive score to each emotion category, the system can quantify and measure the magnitude of changes in user emotions (e.g., an increase in score from "expectation" to "pleasure" indicates improved emotion, while a decrease in score from "calm" to "dissatisfaction" indicates worsened emotion). Based on this quantitative result, the system can then determine whether to correct the semantic understanding of text A or adjust the response strategy. This quantification method provides a unified numerical standard for emotion analysis, enabling the system to more accurately capture subtle changes in user emotions. This lays a data foundation for subsequent text rewriting based on emotion keywords and generating emotion-appropriate responses, ultimately helping the system achieve dynamic responses to user emotional states and improving the naturalness and user satisfaction of intelligent voice interaction in cultural and tourism scenarios.

[0029] In another preferred embodiment of the present invention, the specific process for obtaining the association weight in the text analysis module is as follows: Obtain the semantic feature vector of text A and the sentiment keyword vector of text B. Concatenate the semantic feature vector and the sentiment keyword vector to form a joint vector. Input the joint vector into a preset multilayer perceptron model, which includes a first fully connected layer, a first ReLU activation layer, a second fully connected layer, a second ReLU activation layer, and a third fully connected layer connected in sequence. Perform a linear transformation on the joint vector through the first fully connected layer and a non-linear transformation through the first ReLU activation layer to obtain the output of the first hidden layer. Perform a linear transformation on the output of the first hidden layer through the second fully connected layer and a non-linear transformation through the second ReLU activation layer to obtain the output of the second hidden layer. Perform a linear transformation on the output of the second hidden layer through the third fully connected layer to obtain an unnormalized weight vector. Input the unnormalized weight vector into a Softmax function for normalization to obtain an associated weight vector. Each element in the associated weight vector represents the influence weight of the corresponding sentiment keyword on the semantic understanding of text A, and the sum of all elements is 1.

[0030] Suppose text A is "Is the glass walkway at scenic spot A open today? Is it affected by the weather?", its semantic feature vector is extracted using the BERT model. Text B is "I went yesterday but it was closed. What a waste of time, so disappointing!", the sentiment keywords "disappointing" and "wasted time" are extracted and mapped to sentiment keyword vectors. The semantic feature vector and sentiment keyword vector are then concatenated in sequence to form an 868-dimensional joint vector. The joint vector is input into the multilayer perceptron model: the first fully connected layer performs a linear transformation on the joint vector, outputting a 512-dimensional vector, which is then passed through the first ReLU activation layer to obtain the output of the first hidden layer; the second fully connected layer performs a linear transformation on the output of the first hidden layer, outputting a 256-dimensional vector, which is then passed through the second ReLU activation layer to obtain the output of the second hidden layer; the third fully connected layer performs a linear transformation on the output of the second hidden layer to obtain an unnormalized weight vector (e.g., [3.2, -1.5]), which is then input into the Softmax function to calculate the probability and obtain the associated weight vector (e.g., [0.82, 0.18]), indicating that the weight of "disappointment" on the semantic understanding of text A is 0.82, and the weight of "wasted trip" is 0.18.

[0031] This mechanism deeply integrates emotional keywords from subsequent user conversations with original semantic features. By leveraging the nonlinear transformation capabilities of a multilayer perceptron, it captures the complex relationships between emotion and semantics, enabling the model to learn the degree of influence of different emotional keywords on the semantic understanding of the original text. The association weights obtained after Softmax normalization can be directly used to guide the rewriting process of text A (e.g., semantic components corresponding to emotional keywords with high weights will be adjusted more significantly), avoiding the neglect of the impact of user emotional changes on demand understanding due to reliance solely on semantic features. This mechanism allows the system to dynamically perceive the need for semantic understanding correction based on emotional fluctuations during interaction. For example, when a user expresses dissatisfaction, the model strengthens the recognition of semantics related to negative emotions based on association weights, thereby more accurately adjusting the response strategy. Ultimately, this helps the system generate personalized interactive content that meets both semantic requirements and adapts to user emotions, improving the naturalness and demand fulfillment efficiency of human-computer dialogue in cultural and tourism scenarios.

[0032] In another preferred embodiment of the present invention, the specific process of obtaining the corrected text A in the text analysis module is as follows: The semantic feature vector of text A is decomposed into multiple semantic components, and the original weights of each semantic component are obtained. By calculating the cosine similarity between the sentiment keywords corresponding to text B and each semantic component, a similarity matrix is ​​constructed to establish a mapping relationship between sentiment keywords and semantic components. Based on the mapping relationship and the association weights, each semantic component is weighted and adjusted to obtain the adjusted semantic components. According to the original weights and the association weights, the new weights of each semantic component are calculated. The adjusted semantic components are linearly combined according to the new weights to obtain the corrected semantic feature vector. All keywords and their corresponding weights are extracted from the corrected semantic feature vector, sorted in descending order of weight, and a word vector sequence is generated. The word vector sequence is integrated using natural language processing techniques to obtain the corrected text A.

[0033] Suppose text A, extracted using the BERT model, yields a 768-dimensional semantic feature vector. Its content is "Is attraction B open tomorrow? Do I need to make a reservation?". This vector is semantically decomposed into four components: "Attraction Name (Attraction B)," "Time Inquiry (Tomorrow)," "Opening Status (Open)," and "Reservation Requirement (Reservation)," with original weights of [0.3, 0.25, 0.25, 0.2]. Text B is "I went before but couldn't get in without a reservation, what a waste of time!". The sentiment keywords "trouble" and "waste of time" are extracted and mapped to a 100-dimensional vector. The cosine similarity between the sentiment keyword vector and each semantic component vector is calculated. For example, "trouble" has a similarity of 0.8 with "reservation requirement" and 0.3 with "opening status," while "waste of time" has a similarity of 0.7 with "time inquiry" and 0.6 with "reservation requirement." A 2×4 similarity matrix is ​​constructed. A mapping relationship is established based on the matrix. For example, "trouble" is mainly associated with "reservation request," and "waste of time" is mainly associated with "time query" and "reservation request." Combining the previously obtained association weight vector [0.7, 0.3] (assuming "trouble" has a weight of 0.7 and "waste of time" has a weight of 0.3), the semantic components are weighted and adjusted: the "reservation request" component is adjusted to the original value × (0.7 × 0.8 + 0.3 × 0.6), and the "time query" component is adjusted to the original value × (0.3 × 0.7). New weights are calculated based on the original weights and association weights, such as the new weight of "reservation request" = 0.2 × (0.7 × 0.8 + 0.3 × 0.6). The adjusted components are linearly combined according to the new weights to obtain the corrected semantic feature vector. The keywords "attraction B," "tomorrow," "open," and "reservation" and their adjusted weights are extracted from this vector to generate the corrected text A: "Is attraction B open tomorrow? Is it necessary to make a reservation in advance? Have any tourists been hindered by not making a reservation?"

[0034] Because emotional keywords in user dialogues often implicitly supplement or correct the original semantic needs, by decomposing the semantic feature vector into fine-grained semantic components, the specific semantic dimensions affected by emotional keywords can be accurately located. Furthermore, cosine similarity calculation and the construction of a similarity matrix can quantify the correlation strength between emotional keywords and each semantic component. After establishing a mapping relationship, weighted adjustments based on correlation weights can make the expression of semantic components more closely match the user's true emotional inclination. Recalculating new weights for semantic components and linearly combining them to generate a correction vector ensures that the corrected semantic features retain the core needs of the original text while incorporating the semantic shifts brought about by emotions. Finally, extracting keywords and integrating the generated corrected text A can more accurately reflect the user's true intentions driven by emotions, providing a more accurate semantic basis for subsequent response generation. This helps the system achieve emotional interaction, enabling the intelligent voice interaction system to more accurately understand user needs and emotional states in cultural and tourism scenarios, thereby improving interaction efficiency and user experience.

[0035] In another preferred embodiment of the present invention, the specific process of replying to text B in the result generation module is as follows: S1. Based on the set of missing requirement keywords, a graph traversal algorithm is used to retrieve first-level nodes and second-level nodes that are directly related to the missing requirement keywords from the domain knowledge graph. The retrieved knowledge node content is then processed in a structured manner to generate a set of candidate response content. S2, concatenate the set of missing demand keywords with the semantic feature vector of text B to obtain the target joint vector, calculate the correlation between the target joint vector and each candidate response content, and select the candidate response content with the highest correlation as the response text of text B.

[0036] If the missing keyword set includes "A Scenic Area Glass Skywalk Opening Hours" and "Reservation Policy", a graph traversal algorithm is used to first retrieve the first-level nodes directly associated with "A Scenic Area Glass Skywalk Opening Hours", such as "Opening Hours" and "Closing Time". Then, second-level nodes connected by relational paths are retrieved, such as "Seasonal Adjustment of Opening Hours". At the same time, the first-level nodes "Reservation Method" and "Reservation Time Limit" and the second-level nodes "Reservation Cancellation Rules" directly associated with "Reservation Policy" are retrieved. The retrieved knowledge node content is parsed to extract the entity "A Scenic Area Glass Skywalk" and its attributes such as "Opening Hours are 8:00-18:00" and "Online Reservation Required 1 Day in Advance" are extracted and converted into a structured data format. The association information of "Opening Hours" and "Weather-Affected Closure Rules" in the first-level and second-level nodes is integrated. Duplicate descriptions of "Reservation Platform" information are removed. Finally, a set of candidate response content is generated, such as "A Scenic Area Glass Skywalk Opening Hours are 8:00-18:00, Reservations Required 1 Day in Advance on the Official Platform" and "Heavy rain may cause temporary closure of the glass skywalk, it is recommended to check the weather before traveling". The missing keyword set is concatenated with the semantic feature vector of text B to obtain the target joint vector. The correlation between the target joint vector and each candidate response content is calculated, and the candidate response content with the highest correlation is selected as the response text of text B. Specifically, if the missing keyword set is "opening hours" and "reservation policy", the semantic feature vector of text B, after being extracted by the BERT model, contains the semantic feature of "not open yesterday". The two are concatenated to form the target joint vector. The semantic feature vector of each candidate response content is extracted. The correlation between the target joint vector and "the opening hours of the glass walkway in Scenic Area A are 8:00-18:00, and reservations must be made one day in advance on the official platform" and "the opening status of other attractions in the scenic area" are calculated using cosine similarity. The former with the highest correlation is selected as the response text.

[0037] Leveraging the structured retrieval capabilities of knowledge graphs, information directly or indirectly related to missing keywords is quickly located, ensuring the comprehensiveness and accuracy of responses and avoiding overlooking potential user needs. The text content of knowledge nodes is parsed, and core entities, attributes, and relationships are extracted using natural language processing techniques. Entities are the specific objects described by the knowledge nodes, attributes are the characteristics or states of the entities, and relationships are the types of associations between the entities and other nodes. Unstructured text information is converted into a unified structured data format, including but not limited to JSON, XML, or key-value pairs. During the standardization process, entities, attributes, and relationships are defined in a formatted manner to ensure the standardization and consistency of the data structure. Information from first-level and second-level nodes is linked and integrated, and contextual information is supplemented through the relationship paths between nodes, forming a logically coherent information set. The secondary nodes are derived nodes connected to the primary nodes through relational paths. They deduplicate the integrated knowledge node content, removing duplicate information entries and retaining core data highly relevant to the missing requirement keywords, thus generating a set of candidate response content to provide a basis for subsequent screening. The missing requirement keywords are concatenated with the semantic feature vector of text B to form a target joint vector, which can simultaneously take into account the user's explicit missing requirements and the semantic information implicit in text B. By calculating the correlation with the candidate response content and selecting the highest correlation, it can be ensured that the generated response text covers the user's explicit needs, making the response more accurate and in line with the user's current dialogue intent. This helps the system achieve personalized interaction based on emotion and semantic fusion, improves the efficiency and naturalness of intelligent voice interaction, and better meets the user's needs in cultural and tourism scenarios.

[0038] In another preferred embodiment of the present invention, the text analysis module further includes, if the difference in sentiment scores is greater than or equal to a preset threshold, identifying demand keywords in the response text a to obtain a set of demand keywords, using a graph traversal algorithm to retrieve first-level nodes and second-level nodes connected by relational paths from the domain knowledge graph that have a direct relationship with the demand keywords, performing structured processing on the retrieved knowledge node content to generate a set of candidate response content, concatenating the set of demand keywords with the semantic feature vector of text B to obtain a target joint vector, calculating the correlation between the target joint vector and each candidate response content, and selecting the candidate response content with the highest correlation as the response text of text B.

[0039] In another preferred embodiment of the present invention, the text generation module further includes a function that directly replies to text B if the relevance is less than a preset threshold.

[0040] For example, if text A is "How do I book tickets for Museum C?" and text B is "Is there a lotus exhibition at Park D?", the semantic feature vector calculation shows a correlation of 0.15, which is less than the threshold. In this case, we can directly process text B, "Lotus exhibition at Park D": by searching the knowledge base in the cultural tourism field, we can extract information such as "The lotus exhibition at Park D is usually held from June to August, and it is currently June, which is the peak blooming season", or call a pre-trained language model to generate a corresponding reply, forming a direct reply text, "Park D is currently holding a lotus exhibition, which runs until June 30. It is recommended to visit in the morning."

[0041] When the relevance of adjacent dialogue texts falls below a preset threshold, it indicates a significant shift in the user's dialogue topic. Forcibly integrating unrelated texts A and B can lead to semantic misunderstandings and confusing responses. Directly responding to text B allows for a rapid response to the user's current needs, avoiding interaction delays caused by topic confusion and ensuring a smooth and real-time interaction flow. In dialogue scenarios with sudden topic shifts, accurately capturing the user's latest focus and skipping unnecessary text integration and correction processes, directly utilizing relevant domain knowledge to generate responses, effectively improves the adaptability of the intelligent voice interaction system to dynamic dialogue topics, contributing to the goal of efficient and natural human-computer emotional interaction.

[0042] The foregoing has provided a detailed description of one embodiment of the present invention, but this description is merely a preferred embodiment and should not be construed as limiting the scope of the invention. All equivalent variations and modifications made within the scope of the claims of this invention should still fall within the patent coverage of this invention.

Claims

1. A smart voice interaction system based on a large-scale cultural tourism model, characterized in that, Includes the following modules: Data acquisition module: Acquires the dialogue text of two consecutive inputs from the user, denoted as text A and text B, and obtains the semantic feature vectors corresponding to text A and text B respectively; Text generation module: Calculates the correlation between text A and text B based on semantic feature vectors. If the correlation is greater than or equal to a preset threshold, text A and text B are integrated to obtain a composite text. Text analysis module: Input text A and the comprehensive text into the preset sentiment analysis model respectively to obtain the sentiment category and the corresponding sentiment score and calculate the sentiment score difference. If the sentiment score difference is less than the preset score threshold, obtain the sentiment keywords corresponding to text B and fuse them with the semantic feature vector of text A to obtain the association weight. Based on the association weight, rewrite text A to obtain the corrected text A. Result generation module: Based on the cultural tourism big data model, obtain the response text a corresponding to text A, identify the demand keywords for the corrected text A and the response text a respectively, compare the identification results to obtain the set of missing demand keywords, and determine the response text b corresponding to text B based on the set of missing demand keywords. Emotion Injection Module: Obtain the emotion category corresponding to the comprehensive text and label it as the current user's emotion category. Input the emotion category and the reply text b into the pre-trained large language model to generate a personalized interactive conversation for the corresponding emotion category.

2. The intelligent voice interaction system based on a large-scale cultural tourism model according to claim 1, characterized in that, The data acquisition module also includes preprocessing the dialogue text. The specific preprocessing process is as follows: Remove special characters from the dialogue text using regular expressions; remove stop words from the dialogue text using a predefined stop word list and perform sentence segmentation on the dialogue text.

3. The intelligent voice interaction system based on a large-scale cultural tourism model according to claim 1, characterized in that, The text analysis module also includes a pre-built quantitative scaling table for emotion categories, which maps a unique positive score to each emotion category. The higher the degree of negativity, the smaller the positive score, and the higher the degree of positivity, the larger the positive score.

4. The intelligent voice interaction system based on a large-scale cultural tourism model according to claim 1, characterized in that, The specific process for obtaining the association weights in the text analysis module is as follows: Obtain the semantic feature vector of text A and the sentiment keywords of text B. Concatenate the semantic feature vector and the sentiment keywords to form a joint vector, which is then input into a preset multilayer perceptron model. The multilayer perceptron model includes a first fully connected layer, a first ReLU activation layer, a second fully connected layer, a second ReLU activation layer, and a third fully connected layer connected in sequence. Perform a linear transformation on the joint vector through the first fully connected layer and a non-linear transformation through the first ReLU activation layer to obtain the output of the first hidden layer. Perform a linear transformation on the output of the first hidden layer through the second fully connected layer and a non-linear transformation through the second ReLU activation layer to obtain the output of the second hidden layer. Perform a linear transformation on the output of the second hidden layer through the third fully connected layer to obtain an unnormalized weight vector. Input the unnormalized weight vector into a Softmax function for normalization to obtain the association weights.

5. The intelligent voice interaction system based on a large-scale cultural tourism model according to claim 4, characterized in that, In the text analysis module, the specific process of obtaining the corrected text A is as follows: The semantic feature vector of text A is decomposed into multiple semantic components, and the original weights of each semantic component are obtained. By calculating the cosine similarity between the emotional keywords corresponding to text B and each semantic component, a similarity matrix is ​​constructed to establish the mapping relationship between emotional keywords and semantic components. Based on the mapping relationship and association weights, the semantic components are weighted and adjusted to obtain the adjusted semantic components; the new weights of each semantic component are calculated based on the original weights and association weights. The adjusted semantic components are linearly combined according to the new weights to obtain the corrected semantic feature vector; all keywords and their corresponding weights are extracted from the corrected semantic feature vector, sorted in descending order of weight to generate a word vector sequence, and the word vector sequence is integrated using natural language processing technology to obtain the corrected text A.

6. The intelligent voice interaction system based on a large-scale cultural tourism model according to claim 1, characterized in that, In the result generation module, the specific generation process of the response text b is as follows: S1. Based on the set of missing requirement keywords, a graph traversal algorithm is used to retrieve first-level nodes and second-level nodes that are directly related to the missing requirement keywords from the domain knowledge graph. The retrieved knowledge node content is then processed in a structured manner to generate a set of candidate response content. S2, concatenate the set of missing demand keywords with the semantic feature vector of text B to obtain the target joint vector, calculate the correlation between the target joint vector and each candidate response content, and select the candidate response content with the highest correlation as the response text of text B.

7. The intelligent voice interaction system based on a large-scale cultural tourism model according to claim 1, characterized in that, The text analysis module further includes, if the difference in sentiment scores is greater than or equal to a preset threshold, identifying demand keywords in the response text a to obtain a set of demand keywords. A graph traversal algorithm is then used to retrieve first-level nodes and second-level nodes connected via relational paths from the domain knowledge graph that are directly related to the demand keywords. The retrieved knowledge node content is then structured to generate a set of candidate response content. The set of demand keywords is concatenated with the semantic feature vector of text B to obtain a target joint vector. The correlation between the target joint vector and each candidate response content is calculated, and the candidate response content with the highest correlation is selected as the response text for text B.

8. The intelligent voice interaction system based on a large-scale cultural tourism model according to claim 1, characterized in that, The text generation module also includes a function that, if the relevance is less than a preset threshold, directly responds to text B.

Citation Information

Patent Citations

  • Intelligent customer service question and answer optimization method and system based on sentiment analysis

    CN119626264A

  • Replay generation method and device based on knowledge graph application intelligent question and answer

    CN120316229A

  • AI customer service robot automatic reply system based on knowledge base

    CN120336473A