Interaction optimization method, system and equipment based on sentiment analysis and storage medium

By adopting an interaction optimization method based on sentiment analysis in the intelligent question-and-answer system, the existing system is solved in dealing with problems in multiple media forms, improving the accuracy and user experience of answers.

CN120163166APending Publication Date: 2025-06-17SHANDONG INSPUR BROADCOM INFORMATION TECHNOLOGY CO LTD +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510064484.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The existing intelligent question-and-answer system has limited practicality and user experience when dealing with user questions in various media forms, and the accuracy of the answers is difficult to guarantee.

Method used

An interaction optimization method based on sentiment analysis is adopted to collect input data from multiple channels, extract multi-source feature data, extract entities and relationships, generate query statements, and score and sort answers based on sentiment labels.

Benefits of technology

It improves the accuracy and user experience of the answers, provides the answer that best matches user needs, and enhances the practicality of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163166A_ABST
    Figure CN120163166A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and particularly provides an interaction optimization method, system and device based on sentiment analysis and a storage medium, and the method comprises the steps: collecting input data of an interaction object from a plurality of channels in an interaction process, and obtaining multi-source data; extracting multi-source feature data from the multi-source data; entities and relations are extracted from the multi-source feature data, and query statements are generated based on the entities and the relations; identifying an emotion label corresponding to an entity in the query statement from the multi-source feature data; and based on the query statement and the corresponding emotion tag, scoring a plurality of answers generated based on the query statement and a pre-constructed knowledge graph, and sorting and displaying the plurality of answers according to scores from high to low. According to the invention, multi-source input data can be identified, and answers most matched with user demands can be provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence, and particularly relates to an interaction optimization method, system, device and storage medium based on sentiment analysis. Background Art

[0002] An Intelligent Question Answering System (IQAS) is a system based on natural language processing (NLP) and artificial intelligence technologies that can understand and answer questions raised by users. Currently, intelligent question answering has broad application prospects in businesses such as building artificial intelligence, intelligent customer service, medical and health, banking business consulting, and legal consulting.

[0003] However, most of the existing intelligent question answering systems focus on the processing of single-modal text, ignoring the actual situation that users may ask questions through various media forms such as voice, pictures, and videos, which limits the practicality of the system and the user experience. In addition, when generating answers based on a knowledge graph, multiple answers may be generated, and the answers may not exactly correspond to the user's questions. How to improve the accuracy of the answers is also a technical problem that needs to be solved. Summary of the Invention

[0004] In view of the above deficiencies of the prior art, the present invention provides an interaction optimization method, system, device and storage medium based on sentiment analysis to solve the above technical problems.

[0005] In a first aspect, the present invention provides an interaction optimization method based on sentiment analysis, including: Collecting input data of an interaction object from multiple channels during the interaction process to obtain multi-source data; Extracting multi-source feature data from the multi-source data; Extracting entities and relationships from the multi-source feature data and generating a query statement based on the entities and relationships; Identifying sentiment labels corresponding to the entities in the query statement from the multi-source feature data; Based on the query statement and the corresponding sentiment labels, scoring multiple answers generated based on the query statement and a pre-constructed knowledge graph, and sorting and displaying the multiple answers in descending order of the scores.

[0006] In an optional embodiment, collecting input data of an interaction object from multiple channels during the interaction process to obtain multi-source data includes: Obtaining the input data of the interaction object, where the input data includes text data, voice data, and image data.

[0007] In an optional embodiment, extracting multi-source feature data from the multi-source data includes: Use a word segmentation tool to perform word segmentation on the text data, and use the Glove model to convert the segmented text data into text features; Use the LibROSA speech package to extract frame-level speech features from the speech data; Use OpenCV to crop the image data, and use the ResNet-50 network model to extract facial features from the cropped image data.

[0008] In an alternative embodiment, entities and relationships are extracted from the multi-source feature data, and a query statement is generated based on the entities and relationships, including: Use a pre-trained language model to extract entities and relationships from the text features; Determine the specific format of the query statement. According to the results of entity relationship extraction, the identified entities and relationships are combined according to the defined query statement format to construct a query statement; Deduplicate and filter the constructed query statement to remove duplicate and invalid query statements.

[0009] In an alternative embodiment, identifying an emotional label corresponding to an entity in the query statement from the multi-source feature data includes: Arrange the text features, speech features, and image features in the order of generation time into a text feature sequence, a speech feature sequence, and an image feature sequence; Use a convolutional network to align the lengths of the text feature sequence, the speech feature sequence, and the image feature sequence; Input the aligned text feature sequence, speech feature sequence, and image feature sequence into a bidirectional recurrent neural network, and use the output parameters of the bidirectional recurrent neural network as the input of a cross-modal attention mechanism to obtain weight matrices for the text feature sequence, the speech feature sequence, and the image feature sequence; Use an attention mechanism to integrate the text feature sequence, the speech feature sequence, and the image feature sequence into a joint feature vector based on the weight matrices of the text feature sequence, the speech feature sequence, and the image feature sequence; Use a pre-trained deep learning model to identify the emotional label of the joint feature vector; Based on the multi-source data collection time corresponding to the emotional label and the multi-source data collection time corresponding to multiple entities in the query statement, establish a correspondence between the emotional label and the multiple entities.

[0010] In an alternative embodiment, the method further includes: Use deep reinforcement learning technology to train the attention mechanism.

[0011] In an alternative embodiment, based on the query statement and the corresponding sentiment label, score multiple answers generated based on the query statement and the pre-constructed knowledge graph, and sort and display the multiple answers in descending order of the scores, including: Preset the mood scores corresponding to the sentiment labels; Calculate the relevance of the answer to multiple entities, and convert the relevance into weights; Set the weights of the user scores; Based on the mood scores, weights of the multiple entities corresponding to the answer, the user scores, and the corresponding weights, calculate the comprehensive score of the answer; Sort and display the multiple answers in descending order of the comprehensive scores.

[0012] In a second aspect, the present invention provides an interaction optimization system based on sentiment analysis, including: A data acquisition module, configured to collect input data of an interaction object from multiple channels during the interaction process to obtain multi-source data; A feature extraction module, configured to extract multi-source feature data from the multi-source data; A content extraction module, configured to extract entities and relationships from the multi-source feature data, and generate a query statement based on the entities and relationships; A sentiment analysis module, configured to identify sentiment labels corresponding to the entities in the query statement from the multi-source feature data; An answer sorting module, configured to score multiple answers generated based on the query statement and the pre-constructed knowledge graph based on the query statement and the corresponding sentiment label, and sort and display the multiple answers in descending order of the scores.

[0013] In a third aspect, there is provided a device, including: A memory, configured to store an interaction optimization program based on sentiment analysis; A processor, configured to implement the steps of the interaction optimization method based on sentiment analysis provided in the first aspect when executing the interaction optimization program based on sentiment analysis.

[0014] In a fourth aspect, there is provided a computer-readable storage medium, on which an interaction optimization program based on sentiment analysis is stored. When the interaction optimization program based on sentiment analysis is executed by a processor, the steps of the interaction optimization method based on sentiment analysis provided in the first aspect are implemented.

[0015] The beneficial effects of the present invention are as follows. The interaction optimization method, system, device and storage medium provided by the present invention perform text content recognition and emotion recognition on multi-source input data, generate multiple relevant answers based on the recognized text content, and use the recognized emotion tags to screen and sort the answers, so as to provide the answer that best matches the user's needs.

[0016] In addition, the design principle of the present invention is reliable, the structure is simple, and it has a very broad application prospect. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0018] Figure 1 It is a schematic flowchart of the method according to an embodiment of the present invention.

[0019] Figure 2 It is another schematic flowchart of the method according to an embodiment of the present invention.

[0020] Figure 3 It is a schematic flowchart of multi-modal feature fusion of the method according to an embodiment of the present invention.

[0021] Figure 4 It is a schematic block diagram of the system according to an embodiment of the present invention.

[0022] Figure 5 It is a schematic structural diagram of a device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] In order to enable those skilled in the art to better understand the technical solutions in the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field of the present invention. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments, and are not intended to limit the present invention.

[0025] The following explains the key terms appearing in the present invention.

[0026] Enter data: Text data is one of the most common data types in tourism data analysis. Text data comes from user-generated text content, such as tourist reviews on online travel websites. Text data can be collected from multiple channels, including but not limited to books, websites, and travel blogs. By training a knowledge graph, extracting the historical context of each scenic spot, and analyzing users' hot issues and hot words, consumer behavior and decisions can be evaluated, and users' behavior habits and preferences can be understood.

[0027] Image data refers to the facial expressions of tourists captured in real time by a camera. During human-computer interaction, multimodal input can be achieved through video, voice dialogue, and text recognition. By analyzing tourists' facial expressions, their emotional states can be captured and analyzed in real time.

[0028] Audio data mainly includes the voices of tourists in conversations. Tourist recordings mainly refer to those during the conversation between users and the intelligent question-and-answer system. Through natural language processing (NLP) technology and voice tone recognition technology, emotional analysis can be performed on a large number of tourist recordings to understand tourists' emotions and needs.

[0029] Multimodality refers to a way of completing tasks by simultaneously or sequentially using two or more different types of information carriers (or modalities) in the fields of information processing, human-computer interaction, perception, and understanding. Common modalities include but are not limited to text, image, video, audio, gesture, etc. Each modality provides a unique perspective and information on different aspects of the real world. For example, text is good at expressing logic and details precisely, images and videos are good at presenting visual scenes, and audio and voice carry information about emotions and instant communication.

[0030] The importance of multimodal fusion: Enhance information complementarity: There is often complementarity between different modalities. By fusing them, the advantages of multiple modalities can be combined to improve the integrity and richness of information. For example, in the scenario of intelligent customer service, combining the tone of voice and the semantics of text can help understand users' emotions more accurately.

[0031] Enhance the user experience: In the field of human-computer interaction, multimodal technology enables machines to better simulate human communication methods, providing a more natural, smooth, and intuitive interaction experience. For example, a smart home system that combines visual and voice feedback can better meet users' living habits and needs.

[0032] Adapting to complex scenarios: In certain specific fields, such as medical diagnosis, security monitoring, autonomous driving, etc., multimodal fusion is particularly important. In these scenarios, it is often necessary to process highly complex and ever-changing information, and a single modality is difficult to meet the requirements, while multimodal fusion can provide a more comprehensive environmental perception and understanding ability.

[0033] The interactive optimization method based on sentiment analysis provided by the embodiments of the present invention is executed by a computer device. Correspondingly, the interactive optimization system based on sentiment analysis runs in the computer device.

[0034] Figure 1 It is a schematic flowchart of the method of an embodiment of the present invention. Among them, Figure 1 The execution subject can be an interactive optimization system based on sentiment analysis. According to different requirements, the order of the steps in this flowchart can be changed, and some can be omitted.

[0035] Such as Figure 1 shown, the method includes: S1. During the interaction process, collect the input data of the interaction object from multiple channels to obtain multi-source data.

[0036] In a complex interaction scenario, the system not only obtains data from channels such as text, voice, and images directly input by the user, but also collects the user's behavior information, preference settings, historical interaction records, etc. in an all-round way through various methods such as log records, third-party interfaces, and sensor data to ensure the comprehensiveness and diversity of the data. In this way, the multi-source data not only includes direct user input, but also covers various types of information that indirectly reflect user needs and environmental backgrounds.

[0037] S2. Extract multi-source feature data from the multi-source data.

[0038] With the help of machine learning algorithms and deep learning models, the system deeply analyzes the collected multi-source data and extracts key feature data. These feature data may include keywords and semantic concepts in text, object recognition and scene classification in images, intonation and speech rate features in speech, as well as user behavior patterns and time preferences. The feature extraction process aims to retain the most valuable information in the data and provide a solid foundation for subsequent processing.

[0039] S3. Extract entities and relationships from the multi-source feature data, and generate query statements based on the entities and relationships.

[0040] Through natural language processing techniques such as named entity recognition and dependency parsing, the system can identify specific entities (such as person names, place names, events, etc.) and their relationships from the feature data. Subsequently, using knowledge graph technology, these entities and relationships are organized into a structured query statement form (such as "<Entity A, relationship, Entity B>") for subsequent semantic understanding and reasoning.

[0041] S4. Identify the sentiment labels corresponding to the entities in the query statement from the multi-source feature data.

[0042] Combining various means such as text sentiment analysis, image sentiment recognition, and user behavior pattern analysis, the system can automatically judge and assign corresponding sentiment labels (such as positive, negative, neutral) for each entity or related context. This process not only considers directly expressed sentiment words but also synthesizes various factors such as context, intonation, and user historical emotions to achieve more accurate sentiment recognition.

[0043] S5. Based on the query statement and the corresponding sentiment labels, score the multiple answers generated based on the query statement and the pre-constructed knowledge graph, and sort and display the multiple answers in descending order of the scores.

[0044] According to factors such as the semantic accuracy of the query statement, the matching degree of the sentiment labels, the integrity and relevance of the answers, and the user's personalized preferences, a comprehensive scoring model is designed. This model scores each candidate answer and sorts them in descending order according to the scores. The sorted answer list can be intuitively presented to the user to help the user quickly find the information that best meets the needs and has the most reference value. At the same time, the system can also continuously optimize the scoring model according to the user's feedback to improve the accuracy and personalization of answer recommendations.

[0045] For the convenience of understanding the present invention, hereinafter, the principle of the interactive optimization method based on sentiment analysis of the present invention will be combined with Figure 2 to further describe the interactive optimization method based on sentiment analysis provided by the present invention.

[0046] In an embodiment of the present invention, based on step S1, the following will give a possible embodiment to non-restrictively elaborate on its specific implementation scheme.

[0047] The data is divided into text data, image data, and audio data. The text data is collected through multiple channels such as tourist reviews on online travel websites, social media, and Q&A engines; the image data is collected through the facial expressions and body movements of tourists captured in real time on social platforms or by cameras; the audio data is collected through the recordings of tourists during online communication with the intelligent Q&A system.

[0048] The data storage of the intelligent question-answering system comes from the knowledge graph constructed based on pre-collected data. After entity recognition, it is stored as an RDF query statement. After entity recognition based on the input question, the knowledge graph is queried, and entity linking is performed on the relevant answers to generate answers.

[0049] The intelligent question-answering system supports users to have a video conversation with the system and conducts intelligent question-answering in the form of "video chatting" with the virtual human.

[0050] During the processing of user input data by the intelligent question-answering system through "video chatting", it supports users to communicate with the system in real time via video, capturing user facial expression data, user voice data, and text data obtained by converting the user's question from speech to text.

[0051] After collecting the data, preprocessing is performed on the text, voice, and image data. First, audio is processed, and the SpeechRecognition tool is used to convert the microphone input into text.

[0052] In an embodiment of the present invention, based on step S2, a possible embodiment will be given below to non-restrictively elaborate on its specific implementation scheme.

[0053] S201. Use a word segmentation tool to perform word segmentation on the text data, and use the Glove model to convert the word-segmented text data into text features.

[0054] Word segmentation and part-of-speech tagging: Use a word segmentation tool (jieba) to cut continuous text into individual lexical units. Perform part-of-speech tagging to assign corresponding parts of speech (such as nouns, verbs, adjectives, etc.) to each word.

[0055] Stop word filtering: Create a stop word list containing common meaningless words (such as "de", "le", "is", "the", etc.). Remove these stop words from the word segmentation results to reduce data sparsity.

[0056] Load the pre-trained Glove model: Download the pre-trained Glove word vector model from the Internet or open-source resources. Ensure that the model is compatible with the vocabulary used by the word segmentation tool.

[0057] Build a vocabulary mapping: Create a mapping table from vocabulary to word vectors. For each word in the word segmentation result, find its corresponding word vector in the Glove model.

[0058] Convert to text features: Replace each word in the text data with its corresponding Glove word vector. For a sentence or paragraph containing multiple words, it can be represented as the average or weighted sum of the word vectors (according to weights such as part of speech, TF-IDF, etc.).

[0059] S202. Extract frame-level speech features from the speech data using the LibROSA speech package.

[0060] Load audio data: Use the librosa.load function in LibROSA to load the audio data, which will return the time series of the audio signal and the sampling rate.

[0061] Extract frame-level speech features: Use the functions provided by LibROSA to extract various frame-level speech features, such as Mel spectrogram, short-time Fourier transform (STFT) results, pitch trajectory, spectral centroid, etc.

[0062] Feature vectorization: Vectorize the extracted features for subsequent input into the speech recognition model for recognition. You can use the librosa.util.frame function provided by LibROSA to convert the time series data into frame-level data, and use libraries such as NumPy for feature concatenation and normalization processing.

[0063] S203. Use OpenCV to crop the image data and use the ResNet-50 network model to extract facial features from the cropped image data.

[0064] I. Environment Preparation Install OpenCV: OpenCV is an open-source computer vision library that provides rich image processing functions. You can install it using pip: pip install opencv-python Install the deep learning framework and ResNet-50 model: You can choose deep learning frameworks such as TensorFlow or PyTorch. The ResNet-50 model can usually be found in the official model libraries of these frameworks and can be directly loaded.

[0065] II. Image Preprocessing Read the image: Use the cv2.imread function in OpenCV to read the image file.

[0066] import cv2 image = cv2.imread('path_to_image.jpg').

[0067] Denoising: You can use the filtering functions provided by OpenCV, such as Gaussian filtering, mean filtering, etc., to remove the noise in the image.

[0068] denoised_image = cv2.GaussianBlur(image, (5, 5), 0) Grayscale conversion: Convert the color image to a grayscale image to reduce the amount of data and simplify subsequent processing.

[0069] gray_image = cv2.cvtColor(denoised_image, cv2.COLOR_BGR2GRAY).

[0070] Crop the image: Crop the image as needed to remove irrelevant background information.

[0071] You can use the cv2.rectangle and cv2.crop functions in OpenCV (note: OpenCV does not have a direct cv2.crop function, but cropping can be achieved through slicing operations) to crop the image.

[0072] Assume that the coordinates of the upper left and lower right corners of the cropping area are already known, x, y, w, h = 50, 50, 200, 200 # Example coordinates; cropped_image = gray_image[y:y+h, x:x+w].

[0073] III. Extract image features using ResNet-50 Load the ResNet-50 model: Use a deep learning framework to load the pre-trained ResNet-50 model.

[0074] from tensorflow.keras.applications import ResNet50; from tensorflow.keras.preprocessing import image as keras_image; From tensorflow.keras.applications.resnet50 import preprocess_input, decode_predictions; import numpy as np; model = ResNet50(weights='imagenet').

[0075] Preprocess the cropped image: Resize the cropped image to the size required for input to the ResNet-50 model (usually 224x224). Normalize the image to meet the input requirements of the model.

[0076] # Resize the image and normalize: input_shape = (224, 224); resized_image = keras_image.load_img(np.uint8(cropped_image), target_size=input_shape); input_image = keras_image.img_to_array(resized_image); input_image = np.expand_dims(input_image, axis=0); input_image = preprocess_input(input_image).

[0077] Extract features: Use the ResNet-50 model to extract image features. You can choose to extract the output of a certain layer of the model as features, such as the output of the last convolutional layer or the input of the fully connected layer.

[0078] # Extract the output of a certain layer of the ResNet50 model as features (example: extract the output of the last convolutional layer): from tensorflow.keras.models import Model; layer_name = 'conv5_block3_3_conv' # Example layer name; intermediate_layer_model = Model(inputs=model.input, outputs=model.get_layer(layer_name).output); features = intermediate_layer_model.predict(input_image).

[0079] Feature post-processing: Post-process the extracted features, such as dimensionality reduction, normalization, etc., to meet the requirements of subsequent tasks.

[0080] # Example: Perform global average pooling on the features (select the post-processing method according to actual needs): global_average_pooling = np.mean(features, axis=(2, 3)).

[0081] In an embodiment of the present invention, based on step S3, a possible embodiment will be given below to non - restrictively elaborate on its specific implementation.

[0082] S301. Use a pre - trained language model for entity and relation extraction Pre - trained language models, such as BERT, GPT series, etc., have achieved remarkable success in the field of natural language processing. These models have learned rich language knowledge and context understanding capabilities through pre - training on large - scale corpora. Therefore, they can be used to perform various downstream tasks, including entity recognition, relation extraction, etc.

[0083] Entity recognition: Entities are key noun phrases in the text, which usually represent specific people, places, organizations, or concepts. Pre - trained language models can identify these entities by analyzing the context information of the text. This usually involves inputting the text into the model and obtaining the prediction results of the model for each word. Then, these prediction results can be used to determine which words are entities.

[0084] Relation extraction: Relations refer to certain connections or associations existing between entities. Pre - trained language models can also be used to identify these relations. A common method is to use the model to analyze the sentence structure in the text and determine which entities have specific relations. This usually involves further processing of the model output to extract useful relation information.

[0085] S302. Determine the specific format of the query statement A query statement is a common form of information representation, which consists of three parts: subject, predicate, and object. In the task of entity - relation extraction, query statements are usually used to represent the relations between entities.

[0086] Define the query statement format: Before starting to extract the query statement, it is necessary to clarify the specific format of the query statement. This usually involves determining the representation methods of the subject, predicate, and object, as well as their association methods. For example, strings can be used to represent entities and relations, and they can be combined into a structured query statement representation.

[0087] Construct the query statement: Once the format of the query statement is determined, the query statement can be constructed based on the results of entity - relation extraction. This usually involves traversing the identified entities and relations and combining them according to the defined query statement format. When constructing the query statement, it is necessary to ensure that each query statement is unique and correctly represents the information in the text.

[0088] S303. Perform deduplication and filtering processing on the constructed query statements During the process of constructing query statements, duplicate or invalid query statements may be generated. Therefore, it is necessary to remove duplicates and filter these query statements to ensure that the final results are useful and accurate.

[0089] Duplicate removal process: The duplicate removal process aims to eliminate duplicate query statements. This usually involves comparing the constructed query statements and deleting those that are exactly the same as the existing query statements. When comparing query statements, techniques such as string matching and hash functions can be used to ensure accuracy.

[0090] Filtering process: The filtering process aims to delete invalid or non-compliant query statements. This usually involves further analyzing and validating the query statements to ensure that they meet specific conditions or criteria. For example, query statements that contain invalid entities or relationships can be deleted, or those that do not conform to specific domain knowledge or rules can be deleted.

[0091] In one embodiment of the present invention, based on step S4, a possible embodiment will be given below to non-restrictively elaborate on its specific implementation scheme. Please refer to Figure 3 .

[0092] S401. Organize the text features, speech features, and image features into text feature sequences, speech feature sequences, and image feature sequences according to the generation time.

[0093] In multimedia data processing, data of different modalities (such as text, speech, and image) are often generated at different time points. To effectively fuse this information, these features first need to be organized into sequences according to the generation time.

[0094] Text feature sequence: According to the generation or recording time of the text data, the text features after word segmentation and conversion into word embeddings are arranged in chronological order to form a text feature sequence.

[0095] Speech feature sequence: The speech features (such as Mel spectrogram, STFT results, etc.) extracted by LibROSA are arranged in chronological order according to the recording time of the speech data to form a speech feature sequence.

[0096] Image feature sequence: For image data, although it is usually static, in some cases (such as video frames, continuous photography, etc.), images are also generated in chronological order. According to the capture time of the image, the image features (such as facial features extracted by ResNet-50) after OpenCV processing are arranged in chronological order to form an image feature sequence.

[0097] S402. Use a convolutional network to align the lengths of the text feature sequence, speech feature sequence, and image feature sequence.

[0098] Since there may be differences in the generation rates of data in different modalities (for example, text may be input at a faster rate while images may be captured at a slower rate), the lengths of the text feature sequence, speech feature sequence, and image feature sequence may not be consistent. To address this issue, a convolutional network (CNN) can be used to align the lengths of these feature sequences.

[0099] Role of the convolutional network: The CNN can extract local information in the feature sequence through convolutional operations and reduce the dimension of the feature sequence through pooling operations. By adjusting the structure and parameters of the CNN, the length of the output feature sequence can be controlled to make it consistent.

[0100] Alignment method: For each feature sequence, it can be input into an independent CNN, and the output dimension of the CNN can be adjusted so that the lengths of all feature sequences are the same. In this way, the aligned text feature sequence, speech feature sequence, and image feature sequence can be obtained.

[0101] In a specific example, aligning the lengths of the feature sequences includes: Represent the feature sequences of the three modalities as: ; T represents the text modality, A represents the speech modality, V represents the visual modality, I represents the length of the discourse segment sequence of different modalities, and D represents the feature dimension of the corresponding modality.

[0102] Then the initial feature information of the three modalities is: = ; where represents the model parameters for feature extraction of different modalities. Since the feature dimensions of vision and speech are significantly smaller than those of text features, 1D temporal convolution is used to convert the extracted unimodal features into the same dimension.

[0103] , , m ; where, represents the convolutional kernel size of different modalities, represents the length after alignment of the discourse segment feature sequences of different modalities, represents the dimension size corresponding to the features of different modalities.

[0104] S403. Input the aligned feature sequences into a bidirectional recurrent neural network, and use the output parameters of the bidirectional recurrent neural network as the input to the cross-modal attention mechanism.

[0105] The Bidirectional Recurrent Neural Network (Bi-RNN) can consider both forward and backward context information simultaneously, so it is very suitable for processing sequence data. The aligned text feature sequence, speech feature sequence, and image feature sequence are input into the Bi-RNN, and its output parameters are used as the input of the cross-modal attention mechanism. A Bi-RNN is trained for each modality of data. For the text feature sequence, there is a Bi-RNN_text; for the speech feature sequence, there is a Bi-RNN_audio; for the image feature sequence, there is a Bi-RNN_image. Each Bi-RNN generates a hidden state sequence of length T, where each hidden state contains the context information at that time step. For each modality of data, its feature sequence is respectively input into the corresponding Bi-RNN and forward propagation is performed. In this way, the hidden state at each time step can be obtained.

[0106] The role of Bi-RNN: Bi-RNN can capture the context information in sequence data and generate the hidden state at each time step. These hidden states contain the global and local information of the sequence data.

[0107] Cross-modal attention mechanism: The cross-modal attention mechanism can dynamically allocate weights according to the importance of different modality data. After obtaining the hidden state sequence of each modality, a cross-modal attention mechanism is introduced to calculate the weights of different modality data at each time step. This attention mechanism can be a simple fully connected network, which takes the hidden state of the Bi-RNN as input and outputs a weight vector. Each element in the weight vector corresponds to the weight at a time step, indicating the importance of the data at that time step for the final sentiment recognition.

[0108] Assume that the feature vectors of text, speech, and facial expressions are respectively 、 、 , then the cross-modal attention mechanism is expressed as:

[0109] where 、 、 are learnable weight matrices, is a query matrix composed of multiple entities, is the feature dimension, is the number of pre-generated answer groups after the text question queries through the knowledge graph. The query matrix is used to retrieve the pre-generated answer groups related to the text question from the knowledge graph, and these answer groups can be further used as context information to enhance the effect of the cross-modal attention mechanism. The feature dimension refers to the length of the feature vector, which determines the size and computational complexity of the weight matrix.

[0110] In the cross-modal attention mechanism, these entities can be used as additional input features to enhance the model's understanding and modeling ability of the relationships between different modal data. For example, in the sentiment recognition task, if the text data mentions a specific entity (such as a person's name, a place name, etc.), and this entity has related sentiment attributes or events in the knowledge graph, then this entity information can be used to enhance the model's understanding of the text sentiment.

[0111] To integrate this entity information into the cross-modal attention mechanism, the following steps can be taken: Feature extraction: First, each entity needs to be converted into the form of a feature vector. This can be achieved by using embedding techniques (such as Word2Vec, BERT, etc.) to map the entity and relationship to a vector in a high-dimensional space.

[0112] Construct the query matrix: Then, combine the feature vectors of all entities into a query matrix.

[0113] Attention calculation: After obtaining the query matrix, it can be interacted with the hidden state sequences of different modalities to calculate the cross-modal attention weights. This can be achieved by using dot product attention, bilinear attention, or other attention mechanisms.

[0114] Fusion and output: Finally, fuse the calculated attention weights with the hidden states of different modalities to generate the final output representation. This output representation can be further used for sentiment recognition, information retrieval, or other related tasks.

[0115] S404. Integrate into a joint feature vector using the attention mechanism.

[0116] After obtaining the weight matrix, the text feature sequence, speech feature sequence, and image feature sequence can be integrated into a joint feature vector using the attention mechanism.

[0117] Integration method: According to the weight matrix, the features in each feature sequence can be weighted and summed to obtain the weighted feature vector of each modality. Then, these weighted feature vectors are concatenated or summed, etc., to obtain the joint feature vector.

[0118] Specifically, introduce the attention mechanism to dynamically allocate the weights of each modality. The feature vectors of text, speech, and facial expressions are respectively , , , then the joint feature vector can be calculated through the attention mechanism: ; where, is an activation function (such as ReLU), and are learnable parameters.

[0119] Substitute the weight matrix of each feature vector obtained in S403 into the initial value of W, and use deep reinforcement learning to optimize the weights. The specific steps are as follows: The state space refers to the joint feature vector of multi-modal data at the current moment. This vector contains the feature information extracted from different modalities such as text, speech, and facial expressions. After appropriate preprocessing and feature extraction steps, these feature information are combined into a high-dimensional vector to represent the complete state at the current moment. The design of the state space is crucial for the performance of the model because it directly affects whether the model can accurately capture and represent the complex relationships between multi-modal data. The state space includes: the joint feature vector of multi-modal data at the current moment .

[0120] The action space is defined as the weight vectors of each modality, and these weight vectors satisfy the condition that their sum is 1. Each element in the weight vector corresponds to the weight of a modality, indicating the importance of that modality for the emotion recognition task in the current state. By adjusting these weights, the model can dynamically balance the contributions of different modalities, thereby optimizing the accuracy of emotion recognition. The design of the action space allows the model to learn how to adaptively adjust the modality weights according to the current state during training. Specifically, the action space includes: the weight vectors of each modality , satisfying = 1.

[0121] The reward function is used to measure the accuracy of emotion recognition, and usually the negative value of the cross-entropy loss is used to represent it. The cross-entropy loss is a commonly used method to measure the difference between the model's predicted distribution and the true distribution, and its negative value is used as the reward signal to encourage the model to make more accurate predictions. In deep reinforcement learning, the reward function is the key to guiding the model's learning, and it determines what kind of behavior the model should pursue during training. Specifically, the reward function is represented by the negative value of the cross-entropy loss.

[0122] After performing deep reinforcement learning, perform modality fusion on the three-dimensional vector data.

[0123] The update of the policy network adopts the Actor-Critic framework, where the Actor network is responsible for generating actions (weight vectors), and the Critic network is responsible for evaluating the value of the current state.

[0124] Perform feature vector conversion on the data of the three modalities through the above process.

[0125] The deep reinforcement learning process includes: Feature vector transformation: First, extract features from the data of the three modalities (text, speech, facial expression) respectively to obtain their respective feature vectors. These feature vectors are the basis for subsequent processing.

[0126] State representation: Combine the feature vectors of the three modalities to form a joint feature vector at the current moment, that is, the state in the state space.

[0127] Action generation: Use the Actor network (policy network) to generate actions (weight vectors) according to the current state. These weight vectors are used to adjust the contributions of different modalities.

[0128] Modal fusion: According to the generated weight vectors, perform weighted fusion on the feature vectors of the three modalities to obtain the fused feature vectors.

[0129] Emotion recognition: Input the fused feature vectors into the emotion recognition model to obtain the results of emotion recognition.

[0130] Reward calculation: Calculate the reward value according to the accuracy of emotion recognition, that is, the negative value of the cross-entropy loss.

[0131] Policy network update: Use the Actor-Critic framework to update the policy network. The Actor network adjusts its weights according to the reward value to generate better actions; the Critic network evaluates the value of the current state and provides feedback to the Actor network.

[0132] Iterative training: Repeat the above steps until the model achieves satisfactory performance on the validation set.

[0133] S405. Identify the emotion label of the joint feature vector using a pre-trained deep learning model.

[0134] After obtaining the joint feature vector, a pre-trained deep learning model (such as an emotion classification model) can be used to identify its emotion label.

[0135] Selection of pre-trained model: A deep learning model that has been trained on a large number of emotion datasets can be selected, such as a convolutional neural network (CNN), a recurrent neural network (RNN), or a Transformer, etc.

[0136] Recognition of emotion label: Input the joint feature vector into the pre-trained model, and the model will output the probability distribution of the emotion label. The emotion label with the highest probability can be selected as the final recognition result.

[0137] S406. Establish the correspondence between the emotion label and the entity based on the multi-source data collection time corresponding to the emotion label and the multi-source data collection time corresponding to the entity of the query statement.

[0138] Finally, it is necessary to establish the correspondence between the sentiment label and the query statement according to the multi-source data collection time corresponding to the sentiment label and the multi-source data collection time corresponding to the query statement (composed of an entity, a relationship, and another entity).

[0139] Time matching: For each sentiment label, it is necessary to find the entity whose collection time is closest to it. This can be achieved by comparing timestamps or time windows.

[0140] Correspondence establishment: Once the matching entity is found, the sentiment label can be associated with this entity. In this way, entity data with sentiment information can be obtained, providing support for subsequent analysis and applications.

[0141] In an embodiment of the present invention, based on step S5, a possible embodiment will be given below to non-restrictively elaborate on its specific implementation.

[0142] S501. Preset the emotion scores corresponding to the sentiment labels.

[0143] To quantify these sentiment labels, an emotion score can be set for each label. For example, negative is -1, positive is 1, and no attitude is 0.

[0144] S502. Calculate the relevance between the answer and multiple entities and convert the relevance into weights.

[0145] In the multi-modal sentiment recognition model, multiple sentiment-related entities may be extracted from text, speech, and images. To evaluate the relevance of each entity to the given answer, various techniques can be used, such as cosine similarity, Jaccard similarity, or deep learning-based matching algorithms. Once the relevance is calculated, it can be converted into a weight, and the higher the relevance, the greater the weight.

[0146] S503. Set the weight of the user rating.

[0147] In some cases, users will be asked to rate the results of sentiment recognition to obtain their feedback on the model performance. To incorporate these user ratings into the calculation of the comprehensive score, a weight needs to be set for the user ratings. This weight can be determined based on factors such as the credibility of the user, the consistency of historical ratings, or other factors.

[0148] S504. Calculate the comprehensive score of the answer based on the emotion scores, weights of the multiple entities corresponding to the answer, the user rating, and the corresponding weights.

[0149] Now, for each entity, we have the sentiment score, weight, user rating, and its weight. To obtain a comprehensive score, these values can be summed up with weights. Specifically, the sentiment score of each entity can be multiplied by its weight, and then all the weighted scores are added together. Next, the user rating is multiplied by its weight and added to the previous sum. In this way, a comprehensive score is obtained, which reflects the overall performance of the answer in the sentiment recognition task.

[0150] S505. Sort and display multiple answers in descending order based on the comprehensive score.

[0151] Finally, multiple answers can be sorted according to the comprehensive score. The higher the score, the better the performance of the answer in the sentiment recognition task, so it should be preferentially displayed to the user. This sorting method can help users find the answers they are interested in more quickly and improve the overall user experience.

[0152] The above embodiments have the following beneficial effects: Improved emotional perception accuracy: Through the multi-modal emotional perception interface, using computer vision, speech recognition, and natural language processing technologies, emotional features are extracted from facial expressions in videos, the pitch and rhythm of speech, and the semantics of text respectively. The cross-modal fusion emotional recognition front-end can achieve a comprehensive judgment of the user's emotional state, significantly improving the accuracy and robustness of emotional recognition.

[0153] Optimization of emotion-oriented recommended content: Combining the results of user emotion analysis with cultural tags in the knowledge graph, an algorithm for dynamically adjusting the recommended tag set is designed. According to the user's emotional tendency and attitude, the recommended content is dynamically adjusted to avoid displaying content that may cause negative emotions and ensure that the recommended content is positive, thus improving the user experience.

[0154] More considerate personalized recommendation: Using a machine learning model to learn the user's emotional feedback to form a personalized recommendation strategy. The developed emotion-interest dual-modal personalized recommendation engine can predict the user's preference for different cultural tags and quickly adjust the recommended content according to the change of the emotional state, providing more considerate personalized services.

[0155] Improved user experience: Through the dynamic adjustment of user feedback, the system can display results that arouse the user's positive emotions, making the human-computer interaction more natural and smooth, and enhancing the user experience.

[0156] Enhanced intelligent service: The intelligent customer service system can more accurately understand the customer's emotions by analyzing the customer's voice, text, and even facial expressions, providing more user-friendly services and solutions, and enhancing the service quality.

[0157] Enhanced Market Insights: Social media analysis tools integrate sentiment analysis of text, images, and video content to help businesses and market research companies gain insights into consumer sentiment and optimize marketing strategies.

[0158] In another embodiment of the present invention, an interest fusion gating unit can be used to process multi-source feature data, including: 1. Design of the interest fusion gating unit (1) Input layer Video input: Capture the video conversation between the user and the customer service, and extract visual features such as facial expressions and body movements.

[0159] Voice input: Extract the voice features of the user, including pitch, speech rate, volume, etc.

[0160] Text input: Use natural language processing technology to parse the questions and entities mentioned by the user, and extract keywords and semantic information.

[0161] (2) Interest fusion gating mechanism Design a gating unit that can dynamically adjust the weights of different modalities (visual, voice, text) according to the user's interests (such as mentioned entities, question types, etc.).

[0162] Use the attention mechanism to capture the user's attention degree to specific information in different modalities, so as to achieve effective fusion of the user's interests.

[0163] (3) Emotion scoring module Apply an emotion recognition model to the fused multi-modal features, which can identify the user's emotion (positive, negative, no attitude).

[0164] According to the preset load parameters (such as the weights of different modalities, the thresholds of emotion recognition, etc.), perform weighted calculation on the emotion score to obtain the final emotion score.

[0165] 2. Emotion label scoring Definition of emotion labels: Negative (-1), Positive (1), No attitude (0).

[0166] Calculation of emotion score: According to the probability distribution output by the emotion recognition model, combined with the definition of emotion labels, calculate the user's emotion score.

[0167] .

[0168] Model improvement: Continuously iterate and optimize the model parameters to improve the accuracy and stability of emotion recognition.

[0169] .

[0170] 3. User Feedback and Answer Ranking (1)User Feedback Collection After the conversation ends, show the answer generated by the system to the user and invite the user to provide feedback on the accuracy and satisfaction of the answer.

[0171] Collect user feedback data for subsequent training and optimization of the model.

[0172] (2)Answer Ranking Coefficient Model Evaluation Coefficient: Use the likelihood function that maximizes the data as the answer evaluation coefficient, which is equivalent to minimizing the loss function. This can be achieved by training a ranking model that can rank answers based on user feedback and answer quality.

[0173] 。

[0174] Emotional Index Composite Operation: Perform a composite operation on the emotional score (Score) and the result (tanL) after processing by the loss function to comprehensively consider the accuracy of the answer and the user's emotional state.

[0175] Process the loss function tanL, perform a composite operation with the emotional index Score: 。

[0176] (3)Answer Ranking Rank the answers generated by the system according to the results of the composite operation. Show the ranked answers to the user, and give priority to showing the answers with higher scores.

[0177] 4. Implementation Steps Data Preparation: Collect and preprocess video conversation data, including video frame extraction, voice signal conversion, text content parsing, etc.

[0178] Model Training: Train the interest fusion gated unit, emotion recognition model and ranking model, and optimize them using historical conversation data and user feedback data.

[0179] System Deployment: Deploy the trained model to the video customer service system to achieve real-time emotional scoring and answer ranking functions.

[0180] Monitoring and Optimization: Continuously monitor the system performance, collect user feedback data, and iteratively optimize the model.

[0181] In some embodiments, the interaction optimization system based on sentiment analysis may include multiple functional modules composed of computer program segments. The computer programs of each program segment in the interaction optimization system based on sentiment analysis may be stored in the memory of a computer device and executed by at least one processor to perform (see Figure 1 description) the functions of interaction optimization based on sentiment analysis.

[0182] In this embodiment, according to the functions it performs, the interaction optimization system based on sentiment analysis may be divided into multiple functional modules, as Figure 4 shown. The functional modules of system 400 may include: a data acquisition module 410, a feature extraction module 420, a content extraction module 430, a sentiment analysis module 440, and an answer ranking module 450. The module referred to in the present invention means a series of computer program segments that can be executed by at least one processor and can complete fixed functions, and are stored in the memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments.

[0183] The data acquisition module is used to collect input data of an interaction object from multiple channels during the interaction process to obtain multi-source data; The feature extraction module is used to extract multi-source feature data from the multi-source data; The content extraction module is used to extract entities and relationships from the multi-source feature data and generate query statements based on the entities and relationships; The sentiment analysis module is used to identify sentiment labels corresponding to the entities in the query statement from the multi-source feature data; The answer ranking module is used to score multiple answers generated based on the query statement and a pre-constructed knowledge graph based on the query statement and the corresponding sentiment labels, and rank and display the multiple answers from high to low according to the scores.

[0184] Optionally, as an embodiment of the present invention, collecting input data of an interaction object from multiple channels during the interaction process to obtain multi-source data includes: Obtaining the input data of the interaction object, where the input data includes text data, voice data, and image data.

[0185] Optionally, as an embodiment of the present invention, extracting multi-source feature data from the multi-source data includes: Using a word segmentation tool to perform word segmentation on the text data, and using the Glove model to convert the segmented text data into text features; Using the LibROSA speech package to extract frame-level speech features from the voice data; Crop the image data using OpenCV, and extract facial features from the cropped image data using the ResNet-50 network model.

[0186] Optionally, as an embodiment of the present invention, extracting entities and relationships from the multi-source feature data, and generating a query statement based on the entities and relationships, includes: Extracting entities and relationships from the text features using a pre-trained language model; Determine the specific format of the query statement, and according to the results of entity relationship extraction, combine the identified entities and relationships according to the defined query statement format to construct a query statement; Perform deduplication and filtering processing on the constructed query statement to remove duplicate and invalid query statements.

[0187] Optionally, as an embodiment of the present invention, identifying an emotion label corresponding to the entity in the query statement from the multi-source feature data, includes: Organize the text features, speech features, and image features according to the generation time into a text feature sequence, a speech feature sequence, and an image feature sequence; Use a convolutional network to align the lengths of the text feature sequence, the speech feature sequence, and the image feature sequence; Input the aligned text feature sequence, speech feature sequence, and image feature sequence into a bidirectional recurrent neural network, and use the output parameters of the bidirectional recurrent neural network as the input of a cross-modal attention mechanism to obtain the weight matrices of the text feature sequence, the speech feature sequence, and the image feature sequence; Use an attention mechanism to integrate the text feature sequence, the speech feature sequence, and the image feature sequence into a joint feature vector based on the weight matrices of the text feature sequence, the speech feature sequence, and the image feature sequence; Use a pre-trained deep learning model to identify the emotion label of the joint feature vector; Based on the multi-source data collection time corresponding to the emotion label, and the multi-source data collection time corresponding to multiple entities in the query statement, establish the corresponding relationship between the emotion label and the multiple entities.

[0188] Optionally, as an embodiment of the present invention, further includes: Use deep reinforcement learning technology to train the attention mechanism.

[0189] Optionally, as an embodiment of the present invention, score multiple answers generated based on the query statement and a pre-constructed knowledge graph based on the query statement and the corresponding emotion label, and sort and display the multiple answers from high to low according to the scores, includes: Pre-set the emotion score corresponding to the emotion label; Calculate the relevance between the answer and multiple entities, and convert the relevance into weights; Set the weights of user ratings; Calculate the comprehensive score of the answer based on the sentiment scores, weights of the multiple entities corresponding to the answer, user ratings, and their corresponding weights; Sort and display multiple answers in descending order based on the comprehensive scores.

[0190] Figure 5 The interactive optimization method based on sentiment analysis provided by the embodiments of the present application can be applied to devices. Those skilled in the art can understand that the device structure involved in the embodiments of the present invention does not constitute a limitation on the device. The device may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements. In the embodiments of the present invention, the device includes, but is not limited to, laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments of the present application described herein and / or claimed.

[0191] Among them, the device 500 may include: a processor 510, a memory 520, and a communication unit 530. These components communicate through one or more buses. Those skilled in the art can understand that the structure of the server shown in the figure does not constitute a limitation on the present invention. It can be a bus structure, a star structure, or may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.

[0192] Among them, the memory 520 can be used to store the execution instructions of the processor 510. The memory 520 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disks, or optical disks. When the execution instructions in the memory 520 are executed by the processor 510, the device 500 can execute some or all of the steps in the above method embodiments.

[0193] The processor 510 is the control center of the storage device, connecting various parts of the entire electronic device through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 520, and by invoking data stored in the memory, it performs various functions of the electronic device and / or processes data. The processor may be composed of an integrated circuit (IC), for example, it may be composed of a single packaged IC, or it may be composed of multiple packaged ICs with the same or different functions connected together. For example, the processor 510 may include only a central processing unit (CPU). In the embodiments of the present invention, the CPU may be a single-core processor or may include multiple cores.

[0194] The communication unit 530 is used to establish a communication channel so that the storage device can communicate with other devices. It receives user data sent by other devices or sends user data to other devices.

[0195] The present invention also provides a computer storage medium. The computer storage medium can store a program which, when executed, may include some or all of the steps in the various embodiments provided by the present invention. The storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), or the like.

[0196] Those skilled in the art can clearly understand that the technology in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc., which can store program codes, and includes several instructions to enable a computer device (which may be a personal computer, a server, or a second device, a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.

[0197] For the same or similar parts among the various embodiments in this specification, reference can be made to each other. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the descriptions in the method embodiments.

[0198] In several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the system or module can be in electrical, mechanical or other forms.

[0199] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they can be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0200] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module.

[0201] Although the present invention has been described in detail by referring to the accompanying drawings and in combination with the preferred embodiments, the present invention is not limited thereto. Without departing from the spirit and essence of the present invention, those of ordinary skill in the art can make various equivalent modifications or substitutions to the embodiments of the present invention, and these modifications or substitutions should all be within the scope of the present invention. / Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention.

Claims

1. An interactive optimization method based on sentiment analysis, characterized in that: include: During the interaction process, input data of the interaction object is collected from multiple channels to obtain multi-source data; Extracting multi-source feature data from the multi-source data; Extracting entities and relationships from the multi-source feature data, and generating query statements based on the entities and relationships; Identifying sentiment tags corresponding to entities in the query sentence from the multi-source feature data; Based on the query statement and the corresponding sentiment tag, multiple answers generated based on the query statement and the pre-built knowledge graph are scored, and the multiple answers are sorted and displayed from high to low according to the scores.

2. The method according to claim 1, characterized in that During the interaction process, input data of the interaction object is collected from multiple channels to obtain multi-source data, including: The input data of the interactive object is obtained, where the input data includes text data, voice data and image data.

3. The method according to claim 2, characterized in that Extracting multi-source feature data from the multi-source data includes: Using a word segmentation tool to perform word segmentation on the text data, and using a GloVe model to convert the word segmented text data into text features; Extracting frame-level speech features from the speech data using the LibROSA speech package; OpenCV is used to crop the image data, and the ResNet-50 network model is used to extract facial features from the cropped image data.

4. The method according to claim 3, characterized in that Extracting entities and relationships from the multi-source feature data, and generating query statements based on the entities and relationships, including: Use pre-trained language models to extract entities and relations from text features; Determine the specific format of the query statement, and based on the results of entity relationship extraction, combine the identified entities and relationships according to the defined query statement format to construct the query statement; The constructed query statements are deduplicated and filtered to remove duplicate and invalid query statements.

5. The method according to claim 4, characterized in that Identifying a sentiment tag corresponding to an entity in the query sentence from the multi-source feature data includes: Arrange the text features, voice features and image features into a text feature sequence, a voice feature sequence and an image feature sequence according to the generation time; Use convolutional networks to align the lengths of text feature sequences, speech feature sequences, and image feature sequences; Input the aligned text feature sequence, speech feature sequence and image feature sequence into a bidirectional recurrent neural network, and use the output parameters of the bidirectional recurrent neural network as the input of the cross-modal attention mechanism to obtain the weight matrices of the text feature sequence, speech feature sequence and image feature sequence; By using the attention mechanism, based on the weight matrices of the text feature sequence, the speech feature sequence and the image feature sequence, the text feature sequence, the speech feature sequence and the image feature sequence are integrated into a joint feature vector; Identifying the sentiment label of the joint feature vector using a pre-trained deep learning model; Based on the multi-source data collection time corresponding to the sentiment tag and the multi-source data collection time corresponding to the multiple entities in the query statement, a corresponding relationship between the sentiment tag and the multiple entities is established.

6. The method according to claim 5, characterized in that The method further comprises: The attention mechanism is trained using deep reinforcement learning technology.

7. The method according to claim 5, characterized in that Based on the entities of the query and the corresponding sentiment tags, multiple answers generated based on the query and the pre-built knowledge graph are scored, and the multiple answers are sorted and displayed from high to low according to the scores, including: Pre-set the emotion scores corresponding to the emotion labels; Calculate the relevance of the answer to multiple entities and convert the relevance into weights; Set the weight of user ratings; Calculate the comprehensive score of the answer based on the sentiment scores, weights, user ratings and corresponding weights of multiple entities corresponding to the answer; Multiple answers are sorted and displayed from high to low based on the comprehensive scores.

8. An interactive optimization system based on sentiment analysis, characterized in that: include: A data acquisition module is used to collect input data of the interaction object from multiple channels during the interaction process to obtain multi-source data; A feature extraction module, used to extract multi-source feature data from the multi-source data; A content extraction module, used to extract entities and relationships from the multi-source feature data, and generate query statements based on the entities and relationships; A sentiment analysis module, configured to identify sentiment tags corresponding to entities in the query sentence from the multi-source feature data; The answer sorting module is used to score multiple answers generated based on the query statement and the pre-built knowledge graph based on the query statement and the corresponding sentiment tag, and sort and display the multiple answers from high to low according to the score.

9. A device, characterized in that: include: A memory, used for storing an interactive optimization program based on sentiment analysis; A processor, used to implement the steps of the interactive optimization method based on sentiment analysis as described in any one of claims 1 to 7 when executing the interactive optimization program based on sentiment analysis.

10. A computer-readable storage medium storing a computer program, characterized in that: The readable storage medium stores an interaction optimization program based on sentiment analysis, and when the interaction optimization program based on sentiment analysis is executed by a processor, the steps of the interaction optimization method based on sentiment analysis as described in any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Elderly emotion bidirectional interaction method based on artificial intelligence technology

    CN120632431A

  • An old person emotion two-way interaction method based on artificial intelligence technology

    CN120632431B