Response method, device and equipment for navigation dialogue system and storage medium
By integrating the multimodal information of the guide dialogue system features and emotional recognition, combining user portraits and context information, personalized and emotional reply information is generated, the shortcomings of the existing guide system in personalized experience and interactivity are solved, and deeper user interaction and emotional communication are achieved.
Patent Information
- Application Number
- CN202510097597.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-23
AI Technical Summary
The existing guided system has shortcomings in personalized experience and interactivity. AR-based systems are difficult to achieve in-depth interaction, and rules-based intelligent Q&A systems are difficult to deal with complex problems and ignore emotional interactions.
By fusing the multimodal information (text, voice, image) of the guide dialogue system features, a multimodal representation is generated, and emotional recognition is performed based on this representation. Combining user portraits and context information, we generate recommendation information, retrieve target knowledge information from the knowledge base, and output personalized and emotional reply information.
It realizes a more comprehensive and rich guided interactive experience, enhances personalized emotional response and emotional communication with users, and improves user personalized experience and interactive satisfaction.
Smart Images

Figure CN120030122A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a reply method, device, equipment and storage medium for a navigation dialogue system. Background Art
[0002] In guided tour scenarios, such as museum guided tours, exhibition guided tours, etc., there are currently many deficiencies in experience and services, especially in terms of personalized experience and interactivity.
[0003] Specifically, the current tour guide systems mainly include tour guide systems based on augmented reality (AR) or intelligent question-and-answer systems based on rules. However, the AR-based tour guide system can only enrich the user's visual experience to a certain extent, and it has not really achieved deep interaction with the user. During the visit, the user mainly watches and reads materials, and it is difficult to obtain personalized tour guide services like real tour guides based on personal interests, backgrounds and needs, resulting in insufficient personalized experience. The rule-based intelligent question-and-answer system can only answer some preset common questions, and it is often limited to preset questions and answers, and it is difficult to deal with complex or open questions. In addition, it ignores the emotional interaction with the user, resulting in a tour guide experience that lacks emotional communication and humanistic care. Summary of the invention
[0004] In view of this, the present application provides a reply method, device, equipment and storage medium of a navigation dialogue system to solve the deficiencies in the related art.
[0005] In a first aspect of the present application, a reply method of a navigation dialogue system is provided, the method comprising:
[0006] Performing feature fusion on multimodal information input into the guide dialogue system to obtain a multimodal representation, and obtaining an emotion recognition result according to the multimodal representation, wherein the multimodal information includes text, voice and image information in the guide scenario in which the guide dialogue system is applied;
[0007] Convert the multimodal information into a natural language description, and obtain a user portrait according to the natural language description;
[0008] Generate recommendation information according to the user portrait and context information of the navigation dialogue system, and determine target knowledge information matching the recommendation information from a knowledge base preset under the navigation scenario;
[0009] Output reply information according to the target knowledge information, the natural language description, the emotion recognition result and the user portrait.
[0010] According to an embodiment of the present application, the step of performing feature fusion on the multimodal information input into the guide dialogue system to obtain a multimodal representation includes:
[0011] Performing feature extraction on the text, voice and image information respectively;
[0012] The extracted features are aligned and fused using an attention mechanism to obtain the multimodal representation.
[0013] According to an embodiment of the present application, obtaining an emotion recognition result according to the multimodal representation includes:
[0014] Inputting the multimodal representation into a pre-trained emotion recognition model so that the emotion recognition model outputs the emotion recognition result;
[0015] Wherein, the emotion recognition model is used for:
[0016] Analyzing the text features, speech features, and image features included in the multimodal representation respectively to obtain the corresponding user emotional state;
[0017] The emotion recognition result is output according to the obtained user emotion state and the preset weight.
[0018] According to an embodiment of the present application, obtaining a user portrait according to the natural language description includes:
[0019] Performing semantic analysis on the natural language description;
[0020] Extracting features from the result of semantic analysis to obtain target features, wherein the target features at least include user interest information and user background information;
[0021] The user portrait is constructed according to the target features.
[0022] According to one embodiment of the present application, the knowledge base includes a knowledge graph, a vector database and a structured database, wherein the knowledge graph is used to store the relationship information between entities in the navigation scenario, the vector database is used to store the information of the entities in the navigation scenario, and the structured database is used to store the information of the navigation scene itself.
[0023] According to an embodiment of the present application, determining the target knowledge information matching the recommended information from the knowledge base preset in the navigation scenario includes:
[0024] Inputting the recommendation information into a pre-trained retrieval model so that the retrieval model outputs the target knowledge information;
[0025] Wherein, the retrieval model is used for:
[0026] respectively determining the similarity between each piece of knowledge information in the knowledge base and the recommended information;
[0027] The top N pieces of knowledge information ranked by similarity are used as the target knowledge information, and the target knowledge information is output.
[0028] According to an embodiment of the present application, outputting reply information according to the target knowledge information, the natural language description, the emotion recognition result, and the user portrait includes:
[0029] Obtaining preliminary response information according to the target knowledge information and the natural language description;
[0030] According to the preliminary reply information, the emotion recognition result and the user portrait, final reply information is output, and the final reply information includes text, voice and image information.
[0031] According to an embodiment of the present application, outputting final reply information according to the preliminary reply information, the emotion recognition result and the user portrait includes:
[0032] The preliminary reply information is processed according to the user's emotional state, the user portrait and the emotional interaction strategy to generate and output the final reply information, wherein the emotional interaction strategy is generated by a pre-trained emotion enhancement model.
[0033] According to one embodiment of the present application, the method further includes:
[0034] The user's feedback information input into the navigation dialogue system in response to the final reply information is input into the emotion enhancement model, so that the emotion enhancement model updates the emotion interaction strategy according to the feedback information.
[0035] In a second aspect of the present application, a reply device of a guide dialogue system is provided, the device comprising:
[0036] An emotion recognition unit, used for performing feature fusion on multimodal information input into the guide dialogue system to obtain a multimodal representation, and obtaining an emotion recognition result according to the multimodal representation, wherein the multimodal information includes text, voice and image information in the guide scenario applied by the guide dialogue system;
[0037] A determination unit, configured to convert the multimodal information into a natural language description and obtain a user portrait according to the natural language description;
[0038] A recommendation unit, configured to generate recommendation information according to the user portrait and context information of the navigation dialogue system, and determine target knowledge information matching the recommendation information from a knowledge base preset under the navigation scenario;
[0039] A reply unit is used to output reply information based on the target knowledge information, the natural language description, the emotion recognition result and the user portrait.
[0040] In a third aspect of the present application, an electronic device is provided, including a processor and a memory, wherein the memory stores machine executable instructions that can be executed by the processor, and the processor is used to execute the machine executable instructions to implement the steps of the method proposed in the above embodiment.
[0041] In a fourth aspect of the present application, a machine-readable storage medium is provided, wherein the machine-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are executed by a processor, the steps of the method proposed in the above embodiment are implemented.
[0042] It can be seen from the above technical solutions that the input of the guided tour dialogue system is multimodal information, including text, voice and image information in the guided tour scenario used by the guided tour dialogue system. By combining multiple information sources, a more comprehensive and rich guided tour interaction experience can be provided.
[0043] In addition, emotion recognition results are obtained based on multimodal information, and the response information is emotion optimized using the emotion recognition results and user portraits, so that the navigation dialogue system can form personalized emotional responses and enhance emotional communication with users.
[0044] In addition, the multimodal information is converted into a natural language description, and a user profile is obtained based on the natural language description. Recommended information is generated based on the user profile and the context information of the navigation dialogue system, and target knowledge information matching the recommended information is determined from the knowledge base. By combining the recommendation mechanism of user profile and context information, the reply information output by the navigation dialogue system can meet the interests and needs of different users, thereby enhancing the personalized experience of users.
[0045] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 It is a flowchart of a reply method of a navigation dialogue system provided in an embodiment of the present application;
[0047] Figure 2 It is a flowchart of a preliminary response model training process provided by an embodiment of the present application;
[0048] Figure 3 It is a flowchart of a reply method of a navigation dialogue system provided in an embodiment of the present application;
[0049] Figure 4 It is a structural schematic diagram of a reply device of a navigation dialogue system provided in an embodiment of the present application;
[0050] Figure 5 It is a schematic diagram of the hardware structure of an electronic device shown in an exemplary embodiment of the present application. DETAILED DESCRIPTION
[0051] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0052] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms "a", "said" and "the" used in this application and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings.
[0053] In order to enable those skilled in the art to better understand the technical solutions provided by the embodiments of the present application and to make the above-mentioned purposes, features and advantages of the embodiments of the present application more obvious and understandable, the technical solutions in the embodiments of the present application are further described in detail below in conjunction with the accompanying drawings.
[0054] In guided tour scenarios, such as museum guided tours, exhibition guided tours, etc., there are currently many deficiencies in experience and services, especially in terms of personalized experience and interactivity.
[0055] Specifically, the current tour guide systems mainly include tour guide systems based on augmented reality (AR) or intelligent question-and-answer systems based on rules. However, the AR-based tour guide system can only enrich the user's visual experience to a certain extent, and it has not really achieved deep interaction with the user. During the visit, the user mainly watches and reads materials, and it is difficult to obtain personalized tour guide services like real tour guides based on personal interests, backgrounds and needs, resulting in insufficient personalized experience. The rule-based intelligent question-and-answer system can only answer some preset common questions, and it is often limited to preset questions and answers, and it is difficult to deal with complex or open questions. In addition, it ignores the emotional interaction with the user, resulting in a tour guide experience that lacks emotional communication and humanistic care.
[0056] In view of this, an embodiment of the present application discloses a reply method of a navigation dialogue system to address the deficiencies in the related art.
[0057] like Figure 1 As shown, Figure 1 It is a flowchart of a reply method of a navigation dialogue system provided in an embodiment of the present application.
[0058] The reply method of the navigation dialogue system may include the following steps:
[0059] S101: Perform feature fusion on multimodal information input into the guide dialogue system to obtain a multimodal representation, and obtain an emotion recognition result based on the multimodal representation, wherein the multimodal information includes text, voice and image information in a guide scenario applied by the guide dialogue system.
[0060] In an embodiment of the present application, the navigation dialogue system can be installed on an electronic device to provide a convenient interactive experience for the user. The electronic device can include a smart phone, a tablet computer, a dedicated terminal, and the like.
[0061] The navigation dialogue system has an interactive interface, and the user can input multimodal information in the interactive interface of the navigation dialogue system. The multimodal information here includes text information, voice information and image information in the navigation scenario applied by the navigation dialogue system.
[0062] The embodiments of the present application do not specifically limit the tour-guiding scenarios to which the tour-guiding dialogue system is applied. For example, the tour-guiding scenarios to which the tour-guiding dialogue system is applied may include museum tour-guiding scenarios, exhibition tour-guiding scenarios, and historical site tour-guiding scenarios.
[0063] In some embodiments, the image information in the guided tour scene includes image information of entities in the guided tour scene, such as an image of a four-legged tripod in a museum guided tour scene, and also includes personal image information, such as the user's own body posture, facial expression, etc.
[0064] In some embodiments, the user can trigger a preset control on the interactive interface of the navigation dialogue system for indicating input of text information, and call the text box function to type in text information related to the navigation scene.
[0065] The user can trigger a preset control in the interactive interface of the navigation dialogue system to indicate the input of voice information, and call the audio acquisition module (such as a microphone, sound sensor, etc.) of the electronic device to collect and input voice information in the navigation scenario.
[0066] The user can trigger a preset control in the interactive interface of the navigation dialogue system to indicate the input of image information, and call the image acquisition module (such as a camera, image sensor, etc.) of the electronic device to collect and input image information in the navigation scene.
[0067] In an embodiment of the present application, for multimodal information input into the navigation dialogue system, feature fusion of the multimodal information may be performed to obtain a multimodal representation, and an emotion recognition result may be obtained based on the multimodal representation.
[0068] Multimodal representation refers to converting information from different modalities (text, speech, and images) into a unified form. Each modality provides different perspectives and information about the same data or scene, and the goal of multimodal representation is to fuse the information from these different perspectives together to form a more comprehensive and rich data representation.
[0069] In some embodiments, the multimodal information input into the navigation dialogue system is subjected to feature fusion to obtain a multimodal representation, including:
[0070] S1011: Extract features of text, voice and image information respectively.
[0071] Specifically, the information of each modality can be used to extract feature vectors through specific feature extraction methods.
[0072] For text information, feature vectors can be extracted from text information through natural language processing technologies such as word embedding and bidirectional encoder representation BERT (Bidirectional Encoder Representations from Transformers).
[0073] For speech information, feature vectors can be extracted from speech information through speech recognition and processing technologies, such as Mel-Frequency Cepstral Coefficients (MFCCs) and spectrum features.
[0074] For image information, feature vectors can be extracted from image information through computer vision techniques, such as convolutional neural networks (CNNs) and visual geometry group (VGG).
[0075] S1012: Use the attention mechanism to align and fuse the extracted features to obtain a multimodal representation.
[0076] Specifically, feature vectors of different modalities (text, speech, and image) may have different dimensions and semantics. First, the attention mechanism is used to align feature vectors of different modalities (text, speech, and image) to the same space to ensure that they are comparable when fused. The alignment process can include time alignment, space alignment, and semantic alignment.
[0077] Then, the aligned feature vectors of different modalities (text, speech, and image) are fused into a unified multimodal representation. The fusion stages of feature vectors can include: early fusion, i.e., fusion at the feature level; late fusion, i.e., fusion at the decision level; or intermediate fusion, i.e., fusion at the intermediate layer of the model. The fusion methods of feature vectors can include simple vector concatenation, weighted summation, or more complex deep learning models.
[0078] In an embodiment of the present application, the input of the navigation dialogue system is multimodal information, including text information, voice information and image information in the navigation scenario used by the navigation dialogue system. By combining multiple information sources, a more comprehensive and rich navigation interaction experience can be provided.
[0079] In addition, the embodiments of the present application do not specifically limit the method of "obtaining emotion recognition results based on multimodal representation". For example, the fused multimodal representation can be input into a pre-trained deep learning-based model so that the model outputs emotion recognition results; or, the fused multimodal representation can be input into a classifier such as a support vector machine (SVM) or a multilayer perceptron (MLP) so that the classifier classifies the fused multimodal representation and finally outputs emotion recognition results.
[0080] In an embodiment of the present application, emotion recognition results are obtained based on multimodal information, so that in subsequent steps, the emotion recognition results and user portraits can be used to optimize the emotion of the reply information, so that the navigation dialogue system can form a personalized emotional response, improve the emotional interaction capabilities, and enhance the emotional communication with users.
[0081] S102: Convert the multimodal information into a natural language description, and obtain a user portrait according to the natural language description.
[0082] Natural language description refers to the process of using natural language (such as English, Chinese, etc.) to describe or express information, events, objects, concepts or data. Converting the multimodal information input into the navigation dialogue system into natural language description is actually a cross-modal conversion, that is, integrating information from different modalities into a unified text modality description.
[0083] In some embodiments, for the text information in the navigation scene in the multimodal information, since it is in the form of natural language description, no additional conversion processing is required, and the text information can be directly used for integration;
[0084] For the non-text information in the navigation scenario in the multimodal information, including voice and image information, they can be converted into text information through specific conversion methods. For example, for voice information, it can be converted into text information through voice recognition technology; for image information, it can be input into a deep learning-based model (such as a generative adversarial network, a sequence-to-sequence model, etc.) so that the model can analyze and understand the image information and then output a text description of the image content, that is, text information;
[0085] The original text information in the multimodal information is integrated with the converted text information to generate a natural language description.
[0086] After obtaining the natural language description, the user portrait and context information of the navigation dialogue system will be obtained based on the natural language description.
[0087] In some embodiments, obtaining a user profile according to a natural language description includes:
[0088] S1021: Performing semantic analysis on the natural language description;
[0089] Semantic analysis aims to extract semantic information from natural language descriptions so that computers can understand the meaning of the text. Semantic analysis can include lexical analysis, syntactic analysis, pragmatic analysis, and contextual analysis.
[0090] S1022: Extracting features from the result of semantic analysis to obtain target features, where the target features at least include user interest information and user background information;
[0091] After completing the semantic analysis, it is necessary to extract target features from the results of the semantic analysis. The target features include at least the user's interest information, such as specific topics or fields (such as music, sports, technology, etc.), preferences for certain activities (such as reading, travel, fitness, etc.), and the user's background information, such as personal information (such as age, gender, occupation, educational background, etc.), geographic information (such as residence, frequently visited places, etc.). Of course, the target features may also include the user's tour information in the guided scene (such as tour time, tour mode, tour path, etc.), etc., and the embodiments of the present application are not specifically limited to this.
[0092] In addition, the embodiments of the present application do not specifically limit the method of feature extraction. For example, statistical methods, machine learning, deep learning and other methods can be used to perform feature extraction.
[0093] S1023: Construct the user portrait according to the target features.
[0094] Based on the extracted target features, a user profile is constructed. A user profile is a model that summarizes user attributes, interests, and behaviors. Its construction steps may include:
[0095] Integrate the extracted target features to generate a user feature set;
[0096] Quantify each feature in the user feature set, such as using a score, frequency, or weight to indicate the strength or importance of each feature;
[0097] Build a user profile based on the quantitative features. The user profile can be a multi-dimensional data structure, such as a vector, graph, or database.
[0098] In some embodiments, the user profile will be updated based on new multimodal information input into the guided dialogue system to maintain its accuracy and timeliness.
[0099] In some embodiments, the constructed user profile may include a long-term user profile and a short-term user profile.
[0100] Among them, long-term user portraits are constructed based on users’ long-term characteristics, preferences, and behaviors. They are highly stable and can reflect users’ long-term habits and inherent attributes. Long-term user portraits can include age, gender, educational background, long-term hobbies, etc.
[0101] Short-term user profiles are built based on the user's recent behavior and short-term interaction data, focusing on the user's current intentions and temporary preferences. Short-term user profiles can include current emotional state, recent search history, and temporary interests.
[0102] In some embodiments, a knowledge graph may be used to store long-term user portraits, and a neural Turing machine or a memory network may be used to store short-term user portraits.
[0103] In the embodiment of the present application, the multimodal information is converted into a natural language description, and a user portrait is obtained based on the natural language description, so that in the subsequent steps, recommendation information can be generated based on the user portrait and the context information of the navigation dialogue system, and then the target knowledge information matching the recommendation information is determined from the knowledge base. By combining the recommendation mechanism of the user portrait and the context information, the reply information output by the navigation dialogue system can meet the interests and needs of different users, thereby enhancing the personalized experience of users.
[0104] S103: Generate recommendation information according to the user portrait and the context information of the navigation dialogue system, and determine target knowledge information matching the recommendation information from a knowledge base preset in the navigation scenario.
[0105] In some embodiments, context information may be extracted from the conversation content of the navigation dialogue system to obtain context information of the navigation dialogue system.
[0106] The conversation content of the navigation dialogue system includes the currently input multimodal information, i.e., "the multimodal information input into the navigation dialogue system" in S101, and the historical conversation content of the navigation dialogue system, including the historically input multimodal information and the historically output reply information.
[0107] In some embodiments, the historical conversation content of the navigation dialogue system includes all historical conversation content contained in the current conversation between the user and the navigation dialogue system, or the most recent N rounds of historical conversation content of the current conversation between the user and the navigation dialogue system, which can be set according to actual needs.
[0108] When the historical conversation content of the navigation dialogue system includes all historical conversation content contained in the current conversation between the user and the navigation dialogue system, the obtained context information of the navigation dialogue system can be called global context information;
[0109] When the historical conversation content of the navigation dialogue system includes the most recent N rounds of historical conversation content of the current conversation between the user and the navigation dialogue system, the obtained context information of the navigation dialogue system may be referred to as local context information.
[0110] In some embodiments, extracting context information from the conversation content of the navigation dialogue system to obtain the context information of the navigation dialogue system includes:
[0111] S1031: Converting the dialogue content of the tour guide dialogue system into a natural language description;
[0112] Since the dialogue content of the navigation dialogue system includes the currently input multimodal information (text, voice and image information), the historically input modal information and the historically output reply information, it can be converted first and uniformly converted into natural language description.
[0113] S1032: extracting keywords from the converted natural language description;
[0114] These keywords can include nouns, verbs, adjectives, etc., which can represent the theme and focus of the conversation.
[0115] S1033: performing semantic association analysis on the converted natural language description to obtain semantic association information between different parts of the natural language description;
[0116] Semantic association information can include causal relationships, transitional relationships, and progressive relationships.
[0117] S1034: Integrate the extracted keywords and semantic association information to obtain context information of the navigation dialogue system.
[0118] In some embodiments, generating recommendation information according to the user portrait and context information of the navigation dialogue system includes:
[0119] Recommendation information is determined from the content library based on the user portrait, contextual information of the navigation dialogue system, and a recommendation strategy, wherein the recommendation strategy is generated by a pre-trained recommendation model.
[0120] That is, based on the recommendation strategy generated by the recommendation model, recommended content that matches the user portrait and context information is filtered out from the pre-set content library.
[0121] In some embodiments, the user's feedback information input into the navigation dialogue system for the final reply information can be input into the recommendation model, so that the recommendation model updates the recommendation strategy according to the feedback information. The personalized recommendation strategy is optimized through the feedback information input by the user, so that the reply information output by the navigation dialogue system can meet the interests and needs of different users, thereby improving the personalized experience of users.
[0122] In an embodiment of the present application, target knowledge information matching the recommended information is determined from a knowledge base in a preset navigation scenario.
[0123] A plurality of pieces of knowledge information related to the navigation scene are stored in a preset knowledge base, and target knowledge information matching the recommended information is determined from the plurality of pieces of knowledge information related to the navigation scene.
[0124] In some embodiments, the knowledge information related to the navigation scene may include information about the navigation scene itself and information related to entities in the navigation scene. The information related to entities in the navigation scene may include relationship information between entities in the navigation scene and information about entities in the navigation scene.
[0125] Assume that the tour scene is a museum tour scene. Information about the tour scene itself, such as the layout of the museum's exhibition hall, opening hours, and service facilities. Relationship information between entities in the tour scene, such as the relationship between exhibits and creators, and the relationship between artists and schools. Information about entities in the tour scene, such as the name of the exhibits, historical background, and materials.
[0126] In addition, the embodiments of the present application do not specifically limit the method of "determining target knowledge information matching the recommended information from the knowledge base in the preset navigation scenario". For example, the target knowledge information matching the recommended information can be determined from the knowledge base by similarity calculation, rule engine or keyword matching.
[0127] In the embodiment of the present application, after obtaining the user portrait in the manner described in S102, recommendation information can be generated based on the user portrait and the context information of the navigation dialogue system, and target knowledge information matching the recommendation information can be determined from the knowledge base. By combining the recommendation mechanism of the user portrait and the context information, the reply information output by the navigation dialogue system can meet the interests and needs of different users, thereby enhancing the personalized experience of users.
[0128] S104: Outputting reply information according to the target knowledge information, the natural language description, the emotion recognition result and the user portrait.
[0129] The target knowledge information is the knowledge information in the knowledge base that matches the recommended information and fits the user's interests and background; the natural language description is related to the multimodal information input by the user into the navigation dialogue system; the emotion recognition result reflects the user's current emotional state; the user portrait integrates the user's attributes, interests and behaviors.
[0130] Therefore, the response information output by combining target knowledge information, natural language description, emotion recognition results and user portraits can not only meet the interests, backgrounds and needs of different users, thereby enhancing the user's personalized experience, but also have emotional interaction, thereby enhancing emotional communication with users.
[0131] In some embodiments, outputting reply information according to the target knowledge information, the natural language description, the emotion recognition result, and the user portrait includes:
[0132] S1041: Obtaining preliminary response information based on target knowledge information and natural language description;
[0133] The target knowledge information, i.e., the knowledge information in the knowledge base that matches the recommended information, fits the user's interests, background, and needs, and can be called "retrieved knowledge information". The natural language description is associated with the multimodal information that the user inputs into the navigation dialogue system, and can be called "user query".
[0134] The retrieved knowledge information is integrated with the user query to generate a preliminary, personalized response information that meets the user's interests, background and needs, thereby enhancing the user's personalized experience.
[0135] S1042: Outputting final reply information based on the preliminary reply information, the emotion recognition result and the user portrait.
[0136] The preliminary reply information generated in S1041 is optimized at the emotional level. The current emotional state of the user is obtained through the emotion recognition result, and the preliminary reply information is adjusted and enriched at the emotional level in combination with the user attributes, interests, and behaviors contained in the user portrait. In this way, the final reply information output can not only meet the personalized needs of the user, but also integrate emotional care, thereby enhancing the emotional interaction and communication with the user.
[0137] In some embodiments, the final reply information may include text, voice, and image information.
[0138] That is, the final reply information output to the user is multimodal information, which not only includes text information, but also can convert the text information into voice information through speech synthesis, and display relevant pictures on the interactive interface at the same time.
[0139] In an embodiment of the present application, the input of the navigation dialogue system is multimodal information, including text, voice and image information in the navigation scenario used by the navigation dialogue system. By combining multiple information sources, a more comprehensive and rich navigation interaction experience can be provided. In addition, the emotion recognition result is obtained based on the multimodal information, and the emotion recognition result and the user portrait are used to optimize the response information, so that the navigation dialogue system can form a personalized emotional response and enhance the emotional communication with the user. In addition, the multimodal information is converted into a natural language description, and the user portrait is obtained based on the natural language description. Recommendation information is generated based on the user portrait and the context information of the navigation dialogue system, and the target knowledge information matching the recommended information is determined from the knowledge base. By combining the recommendation mechanism of user portraits and context information, the response information output by the navigation dialogue system can meet the interests and needs of different users, thereby enhancing the personalized experience of users.
[0140] In some embodiments, obtaining an emotion recognition result according to the multimodal representation includes:
[0141] Input the multi-modal representation into a pre-trained sentiment recognition model so that the sentiment recognition model outputs a sentiment recognition result;
[0142] Among them, the sentiment recognition model is used to:
[0143] S1013: Analyze the text features, speech features, and image features included in the multi-modal representation respectively to obtain the corresponding user sentiment state;
[0144] Specifically, the sentiment recognition model will analyze each part (text feature vector, speech feature vector, and image feature vector) in the input multi-modal representation separately. That is, the sentiment recognition model will consider the user sentiment state contained in each modality separately. For example, the text feature vector may contain the sentiment clues in the text information input by the user, the speech feature vector may contain the sentiment clues such as the pitch, rhythm, and intensity of the user's speech, and the image feature vector may contain the visual sentiment clues such as the user's facial expression and body language.
[0145] By analyzing the feature vectors of each modality, the sentiment recognition model can infer the corresponding user sentiment state, such as happy, sad, angry, etc.
[0146] S1014: Output a sentiment recognition result according to the obtained user sentiment state and the pre-set weights.
[0147] Specifically, each modality (text, speech, and image) is assigned a pre-set weight, and these weights reflect the relative importance of different modalities in the sentiment recognition process. For example, in some cases, speech and facial expressions may reflect the user's true sentiment better than text. Therefore, the weights assigned to the speech modality and the image modality are both 0.4, while the weight assigned to the text modality is 0.2.
[0148] The sentiment recognition model will comprehensively output a sentiment recognition result according to the user sentiment state obtained by analyzing each modality and their weights. This sentiment recognition result is the sentiment state that the sentiment recognition model believes the user is most likely to be in.
[0149] In this embodiment, the sentiment recognition model can comprehensively consider the user sentiment state contained in different modalities, use the pre-set weights, and output a comprehensive sentiment recognition result, thereby improving the accuracy and robustness of sentiment recognition.
[0150] In some embodiments, a plurality of knowledge information related to the navigation scene is stored in a pre-set knowledge base. The knowledge information related to the navigation scene may include information about the navigation scene itself and information related to entities in the navigation scene. Among them, the information related to entities in the navigation scene may include relationship information between entities in the navigation scene and information about entities in the navigation scene.
[0151] The above knowledge information is stored in different types of databases in the knowledge base. Specifically, the knowledge base may include knowledge graphs, vector databases, and structured databases.
[0152] Among them, since the knowledge graph has significant advantages in expressing complex relationships between entities, it is very suitable for storing relationship information between entities in navigation scenarios.
[0153] Taking the museum tour scenario as an example, the relationship information in the knowledge graph can be organized according to the following structure:
[0154] (Exhibit) — [belongs to] — (exhibition hall);
[0155] (Exhibit) — [Created in] — (year);
[0156] (Exhibit) — [Creator] — (Artist);
[0157] (artist) — [specializes in] — (genre);
[0158] ....
[0159] For example, graph databases such as Neo4j and JanusGraph can be used to store data in the knowledge graph.
[0160] The vector database is used to store the information of entities in the navigation scene. The vector database can be used to store vector representations of text information such as the name, historical background, and materials of museum exhibits.
[0161] For example, vector databases such as Faiss or Milvus may be used.
[0162] The structured database is used to store the information of the guided scene itself. For structured information, such as the layout of the museum's exhibition halls, opening hours, and service facilities, a structured database can be used to store it.
[0163] Exemplarily, a structured database such as PostgreSQL or MySQL may be used.
[0164] In some embodiments, the knowledge information related to the tour guide scenario stored in the knowledge base includes two parts: core knowledge information and dynamic knowledge information. Among them, the core knowledge information refers to relatively stable content, which is stored in the main database, and the dynamic knowledge information refers to content that often changes, which is stored in the distributed cache system.
[0165] Taking the "information about the tour guide scenario itself" in the "knowledge information related to the tour guide scenario" as an example, it includes two parts: core knowledge information and dynamic knowledge information. The core knowledge information may include relatively stable content such as the regular opening hours of the museum and permanent service facilities, and these contents are stored in the structured database of the main database; the dynamic knowledge information may include content that often changes such as the temporary exhibition information, emergency information, and real-time number of visitors in the museum, and these contents are stored in the structured database of the distributed cache system.
[0166] In some embodiments, an incremental update strategy is adopted to update the core knowledge information. The version control system (such as Git or SVN) is used to manage the core knowledge information, and the version differences are checked at set intervals (such as every 24 hours, every week, etc.), and only the changed core knowledge information is updated.
[0167] For different types of information in the dynamic knowledge information, the update frequency of this type of information can be determined according to the importance level of this type of information.
[0168] For example, for information with a relatively high importance level such as emergency information (such as fire alarms, sudden medical events, etc.), the update is triggered immediately to ensure that relevant personnel can obtain these key information in the first time so as to take corresponding countermeasures quickly; for the real-time number of visitors, the update frequency can be relatively low, such as once every half hour or once every hour, just to meet the needs of daily operation management and tourist inquiries.
[0169] In this way, resources can be reasonably allocated, so that the knowledge information in the knowledge base can be updated in a timely manner to ensure the timeliness and accuracy of the content, and unnecessary resource waste will not be caused due to overly frequent updates.
[0170] In some embodiments, determining the target knowledge information that matches the recommended information from the knowledge base under the preset tour guide scenario includes:
[0171] Inputting the recommended information into a pre-trained retrieval model so that the retrieval model outputs the target knowledge information;
[0172] Among them, the retrieval model is used for:
[0173] S1035: respectively determine the similarity between each knowledge information in the knowledge base and the recommended information;
[0174] When the recommendation information is input into the retrieval model, the retrieval model converts the recommendation information into a vector representation.
[0175] The retrieval model needs to be able to interact with the knowledge base to obtain knowledge information from the knowledge base. Specifically, the retrieval model can send a request to the knowledge base through the application programming interface (API) and receive a response from the knowledge base, thereby obtaining all the knowledge information stored in the knowledge base. After obtaining all the knowledge information stored in the knowledge base, the retrieval model will convert each piece of knowledge information into a vector representation.
[0176] The similarity between each knowledge information vector and the recommendation information vector is calculated respectively. The similarity calculation method may be cosine similarity, Euclidean distance or a self-defined similarity function, which is not specifically limited in the embodiments of the present application.
[0177] S1036: Taking the top N pieces of knowledge information in similarity ranking as the target knowledge information, and outputting the target knowledge information.
[0178] The retrieval model sorts the calculated similarity values in order from high to low.
[0179] In the sorted list, the retrieval model will select the top N pieces of knowledge information with the highest similarity as the target knowledge information and output them. It should be noted that N is a preset parameter, which can be set according to specific application scenarios and requirements. For example, if you want to provide users with more accurate recommendations, N can be set to a smaller value, such as 1, 2, etc.; if you want to provide users with a wider range of recommendations, N can be set to a larger value, such as 5, 6, etc.
[0180] In this embodiment, after obtaining the user portrait in the manner described in S102, recommendation information can be generated based on the user portrait and the context information of the navigation dialogue system, and target knowledge information matching the recommendation information can be determined from the knowledge base. By combining the recommendation mechanism of the user portrait and the context information, the reply information output by the navigation dialogue system can meet the interests and needs of different users, thereby improving the personalized experience of users.
[0181] In some embodiments, preliminary response information is obtained based on the target knowledge information and the natural language description, including:
[0182] Inputting the target knowledge information and the natural language description into the pre-trained preliminary response model so that the preliminary response model outputs preliminary response information;
[0183] Among them, the training process of the preliminary response model is as follows Figure 2 As shown, including:
[0184] S201: pre-training the initial model according to the general corpus to obtain a pre-trained model;
[0185] General corpus refers to a large-scale dataset containing various types of text, such as books, articles, and web content. The initial model can use a transformer architecture similar to ChatGPT (Chat Generative Pre-trained Transformer).
[0186] The initial model is pre-trained on a general corpus to allow the model to deeply understand the deep structure and meaning of the language. After pre-training, the model parameters are optimized to generate a pre-trained model. This pre-trained model can capture the characteristics of general language and has basic language understanding and generation capabilities.
[0187] S202: fine-tuning the pre-trained model according to the professional corpus in the tour guide scenario to obtain a pre-trained-fine-tuned model;
[0188] Professional corpus in the tour guide scenario refers to text datasets related to the tour guide scenario. In the example of a museum tour, professional corpus may include exhibit descriptions, historical background, artistic style, artist introductions, and exhibition reviews.
[0189] The pre-trained model is further fine-tuned using professional corpus in the tour guide scenario. The purpose is to allow the model to further master the proprietary vocabulary, concepts and expressions in the tour guide scenario while maintaining its general language capabilities.
[0190] During fine-tuning, one or more task-specific layers can be added to the output layer of the model and the model can be trained with a smaller learning rate to avoid destroying the general language knowledge learned by the model during the pre-training phase.
[0191] In the example of a museum tour scenario, specific task objectives can be designed, such as exhibit information retrieval, tour route planning, and art style analysis. Through multi-task learning, the model can handle a variety of museum tour queries.
[0192] After fine-tuning, the model parameters were optimized, resulting in a pre-trained-fine-tuned model that combines general language understanding capabilities with expertise in navigation scenarios, providing users with a more accurate and richer navigation experience.
[0193] S203: Perform fusion training on the pre-trained and fine-tuned models to obtain a preliminary response model; the fusion training includes fusing the input knowledge information and the user query to output preliminary response information.
[0194] During the fusion training process, the model learns how to combine knowledge information with user queries to generate accurate, coherent, and informative preliminary response information.
[0195] In an embodiment of the present application, a preliminary response model is used to fuse the target knowledge information (i.e., the retrieved knowledge information) with the natural language description (i.e., the user query) to generate a preliminary, personalized response information that meets the user's interests, background, and needs, thereby enhancing the user's personalized experience.
[0196] In some embodiments, the final reply information is outputted according to the preliminary reply information, the emotion recognition result and the user portrait, including:
[0197] According to the user's emotional state, user portrait and emotional interaction strategy, the preliminary reply information is processed to generate and output the final reply information, wherein the emotional interaction strategy is generated by a pre-trained emotion enhancement model.
[0198] The emotion recognition results comprehensively consider the user's emotional state contained in different modalities; the user portrait contains data such as user attributes, interests and behaviors; the emotion recognition strategy is used to guide how to use the emotion recognition results and user portraits to adjust and enrich the initial response information at the emotional level.
[0199] For example, the user's emotional state contained in the emotion recognition result can be used as the emotional tone to adjust the preliminary reply information. Then, based on the data in the user portrait, the adjusted preliminary reply information is optimized and personalized at the emotional level to make it more suitable for the user's emotional state, communication habits and language style. Finally, after the above emotional optimization and personalized processing, the output final reply information can not only meet the user's personalized needs, but also integrate emotional care, thereby enhancing emotional interaction and communication with the user.
[0200] In some embodiments, the user's feedback information regarding the final reply information input into the guided dialogue system is input into the emotion enhancement model, so that the emotion enhancement model updates the emotion interaction strategy according to the feedback information.
[0201] Exemplarily, the user's feedback information regarding the final reply information may include explicit feedback information and / or implicit feedback information.
[0202] Among them, explicit feedback information includes ratings, comments and / or direct emotional labels (such as "satisfied", "unsatisfied"), etc.; implicit feedback information includes users' follow-up questions, the continuation method of the conversation and / or the interruption of the conversation, etc.
[0203] In some embodiments, the emotion interaction strategy can be continuously optimized in the emotion enhancement model through reinforcement learning methods.
[0204] The specific implementation can be as follows:
[0205] Define the state space: including the user's current emotional state (i.e., emotion recognition results), user portrait, and preliminary response to user queries (i.e., preliminary reply information), etc. The state space can be represented as a multidimensional vector, with each dimension corresponding to an element in the state space.
[0206] Define the action space: including different types of emotional interaction strategies. The action space is usually a finite set that defines all possible actions that the emotion enhancement model can take, and each action is a specific strategy that the model can execute.
[0207] Reward function: A reward function is designed based on the user's feedback information (including explicit feedback information and / or implicit feedback information) regarding the final reply information.
[0208] Policy network: Use algorithms such as Deep Q-Network DQN (Deep Q-Network) or Proximal Policy Optimization PPO (Proximal Policy Optimization) to learn the optimal emotional interaction strategy.
[0209] In this embodiment, the feedback information of the user on the final reply information input into the navigation dialogue system is input into the emotion enhancement model, so that the emotion enhancement model updates the emotion interaction strategy according to the feedback information, thereby dynamically adjusting its emotion interaction strategy according to the personalized reactions of different users. For example, appropriate expressions can be adopted for users of different age groups, or the detailedness of the explanation and the intensity of the emotion expression can be adjusted according to the user's interest in the content, thereby enhancing the emotional communication with the user, optimizing the user's interactive experience, and improving the user's interaction satisfaction.
[0210] like Figure 3 As shown, Figure 3 It is a flowchart of a reply method of a navigation dialogue system provided in an embodiment of the present application.
[0211] The multimodal information input into the navigation dialogue system is subjected to feature fusion to obtain a multimodal representation. The multimodal information includes text, voice and image information in the navigation scenario used by the navigation dialogue system. In some embodiments, feature extraction is performed on the text, voice and image information respectively, and the extracted features are aligned and fused using an attention mechanism to obtain a multimodal representation.
[0212] Obtaining an emotion recognition result based on the multimodal representation. In some embodiments, the multimodal representation can be input into a pre-trained emotion recognition model so that the emotion recognition model outputs an emotion recognition result; wherein the emotion recognition model is used to: respectively analyze the text features, speech features, and image features included in the multimodal representation to obtain the corresponding user emotion state; and output the emotion recognition result based on the obtained user emotion state and the pre-set weight.
[0213] Converting multimodal information into natural language descriptions, that is, performing cross-modal conversion, is to integrate information from different modalities into a unified text modality description.
[0214] The user portrait is obtained based on the natural language description. In some embodiments, the natural language description is subjected to semantic analysis; the result of the semantic analysis is subjected to feature extraction to obtain target features, the target features at least including user interest information and user background information; and the user portrait is constructed based on the target features.
[0215] Context information is extracted from the dialogue content of the navigation dialogue system, thereby obtaining the context information of the navigation dialogue system. The dialogue content of the navigation dialogue system includes the currently input multimodal information, the historically input multimodal information and the historically output reply information.
[0216] Generate recommendation information based on the user profile and the context information of the navigation dialogue system. In some embodiments, the recommendation information is determined from the content library based on the user profile, the context information of the navigation dialogue system, and the recommendation strategy, wherein the recommendation strategy is generated by a pre-trained recommendation model. That is, based on the recommendation strategy generated by the recommendation model, the recommended content matching the user profile and the context information is screened from the pre-set content library.
[0217] Determine the target knowledge information that matches the recommended information from the knowledge base under the preset navigation scenario. The preset knowledge base stores multiple pieces of knowledge information related to the navigation scenario, and the knowledge base includes a real-time update mechanism that can adaptively collect Internet and offline information for automatic update. In some embodiments, the recommended information is input into a pre-trained retrieval model so that the retrieval model outputs the target knowledge information; wherein the retrieval model is used to: respectively determine the similarity between each piece of knowledge information in the knowledge base and the recommended information; use the top N pieces of knowledge information in similarity ranking as the target knowledge information, and output the target knowledge information.
[0218] The target knowledge information and natural language description are input into the pre-trained preliminary response model so that the preliminary response model outputs preliminary response information. The training process of the preliminary response model includes: pre-training the initial model according to the general corpus to obtain the pre-trained model; fine-tuning the pre-trained model according to the professional corpus in the tour scene to obtain the pre-trained-fine-tuned model; performing fusion training on the pre-trained-fine-tuned model to obtain the preliminary response model; the fusion training includes fusing the input knowledge information and the user query to output the preliminary response information.
[0219] According to the user's emotional state, user portrait and emotional interaction strategy, the preliminary reply information is processed to generate and output the final reply information, wherein the emotional interaction strategy is generated by a pre-trained emotion enhancement model.
[0220] In an embodiment of the present application, the input of the navigation dialogue system is multimodal information, including text, voice and image information in the navigation scenario used by the navigation dialogue system. By combining multiple information sources, a more comprehensive and rich navigation interaction experience can be provided. In addition, the emotion recognition result is obtained based on the multimodal information, and the emotion recognition result and the user portrait are used to optimize the response information, so that the navigation dialogue system can form a personalized emotional response and enhance the emotional communication with the user. In addition, the multimodal information is converted into a natural language description, and the user portrait is obtained based on the natural language description. Recommendation information is generated based on the user portrait and the context information of the navigation dialogue system, and the target knowledge information matching the recommended information is determined from the knowledge base. By combining the recommendation mechanism of user portraits and context information, the response information output by the navigation dialogue system can meet the interests and needs of different users, thereby enhancing the personalized experience of users.
[0221] The above content describes the method provided by the present application. The following describes the device provided by the present application:
[0222] See also Figure 4 , which is a structural diagram of a reply device of a guided dialogue system provided in an embodiment of the present application.
[0223] like Figure 4 As shown, the device may include:
[0224] The emotion recognition unit 410 is used to perform feature fusion on the multimodal information input into the guide dialogue system to obtain a multimodal representation, and obtain an emotion recognition result according to the multimodal representation, wherein the multimodal information includes text, voice and image information in the guide scenario applied by the guide dialogue system;
[0225] A determination unit 420, configured to convert the multimodal information into a natural language description, and obtain a user portrait according to the natural language description;
[0226] A recommendation unit 430, configured to generate recommendation information according to the user portrait and the context information of the navigation dialogue system, and determine target knowledge information matching the recommendation information from a knowledge base preset under the navigation scenario;
[0227] The reply unit 440 is used to output reply information according to the target knowledge information, the natural language description, the emotion recognition result and the user portrait.
[0228] In some embodiments, the emotion recognition unit 410 is specifically used to:
[0229] Performing feature extraction on the text, voice and image information respectively;
[0230] The extracted features are aligned and fused using an attention mechanism to obtain the multimodal representation.
[0231] In some embodiments, the emotion recognition unit 410 is specifically used to:
[0232] Inputting the multimodal representation into a pre-trained emotion recognition model so that the emotion recognition model outputs the emotion recognition result;
[0233] Wherein, the emotion recognition model is used for:
[0234] Analyzing the text features, speech features, and image features included in the multimodal representation respectively to obtain the corresponding user emotional state;
[0235] The emotion recognition result is output according to the obtained user emotion state and the preset weight.
[0236] In some embodiments, the determining unit 420 is specifically configured to:
[0237] Performing semantic analysis on the natural language description;
[0238] Extracting features from the result of semantic analysis to obtain target features, wherein the target features at least include user interest information and user background information;
[0239] The user portrait is constructed according to the target features.
[0240] In some embodiments, the knowledge base includes a knowledge graph, a vector database and a structured database, wherein the knowledge graph is used to store the relationship information between entities in the navigation scenario, the vector database is used to store the information of the entities in the navigation scenario, and the structured database is used to store the information of the navigation scene itself.
[0241] In some embodiments, the recommendation unit 430 is specifically configured to:
[0242] Inputting the recommendation information into a pre-trained retrieval model so that the retrieval model outputs the target knowledge information;
[0243] Wherein, the retrieval model is used for:
[0244] respectively determining the similarity between each piece of knowledge information in the knowledge base and the recommended information;
[0245] The top N pieces of knowledge information ranked by similarity are used as the target knowledge information, and the target knowledge information is output.
[0246] In some embodiments, the reply unit 440 is specifically configured to:
[0247] Obtaining preliminary response information according to the target knowledge information and the natural language description;
[0248] According to the preliminary reply information, the emotion recognition result and the user portrait, final reply information is output, and the final reply information includes text, voice and image information.
[0249] In some embodiments, the reply unit 440 is specifically configured to:
[0250] The preliminary reply information is processed according to the user's emotional state, the user portrait and the emotional interaction strategy to generate and output the final reply information, wherein the emotional interaction strategy is generated by a pre-trained emotion enhancement model.
[0251] In some embodiments, the apparatus further comprises an updating unit, wherein the updating unit is configured to:
[0252] The user's feedback information input into the navigation dialogue system in response to the final reply information is input into the emotion enhancement model, so that the emotion enhancement model updates the emotion interaction strategy according to the feedback information.
[0253] The implementation process of the functions and effects of each unit in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.
[0254] The present application embodiment also provides a hardware structure. Figure 5 , Figure 5 The hardware structure diagram of an electronic device provided in an embodiment of the present application is shown in FIG. Figure 5As shown, the hardware structure may include: a processor and a machine-readable storage medium, the machine-readable storage medium storing machine-executable instructions that can be executed by the processor; the processor is used to execute the machine-executable instructions to implement the method disclosed in the above example of this application.
[0255] Based on the same application concept as the above method, an embodiment of the present application also provides a machine-readable storage medium, on which a number of computer instructions are stored. When the computer instructions are executed by a processor, the method disclosed in the above example of the present application can be implemented.
[0256] Exemplarily, the above-mentioned machine-readable storage medium can be any electronic, magnetic, optical or other physical storage device that can contain or store information, such as executable instructions, data, etc. For example, the machine-readable storage medium can be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drive (such as hard disk drive), solid state drive, any type of storage disk (such as optical disk, DVD, etc.), or similar storage medium, or a combination thereof.
[0257] It should be noted that, in this article, relational terms such as target and target are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0258] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A reply method for a navigation dialogue system, characterized in that: The method includes: Performing feature fusion on multimodal information input into the guide dialogue system to obtain a multimodal representation, and obtaining an emotion recognition result according to the multimodal representation, wherein the multimodal information includes text, voice and image information in the guide scenario in which the guide dialogue system is applied; Convert the multimodal information into a natural language description, and obtain a user portrait according to the natural language description; Generate recommendation information according to the user portrait and context information of the navigation dialogue system, and determine target knowledge information matching the recommendation information from a knowledge base preset under the navigation scenario; Output reply information according to the target knowledge information, the natural language description, the emotion recognition result and the user portrait.
2. The method according to claim 1, characterized in that The step of fusing features of the multimodal information input into the guide dialogue system to obtain a multimodal representation includes: Performing feature extraction on the text, voice and image information respectively; The extracted features are aligned and fused using an attention mechanism to obtain the multimodal representation.
3. The method according to claim 1, characterized in that Obtaining the emotion recognition result according to the multimodal representation includes: Inputting the multimodal representation into a pre-trained emotion recognition model so that the emotion recognition model outputs the emotion recognition result; Wherein, the emotion recognition model is used for: Analyzing the text features, speech features, and image features included in the multimodal representation respectively to obtain the corresponding user emotional state; The emotion recognition result is output according to the obtained user emotion state and the preset weight.
4. The method according to claim 1, characterized in that: The obtaining of the user portrait according to the natural language description includes: Performing semantic analysis on the natural language description; Extracting features from the result of semantic analysis to obtain target features, wherein the target features at least include user interest information and user background information; The user portrait is constructed according to the target features.
5. The method according to claim 1, characterized in that The knowledge base includes a knowledge graph, a vector database and a structured database, wherein the knowledge graph is used to store the relationship information between entities in the navigation scenario, the vector database is used to store the information of the entities in the navigation scenario, and the structured database is used to store the information of the navigation scenario itself.
6. The method according to claim 1, characterized in that The determining target knowledge information matching the recommended information from the knowledge base preset in the navigation scenario includes: Inputting the recommendation information into a pre-trained retrieval model so that the retrieval model outputs the target knowledge information; Wherein, the retrieval model is used for: respectively determining the similarity between each piece of knowledge information in the knowledge base and the recommended information; The top N pieces of knowledge information ranked by similarity are used as the target knowledge information, and the target knowledge information is output.
7. The method according to claim 1, characterized in that The outputting reply information according to the target knowledge information, the natural language description, the emotion recognition result and the user portrait includes: Obtaining preliminary response information according to the target knowledge information and the natural language description; According to the preliminary reply information, the emotion recognition result and the user portrait, final reply information is output, and the final reply information includes text, voice and image information.
8. The method according to claim 7, characterized in that The outputting final response information according to the preliminary response information, the emotion recognition result and the user portrait includes: The preliminary reply information is processed according to the user's emotional state, the user portrait and the emotional interaction strategy to generate and output the final reply information, wherein the emotional interaction strategy is generated by a pre-trained emotion enhancement model.
9. The method according to claim 8, characterized in that The method further comprises: The user's feedback information input into the navigation dialogue system in response to the final reply information is input into the emotion enhancement model, so that the emotion enhancement model updates the emotion interaction strategy according to the feedback information.
10. A reply device for a guide dialogue system, characterized in that: The device includes: An emotion recognition unit, used for performing feature fusion on multimodal information input into the guide dialogue system to obtain a multimodal representation, and obtaining an emotion recognition result according to the multimodal representation, wherein the multimodal information includes text, voice and image information in the guide scenario applied by the guide dialogue system; A determination unit, configured to convert the multimodal information into a natural language description and obtain a user portrait according to the natural language description; A recommendation unit, configured to generate recommendation information according to the user portrait and context information of the navigation dialogue system, and determine target knowledge information matching the recommendation information from a knowledge base preset under the navigation scenario; A reply unit is used to output reply information based on the target knowledge information, the natural language description, the emotion recognition result and the user portrait.
11. An electronic device, characterized in that: The invention comprises a processor and a memory, wherein the memory stores machine executable instructions that can be executed by the processor, and the processor is used to execute the machine executable instructions to implement the method according to any one of claims 1 to 9.
12. A machine-readable storage medium, characterized in that: The machine-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are executed by a processor, the method according to any one of claims 1 to 9 is implemented.
Citation Information
Cited By
Vector database-based science and technology museum exhibition item intelligent retrieval and recommendation system and method
CN121166779A