Personalized interaction and scene adaptation system of digital human AI tour guide assistant

By combining multimodal data fusion and knowledge graphs, personalized interaction strategies are generated, which solves the problem of accurate answers of AI digital humans in complex emotions and multiple scenarios, and improves the user experience.

CN120595952APending Publication Date: 2025-09-05WUHAN YINQIAO NANHAI PHOTOELECTRIC CO LTD

Patent Information

Application Number
CN202511001203.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing AI digital humans face challenges in complex sentiment analysis and understanding real-time multimodal data in multiple scenarios, resulting in mechanical responses or inability to accurately answer user questions, affecting user experience.

Method used

It adopts user portrait data acquisition module, multimodal feature extraction module, feature fusion module, scenario adaptation module and personalized interaction module to generate intent categories and emotional states through multimodal data fusion, and generates personalized interactive answers in combination with knowledge graph.

Benefits of technology

It achieves accurate analysis of complex emotions and personalized interaction in multiple scenarios, improves user experience, and solves the problem of digital humans accurately answering questions with complex emotions and across scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120595952A_ABST
    Figure CN120595952A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of AI digital human interaction, and particularly discloses a personalized interaction and scene adaptation system for a digital human AI tour guide assistant, which comprises the following steps: acquiring user portrait multi-modal data at the same moment, extracting features from the multi-modal data, performing multi-modal feature fusion on the extracted features, and generating fusion features; generating an intention category and an emotional state based on the fusion feature; generating a corresponding scene adaptation strategy based on the intention category and the emotional state; and on the basis of a corresponding scene adaptation strategy and in combination with the knowledge graph, generating a personalized interaction answer. According to the digital human system provided by the invention, complex expressions and emotion problems of the user can be processed, and meanwhile, multi-scene adaptability can be realized for cross-scene questions of the user, so that the user experience is improved. A stage in which the digital human can be gradually transited to pure AI driving; and the commercial value and the social influence of the method can be displayed in more application scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of AI digital human interaction technology and relates to a personalized interaction and scene adaptation system for a digital human AI tour guide assistant. Background Art

[0002] With the rapid development of artificial intelligence, AI digital humans are virtual characters that can simulate the appearance, voice, movements, and expressions of real people. They integrate advanced technologies such as computer graphics, speech synthesis, machine learning, and natural language processing. As a cutting-edge technological product, the direction and trends of AI digital humans' technological development are attracting much attention and anticipation.

[0003] The following prior art: Patent publication number CN119916947A discloses an intelligent interaction method and system based on AI digital humans. The intelligent interaction system includes an environmental perception and analysis module, an AI digital human processing module, a control execution output module, and an evaluation and maintenance interaction module, which operate in sequence. The key technical points are: By combining computer graphics and artificial intelligence technology, the intelligent interaction system successfully constructs an AI digital human capable of intelligently controlling lighting strategies, achieving intelligent environmental perception and status judgment, autonomous formulation and execution of lighting strategies, and efficient voice command recognition and response technical effects; based on the movement speed of the infrared signal source, the system builds and runs a dynamic adjustment model to generate a brightness increment value for each adjustment corresponding to lamp A, which realizes dynamic adjustment of the lighting strategy, ensuring that pedestrians can clearly see the road conditions while avoiding the impact of excessive brightness on stage actors.

[0004] Another example is the patent with publication number CN119902625A, which discloses an AI-based virtual digital human interaction system. The system is used to: obtain a user feature data set by extracting features from the interaction data set, calculate the audio change rate and speech speed stability, obtain voice emotion data, calculate the syntactic complexity and count the number of emotional words, obtain text emotion data, calculate the symmetry of facial expressions, obtain facial emotion data, calculate body stability and identify visual stay areas, obtain attention state data, and adjust the virtual digital human's tone and speech speed based on it, optimize the sentence complexity and emotional words in the virtual digital human's text, adjust the changes in the virtual digital human's eyes, eyebrows and mouth corners, and adjust the information display order and visual focus guidance.

[0005] Problems with existing technologies: Although existing AI digital humans can recognize commands such as text and voice, they have difficulty in some complex sentiment analysis, such as identifying complex emotions (such as irony and implicit complaints), resulting in mechanical and general responses; and understanding real-time multimodal data in multiple scenarios is still challenging. For example, in a museum or entertainment scenario, if a user asks a cross-scenario question (such as suddenly asking "where are the good restaurants nearby"), the digital human may not be able to answer accurately or may give a chaotic answer due to the fragmentation of the knowledge graph, causing user dissatisfaction. Summary of the Invention

[0006] In view of the above problems existing in the prior art, the present invention provides a personalized interaction and scene adaptation system of a digital human AI tour guide assistant, which is used to solve the above technical problems.

[0007] In order to achieve the above-mentioned and other purposes, the technical solutions adopted by the present invention are as follows: The present invention provides a personalized interaction and scene adaptation system for a digital human AI tour guide assistant, which includes: a user portrait data acquisition module, a multimodal feature extraction module, a feature fusion module, a scene adaptation module, and a personalized interaction module; The above modules are connected via wired and / or wireless connections to achieve data transmission between modules; User portrait data acquisition module, used to obtain multimodal data of user portraits at the same time; Multimodal feature extraction module, used to extract features from multimodal data respectively; The feature fusion module is used to perform multimodal feature fusion on the extracted features to generate fused features; and generate intent categories and emotional states based on the fused features; The scenario adaptation module generates the corresponding scenario adaptation strategy based on the intent category and emotional state; The personalized interaction module generates personalized interactive responses based on the corresponding scenario adaptation strategy and combined with the knowledge graph.

[0008] Furthermore, the multimodal data includes: voice data, text data, facial image data and body movement data.

[0009] Furthermore, the features include: voice features, text features, facial features, and body features.

[0010] Furthermore, the speech features include: speaking speed, pitch, and timbre; the text features include: emotional words, negative words, keywords, degree adverbs, pronouns, punctuation marks, and sentences; the facial features include: eye key points, eyebrow key points, mouth key points, and cheek muscle key points; and the limb features include: shoulder joints, elbow joints, wrist joints, finger joints, and hip joints.

[0011] Furthermore, the process of multimodal feature fusion is as follows: The speech features, text features, facial features and body features are input into the multimodal feature fusion model, and the speech features, text features, facial features and body features are mapped to a common latent space through the encoder layer; the decoder layer interactively fuses them through the cross-attention mechanism, uses the gradient descent method to minimize the loss function, and uses the backpropagation algorithm to train the multimodal feature fusion model and output fusion features; the multimodal feature fusion model includes: a Transformer model or an unsupervised learning model.

[0012] Furthermore, the process of generating intent categories is as follows: inputting the fusion features into the intent classification model, and the intent classification model outputs intent categories as the results. The intent categories include: information query, entertainment, life service, education and learning, feedback and evaluation; The process of generating the emotional state is: inputting the fusion features into the emotional state model, and the emotional state model uses the emotional state as the output result. The emotional state includes: happiness, sadness, anger, surprise, smile and irony.

[0013] Furthermore, the process of generating the corresponding scenario adaptation strategy is as follows: writing a query statement based on the scenario knowledge graph database software to query the scenario strategy corresponding to the intent category, and generating the corresponding scenario adaptation strategy based on the emotional state association rules and the user historical behavior association rules.

[0014] Furthermore, the process of generating personalized interactive answers is as follows: Generate n groups of initial interactive responses based on the corresponding scenario adaptation strategy and knowledge graph; through the fitness function: ,in, represents the fitness value, Represents the weight to evaluate the fitness value of n groups of initial interactive answers, screen out the initial interactive answers corresponding to higher fitness values ​​and perform i genetic operations, the genetic operations include crossover and mutation operations, and record the fitness values ​​after i genetic operations, and take the interactive answer corresponding to the highest fitness value after i genetic operations as the personalized interactive answer, the personalized interactive answer includes voice, text and digital human actions.

[0015] As described above, the personalized interaction and scene adaptation system of a digital human AI tour guide assistant provided by the present invention has at least the following beneficial effects: By collecting multimodal data of user portraits at the same moment, extracting features from the multimodal data separately, fusing the extracted features into multimodal features, generating fused features, and generating intent categories and emotional states based on the fused features; some complex emotions of users, such as difficult-to-identify complex emotions such as irony or implicit complaints, are analyzed and resolved; based on intent categories and emotional states, corresponding scenario adaptation strategies are generated, which solves the problem that when users ask questions across scenarios, digital humans may not be able to answer accurately or give chaotic answers due to the fragmentation of knowledge graphs, leading to user dissatisfaction, and can achieve adaptability in multiple scenarios; based on the corresponding scenario adaptation strategy and combined with the knowledge graph, personalized interactive answers are generated to enhance user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0017] Figure 1 This is a schematic diagram showing the connections between the modules of the personalized interaction and scene adaptation system of a digital human AI tour guide assistant according to the present invention. DETAILED DESCRIPTION

[0018] The above contents described below in conjunction with the implementation of the present invention are merely examples and explanations of the concept of the present invention. Those skilled in the art may make various modifications or additions to the described specific embodiments or replace them in a similar manner. As long as they do not deviate from the concept of the invention or exceed the scope defined by the claims, they shall fall within the scope of protection of the present invention.

[0019] Example 1 See also Figure 1 As shown, a personalized interaction and scene adaptation system for a digital human AI tour guide assistant includes: a user portrait data acquisition module, a multimodal feature extraction module, a feature fusion module, a scene adaptation module, and a personalized interaction module; The above modules are connected via wired and / or wireless connections to achieve data transmission between modules; A user portrait data acquisition module is used to acquire multimodal data of the user portrait at the same moment, wherein the multimodal data includes: voice data, text data, facial image data and body movement data; Based on the above embodiment, the voice data acquisition method is as follows: the built-in noise reduction 360-degree sound pickup array microphone acquires the user's interactive voice in real time to obtain voice data; The text data acquisition method is: the text content input by the user is collected and recorded in real time through the keyboard or handwritten text input interface to obtain text data; The facial image data is obtained by real-time tracking and collecting facial image data using a built-in RGB camera or a depth camera to obtain facial image data; The body movement data is obtained by using a high-speed camera or a depth camera to track and collect human body movements in real time, such as shaking hands, walking, and clasping fists, to obtain body movement data; It should be added that preprocessing operations such as noise reduction, format normalization and timestamp alignment are performed on the multimodal data of user portraits at the same time. This process is an existing technology and will not be described in detail here. It solves the problems of noise interference, format differences and time asynchrony in the original collected multimodal data of user portraits, improves data quality, and ensures data consistency and accuracy.

[0020] The multimodal feature extraction module extracts features from the multimodal data, including speech features, text features, facial features, and body features. Speech features include speech rate, pitch, and timbre. Text features include emotional words, negative words, keywords, degree adverbs, pronouns, punctuation marks, and sentences. Facial features include key points of the eyes, eyebrows, mouth, and cheek muscles. Body features include shoulder joints, elbow joints, wrist joints, finger joints, and hip joints. On the basis of the above embodiment, it should be supplemented that the speech features of the speech data such as speech rate, pitch and timbre are obtained by using Mel-frequency cepstral coefficients, and the speech data is converted into text content by combining the deep learning large language model Transformer. At the same time, the text data content is analyzed to obtain text features such as sentiment words, negative words, degree adverbs, keywords, punctuation marks and sentences; For example, keywords (such as "complaint" and "recommend" directly reflect intent), sentiment words (such as "disappointed" and "satisfied" directly reflect emotional tendencies), negative words or degree adverbs (such as "not at all" and "very good" to strengthen emotional intensity), sentences (such as interrogative sentences to express query intent and exclamatory sentences to express emotion), pronouns (such as "it" refers to the "subject" in the previous text and needs to be analyzed in context), punctuation marks (such as "!!" to express strong emotion and "?" to express question), etc. By analyzing and obtaining text features, we can identify what the user wants to express and understand the user's intent and emotional expression. With the help of the computer vision target detection model, the eye key points, eyebrow key points, mouth key points, cheek muscle key points, shoulder joint points, elbow joint points, wrist joint points, finger joint points and hip joint points are synchronously detected.

[0021] The feature fusion module performs multimodal feature fusion on the extracted features to generate fused features; and generates intent categories and emotional states based on the fused features; The multimodal feature fusion process is as follows: speech features, text features, facial features, and body features are input into a multimodal feature fusion model, and the speech features, text features, facial features, and body features are mapped to a common latent space through the encoder layer; the decoder layer interactively fuses them through a cross-attention mechanism, uses the gradient descent method to minimize the loss function, and uses the backpropagation algorithm to train the multimodal feature fusion model and output fused features. The multimodal feature fusion model includes but is not limited to: a Transformer model or an unsupervised learning model; The process of generating intent categories is as follows: fusion features are input into the intent classification model, which outputs intent categories. The intent categories include but are not limited to information query, entertainment, life services, education and learning, feedback and evaluation, etc. The process of generating the emotional state is as follows: inputting the fused features into the emotional state model, which outputs the emotional state, including but not limited to happiness, sadness, anger, surprise, smile, and irony; Based on the above embodiment, it should be supplemented that the intent classification model and the emotional state model include but are not limited to: natural language processing model, bidirectional long short-term memory network or gated recurrent unit; information query class, for example, weather query: such as asking about the weather conditions of a specific location, traffic query: such as asking about bus and subway routes, real-time traffic conditions, etc.; Entertainment, such as music playback: requesting to play a specific song, artist, or music genre; Life services, such as restaurant recommendations: finding restaurants and food recommendations; travel planning: requesting recommendations for travel destinations and itineraries; Education and learning, such as question answering: raising academic or knowledge-based questions and seeking answers.

[0022] Feedback and evaluation, such as complaints and dissatisfaction: expressing dissatisfaction with products or services; praise and gratitude: expressing gratitude and praise for high-quality services or products.

[0023] The scenario adaptation module generates the corresponding scenario adaptation strategy based on the intent category and emotional state. The process of generating the corresponding scenario adaptation strategy is as follows: writing a query statement based on the scenario knowledge graph database software to query the scenario strategy corresponding to the intent category, and generating the corresponding scenario adaptation strategy based on the emotional state association rules and the user's historical behavior association rules; In addition to the above examples, it should be noted that a knowledge graph is a technology that uses a graph structure to represent and store large amounts of structured and semi-structured knowledge, enabling knowledge representation, reasoning, and querying. For example, a knowledge graph for tourism scenarios should include detailed information on scenic spots in each region, historical and cultural knowledge, stories about historical events and figures, knowledge about ethnic customs, geography and transportation, and travel safety information. Emotional state association rules: If the user is in a positive emotional state, encouraging language can be added to the policy; if the user is in a negative emotional state, comforting language should be provided. For example, under the information query intent category, if the user is in a positive emotional state, the policy could be, "Your question is very valuable. Here are the detailed information we found for you. We hope it meets your needs!" User historical behavior association rules can analyze historical user interaction data to understand their preferences and habits. This module enables adaptability to multiple scenarios and enhances the user experience.

[0024] The personalized interaction module generates personalized interactive responses based on the corresponding scenario adaptation strategy and combined with the knowledge graph. The process of generating personalized interactive responses is as follows: Generate n groups of initial interactive responses based on the corresponding scenario adaptation strategy and knowledge graph; through the fitness function: ,in, represents the fitness value, Represents the weight to evaluate the fitness value of n groups of initial interactive answers, screen out the initial interactive answers corresponding to higher fitness values ​​and perform i genetic operations, the genetic operations include crossover and mutation operations, and record the fitness values ​​after i genetic operations, and take the interactive answer corresponding to the highest fitness value after i genetic operations as the personalized interactive answer, the personalized interactive answer includes voice, text and digital human actions.

[0025] The effects achieved by this embodiment include collecting multimodal data of user portraits at the same moment, extracting features from the multimodal data respectively, performing multimodal feature fusion on the extracted features, generating fused features, and generating intent categories and emotional states based on the fused features; solving some complex emotions of users, such as difficulty in identifying complex emotions such as irony or implicit complaints; generating corresponding scene adaptation strategies based on intent categories and emotional states, solving the problem that when users ask questions across scenes, the digital human may not be able to answer accurately or give chaotic answers due to the fragmentation of the knowledge graph, leading to user dissatisfaction, and can achieve adaptability in multiple scenes and improve user experience; generating personalized interactive answers based on the corresponding scene adaptation strategy and in combination with the knowledge graph.

[0026] Example 2 The personalized interaction and scene adaptation system of a digital human AI tour guide assistant described in this embodiment can be used in scenarios such as government affairs, culture and tourism, education, and entertainment. It can be equipped with assembly robots and deployed on multiple terminals.

[0027] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0028] It should be understood that determining B based on A does not mean determining B based solely on A. B can also be determined based on A and / or other information.

[0029] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

[0030] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A personalized interaction and scene adaptation system for a digital human AI tour guide assistant, characterized by: The system includes: user portrait data acquisition module, multimodal feature extraction module, feature fusion module, scene adaptation module, and personalized interaction module; User portrait data acquisition module, used to obtain multimodal data of user portraits at the same time; Multimodal feature extraction module, used to extract features from multimodal data respectively; The feature fusion module is used to perform multimodal feature fusion on the extracted features to generate fused features; and generate intent categories and emotional states based on the fused features; The scenario adaptation module generates the corresponding scenario adaptation strategy based on the intent category and emotional state; The personalized interaction module generates personalized interactive responses based on the corresponding scenario adaptation strategy and combined with the knowledge graph.

2. The personalized interaction and scene adaptation system of a digital human AI tour guide assistant according to claim 1 is characterized in that: The multimodal data includes: voice data, text data, facial image data and body movement data.

3. The personalized interaction and scene adaptation system of a digital human AI tour guide assistant according to claim 1 is characterized in that: The features include: voice features, text features, facial features, and body features.

4. The personalized interaction and scene adaptation system of a digital human AI tour guide assistant according to claim 3 is characterized in that: The speech features include: speaking speed, pitch, and timbre; the text features include: emotional words, negative words, keywords, degree adverbs, pronouns, punctuation marks, and sentences; the facial features include: eye key points, eyebrow key points, mouth key points, and cheek muscle key points; the limb features include: shoulder joint points, elbow joint points, wrist joint points, finger joint points, and hip joint points.

5. The personalized interaction and scene adaptation system of a digital human AI tour guide assistant according to claim 1 is characterized in that: The process of multimodal feature fusion is as follows: Input speech features, text features, facial features, and body features into the multimodal feature fusion model, and map the speech features, text features, facial features, and body features into a common latent space through the encoder layer; The decoder layer performs interactive fusion through the cross-attention mechanism, uses the gradient descent method to minimize the loss function, and uses the back-propagation algorithm to train the multimodal feature fusion model and output the fusion features; The multimodal feature fusion model includes: a Transformer model or an unsupervised learning model.

6. The personalized interaction and scene adaptation system of a digital human AI tour guide assistant according to claim 1 is characterized in that: The process of generating intent categories is as follows: the fusion features are input into the intent classification model, and the intent classification model outputs intent categories as the results. The intent categories include: information query, entertainment, life service, education and learning, feedback and evaluation; The process of generating the emotional state is: inputting the fusion features into the emotional state model, and the emotional state model uses the emotional state as the output result. The emotional state includes: happiness, sadness, anger, surprise, smile and irony.

7. The personalized interaction and scene adaptation system of a digital human AI tour guide assistant according to claim 1 is characterized in that: The process of generating the corresponding scenario adaptation strategy is as follows: writing a query statement based on the scenario knowledge graph database software to query the scenario strategy corresponding to the intent category, and generating the corresponding scenario adaptation strategy based on the emotional state association rules and the user historical behavior association rules.

8. The personalized interaction and scene adaptation system of a digital human AI tour guide assistant according to claim 1 is characterized in that: The process of generating personalized interactive answers is as follows: Generate n groups of initial interactive responses based on the corresponding scenario adaptation strategy and knowledge graph; through the fitness function: ,in, represents the fitness value, Represents the weight to evaluate the fitness value of n groups of initial interactive answers, screen out the initial interactive answers corresponding to higher fitness values ​​and perform i genetic operations, the genetic operations include crossover and mutation operations, and record the fitness values ​​after i genetic operations, and take the interactive answer corresponding to the highest fitness value after i genetic operations as the personalized interactive answer, the personalized interactive answer includes voice, text and digital human actions.

Citation Information

Patent Citations

  • Virtual digital human interaction system based on AI

    CN119902625A

  • Intelligent interaction method and system based on AI digital human

    CN119916947A

Cited By

  • Social medium anti-mental semantic recognition method and system, storage medium and electronic equipment

    CN121093970A