Intelligent question answering system and method based on AI digital human interactive museum explanation service
By adopting an intelligent question-and-answer system based on AI digital interactive explanation service in the museum, the problem that traditional explanation services are difficult to meet the audience's personalized and diversified needs is solved, high-quality explanation services and multi-language support are achieved, and visiting experience and management efficiency is improved.
Patent Information
- Application Number
- CN202510140237.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-05-27
AI Technical Summary
Traditional museum explanation services are difficult to meet the personalized and diversified needs of the audience, especially for hearing-impaired people and foreign tourists, and language and cultural differences lead to reduced visit experience and satisfaction.
It adopts an intelligent question-and-answer system based on AI digital human interactive museum explanation service, and realizes real-time and accurate explanation services and intelligent question-and-answer functions through the integration of artificial intelligence technology, digital human technology, multilingual translation and sign language action translation technology.
It improves the quality of museum explanations and audience visiting experience, solves the problems of language and cultural differences, provides multilingual support and sign language action translation, reduces operational costs and improves management efficiency.
Smart Images

Figure CN120045673A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to an intelligent question - answering system and method for an AI digital human interactive museum explanation service. Background Art
[0002] The explanation services of traditional museums mainly rely on human docents or static text and picture displays. Although human docents can provide rich and in - depth explanations, due to limited human resources and time and cost constraints, it is difficult to provide customized services for each viewer. Especially for special groups such as the hearing - impaired, it is difficult for them to obtain information through traditional voice - guided systems, while foreign tourists may face challenges of language and cultural differences. These factors may all lead to a reduction in their visiting experience and satisfaction. With the growing demand of viewers for personalized and diversified experiences, the traditional explanation service model has been difficult to meet these needs.
[0003] In such a background, an intelligent question - answering system for an AI digital human interactive museum explanation service has emerged, which represents an important trend in the digital transformation of museums. This service model is an innovative result jointly driven by artificial intelligence, digital human technology, and market demand. By closely integrating AI technology with digital human images and using natural language processing (NLP) and deep - learning technology, the system can achieve intelligent question - answering and personalized services. It can provide multilingual support and sign - language gesture translation, bringing a real - time interactive digital human explanation experience to viewers. It not only improves the quality of museum explanations, enhances the visiting experience of viewers, but also helps to reduce operating costs and improve management efficiency. The AI digital human interactive explanation service provides strong support for the digital transformation of museums and the improvement of viewer experience, and is a key direction for museum service innovation and future development. Summary of the Invention
[0004] In view of the above situation, embodiments of the present application propose an intelligent question - answering system and method for an AI digital human interactive museum explanation service. By integrating artificial intelligence technology, digital human technology, multilingual translation, and sign - language gesture translation technology, real - time and accurate explanation services and intelligent question - answering functions are realized, so as to meet the diverse and refined needs of viewers, improve service quality, reduce operating costs, and improve management efficiency.
[0005] In a first aspect, embodiments of the present application provide an intelligent question - answering system for an AI digital human interactive museum explanation service, including: a multimodal input receiving unit, a user intention recognition module, a vertical - domain intelligent question - answering unit, a digital human answer video generation module, and a human - machine interaction interface module;
[0006] A multi-modal input receiving unit, configured to collect input information conveyed by a user through input paths such as voice, text, or gesture, and send the collected user input information to a user intention recognition module;
[0007] A user intention recognition unit, which analyzes and processes the user input information through speech recognition, natural language processing, and gesture recognition technologies, and extracts a text description of the user intention;
[0008] A vertical domain intelligent question-answering unit, based on the text description expressing the user intention as input, searches for text answers related to the text description, refines the text answers, and obtains text answers related to the user query;
[0009] A digital human answer video generation module, which converts the text answer into speech and gesture actions, and generates a video in real time through the multi-modal behavior of the virtual digital human.
[0010] In some embodiments, the user intention recognition unit has a speech recognition module, configured to receive the voice audio input by the user, perform transcription processing on the corresponding audio, and extract a text description of the user intention.
[0011] In some embodiments, the user intention recognition unit further has a sign language action recognition unit, which performs real-time detection and classification on the collected sign language video, recognizes the sign language action, and converts the collected sign language action into a text description expressing the user intention.
[0012] In some embodiments, the sign language action recognition unit further includes:
[0013] A video processing module, configured to perform frame processing on the collected sign language action video and extract key frames;
[0014] A human body posture recognition module, which recognizes gesture features based on the key frame images;
[0015] A sign language classification recognition module, configured to classify the gesture features and obtain a gesture classification result;
[0016] A gesture translation module, based on the gesture classification result, converts the gesture into a corresponding text description.
[0017] In some embodiments, there is also a document translation module, which translates the text description of the user intention recognized by the user intention recognition unit.
[0018] In some embodiments, the vertical domain intelligent question-answering unit includes:
[0019] Knowledge base construction, which uses the knowledge of the vertical domain to build a local knowledge base and continuously updates the domain knowledge to ensure the timeliness of the knowledge;
[0020] A knowledge base maintenance module, configured to parse and split knowledge documents in a local knowledge base, convert the content of the knowledge documents into vector data by using a vectorization model, and store the converted vector data in a vector database;
[0021] A knowledge base recall module, configured to retrieve text answers related to a search Q&A question from the vector database based on an input text description expressing the user's intention, and recall these relevant text answers;
[0022] A large model service module, configured to receive the recalled text answers, assemble and merge them, and then input them into a museum vertical domain large model, and use the deep learning ability of the museum vertical domain large model to analyze and summarize the recalled text answers, so as to extract text answers related to the user's query.
[0023] In some embodiments, the digital human answer video generation module further has:
[0024] A text-to-speech module, which converts the text answer into speech;
[0025] A text-to-sign language module, which converts the text answer into sign language actions.
[0026] In some embodiments, the digital human answer video generation module further has a digital human rendering module, which is responsible for real-time rendering of the digital human, generating digital human frame data, and presenting the image and actions of the digital human to tourists.
[0027] In some embodiments, there is also a human-computer interaction interface module, which is used for presenting the image and actions of the digital human, and through a front-end interaction page, presenting a video of digital human interaction to tourists.
[0028] In a second aspect, an intelligent Q&A method based on an AI digital human interactive museum explanation service provided by an embodiment of the present application includes:
[0029] Collect the input information conveyed by the user through voice, text or gesture input paths, and send the collected user input information to a user intention recognition module;
[0030] Parse and process the user input information through voice recognition, natural language processing and gesture recognition technologies, and extract a text description of the user's intention;
[0031] Based on the input text description expressing the user's intention, search for text answers related to this text description, refine the text answers, and obtain text answers related to the user's query;
[0032] Convert the text answer into speech and sign language actions, and generate a video in real time through the multi-modal behavior of the virtual digital human.
[0033] The above at least one technical solution adopted in the embodiments of the present application can achieve the following beneficial effects:
[0034] 1. Transcending the traditional explanation mode: Traditional museum explanations mainly rely on human guides or static text and picture displays. However, this patent combines AI and digital human technologies to achieve intelligent and personalized explanation services, transcending the limitations of the traditional explanation mode.
[0035] 2. Construction and training of a deep model in the vertical field of museums: Through deep learning technology, combined with a large amount of literature and data in the museum field, a deep model specifically for museum content is constructed, providing accurate and rich knowledge support for the intelligent question-answering system.
[0036] 3. Multilingual interaction ability: It can overcome language barriers, support real-time translation and interaction in multiple languages, enabling audiences from different countries and regions to use it easily, greatly expanding the audience scope of the museum. Multilingual interaction not only helps with information transmission but also promotes understanding and communication between different cultures, enhancing the international influence of the museum.
[0037] 4. Sign language action translation technology: Combining human motion capture and machine translation technologies, it realizes real-time translation and generation of sign language actions. The system can convert the knowledge content of the explanation into sign language actions, building a barrier-free communication bridge for the hearing-impaired and solving many inconveniences caused by communication barriers during visits. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0039] Figure 1 Shows a system architecture diagram according to an embodiment of the present application.
[0040] Figure 2 Shows a flowchart according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0042] The following will detail the technical solutions provided by each embodiment of the present application in conjunction with the drawings.
[0043] In this application, the interactive digital human conducts real-time conversations with customers and answers questions. Artificial Intelligence (AI) uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, including theories, methods, technologies, and application systems for perceiving the environment, acquiring knowledge, and using knowledge to obtain the best results. Artificial intelligence technologies mainly include several major directions such as computer vision technology, robotics technology, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning. This embodiment mainly relates to natural language processing technology and deep learning technology.
[0044] The digital human question and answer method provided in the embodiments of this application can be applied in an application environment such as Figure 1 . In this environment, the user communicates with the server through the client. The client includes, but is not limited to, various personal computers, laptop computers, smartphones, tablets, and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers, and is used to generate digital humans and perform data interaction with the client.
[0045] The embodiments of this application propose an intelligent question and answer system for an AI digital human interactive museum explanation service, including: a multimodal input receiving unit, a user intention recognition module, a vertical domain intelligent question and answer unit, and a digital human answer video generation module;
[0046] The multimodal input receiving unit is used to collect the input information conveyed by the user through the input paths of voice, text, or gestures, and send the collected user input information to the user intention recognition module.
[0047] The input information refers to the information entered by the user through the operation page of the client. The operation page is, for example, a user interaction interface. The question input information is related to the user's question and answer needs, and refers to the question text information generated according to the content of the question input by the user. The client provides an interaction interface between the user and the digital human, which can be a web chat window or a mobile application, etc. The target user can input questions through the client, and can input questions in ways such as voice input, text input, or selecting existing question options, gestures, etc., and send the initial question information to the server.
[0048] The user intention recognition unit analyzes and processes the user input information through speech recognition, natural language processing, and gesture recognition technologies, and extracts a text description of the user intention.
[0049] The user intention recognition unit is located on the server. The user intention recognition unit responds to the user's input path, selects the corresponding recognition technology to analyze and process the user input information, and extracts a text description of the user intention.
[0050] The vertical domain intelligent Q&A unit searches for text answers related to the text description expressing the user's intention based on the input text description. The digital human answer video generation module converts the text answer into speech and gesture actions, and generates a video in real time through the multimodal behavior of the virtual digital human.
[0051] The functions of each module are specifically described below.
[0052] The multimodal input receiving unit has a user input interface, providing an information input path for personalized services. The information input path includes voice, text, or gesture; the user clicks to select the information input path option on the user input interface, and collects the input information input by the user through information input paths such as voice, text, or gesture. The input information includes voice audio, text information, and gesture actions.
[0053] In one embodiment, the user intention recognition unit has a speech recognition module, which is used to receive the voice audio input by the user, perform transcribing processing on the corresponding audio, and extract the text description of the user intention.
[0054] Combined with Figure 2 The specific implementation process of the speech recognition module of the present invention is further described:
[0055] The tourist selects one of the input paths of the user input interface, question input. The tourist asks a question through voice. The multimodal input receiving unit obtains the tourist's voice, starts the speech recognition module, performs transcribing processing on the corresponding audio, and extracts the text description of the user intention. Specifically: the user intention recognition unit recognizes that the input path is voice input, and starts the speech recognition module to convert the tourist's voice question into the text description of the user intention.
[0056] The user intention recognition unit also has a sign language action recognition unit. In one embodiment of the present application, special groups (hearing-impaired people) input through sign language actions. The multimodal input receiving unit obtains the tourist's gestures. The sign language action recognition unit performs real-time detection and classification on the collected sign language video, accurately recognizes the sign language actions, and converts the collected sign language actions into the text description expressing the user intention. The specific implementation process of the sign language action recognition unit of the present invention is further described as follows:
[0057] In one embodiment, the sign language action recognition unit has a video acquisition module, which needs to acquire the sign language video to be recognized. Specifically: the terminal detects the start operation triggered by the sign language recognition page; in response to this start operation, the built-in camera is started to capture frame images. The terminal captures the frame images of the target object when using sign language through the built-in camera, and combines the captured frame images into a sign language video. When detecting the stop operation triggered on the user interaction interface; in response to this stop operation, the built-in camera is turned off to stop capturing frame images.
[0058] In one embodiment, the sign language action recognition unit of the present application includes: a video processing module, a human body posture recognition module, a sign language classification recognition module, and a gesture translation module.
[0059] The video processing module is used to perform frame processing on the acquired sign language action video and extract key frames;
[0060] For example, using the video processing tool ffmpeg, perform frame processing on the acquired sign language action video and extract key frame images;
[0061] The human body posture recognition module recognizes gesture features based on the key frame images;
[0062] Exemplarily, the human body posture recognition module of the present application can use the OpenPose model to recognize the human body posture of the person in the video. The human body posture includes the coordinate points of the hand position key points, the angles of adjacent joints, the hand shape features, etc.
[0063] The sign language classification recognition module is used to classify the gesture features and obtain the gesture classification result;
[0064] The sign language classification recognition module is built based on the YOLO model; the YOLO model (You Only Look Once) is an algorithm for object detection using a convolutional neural network. Specifically; it is used to input the feature vector group of the gesture features into a pre-established gesture template library and perform one-to-one matching recognition with the pre-stored gesture data to obtain the gesture classification result.
[0065] The specific training steps of the YOLO model are as follows:
[0066] a) Data collection and preparation, collect and prepare the sign language data set, and the sign language data set includes a training set and a validation set.
[0067] Exemplarily collect and prepare the sign language data set, use the Asl_Videos data set and the Roboflow data set, collect video and picture samples of different sign language gestures, and use 90% of the data as the training set and 10% of the data as the validation set;
[0068] b) Model training: Use GPU to accelerate the training process and tune hyperparameters such as the learning rate, batch size, and number of iterations of the model to obtain the best training effect.
[0069] c) Model evaluation: Evaluate the accuracy and robustness of the model by calculating metrics such as accuracy, recall, and F1-score (the harmonic mean of accuracy and recall), and verify the model's ability to recognize sign language letters.
[0070] d) Model testing: Test the model on a real-time video stream, turn on the camera and detect sign language letters in the video in real-time to verify the model's performance in an actual scenario.
[0071] Test the recognition results of the model on different gestures and evaluate the accuracy and real-time performance of the model.
[0072] The gesture translation module converts the gesture into the corresponding text description based on the gesture classification result.
[0073] Specifically, the category number of the hand movement is converted into the corresponding text description through a predefined mapping relationship.
[0074] The intelligent question-answering system also has a document translation module that translates the text description of the user intention recognized by the user intention recognition unit.
[0075] The intelligent question-answering system also has an optimization module that further processes the translated text description, analyzes grammar and semantics to ensure the accuracy and fluency of the translation, and obtains the optimized text description.
[0076] In one embodiment, the vertical domain intelligent question-answering unit searches for text answers related to the text description based on the input text description expressing the user intention, and refines the text answers to obtain text answers related to the user's query.
[0077] The vertical domain intelligent question-answering unit includes:
[0078] Knowledge base construction: Use the knowledge of the vertical domain to build a local knowledge base and continuously update the domain knowledge to ensure the timeliness of the knowledge.
[0079] The knowledge base maintenance module is configured to parse and split the knowledge documents of the local knowledge base, convert the content of the knowledge documents into vector data using a vectorization model, and store the converted vector data in a vector database.
[0080] Selection of the vectorization model: The unsupervised SimCSE is used to obtain text representations. SimCSE is a simple contrastive learning framework for training sentence embeddings. The training process is as follows: First, the same sentence is input into a pre-trained language model (RoBERTa is used in this embodiment) twice. Since the Dropout layer randomly masks some neurons, two different sentence vectors can be obtained. Then, the two outputs of the same sentence are used as positive examples, and the outputs of other sentences are used as negative examples. The training objective is to make the similarity between positive examples higher and the similarity between negative examples lower. Finally, in this embodiment, the cross-entropy loss is used to optimize the model parameters so that the model can learn better sentence vector representations.
[0081] Selection of the vector database: Since Milvus has absolute advantages in terms of large scale, retrieval performance, community influence, etc., Milvus is selected as the vector database in this embodiment.
[0082] The knowledge base recall module is configured to retrieve text answers related to the search Q&A questions from the vector database based on the input text description expressing the user's intention and recall these relevant text answers.
[0083] The large model service module is configured to receive the recalled text answers, assemble and merge them, and then input them into the museum vertical domain large model. Utilize the deep learning ability of the museum vertical domain large model to analyze and summarize the text answers, so as to extract text answers related to the user's query.
[0084] The recalled text answers are sent into the museum vertical domain large model. The museum vertical domain large model uses advanced language processing models (such as GPT, Baichuan, ChatGLM) to analyze and summarize the recalled content. For example, this module will comprehensively recall the text answers and extract the retrieval content (text answers) about the museum's private domain knowledge base.
[0085] The large model in the museum vertical domain is based on large language models and has stronger semantic understanding: Large models (such as large language models) perform excellently in semantic understanding, being able to better understand the intent of user queries rather than relying solely on keyword matching. This makes search results more accurate and relevant. Personalization and context consideration: Large models can better understand context and user personality, thus providing more personalized and precise search results. It can consider the user's past search history and conversation context to provide information that better meets the user's needs. The superiority of the vector knowledge base: The vector knowledge base can better represent text information, making search and matching more precise. Through vector encoding, documents and queries can be represented in a high-dimensional space, facilitating similarity matching and improving the accuracy of search. Multimodal support: Large models can not only process text but also, to some extent, understand multimodal inputs such as images and speech. This provides a wider range of application scenarios for search and question-and-answer involving multiple information forms. Quick adaptation to new knowledge: Large models have a certain ability of transfer learning and can, to a certain extent, quickly adapt to new knowledge. This is very beneficial for the dynamic update of the knowledge base and adapting to the ever-changing information environment.
[0086] The training steps of the large model in the museum vertical domain described in this application are as follows:
[0087] 1) Set training parameters: Set hyperparameters such as learning rate, batch size, and number of iterations according to the data characteristics;
[0088] 2) Batch training: Use the prepared dataset to perform batch training on the model and observe the numerical changes during the training process;
[0089] 3) Mini-batch training: Divide the dataset into multiple small batches and train one small batch each time to improve training efficiency and help the model learn more detailed local features. By adjusting the batch size, balance the training speed and model performance;
[0090] 4) Transfer learning: Use the museum dataset to fine-tune the pre-trained model to make it adapt to the characteristics of the museum vertical domain and quickly adapt to the museum vertical domain through transfer learning;
[0091] 5) Second training: On the basis of transfer learning, use a larger-scale or more finely annotated museum dataset to further train the model, focusing on the model's understanding and expression ability of museum-specific knowledge and context semantic understanding ability;
[0092] 6) Model fine-tuning: For specific tasks or scenarios (such as detailed explanations of specific exhibits and interactive needs of specific audience groups), fine-tune the model to make the model more accurately meet the actual application requirements;
[0093] 7) Hyperparameter Tuning: Utilize Bayesian statistical methods to dynamically adjust the search direction based on historical training results, quickly converging to the optimal hyperparameter combination to continuously improve the training effect;
[0094] 8) Model Evaluation: Select appropriate evaluation metrics (accuracy, recall, F1-score) according to task requirements. Adopt the K-fold cross-validation method to divide the museum dataset into K subsets. Then the model is trained and validated K times on K different training sets, and finally the mean of the K validation results is used as the evaluation result of the model to ensure the generalization ability of the model on different subsets;
[0095] 9) Iterative Optimization: According to the evaluation results, adjust the model structure, training strategy or hyperparameters, and conduct multiple rounds of iterative training until a satisfactory performance level is achieved.
[0096] The digital human answer video generation module converts the text answer into speech and gesture actions, and generates videos in real time through the multimodal behaviors of the virtual digital human.
[0097] The digital human answer video generation module also has:
[0098] The text-to-speech module converts the text answer into speech;
[0099] The text-to-speech module uses NLP technology and generates natural and fluent text answers based on the Transformer-based model.
[0100] The text-to-sign language module converts the text answer into sign language actions.
[0101] The text-to-sign language module includes:
[0102] 1) Text Preprocessing: Segment the text answer generated by the large model according to speech rules, and at the same time perform part-of-speech tagging and syntactic analysis;
[0103] 2) Sign Language Vocabulary Mapping: Based on the constructed sign language dictionary, map the vocabulary in the text answer to sign language symbols;
[0104] 3) Sign Language Grammar Construction: Based on the grammar rules of sign language, construct the sign language sentence structure, and adjust the tense and body posture of the sign language sentence to conform to the sign language expression habits;
[0105] 4) Action Generation: Encode the sign language symbols into an action sequence, and generate gesture actions according to the action sequence.
[0106] The digital human answer video generation module also has a digital human rendering module, which is responsible for real-time rendering of the digital human, generating digital human frame data, and presenting the image and actions of the digital human to the visitors.
[0107] The human-computer interaction interface module outputs the content of the interactive digital human, integrates the generated content, and through the front-end interactive page, shows the video of the digital human interaction to tourists, ensuring that the digital human can interact with tourists in a consistent and attractive manner.
[0108] Combined with Figure 1 The various modules and their functions of the system architecture of the present invention are further described:
[0109] The system structure framework has:
[0110] Front-end display layer: 1) Digital human module, which is modeled based on the video of a specific person image in the museum, generates a 2.5D AI digital human with highly realistic appearance features and limb movements, simulates the behaviors and expressions of real docents, and provides vivid and natural interpretation services; 2) User interaction interface module, which provides an interface for visiting tourists to interact with the system, including functions such as language selection, sign language selection, question input, etc.
[0111] Business logic layer, which has a multilingual translation module, a sign language translation module, an intelligent question-answering module, and an AI digital human control module.
[0112] Data storage and management layer, including:
[0113] 1) Knowledge base module: Stores knowledge information related to the museum, including exhibits introductions, historical backgrounds, etc., for the intelligent question-answering module to query and use;
[0114] 2) User data module: Stores data such as users' basic information and historical behaviors, for personalized recommendation and service optimization;
[0115] 3) Log management module: Records information such as the running status of the system and users' operations, for fault troubleshooting and performance optimization;
[0116] The system also includes that the back-end service layer includes:
[0117] 1) Service scheduling module: Responsible for the scheduling and management of the AI digital human interpretation service, ensuring the efficient operation of the service; 2) Performance monitoring module: Real-time monitors the performance indicators of the system, such as response time, concurrent user number, etc., ensuring the stable operation of the system; 3) Security management module: Ensures the security of the system, including functions such as data encryption and access control, filters and avoids illegal information at three levels of the problem side, model side, and answer side, ensuring that sensitive information does not leak out and the output information is not illegal;
[0118] It also includes an infrastructure layer: 1) Computing resources: including computing devices such as servers and GPUs, which are used to run each module of the system; 2) Storage resources: including storage devices such as disk arrays and cloud storage, which are used to store the data and logs of the system; 3) Network resources: including network devices such as routers and switches, which are used to implement the network communication of the system.
[0119] This application also provides an intelligent question-answering method for an AI digital human interactive museum explanation service, including:
[0120] Collect the input information conveyed by the user through the input paths of voice, text or gestures, and send the collected user input information to the user intention recognition module;
[0121] The user intention recognition unit parses and processes the user input information through speech recognition, natural language processing and gesture recognition technologies, and extracts the text description of the user intention;
[0122] The vertical domain intelligent question-answering unit searches for text answers related to the text description based on the input text description expressing the user intention, refines the text answers, and obtains the text answers related to the user query;
[0123] The digital human answer video generation module converts the text answers into speech and gesture actions, and generates videos in real time through the multi-modal behaviors of the virtual digital human.
[0124] The above are only the embodiments of this application and are not used to limit this application. For those skilled in the art, various changes and modifications can be made to this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this application shall be included within the scope of the claims of this application.
Claims
1. An intelligent question-answering system based on AI digital human interactive museum explanation service, characterized in that: include: Multimodal input receiving unit, user intention recognition module, vertical field intelligent question and answer unit, digital human answer video generation module and human-computer interaction interface module; A multimodal input receiving unit, used to collect input information conveyed by the user through an input path of voice, text or gesture, and send the collected user input information to the user intention recognition module; The user intention recognition unit analyzes and processes user input information through speech recognition, natural language processing, and sign language recognition technology to extract a text description of the user's intention; The vertical field intelligent question-answering unit searches for text answers related to the text description that expresses the user's intention based on the input, refines the text answers, and obtains text answers related to the user's query; The digital human answer video generation module converts text answers into voice and gesture actions, and generates videos in real time through the multimodal behavior of the virtual digital human.
2. The intelligent question-answering system based on AI digital human interactive museum explanation service as claimed in claim 1, characterized in that: The user intention recognition unit has a speech recognition module for receiving speech audio input by the user, transcribing the corresponding audio, and extracting a text description of the user's intention.
3. The intelligent question-answering system based on AI digital human interactive museum explanation service as claimed in claim 1, characterized in that: The user intention recognition unit also has a sign language action recognition unit, which detects and classifies the collected sign language video in real time, recognizes the sign language action, and converts the collected sign language action into a text description expressing the user's intention.
4. The intelligent question-answering system based on AI digital human interactive museum explanation service as claimed in claim 3, characterized in that: The sign language action recognition unit also includes: The video processing module is used to perform frame processing on the collected sign language action video and extract key frames; Human gesture recognition module, based on key frame images, identifies gesture features; A sign language classification and recognition module is used to classify gesture features and obtain gesture classification results; The gesture translation module converts gestures into corresponding text descriptions based on the gesture classification results.
5. The intelligent question-answering system based on AI digital human interactive museum explanation service as claimed in claim 2, 3 or 4, characterized in that: It also has a document translation module for translating the text description of the user intention identified by the user intention identification unit.
6. The intelligent question-answering system based on AI digital human interactive museum explanation service as claimed in claim 5, characterized in that: The vertical field intelligent question answering unit includes: Knowledge base construction: using vertical domain knowledge to build a local knowledge base and continuously updating domain knowledge; A knowledge base maintenance module is configured to parse and split the knowledge documents of the local knowledge base, convert the knowledge document content into vector data using a vectorization model, and store the converted vector data in a vector database; A knowledge base recall module configured to retrieve text answers related to the search question and answer question from a vector database based on an input text description expressing the user's intention, and recall the related text answers; The big model service module is configured to receive the recalled text answers, assemble and merge them, and input them into the museum vertical field big model, and use the deep learning ability of the museum vertical field big model to analyze and summarize the recalled text answers, so as to extract the text answers related to the user query.
7. The intelligent question-answering system based on AI digital human interactive museum explanation service as claimed in claim 6, characterized in that: The digital human answer video generation module also has: Text-to-speech module, which converts text answers into speech; The text-to-sign language module converts text responses into sign language actions.
8. The intelligent question-answering system based on AI digital human interactive museum explanation service as claimed in claim 7, characterized in that: The digital human answer video generation module also has a digital human rendering module, which is responsible for real-time rendering of the digital human, generating digital human frame data, and presenting the image and movements of the digital human to tourists.
9. The intelligent question-answering system based on AI digital human interactive museum explanation service as claimed in claim 7, characterized in that: It also has a human-computer interaction interface module, which is used to present the image and actions of the digital human, and to show the visitors the video of the digital human interaction through the front-end interactive page.
10. An intelligent question-answering method based on AI digital human interactive museum explanation service, characterized in that: include: Collecting input information conveyed by the user through an input path of voice, text or gesture, and sending the collected user input information to a user intention recognition module; Parse and process user input information through speech recognition, natural language processing, and gesture recognition technologies to extract a text description of the user's intention; Based on the input text description expressing the user's intention, searching for text answers related to the text description, refining the text answers, and obtaining text answers related to the user's query; Convert text answers into speech and gesture actions, and generate videos in real time through the multimodal behavior of virtual digital humans.
Citation Information
Patent Citations
Scenic region explanation method and device based on sign language explanation robot
CN115759083A
Digital human customer service question answering method and device, computer equipment and storage medium
CN117609463A
Search question-answering system and method based on large model and electronic equipment
CN117708274A
Intelligent building intercom system and intercom control method
CN119255128A
Cited By
Customer service man-machine intelligent interaction system and method based on communication object information
CN120780156A