Exhibition hall display method and device based on digital human interaction, equipment and medium

By analyzing the audio of user questions and building display content based on the type of demand, it solves the problem that digital people have difficulty in dealing with multiple user needs, and improves the exhibition hall display efficiency and visiting experience.

CN120234404APending Publication Date: 2025-07-01GUANGZHOU TRF CULTURE & TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510290217.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

Digital people in the existing exhibition halls have difficulty dealing with the various types of problems or needs raised by visitors, resulting in reduced efficiency of exhibition hall display guides and poor visiting experience.

Method used

By obtaining user questions, analyzing the user needs and demand types, selecting different display content construction methods according to different demand types, and driving digital people to display the corresponding content to users.

Benefits of technology

It has enabled digital people to adapt to a variety of user questions and deal with various types of questions or needs, thereby improving the efficiency and quality of digital people displaying in the exhibition hall and improving the visiting experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234404A_ABST
    Figure CN120234404A_ABST
Patent Text Reader

Abstract

The invention discloses an exhibition hall display method and device based on digital human interaction, equipment and a medium, and relates to the field of digital human interaction, and the method comprises the steps: responding to the interaction between a user and a digital human, obtaining a user question audio, and analyzing the user question audio to obtain a first user demand and a first demand type; when the first demand type is a guide question type, matching the first user demand based on a preset knowledge base to obtain a first cue word, and constructing a first display content according to the first cue word; when the first demand type is a picture stylization type, matching to obtain a picture stylization model based on a preset second big model in combination with the first user demand, and constructing second display content according to the picture stylization model; and driving the digital person to display the first display content or the second display content to the user. By implementing the method and the device, the efficiency and the quality of display based on the digital human in the exhibition hall can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of digital human interaction, and in particular to an exhibition hall display method, device, equipment and medium based on digital human interaction. Background Art

[0002] With the development of digital human technology, the application fields of digital humans are also increasing. For example, the introduction of digital humans as exhibition hall guides and exhibition hall customer service in traditional exhibition halls can improve the degree of digitization of exhibition halls, improve the efficiency of exhibition hall display and guidance, and improve the visiting experience. However, the digital humans currently introduced in exhibition halls are usually matched with the question and answer library based on the pre-configured search text to answer questions and conduct exhibition and guidance. Therefore, when visitors ask questions to the digital humans in the exhibition hall, they must narrate the questions in a specific format, and this type of digital human is also difficult to handle the various types of questions or needs raised by visitors, which will reduce the efficiency of the exhibition hall's display and guidance and will also have a negative impact on the visitor's experience. Therefore, under the premise of introducing digital human interaction, how to improve the efficiency and quality of the display in the exhibition hall is still a difficult problem that needs to be solved by existing technologies. Summary of the invention

[0003] The present application provides an exhibition hall display method, device, equipment and medium based on digital human interaction, so as to solve the technical problem that the digital human interaction in the existing exhibition hall has an adverse effect on the exhibition hall display efficiency and quality.

[0004] According to a first aspect of the implementation of the present application, a method for exhibition hall display based on digital human interaction is provided, comprising:

[0005] In response to the interaction between the user and the digital human, the user's question audio is obtained, and the user's question audio is analyzed to obtain a first user demand and a first demand type; wherein the first demand type includes a navigation question type and a picture stylization type;

[0006] When the first demand type is a navigation question type, matching the first user demand based on a preset knowledge base to obtain a first prompt word, and constructing a first display content according to the first prompt word;

[0007] When the first requirement type is a picture stylization type, based on a preset second largest model and in combination with the first user requirement, a corresponding picture stylization model is matched and a second display content is constructed according to the picture stylization model;

[0008] The digital human is driven to display the first display content or the second display content to the user.

[0009] This application first responds to the interaction between the user and the digital human, obtains the user's question audio, and parses it to obtain the first user requirement and the first requirement type. Then, according to different first requirement types, different display content construction methods are selected, and further drives the digital human to display to the user. Compared with the prior art where it is difficult for digital humans to handle various types of user requirements, this application first parses the user's question audio after obtaining it, obtains the first user requirement and the first requirement type, and then selects different processing methods for the first user requirement according to the first requirement type to construct different display contents, and further drives the digital human to display to the user. It can adapt to a variety of user question methods and can also handle various types of questions or requirements raised by users, thereby improving the efficiency and quality of the display based on the digital human in the exhibition hall and enhancing the user's visit experience in the exhibition hall.

[0010] In some embodiments of this application, the step of responding to the interaction between the user and the digital human and obtaining the user's question audio specifically includes:

[0011] Respond to the interaction between the user and the digital human to obtain the user's original audio;

[0012] Calculate the root mean square of the samples of the user's original audio, and filter the user's original audio according to the root mean square of the samples and a preset root mean square threshold to obtain the user's question audio.

[0013] This application first obtains the user's original audio, then calculates its root mean square of the samples, and filters the user's original audio according to the root mean square of the samples to obtain the user's question audio. It can filter out the noise that is not the user's question through the root mean square of the samples, improve the accuracy of the user's question audio, reduce the probability of subsequent processing errors, and thus improve the quality of the display based on the digital human in the exhibition hall.

[0014] In some embodiments of this application, the step of parsing the user's question audio to obtain the first user requirement and the first requirement type specifically includes:

[0015] Based on a speech recognition model, recognize the user's question audio to obtain the user's question text;

[0016] Vectorize the user's question text to obtain the first user requirement;

[0017] Based on a preset requirement keyword, detect the first user requirement according to a similarity algorithm to obtain the first requirement type.

[0018] In this application, first, the user's question audio is recognized based on a speech recognition model to obtain the user's question text, and then the user's question text is vectorized to obtain the first user demand. Vectorization can improve the efficiency of subsequent recognition and processing of the first user demand. Furthermore, based on a similarity algorithm, combined with preset demand keywords, the first user demand is detected to obtain the first demand type. Matching through similarity can improve the accuracy of the detected first demand type, and then improve the efficiency and quality of subsequent processing, thereby improving the efficiency and quality of digital human-based display in the exhibition hall.

[0019] In some embodiments of this application, matching the first user demand based on a preset knowledge base to obtain a first prompt word, and constructing first display content according to the first prompt word specifically includes:

[0020] Performing vector matching on the first user demand based on a preset knowledge base to obtain a first prompt word;

[0021] According to the first prompt word, constructing first display content based on a preset first large model.

[0022] In this application, first, vector matching is performed on the first user demand based on a preset knowledge base to obtain a first prompt word, and according to the first prompt word, first display content is constructed based on the first large model. It can convert the first user demand into a first prompt word that is more suitable for interacting with the first large model through the preset knowledge base, and then obtain first display content that better meets the first user demand, improving the quality of the first display content, thereby improving the quality of digital human-based display in the exhibition hall.

[0023] In some embodiments of this application, constructing first display content according to the first prompt word based on a preset first large model specifically includes:

[0024] When the first prompt word is not empty, obtaining a first answer text corresponding to the first prompt word according to the first large model, and constructing first display content according to the first answer text;

[0025] When the first prompt word is empty, obtaining a second answer text corresponding to the first user demand according to the first large model, and constructing first display content according to the second answer text.

[0026] In this application, according to different situations of the first prompt word obtained by vector matching, choosing to construct first display content based on the first prompt word or the first user demand based on the first large model can adapt to user questions in different situations, and then obtain first display content that better meets the first user demand when constructing first display content based on the first large model, improving the quality of the first display content, thereby improving the quality of digital human-based display in the exhibition hall.

[0027] In some embodiments of the present application, driving the digital human to display the first display content or the second display content to the user specifically includes:

[0028] Based on the speech synthesis technology, synthesizing the first display content into a first display audio;

[0029] Based on the speech driving model, driving the digital human according to the first display audio to generate a first display video, and displaying the first display video to the user.

[0030] In the present application, the first display content is first synthesized into a first display audio based on the speech synthesis technology, and the digital human is driven based on the speech driving model to generate and display the first display video. When the first demand type is a guided tour question type, it can be determined that the display method of the first display content is to synthesize the audio first and then drive the digital human to generate the video, so that the display of the first display content better meets the first user demand, thereby improving the quality of the display based on the digital human in the exhibition hall.

[0031] In some embodiments of the present application, based on a preset second large model, in combination with the first user demand, matching to obtain a corresponding picture stylization model, and constructing the second display content according to the picture stylization model specifically includes:

[0032] Obtaining a first style demand corresponding to the first user demand according to the second large model, and matching to obtain a corresponding picture stylization model according to the first style demand;

[0033] Constructing the second display content based on the first supplementary picture of the user and based on the picture stylization model; wherein, the first supplementary picture is obtained by the user in response to the prompt of the digital human for the first style demand.

[0034] In the present application, the first style demand is first obtained according to the second large model, and then the corresponding picture stylization model is matched, so that a picture stylization model that better meets the first user demand can be obtained. Furthermore, when constructing the second display content in combination with the first supplementary picture of the user, a second display content that better meets the first user demand can be obtained, improving the quality of the second display content, thereby improving the quality of the display based on the digital human in the exhibition hall.

[0035] According to the second aspect of the embodiments of the present application, there is provided an exhibition hall display device based on digital human interaction, including a user question processing module, a first content construction module, a second content construction module, and a digital human display module;

[0036] The user question processing module is used to obtain the user question audio in response to the interaction between the user and the digital human, and parse the user question audio to obtain the first user requirement and the first requirement type; wherein, the first requirement type includes a tour question type and a picture stylization type;

[0037] When the first requirement type is the tour question type, the first content construction module is used to match the first user requirement based on a preset knowledge base to obtain a first prompt word, and construct first display content according to the first prompt word;

[0038] When the first requirement type is the picture stylization type, the second content construction module is used to match a corresponding picture stylization model based on a preset second large model in combination with the first user requirement, and construct second display content according to the picture stylization model;

[0039] The digital human display module is used to drive the digital human to display the first display content or the second display content to the user.

[0040] In some embodiments of the present application, the user question processing module includes a raw audio acquisition unit and an audio filtering and processing unit;

[0041] The raw audio acquisition unit is used to obtain the user raw audio in response to the interaction between the user and the digital human;

[0042] The audio filtering and processing unit is used to calculate the root mean square of the samples of the user raw audio, and filter the user raw audio according to the root mean square of the samples and a preset root mean square threshold to obtain the user question audio.

[0043] In some embodiments of the present application, the user question processing module includes a question audio recognition unit, a text vectorization unit and a requirement type detection unit;

[0044] The question audio recognition unit is used to recognize the user question audio based on a speech recognition model to obtain the user question text;

[0045] The text vectorization unit is used to vectorize the user question text to obtain the first user requirement;

[0046] The requirement type detection unit is used to detect the first user requirement based on a similarity algorithm according to a preset requirement keyword to obtain the first requirement type.

[0047] In some embodiments of the present application, the first content construction module includes a prompt word matching unit and a first display construction unit;

[0048] The prompt word matching unit is used to perform vector matching on the first user requirement based on a preset knowledge base to obtain a first prompt word;

[0049] The first display construction unit is used to construct first display content based on the first prompt word and a preset first large model.

[0050] In some embodiments of the present application, the first display construction unit includes a first processing construction subunit and a second processing construction subunit;

[0051] The first processing construction subunit is used to, when the first prompt word is not empty, obtain a first answer text corresponding to the first prompt word according to the first large model, and construct first display content according to the first answer text;

[0052] The second processing construction subunit is used to, when the first prompt word is empty, obtain a second answer text corresponding to the first user requirement according to the first large model, and construct first display content according to the second answer text.

[0053] In some embodiments of the present application, the digital human display module includes a display audio synthesis unit and a display video synthesis unit;

[0054] The display audio synthesis unit is used to synthesize the first display content into first display audio based on speech synthesis technology;

[0055] The display video synthesis unit is used to drive the digital human according to the first display audio based on a speech-driven model, generate first display video, and display the first display video to the user.

[0056] In some embodiments of the present application, the second content construction module includes a style requirement matching unit and a second display construction unit;

[0057] The style requirement matching unit is used to obtain a first style requirement corresponding to the first user requirement according to the second large model, and match a corresponding picture stylization model according to the first style requirement;

[0058] The second display construction unit is used to construct second display content based on the user's first supplementary picture and the picture stylization model; wherein, the first supplementary picture is input by the user in response to the digital human's prompt for the first style requirement.

[0059] This application first responds to the interaction between the user and the digital human, obtains the user's question audio, and parses it to obtain the first user requirement and the first requirement type. Then, according to different first requirement types, different display content construction methods are selected, and further drives the digital human to display to the user. Compared with the prior art in which it is difficult for digital humans to handle various types of user requirements, this application first parses the user's question audio after obtaining it to obtain the first user requirement and the first requirement type, and then selects different processing methods for the first user requirement according to the first requirement type to construct different display contents, and further drives the digital human to display to the user. It can adapt to a variety of user question methods and can also handle various types of questions or requirements raised by users, thereby improving the efficiency and quality of the display based on digital humans in the exhibition hall and enhancing the user's visiting experience in the exhibition hall.

[0060] According to the third aspect of the embodiments of the present application, a computer device is provided, including: a processor; a memory; a computer program stored in the memory and configured to be executed by the processor; wherein when the processor executes the computer program, it implements the method for exhibition hall display based on digital human interaction described in the present application.

[0061] According to the fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores multiple instructions, and the instructions are suitable for being loaded by the processor to execute the method for exhibition hall display based on digital human interaction described in the present application. Description of the Drawings

[0062] Figure 1 : It is a flowchart showing a method for exhibition hall display based on digital human interaction shown in some embodiments of the present application;

[0063] Figure 2 : It is a module structure diagram of a device for exhibition hall display based on digital human interaction shown in some embodiments of the present application. Detailed Embodiments

[0064] The following details the embodiments of the present application. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by combining the drawings are exemplary and are only used to explain some embodiments of the present application, and cannot be understood as a limitation to the embodiments of the present application. Based on the embodiments shown in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0065] In the description of the present application, it should be understood that the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of the present application, unless otherwise specifically defined, the meaning of "a plurality" or "several" is two or more.

[0066] When the digital human currently applied to the exhibition hall interacts with visitors, it usually matches with a pre-configured retrieval text and Q&A library to answer questions and conduct exhibition navigation. Due to the relatively fixed retrieval format and stored content of its retrieval text and Q&A library and the difficulty in updating, it is difficult to handle various types of questions or demands raised by visitors. At the same time, visitors must describe questions in a specific format when asking questions. This will lead to a reduction in the efficiency of exhibition hall display navigation and also have an adverse impact on the experience of visitors. Therefore, on the premise of introducing digital human interaction, how to improve the efficiency and quality of the display in the exhibition hall is still a difficult problem that the existing technology urgently needs to solve.

[0067] Based on the above technical background, please refer to Figure 1 , an embodiment of the present application provides an exhibition hall display method based on digital human interaction, including steps S101 to S104, and the specific steps are as follows:

[0068] Step S101: In response to the interaction between the user and the digital human, obtain the user's question audio, and parse the user's question audio to obtain the first user demand and the first demand type; wherein, the first demand type includes a tour question type and a picture stylization type.

[0069] In some embodiments of the present application, the obtaining the user's question audio in response to the interaction between the user and the digital human specifically includes:

[0070] In response to the interaction between the user and the digital human, obtain the user's original audio;

[0071] Calculate the root mean square (RMS) of the samples of the user's original audio, and filter the user's original audio according to the RMS and a preset RMS threshold to obtain the user's question audio.

[0072] Specifically, the calculating the root mean square of the samples of the user's original audio is specifically: classifying and summarizing the user's original audio by channel to obtain sample audio data; calculating the root mean square of the sample audio data as the root mean square of the samples of the user's original audio. Generally, the calculation method of the root mean square of the samples is specifically:

[0073]

[0074] Wherein, RMS is the root mean square of the sample audio data, N is the total number of samples, and X i is the i-th sample value.

[0075] In some embodiments of the present application, filtering the user's original audio according to the root mean square of the sample and a preset root mean square threshold to obtain the user's question audio specifically includes:

[0076] Eliminating and filtering the audio in the user's original audio whose root mean square of the sample is not greater than the preset root mean square threshold to obtain the user's question audio.

[0077] In some embodiments of the present application, a preferred embodiment of the preset root mean square threshold is 0.05.

[0078] The present application first obtains the user's original audio, then calculates its root mean square of the sample, and filters the user's original audio according to the root mean square of the sample to obtain the user's question audio, which can filter out the noise that is not the user's question through the root mean square of the sample, improve the accuracy of the user's question audio, reduce the probability of subsequent processing errors, and thus improve the quality of the display based on the digital human in the exhibition hall.

[0079] In some embodiments of the present application, parsing the user's question audio to obtain the first user demand and the first demand type specifically includes:

[0080] Based on a speech recognition model, recognizing the user's question audio to obtain the user's question text;

[0081] Vectorizing the user's question text to obtain the first user demand;

[0082] Detecting the first user demand based on a similarity algorithm according to a preset demand keyword to obtain the first demand type.

[0083] In some embodiments of the present application, the speech recognition model includes but is not limited to the Whisper model, the SpeechBrain model, the PaddleSpeech model, and the iFlytek ASR recognition model, and the preferred embodiment is the iFlytek ASR recognition model.

[0084] In some embodiments of the present application, the similarity algorithm includes but is not limited to the cosine similarity algorithm, the semantic similarity algorithm, the edit distance algorithm, and the Jaccard similarity coefficient algorithm, and the preferred embodiment is the cosine similarity algorithm.

[0085] In this application, first, the user's question audio is recognized based on a speech recognition model to obtain the user's question text, and then the user's question text is vectorized to obtain the first user demand. Vectorization can improve the efficiency of subsequent recognition and processing of the first user demand. Then, based on a similarity algorithm, combined with preset demand keywords, the first user demand is detected to obtain the first demand type. Matching through similarity can improve the accuracy of the detected first demand type, and further improve the efficiency and quality of subsequent processing, thereby improving the efficiency and quality of digital human-based display in the exhibition hall.

[0086] Step S102: When the first demand type is a guided tour question type, match the first user demand based on a preset knowledge base to obtain a first prompt word, and construct first display content according to the first prompt word.

[0087] In some embodiments of this application, the matching of the first user demand based on a preset knowledge base to obtain a first prompt word and constructing first display content according to the first prompt word specifically includes:

[0088] Perform vector matching on the first user demand based on a preset knowledge base to obtain a first prompt word;

[0089] According to the first prompt word, construct first display content based on a preset first large model.

[0090] In some embodiments of this application, the first large model includes but is not limited to GPT, Wenxin Yiyan, iFlytek Spark, Tongyi Qianwen, or an LLM Q&A model fine-tuned by Lora. The preferred embodiment is an LLM Q&A model fine-tuned by Lora.

[0091] This application first performs vector matching on the first user demand based on a preset knowledge base to obtain a first prompt word, and constructs first display content based on the first prompt word and the first large model. It can, through the preset knowledge base, transform the first user demand into a first prompt word that is more suitable for interacting with the first large model, and then obtain first display content that better meets the first user demand, improving the quality of the first display content, thereby improving the quality of digital human-based display in the exhibition hall.

[0092] In some embodiments of this application, the constructing first display content according to the first prompt word and based on a preset first large model specifically includes:

[0093] When the first prompt word is not empty, obtain a first answer text corresponding to the first prompt word according to the first large model, and construct first display content according to the first answer text;

[0094] When the first prompt word is empty, obtain the second response text corresponding to the first user requirement according to the first large model, and construct the first display content according to the second response text.

[0095] In some embodiments of the present application, a simple implementation manner of constructing the first display content according to the first response text or the second response text is to directly use the first response text or the second response text as the first display content.

[0096] According to different situations of the first prompt word obtained by vector matching in the present application, select to construct the first display content based on the first prompt word or the first user requirement based on the first large model, which can adapt to user questions in different situations, and then obtain the first display content that better meets the first user requirement when constructing the first display content based on the first large model, improving the quality of the first display content, thereby improving the quality of the display based on the digital human in the exhibition hall.

[0097] Step S103: When the first requirement type is the picture stylization type, based on a preset second large model, combine the first user requirement to match the corresponding picture stylization model, and construct the second display content according to the picture stylization model.

[0098] In some embodiments of the present application, the method of matching the corresponding picture stylization model based on the preset second large model and combining the first user requirement, and constructing the second display content according to the picture stylization model specifically includes:

[0099] Obtain the first style requirement corresponding to the first user requirement according to the second large model, and match the corresponding picture stylization model according to the first style requirement;

[0100] Construct the second display content based on the first supplementary picture of the user and the picture stylization model; wherein, the first supplementary picture is input by the user in response to the prompt of the digital human for the first style requirement.

[0101] In some embodiments of the present application, the second large model includes but is not limited to GPT, Wenxin Yiyan, iFlytek Spark, Tongyi Qianwen or the LLM Q&A model fine-tuned by Lora, and the preferred implementation manner is the LLM Q&A model fine-tuned by Lora.

[0102] The present application first obtains the first style requirement according to the second large model, and then matches the corresponding picture stylization model, which can obtain a picture stylization model that better meets the first user requirement. Furthermore, when constructing the second display content in combination with the first supplementary picture of the user, a second display content that better meets the first user requirement can be obtained, improving the quality of the second display content, thereby improving the quality of the display based on the digital human in the exhibition hall.

[0103] Step S104: Drive the digital human to display the first display content or the second display content to the user.

[0104] In some embodiments of the present application, the driving the digital human to display the first display content or the second display content to the user specifically includes:

[0105] Based on the speech synthesis technology, synthesize the first display content into the first display audio;

[0106] Based on the speech driving model, drive the digital human according to the first display audio to generate the first display video, and display the first display video to the user.

[0107] In some embodiments of the present application, the speech synthesis technology includes but is not limited to ChatTTS tool, Seed-TTS model, FishSpeech model, and Edge-tts tool, and the preferred embodiment is the Edge-tts tool.

[0108] In some embodiments of the present application, the speech driving model includes but is not limited to VideoReTalking model, Sonic model, and Wav2lip model, and the preferred embodiment is the Wav2lip model.

[0109] The present application first synthesizes the first display content into the first display audio based on the speech synthesis technology, and drives the digital human to generate and display the first display video based on the speech driving model. When the first demand type is the guided tour question type, it can be determined that the display method of the first display content is to synthesize the audio first and then drive the digital human to generate the video, making the display of the first display content more in line with the first user demand, thereby improving the quality of the display based on the digital human in the exhibition hall.

[0110] Compared with the prior art, the present application first responds to the interaction between the user and the digital human, obtains the user's question audio, and parses to obtain the first user demand and the first demand type. Then, according to different first demand types, it selects different display content construction methods, and further drives the digital human to display to the user. Compared with the prior art where it is difficult for the digital human to handle various types of user demands, the present application first parses after obtaining the user's question audio to obtain the first user demand and the first demand type, and then selects different processing methods for the first user demand according to the first demand type to construct different display content, and further drives the digital human to display to the user. It can adapt to a variety of user question methods and can also handle various types of questions or demands raised by the user, thereby improving the efficiency and quality of the display based on the digital human in the exhibition hall and enhancing the user's visit experience in the exhibition hall.

[0111] Corresponding to the foregoing method, please refer to Figure 2 An exhibition hall display device based on digital human interaction provided by an embodiment of the present application includes a user question processing module 210, a first content construction module 220, a second content construction module 230, and a digital human display module 240;

[0112] The user question processing module 210 is configured to obtain a user question audio in response to the interaction between the user and the digital human, and parse the user question audio to obtain a first user demand and a first demand type; wherein, the first demand type includes a guided tour question type and a picture stylization type;

[0113] When the first demand type is the guided tour question type, the first content construction module 220 is configured to match the first user demand based on a preset knowledge base to obtain a first prompt word, and construct first display content according to the first prompt word;

[0114] When the first demand type is the picture stylization type, the second content construction module 230 is configured to match a corresponding picture stylization model based on a preset second large model in combination with the first user demand, and construct second display content according to the picture stylization model;

[0115] The digital human display module 240 is configured to drive the digital human to display the first display content or the second display content to the user.

[0116] In some embodiments of the present application, the user question processing module 210 includes an original audio acquisition unit and an audio filtering and processing unit;

[0117] The original audio acquisition unit is configured to obtain user original audio in response to the interaction between the user and the digital human;

[0118] The audio filtering and processing unit is configured to calculate the root mean square of the samples of the user original audio, and filter the user original audio according to the root mean square of the samples and a preset root mean square threshold to obtain a user question audio.

[0119] In some embodiments of the present application, the user question processing module 210 includes a question audio recognition unit, a text vectorization unit, and a demand type detection unit;

[0120] The question audio recognition unit is configured to recognize the user question audio based on a speech recognition model to obtain a user question text;

[0121] The text vectorization unit is configured to vectorize the user question text to obtain a first user demand;

[0122] The demand type detection unit is configured to detect the first user demand based on a similarity algorithm according to a preset demand keyword, and obtain a first demand type.

[0123] In some embodiments of the present application, the first content construction module 220 includes a prompt word matching unit and a first display construction unit;

[0124] The prompt word matching unit is configured to perform vector matching on the first user demand based on a preset knowledge base to obtain a first prompt word;

[0125] The first display construction unit is configured to construct first display content based on the first prompt word and a preset first large model.

[0126] In some embodiments of the present application, the first display construction unit includes a first processing construction subunit and a second processing construction subunit;

[0127] The first processing construction subunit is configured to, when the first prompt word is not empty, obtain a first answer text corresponding to the first prompt word according to the first large model, and construct first display content according to the first answer text;

[0128] The second processing construction subunit is configured to, when the first prompt word is empty, obtain a second answer text corresponding to the first user demand according to the first large model, and construct first display content according to the second answer text.

[0129] In some embodiments of the present application, the digital human display module 240 includes a display audio synthesis unit and a display video synthesis unit;

[0130] The display audio synthesis unit is configured to synthesize the first display content into first display audio based on speech synthesis technology;

[0131] The display video synthesis unit is configured to drive the digital human according to the first display audio based on a voice driving model, generate first display video, and display the first display video to the user.

[0132] In some embodiments of the present application, the second content construction module 230 includes a style demand matching unit and a second display construction unit;

[0133] The style demand matching unit is configured to obtain a first style demand corresponding to the first user demand according to the second large model, and match a corresponding picture stylization model according to the first style demand;

[0134] The second display construction unit is used to construct second display content based on the first supplementary picture of the user and the picture stylization model; wherein, the first supplementary picture is input by the user in response to the prompt of the digital human for the first style requirement.

[0135] This application first responds to the interaction between the user and the digital human, obtains the user's question audio, and parses it to obtain the first user requirement and the first requirement type. Then, according to different first requirement types, different display content construction methods are selected, and further, the digital human is driven to display to the user. Compared with the prior art in which digital humans are difficult to handle various types of user requirements, this application first parses the user's question audio after obtaining it to obtain the first user requirement and the first requirement type, and then selects different processing methods for the first user requirement according to the first requirement type to construct different display content, and further drives the digital human to display to the user. It can adapt to a variety of user question methods and can also handle various types of questions or requirements raised by users, thereby improving the efficiency and quality of the display based on digital humans in the exhibition hall and enhancing the user's visit experience in the exhibition hall.

[0136] It should be understood that the device provided by the embodiment of this application corresponds to the foregoing method. A display device for an exhibition hall based on digital human interaction provided by the embodiment of this application can implement a method for displaying an exhibition hall based on digital human interaction provided by any embodiment of this application.

[0137] Adaptively, the embodiment of this application also provides a computer device and a computer-readable storage medium.

[0138] The computer device includes: a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor;

[0139] Wherein, when the processor executes the computer program, it implements a method for displaying an exhibition hall based on digital human interaction of this application.

[0140] The computer-readable storage medium stores multiple instructions, and the instructions are suitable for being loaded by the processor to execute a method for displaying an exhibition hall based on digital human interaction of this application.

[0141] The above is part of the embodiments of this application. The purpose, technical solution, and beneficial effects of this application have been further described in detail. It should be clear that the above part of the embodiments of this application should not be understood as a limitation of this application. In particular, for those skilled in the art, any changes, modifications, equivalent replacements, and variations made within the spirit and principle of this application should be included within the protection scope of this application.

Claims

1. A method for exhibition hall display based on digital human interaction, characterized in that: include: In response to the interaction between the user and the digital human, the user's question audio is obtained, and the user's question audio is analyzed to obtain a first user demand and a first demand type; wherein the first demand type includes a navigation question type and a picture stylization type; When the first demand type is a navigation question type, matching the first user demand based on a preset knowledge base to obtain a first prompt word, and constructing a first display content according to the first prompt word; When the first requirement type is a picture stylization type, based on a preset second largest model and in combination with the first user requirement, a corresponding picture stylization model is matched and a second display content is constructed according to the picture stylization model; The digital human is driven to display the first display content or the second display content to the user.

2. The exhibition hall display method based on digital human interaction according to claim 1, characterized in that: The step of obtaining the audio of the user's question in response to the interaction between the user and the digital human specifically includes: Responding to the interaction between the user and the digital human, obtaining the user's original audio; The sample root mean square of the user's original audio is calculated, and the user's original audio is filtered according to the sample root mean square and a preset root mean square threshold to obtain the user question audio.

3. The exhibition hall display method based on digital human interaction according to claim 1, characterized in that: The step of parsing the user question audio to obtain the first user demand and the first demand type specifically includes: Based on the speech recognition model, the user question audio is recognized to obtain the user question text; Vectorizing the user question text to obtain a first user demand; According to the preset demand keywords, the first user demand is detected based on a similarity algorithm to obtain a first demand type.

4. The exhibition hall display method based on digital human interaction according to claim 1, characterized in that: The matching the first user demand based on the preset knowledge base to obtain a first prompt word, and constructing a first display content according to the first prompt word specifically includes: Performing vector matching on the first user demand based on a preset knowledge base to obtain a first prompt word; According to the first prompt word, based on a preset first large model, a first display content is constructed.

5. The exhibition hall display method based on digital human interaction according to claim 4 is characterized in that: The step of constructing the first display content according to the first prompt word and based on a preset first large model specifically includes: When the first prompt word is not empty, obtaining a first answer text corresponding to the first prompt word according to the first large model, and constructing a first display content according to the first answer text; When the first prompt word is empty, a second answer text corresponding to the first user demand is obtained according to the first large model, and a first display content is constructed according to the second answer text.

6. The exhibition hall display method based on digital human interaction according to claim 4, characterized in that: The driving the digital human to display the first display content or the second display content to the user specifically includes: Based on speech synthesis technology, synthesize the first display content into a first display audio; Based on the voice-driven model, the digital human is driven according to the first display audio to generate a first display video, and the first display video is displayed to the user.

7. The exhibition hall display method based on digital human interaction according to claim 1, characterized in that: The matching of the preset second largest model with the first user demand to obtain a corresponding image stylization model, and constructing the second display content according to the image stylization model, specifically includes: Acquire a first style requirement corresponding to the first user requirement according to the second large model, and obtain a corresponding image stylization model according to the first style requirement; According to the user's first supplementary picture and based on the picture stylization model, second display content is constructed; wherein the first supplementary picture is input by the user in response to the digital human's prompt for the first style requirement.

8. An exhibition hall display device based on digital human interaction, characterized in that: It includes a user question processing module, a first content construction module, a second content construction module and a digital human display module; The user question processing module is used to obtain the user question audio in response to the interaction between the user and the digital human, and analyze the user question audio to obtain the first user demand and the first demand type; wherein the first demand type includes the navigation question type and the image stylization type; The first content construction module is used to match the first user demand based on a preset knowledge base to obtain a first prompt word when the first demand type is a navigation question type, and construct a first display content according to the first prompt word; The second content construction module is used for, when the first demand type is a picture stylization type, matching a corresponding picture stylization model based on a preset second large model and in combination with the first user demand, and constructing second display content according to the picture stylization model; The digital human display module is used to drive the digital human to display the first display content or the second display content to the user.

9. A computer device, characterized in that: include: processor; Memory; a computer program stored in the memory and configured to be executed by the processor; Wherein, when the processor executes the computer program, an exhibition hall display method based on digital human interaction as described in any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the exhibition hall display method based on digital human interaction as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Exhibition hall voice interaction method and system

    CN120823832A

  • Exhibition hall voice interaction method and system

    CN120823832B