Virtual avatar interaction method and related apparatus, device, system and medium
Through a virtual avatar interaction system and synthesis engine, video stream synthesis was achieved that synchronizes the mouth and body movements of virtual avatars with the timing of their speech. This supports real-time processing of user interaction requests, solves the problem of poor naturalness in virtual avatar interactions, and improves the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- IFLYTEK CO LTD
- Filing Date
- 2022-12-20
- Publication Date
- 2026-07-31
AI Technical Summary
The existing virtual avatars lack naturalness during interaction, which hinders their promotion and application.
The system employs a virtual avatar interaction system, including an interactive terminal, an interactive response server, and an information processing server. Through voice and image data interaction, combined with a virtual avatar synthesis engine, it achieves video stream synthesis with the virtual avatar's mouth and body movements synchronized with the voice timing, supporting real-time interruption and resumption of user interaction requests.
It enhances the naturalness of virtual avatar interactions, and improves the fun and interactive effects of the user experience.
Smart Images

Figure CN116088675B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a virtual avatar interaction method and related devices, equipment, systems and media. Background Technology
[0002] With the development of artificial intelligence technology, virtual avatars have been applied in many industries such as education and entertainment. For example, in the entertainment industry, virtual avatars are already being used to perform singing, dancing, and other entertainment acts for the public; or, in cultural and heritage fields such as museums and memorial halls, the practical application of virtual avatars is gradually attracting attention.
[0003] However, existing virtual avatars often suffer from poor naturalness in their interactions, hindering their widespread adoption. Therefore, improving the naturalness of virtual avatar interactions has become an urgent problem to be solved. Summary of the Invention
[0004] The main technical problem addressed by this application is to provide a virtual avatar interaction method and related devices, equipment, systems, and media that can improve the naturalness of virtual avatar interaction.
[0005] To address the aforementioned technical problems, this application provides a virtual avatar interaction system, comprising an interactive terminal, an interactive response server, and an information processing server. The interactive terminal is communicatively connected to the interactive response server, and the interactive response server is communicatively connected to the information processing server. The information processing server is equipped with an information system for retrieving information during interactive decision-making. The interactive terminal is used to interact with a user to obtain user input data and to acquire and play a video stream from the interactive response server. The input data includes at least one of voice data and image data. The interactive response server is used to make interactive decisions based on the input data, obtaining an interactive decision result. The interactive decision result includes time-synchronized interactive text and action instructions. Based on the synthesized speech of the interactive text and the action instructions, a video stream is synthesized, and the virtual avatar's mouth movements in the video stream are time-sequentially consistent with the synthesized speech, and its body movements are time-sequentially consistent with the action instructions.
[0006] To address the aforementioned technical problems, a second aspect of this application provides a method for testing an interactive system, used to test the virtual avatar interactive system described in the first aspect. The method includes: inputting test data to a test driver interface of the interactive terminal in the virtual avatar interactive system; wherein, when the test data is video data, it is split into audio data and image data by the test driver interface; acquiring sampled data related to test indicators during the interactive response process of the virtual avatar interactive system based on the test data; obtaining the test values of the virtual avatar interactive system on the test indicators based on the sampled data; and determining whether the virtual avatar interactive system has passed the test based on the test values of the virtual avatar interactive system on each test indicator.
[0007] To address the aforementioned technical problems, a third aspect of this application provides a virtual avatar interaction method, comprising: acquiring and playing a first video stream; wherein an interaction response server generates a first interaction decision in response to a first interaction request issued by a user through an interaction terminal, and synthesizes the first video stream in real time using a virtual avatar synthesis engine based on the first interaction decision, wherein the interaction response server uses a keyword in the first interaction request as a marker to indicate whether playback should resume after interruption; in response to a second interaction request from the user while playing the first video stream, sending an interruption synthesis request and a second interaction request to the interaction response server; wherein the interaction response server suspends the synthesis of the first video stream in response to the interruption synthesis request, and synthesizes the second video stream in real time in response to the second interaction request, and after the second video stream is synthesized, determines, based on the marker, whether to continue synthesizing a new first video stream from the interruption position of the first interaction decision; and acquiring and playing the newly synthesized video stream from the interaction response server.
[0008] To address the aforementioned technical problems, a fourth aspect of this application provides a virtual avatar interaction method, comprising: generating a first interaction decision based on a first interaction request issued by an interactive terminal; synthesizing a first video stream using a virtual avatar synthesis engine based on the first interaction decision; and marking the first video stream with a flag indicating whether playback should resume after interruption based on keywords in the first interaction request; wherein the interactive terminal acquires and plays the first video stream; pausing the synthesis of the first video stream in response to an interruption synthesis request issued by the interactive terminal; synthesizing a second video stream in real time in response to a second interaction request issued by the interactive terminal; and determining, based on the flag, whether to continue synthesizing a new first video stream from the interruption position of the first interaction decision after the second video stream synthesis is completed; wherein the interruption synthesis request is sent by the interactive terminal in response to a second interaction request from a user while playing the first video stream, and the interactive terminal acquires and plays the newly synthesized video stream.
[0009] To address the aforementioned technical problems, a fifth aspect of this application provides an interactive system testing apparatus for testing the virtual avatar interactive system described in the first aspect. The apparatus includes an input module, an acquisition module, a calculation module, and a determination module. The input module is used to input test data to the test driver interface of the interactive terminal in the virtual avatar interactive system; wherein, when the test data is video data, it is split into audio data and image data by the test driver interface. The acquisition module is used to acquire sampled data related to test indicators during the interactive response process of the virtual avatar interactive system based on the test data. The calculation module is used to obtain the test values of the virtual avatar interactive system on the test indicators based on the sampled data. The determination module is used to determine whether the virtual avatar interactive system has passed the test based on the test values of the virtual avatar interactive system on each test indicator.
[0010] To address the aforementioned technical problems, a sixth aspect of this application provides a virtual avatar interaction device, comprising: a first acquisition module, a request sending module, and a second acquisition module. The first acquisition module is used to acquire and play a first video stream; wherein, an interaction response server generates a first interaction decision in response to a first interaction request issued by a user through an interaction terminal, and synthesizes the first video stream in real time through a virtual avatar synthesis engine based on the first interaction decision; the interaction response server uses a keyword in the first interaction request as a marker to indicate whether playback should resume after interruption. The request sending module is used to send an interruption synthesis request and a second interaction request to the interaction response server in response to a second interaction request from the user while playing the first video stream; wherein, the interaction response server pauses the synthesis of the first video stream in response to the interruption synthesis request, and synthesizes the second video stream in real time in response to the second interaction request; and after the second video stream is synthesized, it determines, based on the marker, whether to continue synthesizing a new first video stream from the interruption position of the first interaction decision. The second acquisition module is used to acquire and play the newly synthesized video stream from the interaction response server.
[0011] To address the aforementioned technical problems, a seventh aspect of this application provides a virtual avatar interaction device, comprising: a request processing module and an interruption-resumption module. The request processing module is configured to generate a first interaction decision based on a first interaction request issued by an interactive terminal, and to synthesize a first video stream using a virtual avatar synthesis engine based on the first interaction decision, and to use a flag in the first interaction request to mark the first video stream as an indicator of whether to resume playback after interruption; wherein the interactive terminal acquires and plays the first video stream. The interruption-resumption module is configured to pause the synthesis of the first video stream in response to an interruption synthesis request issued by the interactive terminal, and to synthesize a second video stream in real time in response to a second interaction request issued by the interactive terminal, and to determine, after the second video stream synthesis is completed, whether to continue synthesizing a new first video stream from the interruption position of the first interaction decision based on the flag; wherein the interruption synthesis request is sent by the interactive terminal in response to a second interaction request from the user while playing the first video stream, and the interactive terminal acquires and plays the newly synthesized video stream.
[0012] To address the aforementioned technical problems, the eighth aspect of this application provides an interactive terminal, comprising: a communication circuit, a memory, and a processor. The communication circuit and the memory are respectively coupled to the processor. The memory stores program instructions, and the processor executes the program instructions to implement the virtual avatar interaction method in the second aspect described above.
[0013] To address the aforementioned technical problems, the ninth aspect of this application provides an interactive response server, comprising: a communication circuit, a memory, and a processor. The communication circuit and the memory are respectively coupled to the processor. The memory stores program instructions, and the processor executes the program instructions to implement the virtual avatar interaction method in the third aspect described above.
[0014] To address the aforementioned technical problems, the tenth aspect of this application provides a computer-readable storage medium storing program instructions executable by a processor, the program instructions being used to implement the virtual avatar interaction method described in the first or second aspect above.
[0015] The above scheme acquires and plays a first video stream. The interactive response server responds to a user's first interactive request via an interactive terminal, generates a first interactive decision, and synthesizes the first video stream in real-time using a virtual avatar synthesis engine based on the first interactive decision. The interactive response server uses keywords from the first interactive request to mark the first video stream, indicating whether playback should resume after interruption. Furthermore, in response to a user's second interactive request while playing the first video stream, the server sends an interruption synthesis request and a second interactive request to the interactive response server. The interactive response server pauses the synthesis of the first video stream in response to the interruption request and synthesizes the second video stream in real-time in response to the second interactive request. After the second video stream is synthesized, based on the marker, it determines whether to resume synthesis of a new first video stream from the interruption point of the first interactive decision. This allows the acquisition and playback of the newly synthesized video stream from the interactive response server. Thus, during the real-time synthesis of the video stream by the interactive response server and the playback of the video stream by the interactive terminal, the synthesis can be interrupted by a new user interactive request. A new video stream is synthesized in real-time first, and then the resumption of synthesis of the original video stream is determined based on the resume signal. Therefore, this significantly improves the naturalness of the interaction between the cultural relics virtual avatar and the user. Attached Figure Description
[0016] Figure 1 This is a schematic diagram of the framework of an embodiment of the virtual avatar interaction system of this application;
[0017] Figure 2 This is a schematic diagram of the framework of one implementation method of the virtual character synthesis engine;
[0018] Figure 3 This is a flowchart illustrating an embodiment of the virtual avatar interaction method of this application;
[0019] Figure 4 This is a flowchart illustrating another embodiment of the virtual avatar interaction method of this application;
[0020] Figure 5 This is a flowchart illustrating yet another embodiment of the virtual avatar interaction method of this application;
[0021] Figure 6 This is a schematic diagram of the framework of an embodiment of the virtual avatar interaction device of this application;
[0022] Figure 7 This is a schematic diagram of the framework of another embodiment of the virtual avatar interaction device of this application;
[0023] Figure 8 This is a schematic diagram of the framework of an embodiment of the interactive terminal of this application;
[0024] Figure 9 This is a schematic diagram of the framework of an embodiment of the interactive response server of this application;
[0025] Figure 10 This is a flowchart illustrating an embodiment of the interactive system testing method of this application;
[0026] Figure 11 This is a schematic diagram of the framework of an embodiment of the interactive system testing device of this application;
[0027] Figure 12 This is a schematic diagram of a framework of an embodiment of the computer-readable storage medium of this application. Detailed Implementation
[0028] The embodiments of this application will now be described in detail with reference to the accompanying drawings.
[0029] In the following description, specific details such as particular system architectures, interfaces, and technologies are presented for illustrative purposes rather than for limiting purposes, in order to provide a thorough understanding of this application.
[0030] In this paper, the terms "system" and "network" are often used interchangeably. The term "and / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship. Furthermore, "many" in this paper means two or more.
[0031] Please see Figure 1 , Figure 1 This is a schematic diagram of the framework of an embodiment of the virtual avatar interaction system of this application. Figure 1 As shown, the virtual avatar interaction system includes an interactive terminal, an interactive response server, and an information processing server. The interactive terminal is communicatively connected to the interactive response server, and the interactive response server is communicatively connected to the information processing server. The information processing server is equipped with an information system for the interactive response server to retrieve information. The specific structure of the above equipment can be found in the following related embodiments, and will not be elaborated upon here. Figure 1 The circled numbers indicate the data flow of the cultural heritage virtual avatar system. Specifically:
[0032] (1) The circled number 1 indicates that the interactive terminal has acquired input data. Specifically, the text-based interactive terminal can interact with the user and acquire the user's input data, which may include, but is not limited to, voice data, image data, etc. Image data may include, but is not limited to, facial images, gesture images, etc. In other words, in practical applications, users can interact with the interactive terminal through voice, face, gestures, etc. Furthermore, as... Figure 1As shown, the interactive terminal may include, but is not limited to, a voice wake-up interface, a face wake-up interface, a gesture recognition interface, and a test-driven interface. The voice wake-up interface is used to wake up the interactive terminal when a wake-up word is detected in the voice data, thereby enabling the display of a virtual avatar and interaction with the user. The face wake-up interface is used to wake up the interactive terminal when a registered face is detected, thereby enabling the display of a virtual avatar and interaction with the user. The gesture recognition interface is used to identify gesture categories and provide the identified gesture categories to the interaction decision interface in the interaction response server. For example, the face wake-up interface can specifically perform face detection, preprocessing, feature extraction, and matching and recognition operations on face images. Specifically, face detection can determine the position and size of the face; preprocessing can extract a local image containing the face; feature extraction can then be performed, allowing the extracted face features to be searched and matched with feature templates stored in the database. A threshold is pre-set; if the similarity exceeds this threshold, the face image can be determined to be a registered face, thus waking up the interactive terminal. Furthermore, for the implementation principle of the aforementioned voice wake-up interface, please refer to the technical details of voice wake-up; for the implementation principle of the aforementioned gesture recognition interface, please refer to the technical details of gesture recognition. These details will not be repeated here.
[0033] (2) The circled number 2 indicates that the interactive terminal submits voice data to the speech recognition interface in the interactive response server. The speech recognition interface is used to recognize the voice data, obtain the recognized text, and use it as input data for the semantic understanding interface in the interactive response server. It should be noted that the speech recognition interface can use GMM-HMM, recurrent neural networks, end-to-end models of deep learning, etc., and is not limited here. For the implementation principle of speech recognition, please refer to the technical details of GMM-HMM, recurrent neural networks, end-to-end models of deep learning, etc., which will not be elaborated here.
[0034] (3) The circled number 3 indicates that the speech recognition interface passes the recognized text to the semantic understanding interface. The semantic understanding interface is used to understand the recognized text, extract the user's interaction intent for the cultural heritage scenario, and provide the understood interaction intent to the interaction decision interface in the interaction response server. It should be noted that the semantic understanding interface can use gated recurrent units, long short-term memory networks, etc., and is not limited here. For the implementation principle of semantic understanding, please refer to the technical details of gated recurrent units, long short-term memory networks, etc., which will not be elaborated here.
[0035] (4) The circled number 4 indicates that the semantic understanding interface transmits the parsed interaction intent to the interaction decision interface, and the circled number 7 indicates that the gesture recognition interface in the interaction terminal transmits the recognized gesture category to the interaction decision interface. The circled number 5 indicates that the interaction decision interface submits a query request or personal information operation request to the information system in the information processing server based on the above interaction intent and gesture category. The circled number 6 indicates that the interaction decision interface obtains the response information returned by the information system and performs decision processing based on it to obtain the interaction decision result. More specifically, the interaction decision result may include time-synchronized interaction text and action instructions. It should be noted that the information system collects, stores and processes relevant information. Taking the cultural heritage scenario as an example, the relevant information is cultural heritage information, which may include, but is not limited to: relevant information about cultural heritage exhibits (such as the historical origin, technology and cultural value of cultural heritage exhibits), personal information of registered users, etc., which are not limited here. In addition, in order to facilitate querying in the information system, the above cultural heritage information can be stored in a structured form (such as a knowledge graph). Of course, in practical applications, since users usually interact with the interactive terminal by voice, the interactive decision interface can at least retrieve the response message from the information system in the information processing server based on the interaction intent, and perform decision processing based on the response message to obtain the interactive decision result.
[0036] (5) The circled number 8 indicates that the speech synthesis interface receives interactive text. The speech synthesis interface is used to synthesize speech based on interactive text, obtain synthesized speech, and provide input data to the image synthesis interface in the interactive response server.
[0037] (6) The circled number 9 indicates that the image synthesis interface receives synthesized speech, and the circled number 10 indicates that the image synthesis interface receives action commands. The image synthesis interface integrates a virtual image synthesis engine, which is used to generate a video stream driven by at least one of synthesized speech and action commands. In the video stream, the mouth movements of the virtual image are consistent with the synthesized speech in timing, and the body movements are consistent with the action commands in timing.
[0038] Therefore, the interactive response server can make interactive decisions based on input data and obtain interactive decision results. The interactive decision results can include time-synchronized interactive text and action instructions. Based on the synthesized speech and action instructions of the interactive text, a video stream is synthesized. As mentioned above, the mouth movements of the virtual image in the video stream are consistent with the synthesized speech in time, and the body movements are consistent with the action instructions in time.
[0039] Specifically, please refer to Figure 2 , Figure 2 This is a schematic diagram illustrating the framework of one implementation method of a virtual avatar synthesis engine. For example... Figure 2As shown, sample videos can be pre-collected to train the virtual avatar synthesis engine. For example, 0.5 hours, 1.5 hours, etc., of sample videos can be pre-collected, without limitation. Based on this, deep learning of real human voices and facial expressions (e.g., lip movements, cheek movements) can be performed using artificial intelligence technology to establish the correlation between audio features and facial expressions, thereby obtaining the virtual avatar synthesis engine. Of course, to further improve the applicability of the virtual avatar synthesis engine, after collecting the sample videos, time-synchronized action commands (e.g., the anchor reaching out or making gestures in the sample videos) can be extracted. While performing deep learning of real human voices and facial expressions, further deep learning of action commands and body movements can be performed using artificial intelligence technology to establish the correlation between action features and body movements, thereby further improving the virtual avatar synthesis engine. Based on this, in practical applications, interactive text can be synthesized through a speech synthesis interface to obtain synthesized speech. The sound information of the synthesized speech can be extracted and combined with the virtual avatar synthesis engine to drive facial expressions (e.g., lip movements, cheek movements, etc.). Figure 2 (Including facial expression sequences). Furthermore, when action commands synchronized with the synthesized speech exist, the aforementioned audio information and action commands can be combined with the virtual avatar synthesis engine to drive both facial expressions and body movements. Based on this, the aforementioned image sequence and synthesized speech can be aligned on the timeline, achieving parallel processing of the two channels (i.e., the image channel and the speech channel), thereby generating a video stream where the visuals and audio are consistent with the spoken and spoken actions.
[0040] In one implementation scenario, the aforementioned facial expressions can be combined with 3D face reconstruction. 3D face reconstruction involves displaying and modeling the face, parametrically representing facial pose, ID, and expression, while lip shape generation can be controlled through parameterized facial expression movements. It's important to note that because it's parametric modeling, information such as facial pose, ID, and expression can be controlled using parameters, allowing for further facial editing operations such as beautification, face slimming, and costume changes. Furthermore, due to the 3D nature of the face, its presentation and application scenarios can be combined with CG technology to expand into AR / VR and other fields.
[0041] In one implementation scenario, when the interactive terminal transmits data to the interactive response server, it can also simultaneously transmit the exhibition theme of the exhibition area / hall where the interactive terminal is located. The speech synthesis interface can then further combine this exhibition theme with speech synthesis, ensuring that the final synthesized speech matches the exhibition theme. For example, when the exhibition theme is related to children, the synthesized speech can be lively and childlike; or, when the exhibition theme is related to history, the synthesized speech can be deep and resonant; or, when the exhibition theme is related to daily life, the synthesized speech can be relaxed and straightforward. Other cases can be deduced similarly, and will not be listed here.
[0042] In one implementation scenario, as a possible approach, compared to Face2face, which can only transfer facial expressions from the source video to the target video without controlling head posture, this embodiment uses audio-driven technology to directly predict the position of each vertex of the 3D face via speech, or to drive 3D facial skeletal animation via speech; alternatively, a set of parameters can be used to control the 3D face to produce different expressions. This involves defining a correspondence between phonemes and visual pixels, where visual pixels can include, but are not limited to, lip shapes. This correspondence represents the lip shapes corresponding to different phonemes. Therefore, by using voice-driven technology to drive facial movements, the expressions of virtual avatars can be made natural and fluid.
[0043] (7) The circled number 11 indicates that the interactive response server pushes the video stream to the interactive terminal, or the interactive terminal actively pulls the video stream from the interactive response server. Based on this, the interactive terminal can play the video stream to enable interaction with users through virtual avatars in cultural heritage scenarios, thereby enhancing the interactive experience.
[0044] (8) The circled number 12 indicates that during the testing of the virtual avatar interaction system, test data (such as the aforementioned voice data, face images, gesture images, etc.) is input into the test-driven interface. When the test data is video data, the test-driven interface splits it into audio data and image data, which flow through the virtual avatar interaction system. That is, the audio data will be sent to the voice wake-up interface and then processed by the voice recognition interface, etc., while the image data will be sent to the face wake-up interface, gesture recognition interface, etc., and then processed by the interaction decision interface, etc. For the specific flow process, please refer to the relevant descriptions of the circled numbers 1 to 11 above, so as to realize the testing of the virtual avatar interaction system.
[0045] In an implementation scenario, when testing a virtual avatar interaction system, basic metrics such as speech recognition, speech synthesis, virtual avatar synthesis, interaction success rate, and response time need to be tested. Additionally, metrics such as voice wake-up success rate, face wake-up success rate, voice interruption success rate, and gesture recognition can also be tested.
[0046] In a specific implementation scenario, the following requirements must be met for the "speech recognition" indicator: (1) support for near-field audio processing; (2) support for at least one of command word recognition and continuous speech recognition; (3) in a low-noise environment, the speech recognition sentence recognition accuracy is greater than or equal to 85%; and (4) in a high-noise environment, the speech recognition sentence recognition accuracy is greater than or equal to 80%. It should be noted that the low-noise environment and high-noise environment can be set according to the actual application situation. For example, a high-noise environment can specifically refer to an environment with noise intensity above 60dB, while a low-noise environment can specifically refer to an environment with noise intensity below 45dB.
[0047] In a specific implementation scenario, the following requirements must be met for the indicator "speech synthesis": (1) support for volume, speech rate, and tone adjustment; (2) an average sentence synthesis accuracy rate greater than or equal to 90%, and a synthesis accuracy rate of professional terms for cultural heritage scenarios greater than or equal to 95%; (3) a normalized accuracy rate greater than or equal to 85%, and the normalized accuracy rate is tested in the following two dimensions: symbol pronunciation (i.e., correctly pronouncing the symbols, and the symbols refer to non-native language characters) and number pronunciation (i.e., correctly pronouncing the numbers). It should be noted that the sentence synthesis accuracy rate is used to evaluate the accuracy rate of the system's speech broadcast. The formula for calculating the sentence synthesis accuracy rate is: Sentence synthesis accuracy rate = Number of correctly broadcast speech items / Total speech conditions * 100%. In addition, the formula for calculating the normalized accuracy rate is: Speech conditions with correct pronunciation of numbers and symbols / Total number of speech items * 100%.
[0048] In a specific implementation scenario, the following requirements must be met for the indicator "virtual avatar synthesis": (1) Support for 2D and 3D virtual avatars; (2) Video stream synthesis real-time rate greater than or equal to 1.0. It should be noted that the calculation formula for the video stream synthesis real-time rate is: Video stream synthesis real-time rate = Sum of real-time rates of n video streams / n. Among them, the calculation formula for the real-time rate of a single video stream synthesis is: P = L / T, where P represents the real-time rate of a single video stream synthesis, L represents the duration of a single synthesized voice, and T represents the duration of a single video stream synthesis.
[0049] In a specific implementation scenario, for the metric "interaction success rate," since the virtual avatar interaction system requires that the interaction objective be achieved within a predetermined number of interaction rounds, the interaction can be considered successful; otherwise, it is considered a failure. Therefore, the formula for calculating the interaction success rate can be expressed as: P = S / (S+F)*100%. Here, P represents the interaction success rate, S represents the number of successful interactions, and F represents the number of failed interactions. Furthermore, the metric "interaction success rate" must meet the requirement of being greater than or equal to 90%.
[0050] In a specific implementation scenario, the "response time" metric refers to the entire response time of the virtual avatar during interaction, from the moment the user stops speaking to the moment the virtual avatar begins to respond. The "response time" metric must meet the requirement of being less than or equal to 2 seconds. The formula for calculating response time can be expressed as: T = TE - TS. Where T represents the response time, TS represents the moment the user stops speaking, and TE represents the moment the virtual avatar begins to respond.
[0051] In a specific implementation scenario, the following requirements must be met for the indicator "voice wake-up success rate": In a low-noise environment of a cultural heritage site, the voice wake-up success rate is greater than or equal to 90%, and the false wake-up frequency is less than or equal to 0.2 times / hour; in a high-noise environment of a cultural heritage site, the voice wake-up success rate is greater than or equal to 80%, and the false wake-up frequency is less than or equal to 0.1 times / hour.
[0052] In a specific implementation scenario, the following requirement must be met for the metric "face wake-up success rate": face wake-up success rate greater than or equal to 90%.
[0053] In a specific implementation scenario, regarding the metric of "successful voice interruption rate," the virtual avatar interaction system supports interruption during interaction. After an interruption, the virtual avatar stops speaking, and its lips return to a closed state. For the metric "successful voice interruption rate," the following requirement must be met: the success rate of voice interruption must be greater than or equal to 90%. Furthermore, the formula for calculating the success rate of voice interruption can be expressed as: P = n / N. Where P represents the success rate of voice interruption, N represents the total number of interruption attempts, and n represents the number of times an interruption was successfully performed.
[0054] In a specific implementation scenario, for the metric "gesture recognition", the virtual avatar interaction system needs to support both static gestures (such as a thumbs-up gesture) and dynamic gestures (such as a wave gesture).
[0055] In an implementation scenario, before testing a virtual avatar interaction system, it is necessary to prepare for the test, including but not limited to: test data, test environment, etc.
[0056] In a specific implementation scenario, the input data of a virtual avatar interaction system includes both voice and image. Factors affecting recognition performance at the voice level include, but are not limited to: environmental noise, speaker's age, gender, voice volume, speech rate, and articulation clarity. Factors affecting recognition performance at the image level include, but are not limited to: the number of characters, gesture movement speed, gesture complexity, and image clarity. Factors affecting facial wake-up performance at the image level include, but are not limited to: the character's gender, age, presence of static interference (e.g., images or mask models), facial angle, lighting, presence of occlusion, makeup, and image resolution. Furthermore, still using a cultural heritage site as an example, to ensure the system testing closely resembles the real-world performance of the virtual avatar interaction system in a real cultural heritage site, the above data can be collected on-site in the actual cultural heritage site.
[0057] In a specific implementation scenario, taking the cultural heritage scenario as an example, the voice data can cover the basic terms related to the cultural heritage being tested, and be designed from the perspectives of vocabulary coverage, business coverage, syllable coverage, and common usage of the cultural heritage being tested. Specifically, it can include command words, continuous sentences, etc.
[0058] In a specific implementation scenario, the dataset of voice data can meet the following requirements: (1) The sentence recognition rate test should be recorded by at least 20 male and 20 female speakers, and the voice wake-up function test should be recorded by at least 50 speakers; (2) The environmental noise recording should include at least the actual noise of the cultural and museum environment, such as the machine noise at the entrance of the museum exhibition hall, the environmental noise of tourists talking indoors, etc., which are not limited here. The requirements for audio sampling equipment can be found in Table 1, and will not be described in detail here.
[0059] Table 1 Requirements for Audio Sampling Equipment
[0060]
[0061] As mentioned earlier, during system testing, test data can be input into the test driver interface of the text-based interactive terminal. Alternatively, as a possible implementation, a playback device can be used to replay the test data, placed in front of the interactive terminal to simulate a real-world interaction scenario. The requirements for the playback device are detailed in Table 2 and will not be described in detail here.
[0062] Table 2 Playback Equipment Requirements
[0063]
[0064] In a specific implementation scenario, taking a cultural heritage setting as an example, the image data set can be designed for testing in real cultural heritage settings, such as indoor and outdoor museums or cultural exhibition parks; no specific limitations are imposed here. Additionally, ambient lighting can be required to be between 200 lx and 1500 lx.
[0065] Furthermore, for the facial wake-up test, it can be required that at least 5 men and 5 women record the test, and the test can specifically include the following elements, with no fewer than 20 images of each type:
[0066] (1) Human motion blur can be achieved by using image processing software (such as Photoshop) to add blur to the entire image;
[0067] (2) Horizontal rotation angle, pitch angle, tilt angle;
[0068] (3) Facial features are obscured;
[0069] (4) Makeup and photo editing;
[0070] (5) Light;
[0071] (6) Facial expressions;
[0072] (7) The distance between the person in the picture and the camera should be 0.5 to 3 meters, and the number of people in the picture should be controlled between 1 and 4.
[0073] Similar to the face wake-up test, the gesture recognition test can also require at least 5 men and 5 women to record, and can specifically include the following elements, with no fewer than 20 images of each type:
[0074] (1) Provide at least one set of gestures, each set of gestures containing at least five gestures;
[0075] (2) Provide the name and operation description of each gesture, and the start and end of each gesture. The tester needs to restore the same body posture.
[0076] (3) The similarity between any two gestures in each gesture set should be as low as possible in order to distinguish them;
[0077] (4) The gestures in the gesture set are simple and easy to perform;
[0078] (5) The distance between the people in the picture and the camera should be 0.5 to 3 meters. The number of people in the picture should be controlled between 1 and 4. The main tester should stand in the middle of the picture.
[0079] In addition, the requirements that the image acquisition equipment must meet can be found in Table 3, and will not be described in detail here.
[0080] Table 3 Requirements for Image Acquisition Equipment
[0081]
[0082] In a specific implementation scenario, as mentioned earlier, the input data of a virtual avatar interaction system includes both voice and image data. Accordingly, the virtual avatar interaction system should also ensure that it has voice sampling and image acquisition capabilities. Specifically, the interaction terminal can integrate a microphone, camera, etc.
[0083] In a specific implementation scenario, to ensure the reliability and stability of system data transmission, the virtual avatar interaction system should meet the requirements of an uplink bandwidth of not less than 100Kbps and a downlink bandwidth of not less than 9Mbps, and should maintain a stable connection.
[0084] In a specific implementation scenario, during the testing of a virtual avatar interaction system, the near-field sound pickup distance can be less than 1 meter.
[0085] In a specific implementation scenario, the test environment can be either a low-noise environment or a high-noise environment. Furthermore, it can be required that the noise spectrum remain stable and that the noise does not resemble the pronunciation of the command word; see Table 4 for details.
[0086] Table 4 Recording Scenarios in Typical Noise Environments
[0087]
[0088] In a specific implementation scenario, test data can be created through pre-recording or data collection. Furthermore, multiple test datasets can be created based on different test items. During actual testing, the test dataset can be selected as needed. Referring to Table 5, which shows the types and requirements of voice test datasets, the test data should meet the following requirements:
[0089] (1) At least 2000 voice data entries are required, with the following requirements for the quantity of each type of voice data:
[0090] (a) The number of Category A shall not be less than 70% of the total;
[0091] (b) The number of Category B shall not be less than 15% and not more than 20% of the total;
[0092] (c) The number of Category C shall not be less than 5% and not more than 10% of the total;
[0093] (d) Category D is optional, and its quantity shall not exceed 5% of the total.
[0094] (2) There shall be no fewer than 30 speakers of various speech types;
[0095] (3) Voice data with a duration of 3 to 5 seconds account for more than 80% of the total;
[0096] (4) Voice data includes Chinese, Western languages, and numbers. The tester can set the test content according to the system task and application scenario. Each piece of voice data can meet the following requirements:
[0097] (a) Signal-to-noise ratio greater than or equal to 20 dB;
[0098] (b) The new noise level is less than 5 dB;
[0099] (c) With 16-bit quantization, the sample point value is not less than 10000;
[0100] (d) Voice input greater than 4 words per second.
[0101] Table 5. Types and requirements of speech test datasets
[0102]
[0103] In a specific implementation scenario, please refer to Table 6 for the types and requirements of the face wake-up test dataset. Taking the cultural heritage site scenario as an example, the face wake-up test dataset can be collected in a real-world cultural heritage environment. The requirements are as follows:
[0104] (1) The male-to-female ratio of the test subjects was 1:1;
[0105] (2) 80% are between 16 and 60 years old, 10% are under 16 years old, and 10% are over 60 years old;
[0106] (3) Static images: Photos taken of the test object under normal conditions, with no borders and clear images;
[0107] (4) Angles: The horizontal rotation angle, pitch angle and tilt angle of the face shall not exceed ±20 degrees;
[0108] (5) Lighting: strong light, backlight, dim light, normal light;
[0109] (6) Completeness: The facial contours and features are clear, there is no heavy makeup, the face area of the image has not been edited or modified, the glasses frame does not obstruct the eyes, and the lenses are colorless and non-reflective;
[0110] (7) Paper: Matte coated paper, glossy coated paper, frosted paper, glossy photo paper, cardboard paper, ordinary A4 paper;
[0111] (8) Resolution: The printing resolution of paper photos shall not be less than 300 dpi;
[0112] (9) Cropping method: For paper photos, in the two sets of photos for each object, one set retains the complete paper and the other set cuts out the face. Each set of 4 photos is processed to different degrees of facial feature extraction, with 1 photo not extracted and the other 3 photos having facial features randomly extracted.
[0113] (10) Dynamic Images: The recorded video should be taken in the user's normal state (the background must be the same as the test background of the benevolent user), with the test subject's face within the video area, a frame rate of not less than 25fps, a duration of not less than 10 seconds, and a resolution of not less than 1080p. The tilt and angle of the face should not exceed ±30 degrees. Furthermore, the composite video can refer to the requirements for the recorded video and can be synthesized using captured static images;
[0114] (11) Mask: A wearable three-dimensional face mask made of materials such as plastic, paper or silicone, with the same size as a live human face;
[0115] (12) Head mold: The head mold is made of materials such as foam and resin, and its size is consistent with that of a living human face.
[0116] Table 6. Types and requirements of the face wake-up test dataset
[0117]
[0118]
[0119] In a specific implementation scenario, please refer to the types and requirements of the gesture recognition test dataset shown in Table 7.
[0120] Table 7. Types and requirements of gesture recognition test datasets
[0121]
[0122] In one implementation scenario, for the test item "speech recognition test", the virtual avatar interaction system can be put into standby mode, the speech recognition test data can be input into the test driver interface, or the speech recognition test data can be played back using a playback device at near-field distance, and the following content can be recorded:
[0123] (1) Under low noise environment, the recognition results of the virtual image interaction system are compared with the correct results, the number of successful recognitions and the number of failed recognitions are counted, and the sentence recognition rate is determined;
[0124] (2) Under high noise environment, the recognition results of the virtual image interaction system are compared with the correct results, the number of successful recognitions and the number of failed recognitions are counted, and the sentence recognition rate is determined.
[0125] In one implementation scenario, for the test item "speech synthesis test", the following content can be recorded:
[0126] (1) Sentence synthesis accuracy: According to the test speech set, each speech was input into the virtual image interaction system, the number of correctly broadcast speech was counted, and the sentence synthesis accuracy was calculated based on the aforementioned relevant description;
[0127] (2) Normalized accuracy: According to the test speech set, the test speech is input into the virtual image interaction system one by one. The number of speech sentences in which both numbers and symbols are pronounced correctly is counted, and the normalized accuracy is calculated according to the aforementioned relevant description.
[0128] In one implementation scenario, for the test item "virtual image synthesis", the synthesis time of the video stream containing the virtual image can be counted from the first frame to the last frame and the length of each speech, and the real-time synthesis rate of each video stream can be calculated according to the aforementioned description.
[0129] In one implementation scenario, for the test item "interaction success rate", the interactive functions of the virtual avatar interaction system can be statistically analyzed based on the aforementioned test results, and the interaction success rate can be calculated according to the aforementioned relevant description.
[0130] In one implementation scenario, for the test item "response time", the interaction time of the virtual avatar interaction system can be statistically analyzed based on the aforementioned test results, and the response time can be calculated according to the aforementioned relevant description.
[0131] In one implementation scenario, the test item "voice wake-up test" specifically includes wake-up accuracy and false wake-up frequency, as shown below:
[0132] (1) Wake-up accuracy: Set the virtual character interaction system to standby mode and play the wake-up test data at near-field distance using a playback device. When the sound pressure level is 55dB, calculate the virtual character voice wake-up accuracy in low-noise and high-noise environments respectively. Alternatively, input the wake-up test data into the aforementioned test driver interface and calculate the virtual character voice wake-up accuracy in low-noise and high-noise environments respectively.
[0133] (2) False wake-up frequency: The virtual character interaction system is set to standby mode for a certain period of time (e.g., 6 hours), and the false wake-up frequency of the virtual character in low noise environment and high noise environment is recorded.
[0134] In one implementation scenario, for the "face wake-up test" test item, the virtual avatar interaction system can be put into standby mode. The tester is required to walk into the video capture range of the interaction terminal with their face unobstructed and remain there, then walk out of the video capture range and remain there again. Alternatively, a face wake-up test video recorded / captured in a real environment can be input into the aforementioned test driver interface. Based on this, the face wake-up success rate can be calculated.
[0135] In one implementation scenario, for the "gesture recognition test" item, the virtual avatar interaction system can be put into standby mode. Testers walk into the video capture range of the interactive terminal and remain there, completing static and dynamic gesture tests. Alternatively, gesture recognition test videos recorded / captured in a real environment can be input into the aforementioned test driver interface. Based on this, the gesture recognition success rate is calculated.
[0136] In one implementation scenario, for the test item "voice interruption success rate," the virtual avatar interaction system can be put into standby mode. A playback device can be used at near-field distance to play the speech recognition test data. When the sound pressure level meter reads 55dB, the interactive terminal is awakened. During interaction with the virtual avatar, the playback device plays the speech test data, and the virtual avatar interruption result is recorded. Alternatively, the speech recognition test data can be input into the test driver interface, the interactive terminal is awakened, and new speech test data is input into the test driver interface during interaction with the virtual avatar, and the virtual avatar interruption result is recorded. Based on this, the voice interruption success rate can be calculated according to the aforementioned description.
[0137] It should be noted that the above are merely some possible implementation methods for the test environment and test data involved in the testing process of the virtual avatar interaction system, and do not limit the specific settings during the testing process. Specific settings can be made according to actual needs within the framework of the above system. The following sections describe the cultural relic virtual avatar interaction process from the perspectives of the interactive terminal and the interactive response server, respectively.
[0138] Please see Figure 3 , Figure 3 This is a flowchart illustrating an embodiment of the virtual avatar interaction method of this application. Specifically, it may include the following steps:
[0139] Step S31: Obtain and play the first video stream.
[0140] In this embodiment of the disclosure, the interactive response server generates a first interactive decision in response to a first interactive request issued by the user through an interactive terminal, and synthesizes a first video stream in real time through a virtual image synthesis engine based on the first interactive decision. The interactive response server uses keywords in the first interactive request as markers to indicate whether the first video stream should resume playback after interruption.
[0141] In one implementation scenario, please refer to the following: Figure 1Users can interact with the interactive terminal through gestures, voice, or other means to issue an initial interaction request. For example, when interacting with the interactive terminal via voice, a user can say things like "Where is cultural relic A?" or "Please describe cultural relic B?" to issue an initial interaction request. Similarly, when interacting with the interactive terminal via gestures, a user can make gestures like "Raise the volume" or "Lower the volume" to issue an initial interaction request. Other cases can be deduced similarly, and will not be listed here.
[0142] In one implementation scenario, after a user sends an initial interaction request via an interactive terminal, the interaction response server can process the request and generate an initial interaction decision. Please refer to the following for further details. Figure 1 After a user issues a first interaction request via voice, the interaction response server can recognize the voice data through a voice recognition interface to obtain recognized text, and analyze the recognized text through a semantic understanding interface to obtain the interaction intent. Based on the interaction intent, the server then queries the information processing server through an interaction decision interface to obtain a first interaction decision. Alternatively, after a user issues a first interaction request via gesture, the gesture recognition interface in the interaction terminal can recognize the gesture to obtain a recognition result. The interaction decision interface can then directly query the information processing server and combine the gesture recognition result to obtain the first interaction decision. It should be noted that, as described in the aforementioned embodiments, the information processing server may include an information system, which includes, but is not limited to, information related to cultural relics exhibits and system user information, etc., and is not limited here.
[0143] In a specific implementation scenario, taking a user's initial interactive request via voice, "Please introduce Cultural Relics Exhibit B," as an example, the speech recognition interface can recognize the speech data, obtaining the recognized text "Please introduce Cultural Relics Exhibit B." The semantic understanding interface can analyze this recognized text to obtain the interactive intent "Learn about Cultural Relics Exhibit B." Based on this, the interaction decision interface can query relevant knowledge about "Cultural Relics Exhibit B" from the information system on the information processing server and organize the obtained relevant knowledge into interactive text (i.e., the text to be synthesized), "Cultural Relics Exhibit B was made in XX year, and is…." The speech synthesis interface can perform speech synthesis based on the above interactive text to obtain synthesized speech. Then, the image synthesis interface, driven by the synthesized speech through the virtual image synthesis engine, synthesizes the first video stream. Other cases can be deduced similarly, and will not be listed here.
[0144] In a specific implementation scenario, taking a user's initial interaction request "End Interaction" via gesture as an example, the gesture recognition interface identifies the gesture as "End Interaction" and sends it to the interaction decision interface. Since the recognition result does not contain any cultural heritage-related information, it can be directly processed by the interaction decision interface without further interaction with the information processing server. Responding to the recognition result of "End Interaction," the interaction decision interface can directly act on the image synthesis interface, terminating image synthesis. Other cases can be deduced similarly, and will not be listed here.
[0145] It should be noted that the working principle of the virtual avatar synthesis engine can be found in the relevant description of the virtual avatar synthesis engine in the aforementioned public embodiments, and will not be repeated here.
[0146] In one implementation scenario, while the interactive response server is synthesizing the first video stream in real time, it can also push the stream to the interactive terminal, or the interactive terminal can actively pull the stream from the interactive response server while the interactive response server is synthesizing the first video stream in real time, and play the obtained first video stream, thereby realizing interaction with the user through a virtual avatar.
[0147] In one implementation scenario, a set of relationship mappings can be pre-maintained. This set can contain mappings between words and whether playback resumes after an interruption. For example, the word "end" maps to "do not resume after interruption," the word "termination" maps to "do not resume after interruption," the word "exhibition" maps to "resume after interruption," the word "where" maps to "resume after interruption," and the word "good" maps to "do not resume after interruption." It should be noted that the above examples are merely possible settings in practical applications and do not limit the actual settings.
[0148] In a specific implementation scenario, when a user interacts via voice, the recognition text obtained from the voice data can be detected through a voice recognition interface to determine whether the first interaction request contains words defined in the aforementioned relational mapping set. If a word is detected, the first video stream is further marked with a flag indicating whether playback will resume after an interruption, based on whether the word in the relational mapping set maps to "no resuming playback after interruption" or "resuming playback after interruption". For example, taking the first interactive request "Please introduce cultural relic exhibit B" as an example, since the speech recognition interface detects the word "exhibit" in the recognized text "Please introduce cultural relic exhibit B", and this word is mapped to "interrupt replay" in the relation mapping set, the first video stream marker generated based on this first interactive request can be a flag representing interrupted replay, such as 1, TRUE, etc.; or, taking the first interactive request "Okay, I understand" as an example, since the speech recognition interface detects the word "okay" in the recognized text "Okay, I understand", and this word is mapped to "do not resume playback after interruption" in the relation mapping set, the first video stream marker generated based on this first interactive request can be a flag representing do not resume playback after interruption, such as 0, FALSE, etc. Other cases can be deduced by analogy, and will not be listed here.
[0149] In a specific implementation scenario, when a user interacts via gesture, the recognition results obtained through the gesture recognition interface can be searched for the existence of words defined in the aforementioned relational mapping set. If a word is detected, it is further marked with a flag indicating whether playback will resume after interruption, based on whether the word in the relational mapping set maps to "do not resume playback after interruption" or "resume playback after interruption." For example, taking the "end interaction" gesture, since the word "end" is detected in the recognition results and this word maps to "do not resume playback after interruption," a flag indicating that playback will not resume after interruption can be used to mark the first video stream generated based on the first interaction request, such as 0, FALSE, etc. Other cases can be deduced similarly, and will not be listed here.
[0150] In one implementation scenario, the interactive terminal may be in a dormant state. To wake it up for interaction, it can be switched from dormant to awake using at least one of facial recognition or voice commands. For example, a user can speak a wake-up word into the interactive terminal, causing it to respond and switch to awake mode. Alternatively, a user can pre-register their face on the interactive terminal, and in subsequent applications, simply standing in the terminal's image capture area will trigger the terminal to respond to the registered face and switch to awake mode.
[0151] In one implementation scenario, unlike the aforementioned wake-up methods, a method can also switch to the wake-up state if, in the absence of either a registered face or a wake word, a gaze exceeding a duration threshold is detected. This method, even when neither a registered face nor a wake word is detected, switches to the wake-up state if a gaze exceeding the duration threshold is detected. This eliminates the need to speak a wake word or register a face to wake the interactive terminal, significantly improving the convenience of wake-up, especially for first-time users or those unfamiliar with the interactive terminal, greatly reducing the learning curve and ease of use.
[0152] In a specific implementation scenario, the duration threshold can be set according to the application's needs. For example, to reduce the probability of false wake-ups, the duration threshold can be set to a slightly larger value, such as 5 seconds or 10 seconds; or, to improve the interaction speed, the duration threshold can be set to a smaller value, such as 2 seconds or 3 seconds, etc. There is no limitation here.
[0153] In a specific implementation scenario, to further cater to users who are using the interactive terminal for the first time or are unfamiliar with it, prompts can be output after switching to the wake-up state to guide user interaction. For example, these prompts could be a pre-installed video stream on the interactive terminal, where a virtual avatar could demonstrate how to operate the terminal.
[0154] In a specific implementation scenario, to further reduce the probability of false wake-ups, lip detection can be combined with gaze tracking to determine whether to switch to the wake-up state. Specifically, lip keywords can be detected in each frame of a video captured of the user, and the distance between the upper and lower lips in the image can be determined based on lip key points. Then, the number of frames where the distance between the upper and lower lips is greater than a distance threshold is counted. If the number of frames exceeds a threshold, the user is switched to the wake-up state; otherwise, the user remains in a dormant state. It should be noted that lip key points can be detected using methods such as ASM (Active Shape Model), CPR (Cascaded Pose Regression), and Face++. The specific process can be found in the technical details of these detection methods, which will not be elaborated here. The thresholds in the example above can be set according to the actual application scenario. For example, if high detection accuracy is required, the distance threshold can be set appropriately larger; conversely, if the requirement for detection accuracy is relatively relaxed, the distance threshold can be set appropriately smaller. No limitation is imposed here. Furthermore, to further improve detection accuracy, numerical statistics (such as averaging, weighting, or medianing) can be performed on the distance between the upper and lower lips in each frame to determine a distance threshold. Further, when counting frames, a statistical duration can be set based on the frame rate, and statistics can be performed on each frame within that duration. For example, with a frame rate of 25fps, the statistical duration can be set to 2 seconds, 3 seconds, etc. Further, with a frame rate of 25fps and a statistical duration of 2 seconds, the quantity threshold can be set to 20 frames. Other cases can be deduced similarly and are not limited here. This method, before switching to the wake-up state, first detects the lip keypoints in each frame of the video captured of the user, and determines the distance between the upper and lower lips in the image based on these keypoints. This counts the number of frames where the distance between the upper and lower lips is greater than the distance threshold. If the number of frames exceeds the quantity threshold, the system switches to the wake-up state; otherwise, it remains in sleep mode. Therefore, it can further reduce the probability of false wake-ups.
[0155] It should be noted that after the interactive terminal switches to the wake-up state, the user can interact with the interactive terminal through voice, gestures, and other means.
[0156] Step S32: In response to the user's second interaction request while playing the first video stream, send an interruption synthesis request and a second interaction request to the interaction response server.
[0157] In one implementation scenario, similar to the first interaction request mentioned above, the user can also interact with the interactive terminal via voice, gestures, etc., while the first video stream is playing. The difference is that in this case, the user is considered to have a more urgent interaction need, thus generating a request to interrupt and synthesize the first video stream. This interrupt and synthesize request is then sent together with the second interaction request to the interaction response server. Specifically, the interrupt and synthesize request can be directly sent to the interaction decision interface in the interaction response server for processing. The second interaction request, similar to the first interaction request, can be handled by different interfaces depending on how it is sent. For details, please refer to the description of the first interaction request; it will not be repeated here.
[0158] In a specific implementation scenario, taking the first interactive request "Please introduce Cultural Relics Exhibit B" as an example, as mentioned earlier, the interactive terminal can obtain and play the first video stream related to the introduction of "Cultural Relics Exhibit B". When playing the video to the part about "In the year XXX AD, Privy Councilor XX...", if the user is confused about the historical figure "Privy Councilor XX", they can issue a second interactive request via voice: "Who is Privy Councilor XX?". At the same time, the interactive terminal generates an interruption request and sends the second interactive request and the interruption request to the interactive response server. Other situations can be deduced similarly, and will not be listed in detail here.
[0159] In a specific implementation scenario, taking the first interactive request "Please introduce Cultural Relics Exhibit B" as an example, as mentioned earlier, the interactive terminal can obtain and play the first video stream introducing "Cultural Relics Exhibit B". When playing the video to the part about "In the year XXX AD, Privy Councilor XX...", the user suddenly pauses the introduction of "Cultural Relics Exhibit B" and returns to the homepage to check the exhibition hall's closing time. The user can then send a second interactive request, "Return to Homepage," via gesture. Simultaneously, the interactive terminal generates an interruption request and sends both the second interactive request and the interruption request to the interactive response server. Other scenarios can be deduced similarly, and will not be listed here.
[0160] In this embodiment of the present disclosure, the interactive response server suspends the synthesis of the first video stream in response to the interruption synthesis request, and synthesizes the second video stream in real time in response to the second interactive request. After the synthesis of the second video stream is completed, the server determines, based on a flag, whether to continue synthesizing a new first video stream from the interruption position of the first interactive decision.
[0161] In one implementation scenario, when the synthesis of the first video stream is paused, the second video stream can be synthesized in real time in response to the second interactive request. For details, please refer to the synthesis process of the first video stream mentioned above, which will not be repeated here.
[0162] In one implementation scenario, as mentioned earlier, the interruption request for compositing can be directly handled by the interactive decision interface. In response to the interruption request, the interactive decision interface can directly command the image compositing interface to pause the compositing of the first video stream. Based on this, after the second video stream is completed, it can continue to determine, based on the aforementioned flag, whether to resume compositing of a new first video stream from the interruption point of the first interactive decision.
[0163] In a specific implementation scenario, if the flag indicates that playback was interrupted, it can be determined that a new first video stream will continue to be synthesized from the interruption position of the first interaction decision; conversely, if the flag indicates that playback was not resumed after the interruption, it can be determined that a new first video stream will no longer be synthesized.
[0164] In a specific implementation scenario, as mentioned earlier, the first interactive decision may include interactive text, and the virtual avatar synthesis engine can perform synthesis operations based on the synthesized speech obtained from the interactive text. In this case, the time information corresponding to the interruption position in the synthesized speech can be obtained. Specifically, the frame number of the audio frame corresponding to the interruption position in the synthesized speech can be obtained, and thus the time information can be obtained based on the frame rate and frame number of the synthesized speech. For example, the interval between adjacent frames can be obtained based on the frame rate, and then the frame number can be multiplied by the interval to obtain the time information. Taking the example of playing "In the year XXX AD, Privy Councilor XX...", if the user is confused about the historical figure "Privy Councilor XX", they can make a second interactive request through voice, "Who is Privy Councilor XX?" The frame number of the audio frame corresponding to the interruption position is N, and the frame rate is 25fps, so the time interval between adjacent audio frames is 40ms, and the time information is N*40ms. Other cases can be deduced in the same way, and will not be listed here. At the same time, the phoneme information of the virtual avatar at the interruption position in the first video stream can be obtained. Using the aforementioned example, the phoneme information at the interruption point is the final phoneme of "Privy Councilor XX". Based on this, by combining time information and phoneme information, the text content in the interactive text that has not been broadcast by the virtual avatar in the first video stream can be determined. Again, using the aforementioned example, the text content in the interactive text that has not been broadcast by the virtual avatar in the first video stream is the text information following "In the year XXX AD, Privy Councilor XX". Therefore, the virtual avatar synthesis engine can further perform synthesis operations based on the corresponding parts of the above text content in the synthesized speech to obtain a new first video stream. In the above method, the first interactive decision includes interactive text. The virtual avatar synthesis engine performs synthesis operation based on the synthesized speech of the interactive text. When a new first video stream is determined based on the marker, the time information corresponding to the interruption position in the synthesized speech can be obtained, and the phoneme information of the virtual avatar in the first video stream at the interruption position can be obtained. Based on the time information and phoneme information, the text content in the interactive text that has not been played by the virtual avatar in the first video stream can be determined. Then, through the virtual avatar synthesis engine, the synthesis operation continues based on the corresponding part of the text content in the synthesized speech to obtain a new first video stream. Therefore, the naturalness of interruption resumption can be improved.
[0165] In a specific implementation scenario, unlike the aforementioned implementation methods, the first interactive decision may include time-synchronized action instructions and interactive text. Action instructions are used to guide the limb movements of the virtual character in the synthesized first video stream, specifically including but not limited to reaching out or waving, etc., without limitation here. It should be noted that time synchronization means that the text in the interactive text corresponds to action instructions. For example, the text "Exhibit B" in the aforementioned interactive text "Cultural Relics Exhibit B was made in XX year, is…" can correspond to the action instruction "reach out," and when the first video stream is synthesized, when the text "Cultural Relics Exhibit B" is played in the first video stream, a 3D model of Cultural Relics Exhibit B can be embedded above the virtual character's hand. For details, please refer to the relevant descriptions in the following disclosed embodiments, which will not be repeated here. In other words, the virtual character synthesis engine performs synthesis operations based on the synthesized speech and action instructions of the interactive text. In the case of determining the synthesized new first video stream based on the identifier, the time information corresponding to the interruption position in the synthesized speech can be obtained, as well as the phoneme information of the virtual avatar in the first video stream at the interruption position. Then, based on the time and phoneme information, the text content in the interactive text that has not been broadcast by the virtual avatar in the first video stream can be determined. For details, please refer to the aforementioned descriptions, which will not be repeated here. Based on this, a new first video stream can be synthesized using the virtual avatar synthesis engine, based on the corresponding part of the text content in the synthesized speech and the residual part of the action command after the interruption position. For example, when playing "In the year XXX AD, Privy Councilor XX…", if a user is confused about the historical figure "Privy Councilor XX", they can issue a second interactive request via voice, "Who is Privy Councilor XX?" The text content in the interactive text that has not been broadcast by the virtual avatar in the first video stream is the text information after "In the year XXX AD, Privy Councilor XX", and the residual part of the action command after the interruption position is the action command after "In the year XXX AD, Privy Councilor XX". Other cases can be deduced similarly, and will not be listed here. In the above method, the first interactive decision includes time-synchronized interactive text and action commands. The virtual avatar synthesis engine performs synthesis operations based on the synthesized speech of the interactive text and the action commands. When it is determined to continue synthesizing a new first video stream based on the marker, the engine obtains the time information corresponding to the interruption position in the synthesized speech and the phoneme information of the virtual avatar at the interruption position in the first video stream. Based on the time information and phoneme information, the engine determines the text content in the interactive text that has not been played by the virtual avatar in the first video stream. Then, through the virtual avatar synthesis engine, a new first video stream is synthesized based on the corresponding part of the text content in the synthesized speech and the residual part of the action command after the interruption position, which can improve the naturalness of interruption and resumption.
[0166] Step S33: Obtain and play the newly synthesized video stream from the interactive response server.
[0167] Specifically, the newly synthesized video stream includes at least the aforementioned second video stream. Furthermore, if the flag indicates an interruption of playback, the newly synthesized video stream may further include a new first video stream following the second video stream; conversely, if the flag indicates that playback will not resume after the interruption, the newly synthesized video stream includes only the aforementioned second video stream.
[0168] In one implementation scenario, the interactive terminal can also display 3D models of cultural and museum exhibits. In this case, the interactive terminal can also respond to the recognition of a switching gesture by displaying the next 3D model of the cultural and museum exhibit. For example, the interactive terminal can maintain a list of cultural and museum exhibits, which can be arranged sequentially among the various exhibits displayed in the exhibition hall. For example, the list can be arranged chronologically, or it can be arranged according to popularity; no limitation is made here. The above method, responding to the recognition of a switching gesture by displaying the next 3D model of the cultural and museum exhibit, can enhance the interactive experience.
[0169] In one implementation scenario, the interactive terminal can also recognize a "like" gesture when displaying a 3D model of a cultural and museum exhibit, accumulating a preset score for the currently displayed exhibit. The 3D models of each exhibit are then displayed sequentially based on their accumulated scores. It should be noted that the "like" gesture can include, but is not limited to, a thumbs-up. The preset score can be set to 10 points, 20 points, etc., without limitation. This method, by recognizing a "like" gesture when displaying a 3D model of a cultural and museum exhibit, accumulates a preset score for the currently displayed exhibit, and then displays the 3D models of each exhibit sequentially based on their accumulated scores. Therefore, by recognizing "like" gestures, the popularity information of each exhibit can be collected, allowing the display of 3D models of exhibits according to user popularity, helping users to learn about popular exhibits and improving visit efficiency.
[0170] In one implementation scenario, as mentioned earlier, users can pre-register for a virtual avatar interaction system using their faces. In this case, upon detecting a registered face, the system can obtain the visit route of the user to whom the registered face belongs. Based on the location of the interactive terminal where the registered face is currently detected and the visit route, the system determines the next cultural and museum exhibit the user will visit and displays a third video stream. This third video stream is synthesized in real-time by the interactive response server using a virtual avatar synthesis engine based on the next cultural and museum exhibit to be visited. The virtual avatar in the third video stream indicates the location information of the next cultural and museum exhibit to be visited. It should be noted that interactive terminals can be set up in each exhibition hall / area. For example, interactive terminals can be set up at the entrance of each exhibition hall / area; or, in the case of a large exhibition hall / area, interactive terminals can be further set up within the exhibition hall / area, without limitation. The above method, in response to the detection of a registered face, obtains the visit route of the user to whom the registered face belongs, and determines the next cultural and museum exhibit to be visited by the user based on the location of the interactive terminal that detected the registered face and the visit route, and displays a third video stream. The third video stream is synthesized in real time by the interactive response server based on the next cultural and museum exhibit to be visited through a virtual image synthesis engine. The virtual image in the third video stream indicates the location information of the next cultural and museum exhibit to be visited. Therefore, only cultural and museum interactive systems need to be set up in each exhibition hall / exhibition area to realize multi-terminal interconnection to guide users to visit cultural and museum exhibits, which helps to improve the level of intelligence.
[0171] In a specific implementation scenario, the next cultural relic to be visited can be determined by the visitor route and the location of the interactive terminal where a registered face has been detected. For example, if an interactive terminal is set up at the entrance of an exhibition hall / area, taking the visitor route "Cultural Relic A → Cultural Relic B → Cultural Relic C → Cultural Relic D → Cultural Relic E" as an example, if Cultural Relic B is displayed in Exhibition Hall / Area A where the interactive terminal where a registered face has been detected is located, it can be determined that the next cultural relic to be visited is Cultural Relic B. In this case, the location information of the virtual image in the third video stream indicating the next cultural relic to be visited can be "pointing to Exhibition Hall / Area A"; or, if the interactive terminal where a registered face has been detected is located... If exhibit B in the exhibition hall / area does not display any of the cultural relics exhibits along the aforementioned tour route, then based on the fact that exhibit A has already been visited (e.g., based on previously indicated location information for that exhibit), and that no other exhibits have been visited, the next exhibit to be visited can be determined to be exhibit B. Combining this with the information recorded in the system that "exhibit B is displayed in exhibition hall / area A," the location information in the third video stream indicating the next exhibit to be visited, pointing to exhibition hall / area A, can be determined. Other situations can be deduced similarly, and will not be listed here.
[0172] In a specific implementation scenario, the synthesis process of the third video stream can be referred to the synthesis operation of the first video stream mentioned above, and will not be repeated here.
[0173] In a specific implementation scenario, the interactive terminal can also respond to the viewing request of the user to whom the registered face belongs, displaying the user's viewing progress on their tour route and displaying a fourth video stream. This fourth video stream is synthesized in real-time by the interactive response server using a virtual avatar synthesis engine based on the viewing progress. At least one of the virtual avatar's expression, actions, and voice in the fourth video stream matches the viewing progress. For example, the user can display the tour route by touching the relevant menu on the interactive terminal. Alternatively, the interactive terminal can, upon detecting a registered face, first query whether a tour route for the user to whom the registered face belongs has been generated; if so, it can display the tour route. This is not limited to any particular scenario. Furthermore, when the viewing progress is 50%, at least one of the virtual avatar's expression, actions, and voice in the fourth video stream can be "encouraging"; similarly, when the viewing progress is 90%, at least one of the virtual avatar's expression, actions, and voice in the fourth video stream can be "happy"; similarly, when the viewing progress is 100%, at least one of the virtual avatar's expression, actions, and voice in the fourth video stream can be "excited". Of course, the above are merely possible implementation methods in practical applications and do not limit the specific way of setting up virtual avatars that match the visit progress in actual applications. In response to a viewing request from the user whose face is registered, the above method displays the user's visit progress along their route and shows a fourth video stream. This fourth video stream is synthesized in real-time by the interactive response server based on the visit progress using a virtual avatar synthesis engine. At least one of the virtual avatar's expression, actions, or voice in the fourth video stream matches the visit progress, thus supporting check-ins during the visit and providing support to the user through the virtual avatar, which helps improve the user's visit experience.
[0174] In a specific implementation scenario, users can engage in interactive Q&A sessions with the interactive terminal regarding their cultural heritage preferences. This involves learning about the user's preferred scenes, exhibits, and cultural heritage knowledge through a question-and-answer format. Based on this, the system can respond by ending the interactive Q&A session with the user whose face is registered. A tour route is then generated based on the Q&A, and a fifth video stream is displayed. This fifth video stream is synthesized in real-time by the interactive Q&A server using a virtual avatar synthesis engine, based on the first cultural heritage exhibit visited along the route. The virtual avatar in the fifth video stream indicates the location information of the first visited exhibit. It should be noted that exhibits of interest to the user can be extracted from the interactive Q&A session and sorted according to factors such as the distance of these exhibits from the user and the number of visitors, thus generating the tour route. The above method generates a tour route based on the interactive Q&A session after the session ends, and displays a fifth video stream. This fifth video stream is synthesized in real time by the interactive response server using a virtual avatar synthesis engine based on the first cultural relic exhibit visited in the tour route. The virtual avatar in the fifth video stream indicates the location information of the first cultural relic exhibit visited. Therefore, the interactive Q&A session can guide users through the tour, further enhancing user satisfaction while minimizing visitor time and improving the overall user experience.
[0175] It should be noted that, in the embodiments disclosed in this application, cultural and museum exhibits include, but are not limited to, physical exhibits, and may also include non-physical exhibits displayed by means of photography, projection, etc. For example, some exhibits are susceptible to deterioration due to environmental factors such as exposure and humidity, and only their non-physical exhibits are displayed in the exhibition hall / exhibition area.
[0176] The above scheme acquires and plays a first video stream. The interactive response server responds to a user's first interactive request via an interactive terminal, generates a first interactive decision, and synthesizes the first video stream in real-time using a virtual avatar synthesis engine based on the first interactive decision. The interactive response server uses keywords from the first interactive request to mark the first video stream, indicating whether playback should resume after interruption. Furthermore, in response to a user's second interactive request while playing the first video stream, the server sends an interruption synthesis request and a second interactive request to the interactive response server. The interactive response server pauses the synthesis of the first video stream in response to the interruption request and synthesizes the second video stream in real-time in response to the second interactive request. After the second video stream is synthesized, based on the marker, it determines whether to resume synthesis of a new first video stream from the interruption point of the first interactive decision. This allows the acquisition and playback of the newly synthesized video stream from the interactive response server. Thus, during the real-time synthesis of the video stream by the interactive response server and the playback of the video stream by the interactive terminal, the synthesis can be interrupted by a new user interactive request. A new video stream is synthesized in real-time first, and then the resumption of synthesis of the original video stream is determined based on the resume signal. Therefore, this significantly improves the naturalness of the interaction between the cultural relics virtual avatar and the user.
[0177] It should be noted that the specific steps in the above-described virtual avatar interaction method embodiments can be provided by... Figure 1 The interactive terminal in the virtual avatar interaction system shown above is responsible for the specific composition of the interactive terminal. For details, please refer to the relevant descriptions in the aforementioned virtual avatar interaction system. It will not be repeated here.
[0178] Please see Figure 4 , Figure 4 This is a flowchart illustrating another embodiment of the virtual avatar interaction method of this application. Specifically, it may include the following steps:
[0179] Step S41: Based on the first interaction request issued by the interactive terminal, generate a first interaction decision, and synthesize a first video stream through a virtual image synthesis engine based on the first interaction decision, and mark the first video stream with a flag indicating whether playback will resume after interruption based on the keywords in the first interaction request.
[0180] In this embodiment of the disclosure, the interactive terminal acquires and plays the first video stream. For details, please refer to the relevant descriptions in the foregoing embodiments of the disclosure, which will not be repeated here.
[0181] Step S42: In response to the interruption synthesis request issued by the interactive terminal, pause the synthesis of the first video stream, and in response to the second interaction request issued by the interactive terminal, synthesize the second video stream in real time. After the synthesis of the second video stream is completed, determine, based on the flag, whether to continue synthesizing a new first video stream from the interruption position of the first interaction decision.
[0182] In this embodiment of the disclosure, the interruption request is sent by the interactive terminal in response to the user's second interactive request while playing the first video stream, and the interactive terminal obtains and plays the newly synthesized video stream. For details, please refer to the relevant descriptions in the foregoing embodiments of the disclosure, which will not be repeated here.
[0183] In one implementation scenario, as described in the aforementioned disclosed embodiments, the first interaction decision includes time-synchronized interactive text and action commands. The virtual avatar synthesis engine performs synthesis operations based on the synthesized speech of the interactive text and the action commands. When a new first video stream is determined to be synthesized based on a marker, the time information corresponding to the interruption position in the synthesized speech can be obtained, and the phoneme information of the virtual avatar at the interruption position in the first video stream can be obtained. Based on the time information and phoneme information, the text content in the interactive text that has not been played by the virtual avatar in the first video stream can be determined. Then, through the virtual avatar synthesis engine, a new first video stream is synthesized based on the corresponding part of the text content in the synthesized speech and the remaining part of the action commands after the interruption position. For details, please refer to the relevant descriptions in the aforementioned disclosed embodiments, which will not be repeated here.
[0184] In one implementation scenario, as mentioned above, the first interactive decision includes at least interactive text. During the synthesis of the first video stream using a virtual avatar synthesis engine based on the first interactive decision, in response to the search for matching cultural relics exhibits in the information system of the information processing server using keywords, the matching cultural relics exhibit is used as the target exhibit. Speech synthesis is performed based on the interactive text to obtain synthesized speech. Furthermore, a 3D model of the target exhibit is embedded in the first video stream obtained by synthesizing the image from the synthesized speech using the virtual avatar synthesis engine. In this approach, if a matching cultural relics exhibit is found in the information system based on keywords, a 3D model of that exhibit can be embedded in the synthesized first video stream. This allows for more convenient information interaction through the 3D model, thereby enhancing the user's visitor experience.
[0185] In a specific implementation scenario, taking the first interactive request "Please introduce cultural relic exhibit B" as an example, since the keyword "cultural relic exhibit B" finds a matching cultural relic exhibit in the information system, a 3D model of "cultural relic exhibit B" can be embedded in the first video stream synthesized by the virtual image synthesis engine. Other cases can be deduced similarly, and will not be listed in detail here.
[0186] In a specific implementation scenario, as mentioned above, the first interaction decision may further include action instructions synchronized with the interactive text time. Before or after determining whether to select the matching cultural relic exhibit as the target exhibit, it can be further determined that the first interaction decision also includes action instructions synchronized with the interactive text time, at least including a reaching gesture. Based on this, when embedding the 3D model of the target exhibit into the first video stream obtained by synthesizing the synthesized speech using a virtual avatar synthesis engine, the synthesized speech, action instructions, and the 3D model of the target exhibit can be synthesized using the virtual avatar synthesis engine to obtain the first video stream. That is, the virtual avatar in the first video stream triggers a reaching gesture to display the 3D model of the target exhibit. More specifically, the reaching gesture can be triggered when a keyword matching the target exhibit first appears in the interactive text. For details, please refer to the relevant descriptions in the aforementioned disclosed embodiments, which will not be repeated here. In the above method, before or after selecting the matched cultural and museum exhibits as the target exhibits, it is further determined that the first interactive decision also includes action instructions synchronized with the interactive text time, including at least a reaching gesture. In the process of embedding the three-dimensional model of the target exhibit into the first video stream obtained by performing image synthesis on the synthesized speech through the virtual image synthesis engine, the synthesized speech, action instructions and the three-dimensional model of the target exhibit are image synthesized by the virtual image synthesis engine to obtain the first video stream. The virtual image in the first video stream triggers the reaching gesture to display the three-dimensional model of the target exhibit, which can improve the naturalness of the virtual image.
[0187] It should be noted that this disclosure only describes the parts that were not elaborated in the foregoing disclosure embodiments. For other similar or identical parts, please refer to the foregoing disclosure embodiments, which will not be repeated here.
[0188] The above scheme generates a first interaction decision based on a first interaction request from the interactive terminal, and synthesizes a first video stream through a virtual avatar synthesis engine based on the first interaction decision. It also uses keywords from the first interaction request to mark whether the first video stream should resume playback after an interruption. The interactive terminal acquires and plays the first video stream, pauses the synthesis of the first video stream in response to an interruption synthesis request from the text interaction terminal, and synthesizes a second video stream in real time in response to a second interaction request from the text interaction terminal. After the second video stream synthesis is completed, it determines whether to resume synthesis of a new first video stream from the interruption point of the first interaction decision based on the marker. The interruption synthesis request is sent by the interactive terminal in response to the user's second interaction request while playing the first video stream, and the interactive terminal acquires and plays the newly synthesized video stream. Therefore, during the real-time synthesis of the video stream on the interactive response server and the playback of the video stream on the interactive terminal, the synthesis can be interrupted by a new user interaction request. A new video stream is synthesized in real time first, and then the resumption of playback is determined based on the marker, thus greatly improving the naturalness of the interaction between the cultural relics virtual avatar and the original video stream.
[0189] It should be noted that the specific steps in the above-described virtual avatar interaction method embodiments can be provided by... Figure 1 The interactive response server is executed in the virtual avatar interaction system shown. The specific composition of the interactive response server can be found in the relevant description in the aforementioned virtual avatar interaction system, and will not be repeated here.
[0190] Please see Figure 5 , Figure 5 This is a flowchart illustrating another embodiment of the virtual avatar interaction method of this application. Specifically, it may include the following steps:
[0191] Step S501: The interactive terminal sends an authentication request to the API access layer of the API gateway.
[0192] Specifically, in practical applications, as one possible implementation, an API gateway can be set up between the interactive response server and the interactive terminal to handle authentication, forwarding, etc. The API access layer can use an Nginx+keepalive architecture to achieve high availability load balancing, as well as primary and backup nodes to ensure failover and transfer between services.
[0193] Step S502: The API gateway's API access layer processes the authentication request through the authentication service interface, obtains the authentication result, and returns the authentication result to the interactive terminal through the API gateway's API access layer.
[0194] Specifically, the authentication service interface processes authentication requests to determine whether the interactive terminal can access the interactive response server, thereby improving the security of the virtual avatar interactive system.
[0195] Step S503: The interactive terminal responds to the authentication result, including successful authentication, by sending an initialization request to the API access layer of the API gateway.
[0196] Step S504: The authentication service interface of the API gateway performs authentication verification on the initialization request.
[0197] Step S505: The authentication service interface of the API gateway returns the verification result to the API access layer.
[0198] Step S506: The API access layer returns an error message if the verification result includes a verification failure.
[0199] Step S507: If the verification result includes successful verification, the API access layer returns the video address.
[0200] Specifically, the video address is the network address from which the video stream is subsequently retrieved from the interactive response server. In other words, the interactive terminal can retrieve and play the video stream based on this video address, thereby enabling the user to interact with the virtual avatar through the interactive terminal.
[0201] Step S508: The interactive terminal uploads an interaction request to the API access layer of the API gateway.
[0202] Specifically, the interaction request may include, but is not limited to, text, audio, etc. For details, please refer to the relevant descriptions in the foregoing disclosed embodiments; they will not be repeated here.
[0203] Step S509: In response to the interaction request, including audio, the API access layer directly transmits the audio to the speech recognition interface of the interaction response server for recognition, and obtains the recognized text.
[0204] For details, please refer to the relevant descriptions in the foregoing disclosed embodiments, which will not be repeated here.
[0205] Step S510: The speech recognition interface sends the recognized text to the semantic understanding interface for analysis, interaction intent, and obtains decision text based on the interaction intent.
[0206] Specifically, as described in the foregoing embodiments, after obtaining the interaction intent, the interaction decision interface in the interaction response server can interact with the information system in the information processing server to obtain the decision text. The specific process can be found in the relevant descriptions in the foregoing embodiments, and will not be repeated here.
[0207] Step S511: In response to the interaction request, which includes the text to be analyzed, the API access layer directly transmits the text to be analyzed to the semantic understanding interface of the interaction response server to obtain the interaction intent, and obtains the decision text based on the interaction intent.
[0208] For details, please refer to the relevant descriptions in the foregoing disclosed embodiments, which will not be repeated here.
[0209] Step S512: The semantic understanding interface transmits the decision text to the speech synthesis interface for speech synthesis to obtain synthesized speech.
[0210] For details, please refer to the relevant descriptions in the foregoing disclosed embodiments, which will not be repeated here.
[0211] Step S513: In response to the interaction request, which includes the text to be synthesized, the API access layer directly transmits the text to be synthesized to the speech synthesis interface of the interaction response server for speech synthesis to obtain synthesized speech.
[0212] For details, please refer to the relevant descriptions in the foregoing disclosed embodiments, which will not be repeated here.
[0213] Step S514: The speech synthesis interface inputs the synthesized speech into the image synthesis interface to synthesize a video stream.
[0214] For details, please refer to the relevant descriptions in the foregoing disclosed embodiments, which will not be repeated here.
[0215] Step S515: The image compositing interface transmits the video stream to the push streaming service interface.
[0216] It should be noted that the interactive terminal obtains the synthesized video stream from the streaming service interface through the aforementioned video address and plays it on the interactive terminal.
[0217] Step S516: The interactive terminal sends a new interaction request to the API access layer.
[0218] For details, please refer to the relevant descriptions in the foregoing disclosed embodiments, which will not be repeated here.
[0219] Step S517: The API access layer generates an interruption compositing request and sends it to the image compositing interface to pause the compositing operation.
[0220] Specifically, the interactive response server can first pause the ongoing compositing operation and respond to new interactive requests to perform the compositing operation until the compositing is complete. Based on this, if the identifier of the video stream synthesized in step S514 indicates an interruption of subsequent playback, the compositing operation can be re-executed from the interruption point; otherwise, the compositing operation does not need to be re-executed. It should be noted that the specific meaning of the identifier can be found in the relevant descriptions in the aforementioned disclosed embodiments, and will not be repeated here.
[0221] Step S518: The text-interactive terminal pulls the video stream from the push streaming service interface via the video address.
[0222] For details, please refer to the relevant descriptions in the foregoing disclosed embodiments, which will not be repeated here.
[0223] It should be noted that as long as the network connection between the interactive terminal and the interactive response server remains open, steps S508 to S518 can be executed repeatedly. In other words, as long as the user issues an interactive request through the interactive terminal or issues a new interactive request during the video stream synthesis process, the above-mentioned steps can be re-executed.
[0224] Step S519: The interactive terminal sends a disconnection request to the API access layer.
[0225] Specifically, in scenarios where the interactive terminal needs to be sent for inspection, the interactive terminal needs to disconnect from the interactive response server. In order to minimize the impact on the system, the interactive terminal can send a disconnection request to successfully disconnect from the interactive response server.
[0226] Step S520: The API access layer forwards the disconnect request to the image synthesis interface so that the image synthesis interface stops the synthesis operation.
[0227] Specifically, upon receiving a disconnection request, the image synthesis interface in the interactive response server can stop the synthesis operation.
[0228] Step S521: The image synthesis interface sends a stop streaming command to the streaming service interface.
[0229] Specifically, after the image compositing interface stops the compositing operation, it can command the streaming service interface to stop streaming. At this time, the interactive terminal will no longer obtain video streams from the interactive response server, and there will be no more data interaction between the interactive terminal and the interactive response server.
[0230] The above solution can interrupt the synthesis process when a new user interaction request occurs during the real-time synthesis of the video stream on the interactive response server and the playback of the video stream on the interactive terminal. It first synthesizes a new video stream in real time, and then determines whether to continue synthesizing the original video stream from the interrupted position based on the flag indicating whether to resume playback. Therefore, it can greatly improve the naturalness of the interaction of the cultural heritage virtual image.
[0231] It should be noted that the specific steps in the above-described virtual avatar interaction method embodiments can be provided by... Figure 1 In the virtual avatar interaction system shown, the interactive terminal, interactive response server, and information processing server work together. For the specific composition of the interactive terminal, interactive response server, and information processing server, please refer to the relevant description in the aforementioned virtual avatar interaction system, which will not be repeated here.
[0232] Please see Figure 6 , Figure 6This is a schematic diagram of the framework of an embodiment of the virtual avatar interaction device 60 of this application. The virtual avatar interaction device 60 includes: a first acquisition module 61, a request sending module 62, and a second acquisition module 63. The first acquisition module 61 is used to acquire and play a first video stream; wherein, the interaction response server generates a first interaction decision in response to a first interaction request issued by a user through an interaction terminal, and synthesizes the first video stream in real time through a virtual avatar synthesis engine based on the first interaction decision, and the interaction response server uses a keyword in the first interaction request as a marker to indicate whether playback should resume after interruption; the request sending module 62 is used to send an interruption synthesis request and a second interaction request to the interaction response server in response to a second interaction request from the user while playing the first video stream; wherein, the interaction response server pauses the synthesis of the first video stream in response to the interruption synthesis request, and synthesizes the second video stream in real time in response to the second interaction request, and after the second video stream is synthesized, determines whether to continue synthesizing a new first video stream from the interruption position of the first interaction decision based on the marker; the second acquisition module 63 is used to acquire and play the newly synthesized video stream by the interaction response server.
[0233] The above solution, because the virtual avatar interaction device 60 can implement the steps in the above virtual avatar interaction method embodiment, can interrupt the synthesis when a new interaction request from the user is received during the real-time synthesis of the video stream by the interaction response server and the playback of the video stream by the interaction terminal. It can first synthesize a new video stream in real time, and then determine whether to continue synthesizing the original video stream from the interruption position according to the flag of whether to continue playback. Therefore, it can greatly improve the naturalness of the cultural relics virtual avatar interaction.
[0234] In some disclosed embodiments, the virtual avatar interaction device 60 further includes a state switching module, which is used to switch to the wake-up state in response to the detection of a gaze that exceeds a duration threshold but is not detected when either the registered face or the wake word is not detected; the virtual avatar interaction device 60 also includes a guidance prompt module, which is used to output prompt information to guide the user's interaction.
[0235] Therefore, if neither the registered face nor the wake word is detected, but a gaze exceeding the duration threshold is detected, the system switches to the wake-up state. This eliminates the need to speak a wake word or register a face to switch the interactive terminal to the wake-up state, thus greatly improving the convenience of waking up. This is especially beneficial for first-time users or those unfamiliar with the interactive terminal, significantly reducing the learning curve and ease of use.
[0236] In some disclosed embodiments, the virtual avatar interaction device 60 further includes a key point detection module for detecting lip key points in each frame of a video captured of the user; the virtual avatar interaction device 60 also includes a distance determination module for determining the distance between the upper and lower lips in the image based on the lip key points in the image; the virtual avatar interaction device 60 further includes a frame count module for counting the number of image frames where the distance between the upper and lower lips is greater than a distance threshold; wherein, if the number of frames is greater than a quantity threshold, the device switches to a wake-up state, and if the number of frames is not greater than the quantity threshold, the device maintains a sleep state.
[0237] Therefore, before switching to the wake-up state, the key points of the lips in each frame of the video taken of the user are detected, and the distance between the upper and lower lips in the image is determined based on the key points of the lips in the image. The number of image frames with a distance between the upper and lower lips greater than the distance threshold is counted. Then, if the number of frames is greater than the number threshold, the wake-up state is switched, and if the number of frames is not greater than the number threshold, the sleep state is maintained. Thus, the probability of false wake-up can be further reduced.
[0238] In some disclosed embodiments, the distance threshold is obtained by numerically calculating the distance between the upper and lower lips in each frame of the image.
[0239] Therefore, obtaining the distance threshold by numerically calculating the distance between the upper and lower lips in each frame of the image can further improve the detection accuracy.
[0240] In some disclosed embodiments, the virtual avatar interaction device 60 further includes a switching display module for switching the display of the next cultural and museum exhibit in response to the recognition of a switching gesture.
[0241] Therefore, the above method, in response to the recognition of the switching gesture, switches to display the 3D model of the next cultural and museum exhibit, which can improve the interactive experience.
[0242] In some disclosed embodiments, the virtual avatar interaction device 60 further includes a score accumulation module, which is used to accumulate a preset score for the currently displayed cultural and museum exhibit in response to recognizing a thumbs-up gesture when displaying a 3D model of the cultural and museum exhibit; wherein, the 3D models of each cultural and museum exhibit are displayed sequentially based on the size of their respective accumulated scores.
[0243] Therefore, in response to recognizing a "like" gesture while displaying the 3D model of a cultural and museum exhibit, a preset score is accumulated for the currently displayed exhibit. Thus, the 3D models of each exhibit are displayed sequentially based on the size of their accumulated scores. By recognizing "like" gestures, the popularity information of each exhibit can be collected, thereby supporting the display of the 3D models of each exhibit in order of user popularity. This helps users learn about popular exhibits as much as possible and improves visit efficiency.
[0244] In some disclosed embodiments, the virtual avatar interaction device 60 further includes a route acquisition module, used to acquire the visit route of the user to whom the registered face belongs in response to the detection of a registered face; the virtual avatar interaction device 60 further includes an exhibit determination module, used to determine the next cultural and museum exhibit to be visited by the user based on the location of the interactive terminal where the registered face is currently detected and the visit route; the virtual avatar interaction device 60 further includes a third acquisition module, used to display a third video stream; wherein, the third video stream is synthesized in real time by the interaction response server based on the next cultural and museum exhibit to be visited through a virtual avatar synthesis engine, and the virtual avatar in the third video stream indicates the location information of the next cultural and museum exhibit to be visited.
[0245] Therefore, in response to the detection of a registered face, the system obtains the visit route of the user to whom the registered face belongs. Based on the location of the interactive terminal where the registered face is currently detected and the visit route, it determines the next cultural and museum exhibit that the user will visit and displays a third video stream. The third video stream is synthesized in real time by the interactive response server based on the next cultural and museum exhibit to be visited through a virtual image synthesis engine. The virtual image in the third video stream indicates the location information of the next cultural and museum exhibit to be visited. Therefore, it is only necessary to set up cultural and museum interactive systems in each exhibition hall / exhibition area to realize multi-terminal interconnection to guide users to visit cultural and museum exhibits, which helps to improve the level of intelligence.
[0246] In some disclosed embodiments, the virtual avatar interaction device 60 further includes a progress display module, used to display the visit progress of the user whose registered face belongs on their visit route in response to the user's viewing request; the virtual avatar interaction device 60 also includes a fourth acquisition module, used to display a fourth video stream; wherein, the fourth video stream is synthesized in real time by the interaction response server based on the visit progress through the virtual avatar synthesis engine, and at least one of the virtual avatar's expression, action, and voice in the fourth video stream matches the visit progress.
[0247] Therefore, in response to the viewing request of the user whose face is registered, the system displays the user's tour progress along their tour route and displays a fourth video stream. This fourth video stream is synthesized in real time by the interactive response server based on the tour progress using a virtual avatar synthesis engine. At least one of the expressions, actions, and voices of the virtual avatar in the fourth video stream matches the tour progress, thus enabling users to check in during the tour and receive support from the virtual avatar, which helps improve the user's tour experience.
[0248] In some disclosed embodiments, the virtual avatar interaction device 60 further includes a route generation module, used to generate a tour route based on the interactive Q&A session with the user to whom the registered face belongs, in response to the end of the interactive Q&A session regarding cultural and museum preferences; the virtual avatar interaction device 60 also includes a fifth acquisition module, used to display a fifth video stream; wherein, the fifth video stream is synthesized in real time by the interactive response server based on the first cultural and museum exhibit visited in the tour route through a virtual avatar synthesis engine, and the virtual avatar in the fifth video stream indicates the location information of the first cultural and museum exhibit visited.
[0249] Therefore, after the interactive Q&A session ends, a tour route is generated based on the Q&A, and a fifth video stream is displayed. The fifth video stream is synthesized in real time by the interactive response server based on the first cultural and museum exhibit visited in the tour route using a virtual image synthesis engine. The virtual image in the fifth video stream indicates the location information of the first cultural and museum exhibit visited. Thus, the interactive Q&A can guide users through the tour, which can further enhance user satisfaction while saving users' tour time as much as possible, and helps to improve the user tour experience.
[0250] Please see Figure 7 , Figure 7 This is a schematic diagram of the framework of an embodiment of the virtual avatar interaction device 70 of this application. The virtual avatar interaction device 70 includes: a request processing module 71 and an interruption and resumption module 72. The request processing module 71 is used to generate a first interaction decision based on a first interaction request issued by an interactive terminal, and to synthesize a first video stream through a virtual avatar synthesis engine based on the first interaction decision, and to mark the first video stream with a flag indicating whether to resume playback after interruption based on keywords in the first interaction request; wherein, the interactive terminal acquires and plays the first video stream; the interruption and resumption module 72 is used to pause the synthesis of the first video stream in response to an interruption synthesis request issued by the interactive terminal, and to synthesize a second video stream in real time in response to a second interaction request issued by the interactive terminal, and to determine whether to continue synthesizing a new first video stream from the interruption position of the first interaction decision after the second video stream synthesis is completed, based on the flag; wherein, the interruption synthesis request is sent by the interactive terminal in response to the second interaction request of the user when playing the first video stream, and the interactive terminal acquires and plays the newly synthesized video stream.
[0251] The above solution, because the virtual avatar interaction device 70 can implement the steps in the above virtual avatar interaction method embodiment, can interrupt the synthesis when a new interaction request from the user is received during the real-time synthesis of the video stream by the interaction response server and the playback of the video stream by the interaction terminal. It can first synthesize a new video stream in real time, and then determine whether to continue synthesizing the original video stream from the interruption position according to the flag of whether to continue playback. Therefore, it can greatly improve the naturalness of the cultural relics virtual avatar interaction.
[0252] In some disclosed embodiments, the first interaction decision includes time-synchronized interactive text and action commands. The virtual avatar synthesis engine performs synthesis operations based on the synthesized speech of the interactive text and the action commands. The interruption and resumption module 72 includes a time information acquisition submodule for acquiring the time information corresponding to the interruption position in the synthesized speech; the interruption and resumption module 72 includes a phoneme information acquisition submodule for acquiring the phoneme information of the virtual avatar in the first video stream at the interruption position; the interruption and resumption module 72 includes a text content determination submodule for determining, based on the time information and phoneme information, the text content in the interactive text that has not been played by the virtual avatar in the first video stream; and the interruption and resumption module 72 includes a video stream synthesis submodule for synthesizing a new first video stream through the virtual avatar synthesis engine based on the corresponding part of the text content in the synthesized speech and the residual part of the action commands after the interruption position.
[0253] Therefore, the first interactive decision includes interactive text. The virtual avatar synthesis engine performs synthesis operations based on the synthesized speech of the interactive text. When a new first video stream is determined based on the marker, the time information corresponding to the interruption position in the synthesized speech can be obtained, and the phoneme information of the virtual avatar at the interruption position in the first video stream can be obtained. Based on the time information and phoneme information, the text content in the interactive text that has not been played by the virtual avatar in the first video stream can be determined. Then, through the virtual avatar synthesis engine, the synthesis operation continues based on the corresponding part of the text content in the synthesized speech to obtain a new first video stream, thus improving the naturalness of interruption resumption.
[0254] In some disclosed embodiments, the time information acquisition submodule includes a frame number acquisition unit, used to acquire the frame number of the audio frame corresponding to the interruption position in the synthesized speech; the time information acquisition submodule includes a time analysis unit, used to obtain time information based on the frame rate and frame number of the synthesized speech.
[0255] Therefore, by obtaining the frame number of the audio frame corresponding to the interruption position in the synthesized speech, and then obtaining the time information based on the frame rate and frame number of the synthesized speech, the accuracy of the time information can be improved.
[0256] In some disclosed embodiments, the first interactive decision includes at least interactive text, the request processing module 71 includes a target determination submodule, which is used to select the matching cultural and museum exhibit as the target exhibit in response to the information system in the information processing server retrieving a matching cultural and museum exhibit based on keywords; the request processing module 71 includes a speech synthesis submodule, which is used to perform speech synthesis based on the interactive text to obtain synthesized speech; the request processing module 71 includes a model embedding submodule, which is used to embed a three-dimensional model of the target exhibit into the first video stream obtained by performing image synthesis on the synthesized speech through a virtual image synthesis engine.
[0257] Therefore, if a matching cultural and museum exhibit is found in the information system based on keywords, a 3D model of the exhibit can be embedded in the synthesized first video stream. This allows for more convenient information interaction through the 3D model, thereby enhancing the user's visitor experience.
[0258] In some disclosed embodiments, the request processing module 71 includes an action determination submodule, which is used to determine that the first interaction decision also includes an action instruction synchronized with the interaction text time, including at least a reaching gesture; the model embedding submodule is specifically used to perform image synthesis on the synthesized speech, action instruction and the three-dimensional model of the target exhibit through a virtual image synthesis engine to obtain a first video stream; wherein, the virtual image in the first video stream triggers a reaching gesture to display the three-dimensional model of the target exhibit.
[0259] Therefore, before or after selecting the matched cultural and museum exhibits as the target exhibits, it is further determined that the first interactive decision also includes action instructions synchronized with the interactive text time, including at least a reaching gesture. In the process of embedding the three-dimensional model of the target exhibit into the first video stream obtained by performing image synthesis on the synthesized speech through the virtual image synthesis engine, the synthesized speech, action instructions and the three-dimensional model of the target exhibit are image synthesized by the virtual image synthesis engine to obtain the first video stream. The virtual image in the first video stream triggers the reaching gesture to display the three-dimensional model of the target exhibit, which can improve the naturalness of the virtual image.
[0260] Please see Figure 8 , Figure 8 This is a schematic diagram of the framework of an embodiment of the interactive terminal 80 of this application. The interactive terminal 80 includes a communication circuit 81, a memory 82, and a processor 83. The communication circuit 81 and the memory 82 are respectively coupled to the processor 83. The memory 82 stores program instructions, and the processor 83 is used to execute the program instructions to implement the steps in the above-described virtual avatar interaction method embodiment. Specifically, the interactive terminal 80 may include, but is not limited to, desktop computers, laptops, tablet computers, self-service terminals, etc., and is not limited thereto.
[0261] Specifically, processor 83 controls itself, as well as communication circuit 81 and memory 82, to implement the steps in the above-described virtual avatar interaction method embodiments. Processor 83 can also be called a CPU (Central Processing Unit). Processor 83 may be an integrated circuit chip with signal processing capabilities. Processor 83 can also be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor can be a microprocessor or any conventional processor. Furthermore, processor 83 can be implemented using integrated circuit chips.
[0262] The above solution, since the interactive terminal 80 can implement the steps in the above virtual image interaction method embodiment, can interrupt the synthesis when a new interaction request from the user occurs during the real-time synthesis of the video stream by the interactive response server and the playback of the video stream by the interactive terminal. It can first synthesize a new video stream in real time, and then determine whether to continue synthesizing the original video stream from the interruption position according to the flag of whether to continue playback. Therefore, it can greatly improve the naturalness of the interaction of cultural relics virtual images.
[0263] Please see Figure 9 , Figure 9 This is a schematic diagram of the framework of an embodiment of the interactive response server 90 of this application. The interactive response server 90 includes a communication circuit 91, a memory 92, and a processor 93. The communication circuit 91 and the memory 92 are respectively coupled to the processor 93. The memory 92 stores program instructions, and the processor 93 is used to execute the program instructions to implement the steps in the above-described virtual avatar interaction method embodiment.
[0264] Specifically, processor 93 controls itself, as well as communication circuit 91 and memory 92, to implement the steps in the above-described virtual avatar interaction method embodiments. Processor 93 can also be called a CPU (Central Processing Unit). Processor 93 may be an integrated circuit chip with signal processing capabilities. Processor 93 can also be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor can be a microprocessor or any conventional processor. Furthermore, processor 93 can be implemented using integrated circuit chips.
[0265] The above solution, because the interactive response server 90 can implement the steps in the above virtual image interaction method embodiment, can interrupt the synthesis when a new interaction request from the user occurs during the real-time synthesis of the video stream on the interactive response server and the playback of the video stream on the interactive terminal. It can first synthesize a new video stream in real time, and then determine whether to continue synthesizing the original video stream from the interruption position according to the flag for whether to continue playback. Therefore, it can greatly improve the naturalness of the interaction of cultural relics virtual images.
[0266] Please see Figure 10 , Figure 10 This is a flowchart illustrating an embodiment of the interactive system testing method of this application. It should be noted that the interactive system testing method in this embodiment is applied to the aforementioned virtual avatar interactive system, as detailed in the relevant descriptions in the foregoing embodiments. Specifically, it may include the following steps:
[0267] Step S101: Input test data into the test driver interface of the interactive terminal in the virtual avatar interaction system.
[0268] In this embodiment of the disclosure, when the test data is video data, it is split into audio data and image data by the test driver interface. It should be noted that for details regarding the test driver interface and test data, please refer to the relevant descriptions in the foregoing embodiments of the disclosure, which will not be repeated here.
[0269] Step S102: Obtain sampling data related to test indicators during the interactive response process of the virtual avatar interaction system based on test data.
[0270] In one implementation scenario, if the test metrics include the interaction success rate, the number of successful interactions and the total number of interactions can be obtained.
[0271] In one implementation scenario, when the test metrics include the real-time rate of video stream synthesis, the synthesis duration of the audio and the synthesis duration of the video stream can be obtained when synthesizing each video stream.
[0272] In one implementation scenario, when response time is included as a test metric, the moment in the audio data that represents the user stopping speaking and the moment in the synthesized video stream that the virtual avatar begins to respond can be obtained.
[0273] It should be noted that the above examples are only a few possible implementation methods in actual application. When the test indicators are set to other indicators, the same principle can be applied. No further examples will be given here.
[0274] Step S103: Based on the sampled data, obtain the test value of the virtual avatar interaction system for the test index.
[0275] In one implementation scenario, if the test metric includes the interaction success rate, the percentage of successful interactions out of the total number of interactions can be used as the test value for the interaction success rate.
[0276] In one implementation scenario, when the test metric includes the real-time rate of video stream synthesis, the ratio of the synthesized speech duration to the video stream synthesis duration can be obtained as the real-time rate of a single video stream synthesis, and the average of the real-time rates of multiple video stream synthesis can be obtained as the test value of the real-time rate of video stream synthesis.
[0277] In one implementation scenario, when response time is included as a test metric, the difference between the moment in the audio data when the user stops speaking and the moment in the synthesized video stream when the virtual avatar begins to respond can be used as the test value for response time.
[0278] It should be noted that the above examples are only a few possible implementation methods in actual application. When the test indicators are set to other indicators, the same principle can be applied. No further examples will be given here.
[0279] Step S104: Based on the test values of the virtual avatar interaction system on each test indicator, determine whether the virtual avatar interaction system has passed the test.
[0280] For example, if the test values of the virtual avatar interaction system on all test indicators indicate that the test has passed, then the virtual avatar interaction system can be determined to have passed the test. Conversely, if the test value on at least one test indicator indicates that the test has failed, then the virtual avatar interaction system can be determined to have failed the test.
[0281] The above solution, because the interactive system testing method in this embodiment is applied to the virtual avatar interactive system, and during testing, test data is input to the test driver interface of the interactive terminal in the virtual avatar interactive system, when the test data is video data, it is split into audio data and image data by the test driver interface, and then the sampling data related to the test indicators is obtained during the interactive response of the virtual avatar interactive system based on the test data. Then, based on the sampling data, the test value of the virtual avatar interactive system on the test indicator is obtained, and finally, based on the test value of the virtual avatar interactive system on each test indicator, it is determined whether the virtual avatar interactive system has passed the test, which can help improve the test accuracy.
[0282] Please see Figure 11 , Figure 11 This is a flowchart illustrating an embodiment of the interactive system testing device 1100 of this application. It should be noted that the interactive system testing device 1100 in this embodiment is applied to the aforementioned virtual avatar interactive system, as detailed in the relevant descriptions in the foregoing embodiments. Specifically, it may include an input module 1101, an acquisition module 1102, a calculation module 1103, and a determination module 1104. The input module 1101 is used to input test data to the test driver interface of the interactive terminal in the virtual avatar interactive system; wherein, when the test data is video data, it is split into audio data and image data by the test driver interface. The acquisition module 1102 is used to acquire sampled data related to test indicators during the interactive response process of the virtual avatar interactive system based on the test data. The calculation module 1103 is used to obtain the test values of the virtual avatar interactive system on the test indicators based on the sampled data. The determination module 1104 is used to determine whether the virtual avatar interactive system has passed the test based on the test values of the virtual avatar interactive system on each test indicator.
[0283] The above solution, in this embodiment of the present disclosure, uses the interactive system testing device 1100 applied to the virtual avatar interactive system. During testing, test data is input to the test driver interface of the interactive terminal in the virtual avatar interactive system. When the test data is video data, it is split into audio data and image data by the test driver interface. Then, sampling data related to the test indicators is obtained during the interactive response process of the virtual avatar interactive system based on the test data. Based on the sampling data, the test value of the virtual avatar interactive system on the test indicator is obtained. Finally, based on the test value of the virtual avatar interactive system on each test indicator, it is determined whether the virtual avatar interactive system has passed the test, which helps to improve the test accuracy.
[0284] In some disclosed embodiments, when the test metric includes the interaction success rate, the acquisition module 1102 includes a first acquisition submodule for acquiring the number of successful interactions and the total number of interactions; the calculation module 1103 includes a first calculation submodule for acquiring the percentage of the number of successful interactions in the total number of interactions, as the test value of the interaction success rate.
[0285] In some disclosed embodiments, when the test metric includes the real-time rate of video stream synthesis, the acquisition module 1102 includes a second acquisition submodule, used to acquire the synthesized speech duration and video stream synthesis duration used when synthesizing each video stream; the calculation module 1103 includes a second calculation submodule, used to acquire, for each video stream, the ratio of the synthesized speech duration to the video stream synthesis duration used when synthesizing the video stream as the real-time rate of a single video stream, and to acquire the average of the real-time rates of multiple video streams as the test value of the real-time rate of video stream synthesis.
[0286] In some disclosed embodiments, when the test metric includes response time, the acquisition module 1102 includes a third acquisition submodule for acquiring the moment in the audio data that represents the user stopping speaking and the moment in the synthesized video stream that the virtual avatar begins to respond; the calculation module 1103 includes a third calculation submodule for acquiring the difference between the moment that represents the user stopping speaking and the moment in the synthesized video stream that the virtual avatar begins to respond, as the test value of the response time.
[0287] Please see Figure 12 , Figure 12 This is a schematic diagram of a framework of an embodiment of the computer-readable storage medium 1200 of this application. The computer-readable storage medium 1200 stores program instructions 1201 that can be executed by a processor. The program instructions 1201 are used to implement the steps in any of the above-described embodiments of the virtual avatar interaction method.
[0288] The above-described solution, the computer-readable storage medium 1200, can implement the steps in any of the above-described virtual character interaction method embodiments. Therefore, during the process of real-time synthesis of video streams by the interactive response server and playback of the video streams by the interactive terminal, the synthesis can be interrupted by a new interactive request from the user. A new video stream can be synthesized in real time first, and then the original video stream can be resumed from the interruption position based on the flag indicating whether to continue playback. Therefore, the naturalness of the interaction of cultural relics virtual characters can be greatly improved.
[0289] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.
[0290] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to, and for the sake of brevity, they will not be repeated here.
[0291] In the several embodiments provided in this application, it should be understood that the disclosed methods and apparatus can be implemented in other ways. For example, the apparatus implementations described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0292] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0293] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0294] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0295] If the technical solution of this application involves personal information, the product using this technical solution has clearly informed the user of the personal information processing rules and obtained the user's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using this technical solution has obtained the user's separate consent before processing the sensitive personal information, and also meets the requirement of "express consent". For example, at personal information collection devices such as cameras, clear and prominent signs are set up to inform users that they have entered the scope of personal information collection and that personal information will be collected. If an individual voluntarily enters the collection scope, it is deemed that they have agreed to the collection of their personal information; or on the personal information processing device, with clear signs / information informing users of the personal information processing rules, authorization is obtained from the individual through pop-up information or by asking the individual to upload their personal information; wherein, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.
Claims
1. A virtual avatar interaction system, characterized in that, The system includes a cultural heritage interactive terminal, an interactive response server, and an information processing server. The interactive terminal is communicatively connected to the interactive response server, and the interactive response server is communicatively connected to the information processing server. The information processing server is equipped with an information system for the interactive response server to retrieve information during interactive decision-making. The interactive terminal is used to interact with the user to obtain the user's input data, and to obtain and play a video stream from the interactive response server. The input data includes at least one of voice data and image data. The interactive response server is used to make interactive decisions based on the input data, obtain interactive decision results, and the interactive decision results include time-synchronized interactive text and action instructions. Based on the synthesized speech of the interactive text and the action instructions, a video stream is synthesized, and the mouth movements of the virtual image in the video stream are consistent with the synthesized speech in time, and the body movements are consistent with the action instructions in time. The interactive response server generates a first interactive decision in response to a first interactive request issued by a user through the interactive terminal, and synthesizes a first video stream in real time through a virtual image synthesis engine based on the first interactive decision. The interactive response server also maintains a relational mapping set and uses the relational mapping set as a marker to indicate whether the first video stream is resumed after interruption based on keywords in the first interactive request. The relational mapping set contains the mapping relationship between words and whether the video stream is resumed after interruption. In response to a second interaction request from a user while playing the first video stream, the interactive terminal sends an interruption request and the second interaction request to the interactive response server. The interactive response server responds to the interruption synthesis request by pausing the synthesis of the first video stream, and responds to the second interactive request by synthesizing the second video stream in real time. After the synthesis of the second video stream is completed, it determines, based on the flag, whether to continue synthesizing a new first video stream from the interruption position of the first interactive decision. The interactive terminal acquires and plays the newly synthesized video stream from the interactive response server; wherein, when the flag indicates that subsequent playback has been interrupted, the newly synthesized video stream includes the second video stream and the new first video stream following the second video stream.
2. The system according to claim 1, characterized in that, The interactive terminal is equipped with a test driver interface, which is used to input test data when testing the virtual avatar interactive system, and to split the test data into audio data and image data when the test data is video data.
3. The system according to claim 1 or 2, characterized in that, The interactive response server includes: A speech recognition interface is used to recognize the speech data and obtain the recognized text; A semantic understanding interface is used to understand the identified text and obtain the interaction intent; An interactive decision interface is used to retrieve response information from the information system in the information processing server based at least on the interactive intent, and to perform decision processing based on the response information to obtain the interactive decision result. A speech synthesis interface is used to perform speech synthesis based on the interactive text in the interactive decision results to obtain synthesized speech. The image synthesis interface integrates a virtual image synthesis engine, which is used to generate a video stream driven by at least one of the synthesized voice and action commands.
4. The system according to claim 3, characterized in that, The speech synthesis interface is also used to obtain the exhibition theme of the exhibition area / hall where the interactive terminal is located from the interactive terminal, and to perform speech synthesis based on the exhibition theme and the interactive text in the interactive decision result to obtain synthesized speech that matches the exhibition theme.
5. The system according to claim 1 or 2, characterized in that, The interactive terminal includes: A voice wake-up interface is used to wake up the interactive terminal when a wake-up word is detected in the voice data, so as to display a virtual avatar on the interactive terminal and interact with the user; and / or, A face wake-up interface is used to wake up the interactive terminal when a registered face is detected, so as to display a virtual avatar on the interactive terminal and interact with the user; and / or, A gesture recognition interface is used to identify gesture types and provide the identified gesture categories to the interaction decision interface in the interaction response server, so that the interaction decision interface can retrieve response information from the information system in the information processing server based on the interaction intent and gesture category.
6. A method for testing an interactive system, characterized in that, The method for testing the virtual avatar interaction system according to any one of claims 1 to 5 includes: Input test data into the test driver interface of the interactive terminal in the virtual avatar interaction system; wherein, when the test data is video data, it is split into audio data and image data by the test driver interface; Acquire sampling data related to test indicators during the interactive response process of the virtual avatar interaction system based on the test data; Based on the sampled data, the test values of the virtual avatar interaction system on the test indicators are obtained; Based on the test values of the virtual avatar interaction system on various test indicators, it is determined whether the virtual avatar interaction system has passed the test.
7. The method according to claim 6, characterized in that, When the test metric includes the interaction success rate, the step of obtaining sampled data related to the test metric during the interaction response process of the virtual avatar interaction system based on the test data includes: Get the number of successful interactions and the total number of interactions; The process of obtaining the test value of the virtual avatar interaction system on the test index based on the sampled data includes: The percentage of successful interactions in the total number of interactions is obtained and used as a test value for the interaction success rate.
8. The method according to claim 6, characterized in that, When the test metric includes the real-time rate of video stream synthesis, the step of acquiring sampling data related to the test metric during the interactive response process of the virtual avatar interaction system based on the test data includes: Obtain the synthesized speech duration and video stream synthesis duration used when synthesizing each video stream; The process of obtaining the test value of the virtual avatar interaction system on the test index based on the sampled data includes: For each video stream, the ratio of the synthesized speech duration to the video stream synthesis duration is obtained and used as the real-time synthesis rate of a single video stream. The average real-time rate of multiple video streams is obtained as the test value of the real-time rate of the video stream synthesis.
9. The method according to claim 6, characterized in that, When response time is included as a test metric, the acquisition of sampled data related to the test metric during the interaction response process of the virtual avatar interaction system based on the test data includes: Obtain the moment in the audio data that represents the moment the user stops speaking and the moment in the synthesized video stream that the virtual avatar begins to respond; The process of obtaining the test value of the virtual avatar interaction system on the test index based on the sampled data includes: The difference between the moment when the user stops speaking and the moment when the virtual avatar in the synthesized video stream begins to respond is obtained as the test value of the response time.
10. A virtual avatar interaction method, characterized in that, include: The system acquires and plays a first video stream; wherein, the interactive response server generates a first interactive decision in response to a first interactive request issued by a user through an interactive terminal, and synthesizes the first video stream in real time through a virtual avatar synthesis engine based on the first interactive decision. The interactive response server also maintains a relation mapping set, and refers to the relation mapping set to mark the first video stream with a flag indicating whether playback should resume after interruption based on keywords in the first interactive request. The relation mapping set contains the mapping relationship between words and whether playback should resume after interruption. In response to a second interaction request from a user while playing the first video stream, an interruption synthesis request and the second interaction request are sent to the interaction response server; wherein, the interaction response server pauses the synthesis of the first video stream in response to the interruption synthesis request, and synthesizes the second video stream in real time in response to the second interaction request, and after the synthesis of the second video stream is completed, it determines, based on the flag, whether to continue synthesizing a new first video stream from the interruption position of the first interaction decision; Acquire and play the newly synthesized video stream from the interactive response server; wherein, in the case where the flag indicates an interruption of subsequent playback, the newly synthesized video stream includes the second video stream and the new first video stream following the second video stream.
11. The method according to claim 10, characterized in that, Prior to acquiring and playing the first video stream, the method further includes: In response to a gaze that exceeds the duration threshold but is not detected if either the registered face or the wake word is not detected, the system switches to the wake state and outputs prompts to guide user interaction.
12. The method according to claim 11, characterized in that, Before switching to the wake-up state, the method further includes: Detect lip key points in each frame of a video captured from a user; Based on the key lip points in the image, determine the distance between the upper and lower lips in the image; Count the number of image frames in which the distance between the upper and lower lips is greater than a distance threshold; Specifically, if the number of frames exceeds the quantity threshold, the system switches to the wake-up state; if the number of frames does not exceed the quantity threshold, the system remains in sleep mode.
13. The method according to claim 12, characterized in that, The distance threshold is obtained by statistically analyzing the distance between the upper and lower lips in each frame of the image.
14. The method according to claim 10, characterized in that, The method further includes: In response to the recognition of the switching gesture, the display switches to show the 3D model of the next cultural and museum exhibit; And / or, in response to recognizing a thumbs-up gesture while displaying a 3D model of a cultural and museum exhibit, a preset score is added to the currently displayed cultural and museum exhibit; wherein, the 3D models of each cultural and museum exhibit are displayed sequentially based on the size of their respective accumulated scores.
15. The method according to claim 10, characterized in that, The method further includes: In response to the detection of a registered face, the system obtains the visit route of the user to whom the registered face belongs, and based on the location of the interactive terminal that detected the registered face and the visit route, determines the next cultural and museum exhibit that the user will visit, and displays a third video stream. The third video stream is synthesized in real time by the interactive response server based on the next cultural and museum exhibit to be visited through the virtual image synthesis engine, and the virtual image in the third video stream indicates the location information of the next cultural and museum exhibit to be visited.
16. The method according to claim 15, characterized in that, The method further includes: In response to the viewing request from the user to whom the registered face belongs, the system displays the user's viewing progress along their tour route and displays a fourth video stream. The fourth video stream is synthesized in real time by the interactive response server based on the visit progress using the virtual avatar synthesis engine, and at least one of the virtual avatar's expression, action, and voice in the fourth video stream matches the visit progress.
17. The method according to claim 15, characterized in that, The method further includes: In response to the termination of the interactive Q&A session with the user to whom the registered face belongs regarding cultural heritage preferences, the tour route is generated based on the interactive Q&A session, and the fifth video stream is displayed; The fifth video stream is synthesized in real time by the interactive response server based on the first cultural and museum exhibit visited in the tour route using a virtual image synthesis engine, and the virtual image in the fifth video stream indicates the location information of the first cultural and museum exhibit visited.
18. A virtual avatar interaction method, characterized in that, include: Based on the first interaction request issued by the interactive terminal, a first interaction decision is generated, and based on the first interaction decision, a first video stream is synthesized through a virtual image synthesis engine. A reference relation mapping set is used to mark the first video stream with a flag indicating whether playback should resume after interruption based on the keywords in the first interaction request. The interactive terminal acquires and plays the first video stream, and the interactive response server also maintains a relation mapping set, which contains the mapping relationship between words and whether playback should resume after interruption. In response to an interruption request from the interactive terminal, the synthesis of the first video stream is paused, and in response to a second interaction request from the interactive terminal, a second video stream is synthesized in real time. After the synthesis of the second video stream is completed, based on the flag, it is determined whether to continue synthesizing a new first video stream from the interruption position of the first interaction decision. The interruption request is sent by the interactive terminal in response to a second interaction request from the user while playing the first video stream, and the interactive terminal acquires and plays the newly synthesized video stream. When the flag indicates interruption followed by playback, the newly synthesized video stream includes the second video stream and the new first video stream following the second video stream.
19. The method according to claim 18, characterized in that, The first interactive decision includes time-synchronized interactive text and action commands. The virtual avatar synthesis engine performs a synthesis operation based on the synthesized speech of the interactive text and the action commands. If, based on the identifier, it is determined that the new first video stream will be synthesized, the method further includes: Obtain the time information corresponding to the interruption position in the synthesized speech, and obtain the phoneme information of the virtual image at the interruption position in the first video stream; Based on the time information and the phoneme information, determine the text content in the interactive text that has not been broadcast by the virtual image in the first video stream; The new first video stream is synthesized using the virtual avatar synthesis engine, based on the corresponding part of the text content in the synthesized speech and the remaining part of the action command after the interruption position.
20. The method according to claim 19, characterized in that, The step of obtaining the time information corresponding to the interruption position in the synthesized speech includes: Obtain the frame number of the audio frame corresponding to the interruption position in the synthesized speech; The time information is obtained based on the frame rate and frame number of the synthesized speech.
21. The method according to claim 18, characterized in that, The first interactive decision includes at least interactive text, and the synthesis of the first video stream based on the first interactive decision using a virtual avatar synthesis engine includes: In response to the information system in the information processing server retrieving a matching cultural and museum exhibit based on the keyword, the matching cultural and museum exhibit is taken as the target exhibit, and speech synthesis is performed based on the interactive text to obtain synthesized speech. The three-dimensional model of the target exhibit is embedded in the first video stream obtained by synthesizing the image of the synthesized speech through the virtual image synthesis engine.
22. The method according to claim 21, characterized in that, Before or after selecting the matched cultural relics exhibit as the target exhibit, the method further includes: Determining the first interaction decision also includes action instructions synchronized with the interaction text time, including at least a reaching gesture; The first video stream obtained by synthesizing the synthesized speech using the virtual image synthesis engine, embedding a 3D model of the target exhibit, includes: The first video stream is obtained by synthesizing the synthesized speech, the action commands, and the three-dimensional model of the target exhibit using a virtual avatar synthesis engine. In this process, the virtual image in the first video stream triggers the reaching gesture to display the three-dimensional model of the target exhibit.
23. An interactive system testing device, characterized in that, The apparatus for testing the virtual avatar interaction system according to any one of claims 1 to 5, the apparatus comprising: An input module is used to input test data into the test driver interface of the interactive terminal in the virtual avatar interaction system; wherein, when the test data is video data, it is split into audio data and image data by the test driver interface; The acquisition module is used to acquire sampled data related to test indicators during the interactive response process of the virtual avatar interaction system based on the test data; The calculation module is used to obtain the test value of the virtual avatar interaction system on the test index based on the sampled data; The determination module is used to determine whether the virtual avatar interaction system has passed the test based on the test values of the virtual avatar interaction system on various test indicators.
24. A virtual avatar interaction device, characterized in that, include: The first acquisition module is used to acquire and play a first video stream; wherein, the interactive response server generates a first interactive decision in response to a first interactive request issued by a user through an interactive terminal, and synthesizes the first video stream in real time through a virtual image synthesis engine based on the first interactive decision. The interactive response server also maintains a relation mapping set, and refers to the relation mapping set to mark the first video stream with a flag indicating whether playback should resume after interruption based on keywords in the first interactive request. The relation mapping set contains the mapping relationship between words and whether playback should resume after interruption. The request sending module is configured to, in response to a second interaction request from a user while playing the first video stream, send an interruption synthesis request and the second interaction request to the interaction response server; wherein, the interaction response server suspends the synthesis of the first video stream in response to the interruption synthesis request, and synthesizes the second video stream in real time in response to the second interaction request, and after the synthesis of the second video stream is completed, determines, based on the flag, whether to continue synthesizing a new first video stream from the interruption position of the first interaction decision; The second acquisition module is used to acquire and play the newly synthesized video stream from the interactive response server; wherein, in the case where the flag indicates that the subsequent playback has been interrupted, the newly synthesized video stream includes the second video stream and the new first video stream following the second video stream.
25. A virtual avatar interaction device, characterized in that, include: The request processing module is used to generate a first interaction decision based on a first interaction request issued by the interactive terminal, and to synthesize a first video stream through a virtual image synthesis engine based on the first interaction decision. It also includes a reference relation mapping set that uses keywords from the first interaction request to mark whether the first video stream should resume playback after an interruption. The interactive terminal acquires and plays the first video stream, and the interactive response server maintains a relation mapping set containing mapping relationships between words and whether playback should resume after an interruption. The interruption and resumption module is configured to, in response to an interruption request from the interactive terminal, pause the synthesis of the first video stream, and in response to a second interaction request from the interactive terminal, synthesize a second video stream in real time. After the second video stream is synthesized, based on the flag, determine whether to resume synthesizing a new first video stream from the interruption point of the first interaction decision. The interruption request is sent by the interactive terminal in response to a second interaction request from the user while playing the first video stream, and the interactive terminal acquires and plays the newly synthesized video stream. When the flag indicates interruption and resumption, the newly synthesized video stream includes the second video stream and the new first video stream following the second video stream.
26. An interactive terminal, characterized in that, The device includes a communication circuit, a memory, and a processor. The communication circuit and the memory are respectively coupled to the processor. The memory stores program instructions, and the processor is used to execute the program instructions to implement the virtual avatar interaction method according to any one of claims 10 to 17.
27. An interactive response server, characterized in that, The device includes a communication circuit, a memory, and a processor. The communication circuit and the memory are respectively coupled to the processor. The memory stores program instructions, and the processor is used to execute the program instructions to implement the virtual avatar interaction method according to any one of claims 18 to 22.
28. A computer-readable storage medium, characterized in that, The device stores program instructions that can be executed by a processor, the program instructions being used to implement the virtual avatar interaction method according to any one of claims 6 to 22.