A zero-configuration adaptive virtual digital human real-time interaction method and system thereof

CN122802709APending Publication Date: 2026-09-22JIANGSU NANDA CULTURE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510329051.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

然而,现有的虚拟数字人直播系统仍存在一些不足,导致无法获得更好的用户体验和实时互动多样性

Benefits of technology

[0039]与现有技术相比,本发明所达到的有益效果:本发明提供的零配置自适应的虚拟数字人实时交互方法及其系统,通过获取历史直播数据和用户互动反馈,系统能够为用户提供高度个性化的交互体验;实时采集用户的面部表情和语音数据,分析用户情绪并作出相应调整,虚拟数字人可以根据用户的偏好和历史行为动态调整直播内容和风格,使用户感受到更加贴心的服务。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122802709A_ABST
    Figure CN122802709A_ABST
Patent Text Reader

Abstract

The application discloses a kind of zero configuration adaptive virtual digital human real-time interaction method and system thereof, term e-commerce virtual live broadcast technical field, comprising: obtaining the historical live data of e-commerce live virtual digital person, commodity service image label;Based on time series sampling, the live feedback intention of different terminal users is obtained;The first virtual audio and video stream is generated using the historical live data and commodity service image label;The second virtual audio and video stream is generated by language optimization adjustment of the first virtual audio and video stream using the live feedback intention, and pushed to multi-user terminal.The application generates and optimizes virtual audio and video stream by automatically obtaining and processing historical live data and user feedback intention, realizes adaptive live interaction without manual configuration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of e-commerce virtual live streaming technology, specifically relating to a zero-configuration adaptive real-time interaction method and system for virtual digital humans. Background Technology

[0002] With the rapid development of e-commerce live streaming, virtual digital human technology is increasingly being applied to live streaming scenarios to enhance user experience and interactivity. However, existing virtual digital human live streaming systems still have some shortcomings, resulting in an inability to achieve a better user experience and diverse real-time interactions. For example, the interactivity of virtual digital humans is not natural enough, and they cannot respond to user needs in real time; the appearance and content of virtual digital humans are relatively simple; interaction with users is limited to simple questions and answers, making it difficult to handle complex emotions and contexts; the interaction forms are relatively monotonous, lacking diverse interaction scenarios; and it is difficult to adjust live streaming content and strategies in real time based on user feedback. Summary of the Invention

[0003] The purpose of this invention is to overcome the shortcomings of the prior art and provide a zero-configuration adaptive virtual digital human real-time interaction method and system. By automatically acquiring and processing historical live broadcast data and user feedback intent, it generates and optimizes virtual audio and video streams to achieve adaptive live broadcast interaction without manual configuration.

[0004] To achieve the above objectives, the present invention is implemented using the following technical solution:

[0005] In a first aspect, the present invention provides a zero-configuration adaptive real-time interaction method for virtual digital humans, comprising:

[0006] Acquire historical live streaming data and product / service profile tags of virtual digital humans in e-commerce live streaming;

[0007] The live streaming feedback intent of different terminal users is obtained based on time series sampling;

[0008] A first virtual audio and video stream is generated using the historical live streaming data and product / service profile tags;

[0009] The first virtual audio and video stream is optimized and adjusted using the live feedback intent to generate a second virtual audio and video stream, which is then pushed to a multi-user terminal.

[0010] Furthermore, generating a first virtual audio and video stream using the historical live streaming data and product / service profile tags includes:

[0011] The first virtual live streaming dataset is generated by extracting live streaming content data, live streaming emotion data, and live streaming image data from the historical live streaming data of virtual digital humans, and then preprocessing the data through data classification matrix extraction.

[0012] The system obtains historical product and service profiles from the historical live streaming data of virtual digital humans, as well as product and service profile tags from the current e-commerce live streaming. After data extraction and preprocessing, a second virtual live streaming dataset is generated.

[0013] The third virtual live streaming dataset is generated by extracting and preprocessing user interaction feedback data from the historical live streaming data of virtual digital humans.

[0014] Based on the first virtual live streaming dataset, a deep learning network model is used to perform virtual rendering on the virtual digital human to obtain the first live streaming data packet.

[0015] Based on the second virtual live streaming dataset, the third virtual live streaming dataset, and product / service profile tags, a live streaming script is generated using natural language processing technology. After audio generation, it is fused with the first live streaming data packet to form a first virtual audio / video stream.

[0016] Furthermore, the second virtual audio and video stream is generated by optimizing and adjusting the language of the first virtual audio and video stream using the live broadcast feedback intent, including:

[0017] In response to the triggering operation of the live broadcast feedback intent, obtain the live broadcast emotion feedback tags of different users to the e-commerce live broadcast virtual digital human, and match and confirm the preset live broadcast optimization strategy corresponding to the live broadcast emotion feedback tags;

[0018] The virtual digital human's live streaming behavior is adjusted in real time using the preset live streaming optimization strategy, including at least the live streaming content, live streaming actions, and live streaming emotions, and a new second virtual audio and video stream is rendered and generated.

[0019] Furthermore, based on time-series sampling, the live streaming feedback intent of different terminal users is obtained, including:

[0020] In the e-commerce live streaming screen, at least one live interactive topic and multiple alternative options corresponding to the topic are displayed according to a preset spatial layout arrangement rule;

[0021] In response to different users' triggering operations on the live interactive topics, the alternative options selected by different users are obtained, and the first live feedback tag reflecting the user is matched based on the statistical analysis of the alternative options.

[0022] In response to interactive video commands from end users in e-commerce live streaming, the system collects facial expression data and / or voice data from end users while watching virtual digital human e-commerce live streams in real time, and matches second live stream feedback tags reflecting the user based on statistical analysis of the facial expression data and / or voice data.

[0023] The live feedback intent is generated by matching the first live feedback tag and the second live feedback tag with a preset intent matrix.

[0024] Furthermore, within a preset time period, several consecutive sampling time intervals are divided according to the time sequence, and the alternative options are matched sequentially with at least three priorities according to the order of the sampling time intervals in which the user's response trigger time is located.

[0025] Multi-threading technology is used to process multiple alternative options in different sampling time intervals in parallel and to statistically analyze the number of different priorities. The alternative option with the highest proportion among the first or first two priorities within a preset time period is used to match the first live feedback tag reflecting the user.

[0026] Furthermore, the e-commerce live stream displays methods for triggering interactive commands initiated by end users, including:

[0027] In response to an interactive command triggered by an end user, at least one interactive content list page is displayed. The interactive content list page sequentially displays content including at least e-commerce live streaming scenarios, virtual digital human images, product and service profile tags, and live streaming interactive topics.

[0028] In response to a user's triggering action on the interactive content list page, the first virtual audio and video stream is rendered and adjusted using the options obtained from the interactive content list page.

[0029] Secondly, the present invention provides a zero-configuration adaptive real-time interactive system for virtual digital humans, comprising:

[0030] The adaptive module uses an adaptive control algorithm to adjust the system configuration parameters in real time.

[0031] The data acquisition module obtains historical live streaming data of e-commerce live streaming virtual digital humans, product and service profile tags, and obtains the live streaming feedback intent of different terminal users based on time series sampling;

[0032] The interactive module generates a first virtual audio and video stream using the historical live streaming data and product / service profile tags; and optimizes the language of the first virtual audio and video stream using the live streaming feedback intent to generate a second virtual audio and video stream, which is then pushed to a multi-user terminal.

[0033] Thirdly, the present invention provides an electronic device, comprising:

[0034] processor;

[0035] Memory used to store the processor's executable instructions;

[0036] The processor is configured to execute the instructions to implement the zero-configuration adaptive real-time interaction method for virtual digital humans as described in any one of the first aspects.

[0037] Fourthly, the present invention provides a computer-readable storage medium that, when instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to perform the zero-configuration adaptive real-time interaction method for virtual digital humans as described in any one of the first aspects.

[0038] Fifthly, the present invention provides a computer program product comprising a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps in the zero-configuration adaptive real-time interaction method for virtual digital humans as described in any first aspect.

[0039] Compared with the prior art, the beneficial effects achieved by the present invention are as follows: The zero-configuration adaptive virtual digital human real-time interaction method and system provided by the present invention can provide users with a highly personalized interactive experience by acquiring historical live broadcast data and user interaction feedback; by collecting users' facial expressions and voice data in real time, analyzing users' emotions and making corresponding adjustments, the virtual digital human can dynamically adjust the live broadcast content and style according to users' preferences and historical behavior, so that users can feel more considerate service. Attached Figure Description

[0040] Figure 1 A flowchart of a zero-configuration adaptive real-time interaction method for virtual digital humans provided in an embodiment of the present invention.

[0041] Figure 2 This is a flowchart for generating a first virtual audio and video stream, provided as an embodiment of the present invention.

[0042] Figure 3 This is a flowchart for generating a second virtual audio and video stream, provided as an embodiment of the present invention.

[0043] Figure 4 This is a flowchart for obtaining the live streaming feedback intent of different terminal users, provided as an embodiment of the present invention.

[0044] Figure 5 A flowchart of the first live feedback tag provided in an embodiment of the present invention.

[0045] Figure 6 This is a flowchart illustrating a method for triggering interactive commands initiated by end users, as provided in an embodiment of the present invention. Detailed Implementation

[0046] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and should not be construed as limiting the scope of protection of the present invention. The components of the embodiments of the present disclosure described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present disclosure provided in the accompanying drawings is not intended to limit the scope of the claimed disclosure, but merely represents selected embodiments of the present disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of the present disclosure without inventive effort are within the scope of protection of the present disclosure.

[0047] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0048] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0049] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0050] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0051] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0052] like Figure 1 As shown, this embodiment of the invention provides a zero-configuration adaptive real-time interaction method for virtual digital humans, including the following steps:

[0053] Step S1: Obtain historical live streaming data and product / service profile tags of the e-commerce live streaming virtual digital human.

[0054] This step involves acquiring audio and video content, livestream scripts, user interaction records (such as bullet comments, reviews, gifts, etc.) from the virtual digital human's past e-commerce livestreams, as well as feedback data from viewers regarding the e-commerce livestream content or products. It also involves extracting product feature tags, including category, function, target user group, and price range, and generating a tagging system based on service characteristics (such as after-sales guarantee and delivery method).

[0055] For example, it collects historical data from virtual digital humans' past live streams of smartwatches, including their speech, actions, and expressions; it also collects user feedback data on the smartwatches, including "battery life" and "feature introductions." Furthermore, it extracts past product and service profile tags, including "activity tracking" and "health features."

[0056] In step S1, the acquired historical data of the e-commerce live streaming virtual digital human and product service profile tags are preprocessed to ensure the quality of subsequent virtual audio and video stream generation.

[0057] First, data cleaning removes missing values, outliers, and duplicate records to ensure data integrity and accuracy. Next, the data undergoes formatting, standardization, and normalization to adapt it to subsequent processing procedures.

[0058] Text data needs to undergo operations such as word segmentation and stop word removal to extract key information.

[0059] Voice data needs to undergo noise reduction processing to improve voice quality.

[0060] Secondly, through feature extraction and annotation, product / service profile tags and user feedback are categorized. For example, "battery life" is labeled as "feature introduction." Simultaneously, sentiment data is annotated to support emotional interaction. Data augmentation generates more diverse voice and text samples to improve the model's generalization ability.

[0061] Finally, by using multimodal data fusion, the timing alignment and feature fusion of audio, video and text data are ensured, generating a high-quality first virtual audio and video stream. The preprocessing operation not only improves the usability of the data, but also enhances the personalization and interactivity of the virtual digital human live stream.

[0062] Step S2: As Figure 4 As shown, the live streaming feedback intent of different terminal users is obtained based on time series sampling.

[0063] In this step, the live streaming feedback intent of different terminal users is obtained based on time-series sampling, including:

[0064] Step S2-1: Display at least one live interactive topic and multiple alternative options corresponding to the topic in the e-commerce live broadcast screen according to the preset spatial layout arrangement rules. These alternative options can be used to collect user feedback in real time during the live broadcast.

[0065] In addition, this step can also collect other user feedback during the live stream in real time, such as but not limited to bullet comments, comments, likes, and shares. However, the collected feedback needs to be preprocessed, such as extracting keywords, to meet the requirements of the natural language processing model.

[0066] Step S2-2: In response to different users' triggering operations on live interactive topics, obtain the alternative options selected by different users, and match the first live feedback tag reflecting the user based on the statistical analysis of the alternative options.

[0067] In this step, based on all the options selected by the users, the user's tendency feedback is statistically analyzed, such as positive or negative, like or dislike, etc., which is used as a tendency-based primary live stream feedback tag.

[0068] Steps S2-3: In response to interactive video commands from end users during the e-commerce live stream, real-time facial expression data and / or voice data of end users watching the virtual digital human e-commerce live stream are collected. Based on statistical analysis of the facial expression data and / or voice data, a second live stream feedback tag reflecting the user is matched. This can be achieved through the terminal device's camera, microphone, etc., to acquire the user's facial expression data and / or voice data in real time.

[0069] In this step, facial expression data and / or voice data of users are parsed, extracted, and analyzed to determine the viewing emotional trends of all users. For example, the tone and emotional expression of the virtual anchor are adjusted according to the emotional feedback of users (such as positive or negative). When users show interest, the virtual anchor can use a more enthusiastic tone.

[0070] Step S2-4: Use the first live feedback tag and the second live feedback tag to perform a preset intent matrix match to generate live feedback intent.

[0071] In the steps, an intent preference matrix diagram corresponding to two types of live feedback tags is pre-set so that the corresponding live feedback intent can be adaptively matched and selected.

[0072] Furthermore, natural language processing (NLP) techniques can be used to perform semantic analysis on user textual feedback to identify user intent. For example, a user might have questions about product prices, features, or promotional activities. In this embodiment, pre-trained language models (such as chatGPT, deepseek R1, and other big data models) are used to optimize conversational techniques, ensuring that the language expression is more natural, fluent, and meets user needs.

[0073] Step S3: As Figure 2 As shown, a first virtual audio-visual stream is generated using historical live-stream data and product / service profile tags. For example, a live-stream script can be directly generated by combining product / service profile tags and historical user feedback, such as "This smartwatch has a battery life of up to 7 days, making it perfect for outdoor sports enthusiasts." This live-stream script is then converted into speech and integrated with the movements and expressions of a virtual digital human to form the first virtual audio-visual stream.

[0074] In this step, a first virtual audio and video stream is generated using historical live streaming data and product / service profile tags, including:

[0075] Step S3-1: Obtain live stream content data, live stream emotion data, and live stream image data from the historical live stream data of the virtual digital human. After preprocessing by extracting data classification matrix, generate the first virtual live stream dataset. Furthermore, convert the voice data into text using a feature extraction algorithm and extract emotion feature vectors to generate the first virtual live stream dataset.

[0076] Step S3-2: Obtain historical product and service profiles from the virtual digital human's historical live stream data, as well as product and service profile tags from the current e-commerce live stream. After data extraction and preprocessing, generate a second virtual live stream dataset. For example, use a tagging method to convert product features into structured data, which facilitates subsequent processing to generate the second virtual live stream dataset.

[0077] Step S3-3: Obtain user interaction feedback data from the historical live streaming data of the virtual digital human. After data extraction and preprocessing, extract the key intentions and sentiments in user comments using natural language processing technology to generate the third virtual live streaming dataset.

[0078] Step S3-4: Based on the first virtual live streaming dataset, use a deep learning network model to perform virtual rendering on the virtual digital human to obtain the first live streaming data packet;

[0079] Step S3-5: Based on the second virtual live streaming dataset, the third virtual live streaming dataset, and the product and service profile tags, generate a live streaming script using natural language processing technology. After audio generation, merge the script with the first live streaming data packet to form the first virtual audio and video stream.

[0080] Step S4: As Figure 3As shown, the language of the first virtual audio and video stream is optimized and adjusted using the live feedback intent to generate a second virtual audio and video stream, which is then pushed to the multi-user terminal.

[0081] For example, during a live stream, the system displays the interactive topic "Which smartwatch function are you most interested in?" and provides options such as "Battery Life," "Activity Monitoring," and "Health Features." When the end user selects "Battery Life," the system matches the first live stream feedback tag. Simultaneously, it collects the user's facial expression data in real time, analyzes the user's emotional feedback, and matches a second live stream feedback tag. Based on the feedback tag, the system adjusts the virtual avatar's live stream behavior, adding detailed information about battery life and adjusting expressions and movements to better engage the user. Finally, a second virtual audio and video stream is re-rendered and pushed to the user's device.

[0082] In this step, the first virtual audio and video stream is optimized and adjusted using live feedback intent to generate the second virtual audio and video stream, including:

[0083] Step S4-1: In response to the triggering operation of the live broadcast feedback intent, obtain the live broadcast emotional feedback tags of different users to the e-commerce live broadcast virtual digital human, and match and confirm the preset live broadcast optimization strategy corresponding to the live broadcast emotional feedback tags.

[0084] Step S4-2: Use a preset live streaming optimization strategy to adjust the live streaming behavior of the virtual digital human in real time, including at least the live streaming content, live streaming actions, and live streaming emotions, and render and generate a new second virtual audio and video stream.

[0085] The optimized virtual digital human's live-streaming script is generated using deep learning speech synthesis technology, ensuring lip-sync with speech and automatic lip-sync calibration using a super lip-sync model library to ensure an error rate of less than 3%. The optimized speech is then rendered along with the virtual anchor's actions, expressions, and the surrounding scene to generate a new second virtual audio-visual stream.

[0086] In this step, the virtual audio and video streams can also support multilingual and dialectal speech synthesis to meet the needs of users in different regions.

[0087] In this embodiment, as Figure 5 As shown, within a preset time period, several consecutive sampling time intervals are divided according to the time sequence, and at least three priorities are matched sequentially for the candidate options according to the order of the sampling time intervals in which the user's response trigger time is located.

[0088] Multithreading technology is used to process multiple alternative options in parallel within different sampling time intervals and to statistically analyze the number of different priorities. The alternative option with the highest proportion among the first or first two priorities within a preset time period is used to match the first live feedback tag reflecting the user.

[0089] In this embodiment, as Figure 6 As shown, the e-commerce live streaming screen displays methods for triggering interactive commands initiated by end users, including:

[0090] In response to an interactive command triggered by an end user, at least one interactive content list page is displayed. The interactive content list page sequentially displays content including at least e-commerce live streaming scenarios, virtual digital human figures, product and service profile tags, and live streaming interactive topics.

[0091] In response to the end user's trigger operation on the interactive content list page, the first virtual audio and video stream is rendered and adjusted using the options obtained from the interactive content list page.

[0092] In this step, the virtual digital human anchor can trigger a real-time question-and-answer session based on the interactive commands of the end user, quickly feed back the user's interactive content to the virtual digital human, generate answers in real time based on the interactive content and push them to the current end user, support keyword interaction and atmosphere guidance, and effectively improve end user participation and interaction.

[0093] Secondly, embodiments of the present invention provide a zero-configuration adaptive real-time interactive system for virtual digital humans, which monitors user interaction data in real time and dynamically adjusts live streaming content and interaction methods to ensure a smooth and interactive user experience. Specifically, it includes the following modules:

[0094] The adaptive module uses an adaptive control algorithm to adjust the system configuration parameters in real time.

[0095] The data acquisition module obtains historical live streaming data of e-commerce live streaming virtual digital humans, product and service profile tags, and obtains the live streaming feedback intent of different terminal users based on time series sampling;

[0096] The interactive module generates a first virtual audio and video stream using historical live streaming data and product / service profile tags; and optimizes and adjusts the language of the first virtual audio and video stream using live streaming feedback intent to generate a second virtual audio and video stream, which is then pushed to multiple user terminals.

[0097] In this embodiment, the system adopts an adaptive module design, supporting multiple platforms (such as H5, iOS, and Android) with zero configuration and automatic detection of the user's device's network environment and hardware performance. Specifically, in low-bandwidth environments, it automatically reduces video resolution; on high-performance devices, it enables higher resolution and more complex rendering effects.

[0098] Thirdly, the present invention provides an electronic device, comprising:

[0099] processor;

[0100] Memory used to store processor-executable instructions;

[0101] The processor is configured to execute instructions to implement a zero-configuration adaptive real-time interaction method for virtual digital humans, as described above.

[0102] Fourthly, the present invention provides a computer-readable storage medium that, when the instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to perform a zero-configuration adaptive real-time interaction method for virtual digital humans as described above.

[0103] Fifthly, the present invention provides a computer program product comprising a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps in the zero-configuration adaptive real-time interaction method for virtual digital humans as described above.

[0104] In this embodiment, the terminal device can be a user equipment (UE), a mobile device, a user terminal, a terminal, a personal digital assistant (PDA), a handheld device, a computing device, an in-vehicle device, a wearable device, etc.

[0105] The technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.

[0106] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0107] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the technical principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A zero-configuration adaptive real-time interaction method for virtual digital humans, characterized in that, include: Acquire historical live streaming data and product / service profile tags of virtual digital humans in e-commerce live streaming; The live streaming feedback intent of different terminal users is obtained based on time series sampling; A first virtual audio and video stream is generated using the historical live streaming data and product / service profile tags; The first virtual audio and video stream is optimized and adjusted using the live feedback intent to generate a second virtual audio and video stream, which is then pushed to a multi-user terminal.

2. The zero-configuration adaptive real-time interaction method for virtual digital humans according to claim 1, characterized in that, Generating a first virtual audio and video stream using the historical live streaming data and product / service profile tags includes: The first virtual live streaming dataset is generated by extracting live streaming content data, live streaming emotion data, and live streaming image data from the historical live streaming data of virtual digital humans, and then preprocessing the data through data classification matrix extraction. The system obtains historical product and service profiles from the historical live streaming data of virtual digital humans, as well as product and service profile tags from the current e-commerce live streaming. After data extraction and preprocessing, a second virtual live streaming dataset is generated. The third virtual live streaming dataset is generated by extracting and preprocessing user interaction feedback data from the historical live streaming data of virtual digital humans. Based on the first virtual live streaming dataset, a deep learning network model is used to perform virtual rendering on the virtual digital human to obtain the first live streaming data packet. Based on the second virtual live streaming dataset, the third virtual live streaming dataset, and product / service profile tags, a live streaming script is generated using natural language processing technology. After audio generation, it is fused with the first live streaming data packet to form a first virtual audio / video stream.

3. The zero-configuration adaptive real-time interaction method for virtual digital humans according to claim 1, characterized in that, Using the live stream feedback intent to perform language optimization and adjustment on the first virtual audio and video stream to generate a second virtual audio and video stream, including: In response to the triggering operation of the live broadcast feedback intent, obtain the live broadcast emotion feedback tags of different users to the e-commerce live broadcast virtual digital human, and match and confirm the preset live broadcast optimization strategy corresponding to the live broadcast emotion feedback tags; The virtual digital human's live streaming behavior is adjusted in real time using the preset live streaming optimization strategy, including at least the live streaming content, live streaming actions, and live streaming emotions, and a new second virtual audio and video stream is rendered and generated.

4. The zero-configuration adaptive real-time interaction method for virtual digital humans according to claim 3, characterized in that, The live streaming feedback intent of different terminal users is obtained based on time series sampling, including: In the e-commerce live streaming screen, at least one live interactive topic and multiple alternative options corresponding to the topic are displayed according to a preset spatial layout arrangement rule; In response to different users' triggering operations on the live interactive topics, the alternative options selected by different users are obtained, and the first live feedback tag reflecting the user is matched based on the statistical analysis of the alternative options. In response to interactive video commands from end users in e-commerce live streaming, the system collects facial expression data and / or voice data from end users while watching virtual digital human e-commerce live streams in real time, and matches second live stream feedback tags reflecting the user based on statistical analysis of the facial expression data and / or voice data. The live feedback intent is generated by matching the first live feedback tag and the second live feedback tag with a preset intent matrix.

5. The zero-configuration adaptive real-time interaction method for virtual digital humans according to claim 4, characterized in that, Within a preset time period, several consecutive sampling time intervals are divided according to the time sequence, and the candidate options are matched sequentially according to the order of the sampling time intervals in which the user's response trigger time is located, with at least three priorities. Multi-threading technology is used to process multiple alternative options in parallel within different sampling time intervals and statistically analyze the number of different priorities. The alternative option with the highest proportion among the first or first two priorities within a preset time period is used to match the first live feedback tag reflecting the user.

6. The zero-configuration adaptive real-time interaction method for virtual digital humans according to claim 4, characterized in that, The e-commerce live stream displays methods for triggering interactive commands initiated by end users, including: In response to an interactive command triggered by an end user, at least one interactive content list page is displayed. The interactive content list page sequentially displays content including at least e-commerce live streaming scenarios, virtual digital human images, product and service profile tags, and live streaming interactive topics. In response to a user's triggering action on the interactive content list page, the first virtual audio and video stream is rendered and adjusted using the options obtained from the interactive content list page.

7. A zero-configuration adaptive real-time interactive system for virtual digital humans, characterized in that, include: The adaptive module uses an adaptive control algorithm to adjust the system configuration parameters in real time. The data acquisition module obtains historical live streaming data and product / service profile tags of the virtual digital human in e-commerce live streaming. And obtain the live streaming feedback intent of different terminal users based on time series sampling; The interactive module uses the historical live streaming data and product / service profile tags to generate a first virtual audio / video stream; The system then uses the live stream feedback intent to perform language optimization and adjustment on the first virtual audio and video stream to generate a second virtual audio and video stream, which is then pushed to a multi-user terminal.

8. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the zero-configuration adaptive real-time interaction method for virtual digital humans as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device is able to perform the zero-configuration adaptive real-time interaction method for virtual digital humans as described in any one of claims 1 to 6.

10. A computer program product, characterized in that, The computer program product includes a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps in the zero-configuration adaptive real-time interaction method for virtual digital humans as described in any one of claims 1 to 6.