Agricultural product personalized recommendation digital human live broadcast interaction system based on large model
The digital human live-streaming interactive system for personalized recommendations of agricultural products based on a large model has solved the problem of the disconnect between interaction and knowledge in agricultural product live-streaming, realizing personalized recommendations and professional display, and improving the conversion rate and interaction quality of live-streaming.
Patent Information
- Application Number
- CN202511828607.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-05
- Publication Date
- 2026-02-03
AI Technical Summary
Existing technologies for agricultural product live streaming suffer from problems such as a disconnect between interaction and knowledge, a lack of personalized and real-time support in recommendations, and a lack of professional product presentation. These issues make it difficult to meet the professional needs of agricultural product live streaming, such as product identification, cooking association, and real-time data fusion.
The system employs a large-scale model-based personalized recommendation digital human live streaming interactive system for agricultural products, which includes a configuration module, an interaction acquisition module, a multimodal analysis module, a personalized content generation module, and a synchronous streaming module. By acquiring multimodal user interaction content in real time and combining it with an agricultural product knowledge base and a user profile database, the system generates personalized recommendation suggestions and drives the digital human anchor to perform matching broadcast voice, facial expressions, and body movements.
It achieves deep integration of domain knowledge and real-time data, providing intelligent and personalized recommendations, improving the commercial conversion efficiency and professional performance of live streaming, enhancing the persuasiveness and immersiveness of the presentation, and ensuring smooth intelligent interaction in complex environments.
Smart Images

Figure CN121462784A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of live streaming technology, and in particular to a digital human live streaming interactive system for personalized recommendations of agricultural products based on a large model. Background Technology
[0002] In recent years, with the rapid development of artificial intelligence and computer graphics technologies, digital human anchors, as an innovative interactive medium, have been gradually applied to the field of e-commerce live streaming. Compared with real anchors, digital humans have advantages such as being online 24 / 7, having a stable image, and controllable costs, providing brands with new marketing solutions. However, as the application of digital human technology in e-commerce live streaming becomes increasingly widespread, existing solutions have revealed significant shortcomings in vertical sectors, especially in agricultural product live streaming.
[0003] For example, the existing technology (publication number CN119967197A) discloses a digital human live streaming interaction method and system based on a large model. Although it uses a large model to achieve multimodal interaction and high-quality presentation, its knowledge base and decision-making logic are relatively generalized, making it difficult to meet the professional needs of agricultural product live streaming for product identification, cooking association, and real-time data fusion, resulting in a lack of depth and relevance in recommended content.
[0004] For example, the existing technology (publication number CN116996703A) discloses a digital human live streaming interaction method, system, device and storage medium. Although it can efficiently process bullet comments and achieve stylized interaction, its interaction dimension is single, and the knowledge base and action design are not closely integrated with the physical characteristics of agricultural products. It is difficult to conduct professional product display and in-depth persuasion, resulting in the interaction effect remaining superficial and making it difficult to effectively improve the conversion rate.
[0005] Therefore, existing technologies for live-streaming agricultural products suffer from core problems such as a disconnect between interaction and knowledge, a lack of personalized and real-time support in recommendations, and a lack of professional product presentation in demonstrations. Thus, there is an urgent need for a digital human-based live-streaming interactive system that can deeply integrate domain knowledge, real-time data, and multimodal interaction to achieve intelligent, personalized, and embodied recommendations. Summary of the Invention
[0006] The main objective of this invention is to overcome the shortcomings of the prior art and provide a digital human live interactive system for personalized recommendations of agricultural products based on a large model.
[0007] The technical solution adopted by this invention to achieve its technical objective is: a digital human live-streaming interactive system for personalized recommendations of agricultural products based on a large model, comprising: The configuration module is used to configure the digital human anchor's appearance, voice, and initial behavioral logic related to agricultural product recommendations according to the live marketing goals; The interactive acquisition module is used to acquire multimodal user interaction content in the live broadcast room in real time. The multimodal user interaction content includes at least bullet screen text, video chat screen and voice. The multimodal analysis module connects the agricultural product knowledge base and the user profile database. It is used to analyze the multimodal user interaction content through a large model, extract user intent, emotional state and entity features related to agricultural products, and combine them with user historical behavior data to generate comprehensive interactive instructions that include personalized agricultural product recommendations. The personalized content generation module is used to generate personalized recommended voice text based on the comprehensive interactive instructions, and drive the digital human anchor to generate matching broadcast voice, facial expressions and body movement sequences based on the recommended voice text and the characteristics of agricultural products. The synchronous streaming module is used to perform time-series alignment and synchronous rendering of the generated broadcast voice, facial expressions and body movement sequences, and push the final audio and video streams to the live streaming platform.
[0008] Preferably, the multimodal analysis module is specifically used for: Large models are used to identify agricultural products or cooking scenes in video chat. Analyze the user's intentions in asking for prices, comparing prices, inquiring about taste, origin, cooking methods, or appearance in the bullet screen text and voice; The identified entities, scenarios, and user intentions are matched with product information in the agricultural product knowledge base and combined with preference information in the user profile database to generate the comprehensive interactive instructions. The instructions include at least the specific agricultural product recommended, the reason for the recommendation, and related consumption suggestions.
[0009] Preferably, the personalized content generation module includes: The text generation unit is used to generate natural language text that conforms to the style of a digital human anchor and contains the core selling points and promotional information of the recommended products, based on the comprehensive interactive instructions. A speech generation unit is used to convert the natural language text into speech audio; The action-driven unit is used to call or generate corresponding digital human anchor action instructions from a preset action library based on the semantics and emotion of the natural language text and the physical characteristics of the recommended agricultural products. The actions include, but are not limited to: pointing to a specific product area, simulating tasting actions, and displaying product size or color comparisons.
[0010] Preferably, the system further includes a real-time decision-making module, used for: Based on the products recommended in the comprehensive interactive instructions, query the inventory database in real time; If the inventory is below a preset threshold, a low inventory warning will be added when generating the natural language text, or the multimodal analysis module will be triggered to start the alternative product recommendation process.
[0011] Preferably, the interactive acquisition module further includes a filtering submodule, used for: The obtained bullet screen text is filtered for invalid information. The filtering rules include removing sensitive words, duplicate content, and advertising information that is not related to the current recommended topic. Video chat requests are prioritized based on factors including user follower level, historical spending amount, and the relevance of the current interaction content to the live stream topic.
[0012] Preferably, the configuration module also provides a visual interface for operators to upload or link real-time agricultural product data sources, including product images, prices, origin traceability information, and test reports, and sets recommended sales script templates and key display actions for digital human anchors for different products.
[0013] Preferably, the synchronous streaming module aligns the audio and lip-sync animation sequences using a dynamic time warping algorithm and smooths the transitions between motion frames using an optical flow compensation algorithm to ensure audio-visual synchronization.
[0014] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program as a functional module of the system described above.
[0015] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the functional modules of the system described above.
[0016] The working principle of this large-scale model-based personalized agricultural product recommendation digital human live-streaming interactive system is as follows: The system achieves intelligent and personalized interaction and recommendations from digital human anchors in agricultural product live-streaming scenarios through an integrated process. First, the configuration module allows operators to customize the digital human's image, voice, and initial behavioral logic according to marketing goals, and connects it to real-time agricultural product data sources and script templates through a visual interface. After the live stream begins, the interaction acquisition module captures and preprocesses multimodal interactive content from users in real time, including bullet comments, video feeds, and voice messages, ensuring input quality through filtering and prioritization. Subsequently, the core multimodal analysis module calls upon the large-scale model to deeply analyze this interactive content, identify user intent, emotions, and agricultural product entities in the video, and, combined with the built-in agricultural product knowledge base and user profile database, generates comprehensive interactive instructions that integrate specific product recommendations, reasons, and consumption suggestions. The personalized content generation module then, based on these instructions, synchronously generates natural language recommendation text and corresponding voice messages that match the anchor's style, and drives the digital human to execute a sequence of actions highly matched to the recommended content and agricultural product characteristics (such as simulated tasting and comparative demonstrations). Finally, the synchronous streaming module uses algorithms such as dynamic time warping to ensure precise synchronization and smooth rendering of voice, lip movements, and actions, and pushes the final immersive interactive live stream to the live streaming platform, thereby dynamically responding to user needs and achieving an experience upgrade from watching to personalized shopping guidance.
[0017] Compared with the prior art, the beneficial effects of the present invention are: This digital human live-streaming interactive system for personalized agricultural product recommendations, based on a large model, integrates a dedicated agricultural product knowledge base with real-time data. The system can perform domain-based reasoning and provide professional recommendations with scientific evidence and personalized suggestions, achieving in-depth professional recommendations and solving the problem of vague recommendation content in existing technologies.
[0018] This digital human-based live-streaming interactive system for personalized recommendations of agricultural products, based on a large model, integrates user profiles, real-time intent, and dynamic inventory data. By shifting from passive response to proactive and accurate recommendations and real-time decision-making, the system enhances its accuracy in conversion and effectively improves the commercial conversion efficiency of live-streaming.
[0019] This digital human live-streaming interactive system for personalized agricultural product recommendations, based on a large model, innovatively binds digital human actions with the physical characteristics of agricultural products and the semantics of recommendations. Through a professional action library, it achieves embodied performance, greatly enhancing the persuasiveness and immersiveness of the presentation.
[0020] This digital human live-streaming interactive system for personalized recommendations of agricultural products, based on a large model, maintains smooth, real-time intelligent interaction and high-quality output in complex live-streaming environments through intelligent filtering, priority scheduling, and advanced synchronous rendering technology. Attached Figure Description
[0021] Figure 1 This is a system framework diagram of a digital human live-streaming interactive system for personalized recommendations of agricultural products based on a large model.
[0022] Figure 2 System framework diagram for the personalized content generation module.
[0023] in: Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. However, it should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of the invention. Furthermore, descriptions of well-known structures and technologies are omitted in the following description to avoid unnecessarily obscuring the concept of the invention.
[0025] In the description of this invention, it should be noted that when an element is referred to as being "fixed to" or "set on" another element, it can be directly on or indirectly on the other element. When an element is referred to as being "connected to" another element, it can be directly connected to or indirectly connected to the other element.
[0026] In the description of this invention, it should be noted that the terms "center," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship commonly used when the product of this invention is in use. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified. "Several" means one or more, unless otherwise explicitly specified.
[0027] In the description of this invention, it should also be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," "connect," and "link" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances. Example 1:
[0028] Please see Figure 1 and Figure 2 A digital human live-streaming interactive system for personalized recommendations of agricultural products based on a large model includes a configuration module, an interaction acquisition module, a multimodal analysis module, a personalized content generation module, a synchronous streaming module, and a real-time decision-making module.
[0029] In this implementation, the configuration module is used to configure the digital human anchor's image, voice, and initial behavioral logic related to agricultural product recommendations according to the live-streaming marketing objectives.
[0030] The configuration module also provides a visual interface for operators to upload or link real-time agricultural product data sources, including product images, prices, origin traceability information, and test reports. It also allows setting recommended sales script templates and key display actions for digital human anchors for different products.
[0031] In this implementation, the interaction acquisition module is used to acquire multimodal user interaction content in the live broadcast room in real time. The multimodal user interaction content includes at least bullet screen text, video chat screen, and voice.
[0032] The interactive acquisition module also includes a filtering submodule, which is used to filter invalid information from the acquired bullet screen text. The filtering rules include removing sensitive words, duplicate content, and advertising information that is irrelevant to the current recommended topic. The video chat requests are prioritized, and the ranking criteria include at least the user's fan level, historical spending amount, and the relevance of the current interactive content to the live broadcast topic.
[0033] In this implementation, the multimodal analysis module connects the agricultural product knowledge base and the user profile database. It is used to analyze the multimodal user interaction content through a large model, extract user intent, emotional state, and entity features related to agricultural products, and combine them with user historical behavior data to generate comprehensive interactive instructions that include personalized agricultural product recommendations.
[0034] The multimodal analysis module is specifically used to: identify agricultural product entities or cooking scenes in video chat using a large model; analyze user intentions in the bullet screen text and voice messages, such as inquiring about prices, comparing prices, asking about taste, origin, cooking methods, or appearance; match the identified entities, scenes, and user intentions with product information in the agricultural product knowledge base, and combine them with preference information in the user profile database to generate the comprehensive interactive instructions, which at least include the specific agricultural product recommended, the reason for the recommendation, and related consumption suggestions.
[0035] In this implementation, the personalized content generation module is used to generate personalized recommended voice text based on the comprehensive interactive instructions, and drive the digital human anchor to generate matching broadcast voice, facial expressions and body movement sequences based on the recommended voice text and the characteristics of agricultural products.
[0036] The personalized content generation module includes: The text generation unit is used to generate natural language text that conforms to the style of a digital human anchor and contains the core selling points and promotional information of the recommended products, based on the comprehensive interactive instructions. A speech generation unit is used to convert the natural language text into speech audio; The action-driven unit is used to call or generate corresponding digital human anchor action instructions from a preset action library based on the semantics and emotion of the natural language text and the physical characteristics of the recommended agricultural products. The actions include, but are not limited to: pointing to a specific product area, simulating tasting actions, and displaying product size or color comparisons.
[0037] In this implementation, the synchronous streaming module is used to perform time-series alignment and synchronous rendering of the generated broadcast voice, facial expressions and body movement sequences, and push the final audio and video stream to the live streaming platform.
[0038] The synchronous streaming module aligns the audio and lip-sync animation sequences using a dynamic time warping algorithm and smooths the transitions between motion frames using an optical flow compensation algorithm to ensure audio-visual synchronization.
[0039] In this implementation, the real-time decision-making module is used to query the inventory database in real time based on the products recommended in the comprehensive interactive instructions; if the inventory is lower than a preset threshold, a low inventory prompt is added when generating the natural language text, or the multimodal analysis module is triggered to start the alternative product recommendation process.
[0040] Specifically, the usage process of this large-model-based personalized agricultural product recommendation digital human live-streaming interactive system is as follows, including the following five stages: Phase 1: Personalized battlefield configuration and data fusion.
[0041] The process begins with the configuration module. Operators can not only define the visual image (skin color, clothing, facial features) and voice characteristics (selected from the voice library or customized training) of the digital human anchor through a visual interface (such as a drag-and-drop editor), but more importantly, they need to import and structure the real-time agricultural product data source for this live broadcast.
[0042] This includes linking product SKUs, multi-angle images, dynamic prices, origin traceability information (such as blockchain hash values), quality inspection reports, etc., with a pre-defined domain knowledge graph (including product attributes, cooking methods, nutritional associations, etc.).
[0043] Meanwhile, operators can pre-set differentiated wording templates (such as emphasizing crispness and highlighting the original cut) and corresponding key display action instructions (such as holding it up with both hands and cutting the cross-section with a knife) for different products (such as selenium-enriched apples and chilled steak), laying the style and material foundation for subsequent personalized generation.
[0044] Phase Two: High-concurrency sensing and intelligent filtering of multi-source heterogeneous signals.
[0045] After the live stream starts, the interaction acquisition module continuously monitors the live stream room and processes the following three types of high-concurrency data streams in parallel: Text stream: Real-time crawling of bullet comments, and filtering submodules apply rules based on sensitive word databases, duplicate detection algorithms (such as SimHash) and topic relevance models to remove noise and advertisements; Audio and video request stream: Processes users' video call requests. The priority scheduling submodule sorts them in real time based on a dynamic weight model (combining user fan level, historical spending, and semantic relevance of current speech to the live broadcast topic) to ensure that high-value and highly relevant interactions are given priority. Voice stream: Real-time capture of audio from live chat sessions. This module packages the cleaned and sorted multimodal raw data into time-stamped data packets and sends them to the downstream decision-making center.
[0046] Phase 3: Deep intent insight and real-time decision-making based on domain knowledge enhancement.
[0047] The multimodal analysis module uses a visual large model (such as ViT) to perform entity detection and scene recognition on video frames to determine whether the user has displayed a certain fruit, vegetable, or kitchen scene. Simultaneously, real-time speech recognition (ASR) and NLP parsing of text are performed to extract entities related to purchasing intent, such as inquiries, price comparisons, taste, cooking methods, and appearance. Then, these extracted multimodal feature vectors are used for joint vector retrieval and inference with an agricultural product knowledge base and a user profile database (which records users' historical purchasing preferences and taste preferences).
[0048] For example, if the system recognizes a user holding a "tomato" and asking "stewed beef brisket," and combines this with the user's historical preference for "sour flavors," it will generate a structured, comprehensive interactive instruction. This instruction not only recommends "a sandy-textured tomato from a specific origin," but also includes the reasoning that "it has moderate acidity, doesn't fall apart even after long cooking, and is suitable for stewing beef brisket," as well as personalized consumption suggestions such as "stir-frying it with ginger slices to remove the fishy smell." The real-time decision-making module intervenes at this point, checking the inventory of the recommended product. If it's scarce, it will be marked in the instruction, triggering alternative solutions.
[0049] Phase 4: Stylized and embodied content generation and driving force.
[0050] After receiving the instruction, the personalized content generation module initiates collaborative creation: The text generation unit takes structured instructions, inserts them into a stylized language model (such as playful, professional, friendly) pre-set for the current digital human, and generates a natural and fluent script that includes the core selling points of the product ("This tomato, look at its stem, it's especially fresh...") and possible inventory information. The speech generation unit uses a TTS engine to convert text into voice audio that matches the digital human avatar and is rich in emotional fluctuations; The action-driven unit simultaneously parses the semantics of the text (e.g., "compare the sizes"), the emotion ("surprise"), and the physical characteristics of the product (e.g., "crisp apple"), intelligently calling or parameterizing a series of action instructions from the domain action library. For example, corresponding to "crisp," it might generate an audio effect instruction to "pick up the apple and bring it close to the microphone to simulate a crisp chewing sound," driving the digital human to make corresponding gestures and expressions. The actions in the action library are pre-recorded or physically simulated and are strongly related to the characteristics of agricultural products.
[0051] Phase 5: Immersive rendering and output with frame-level precise synchronization.
[0052] The synchronous streaming module handles the challenging issue of time synchronization by employing a dynamic time warping algorithm. This algorithm precisely aligns keyframes of the phoneme sequence with the digitized human's lip-sync animation at the frame level, ensuring a strict match between "speaking" and "movement." Simultaneously, an optical flow compensation algorithm interpolates and smooths transition frames in the motion sequence, eliminating any sense of mechanical jumps. All elements (background, digitized human model, effects, and audio) are synthesized via a real-time rendering engine and ultimately encoded into ultra-low-latency audio and video streams. These streams are stably pushed to the live streaming platform via protocols such as RTMP / WebRTC, presenting viewers with a "virtual agricultural product expert" who can understand, comprehend, and provide personalized responses and recommendations. Implementation: 2:
[0053] Based on the above embodiments, this invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program as a functional module of the system described above.
[0054] The solution in this embodiment can be selectively combined with solutions in other embodiments. Implementation: 3:
[0055] Based on the above embodiments, this invention also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the functional modules of the system described above.
[0056] The solution in this embodiment can be selectively combined with solutions in other embodiments.
[0057] It should be noted that although the above embodiments have been described herein, this does not limit the scope of patent protection of this invention. Therefore, any changes and modifications made to the embodiments described herein based on the innovative concept of this invention, or equivalent structural, procedural, or functional transformations made using the description and drawings of this invention, directly or indirectly applying the above technical solutions to other related technical fields, are all included within the scope of protection of this invention.
Claims
1. A digital human live-streaming interactive system for personalized recommendations of agricultural products based on a large model, characterized in that: include: The configuration module is used to configure the digital human anchor's appearance, voice, and initial behavioral logic related to agricultural product recommendations according to the live marketing goals; The interactive acquisition module is used to acquire multimodal user interaction content in the live broadcast room in real time. The multimodal user interaction content includes at least bullet screen text, video chat screen and voice. The multimodal analysis module connects the agricultural product knowledge base and the user profile database. It is used to analyze the multimodal user interaction content through a large model, extract user intent, emotional state and entity features related to agricultural products, and combine them with user historical behavior data to generate comprehensive interactive instructions that include personalized agricultural product recommendations. The personalized content generation module is used to generate personalized recommended voice text based on the comprehensive interactive instructions, and drive the digital human anchor to generate matching broadcast voice, facial expressions and body movement sequences based on the recommended voice text and the characteristics of agricultural products. The synchronous streaming module is used to perform time-series alignment and synchronous rendering of the generated broadcast voice, facial expressions and body movement sequences, and push the final audio and video streams to the live streaming platform.
2. The digital human live-streaming interactive system for personalized agricultural product recommendations based on a large model as described in claim 1, characterized in that, The multimodal analysis module is specifically used for: Large models are used to identify agricultural products or cooking scenes in video chat. Analyze the user's intentions in asking for prices, comparing prices, inquiring about taste, origin, cooking methods, or appearance in the bullet screen text and voice; The identified entities, scenarios, and user intentions are matched with product information in the agricultural product knowledge base and combined with preference information in the user profile database to generate the comprehensive interactive instructions. The instructions include at least the specific agricultural product recommended, the reason for the recommendation, and related consumption suggestions.
3. The digital human live-streaming interactive system for personalized agricultural product recommendations based on a large model as described in claim 2, characterized in that, The personalized content generation module includes: The text generation unit is used to generate natural language text that conforms to the style of a digital human anchor and contains the core selling points and promotional information of the recommended products, based on the comprehensive interactive instructions. A speech generation unit is used to convert the natural language text into speech audio; The action-driven unit is used to call or generate corresponding digital human anchor action instructions from a preset action library based on the semantics and emotion of the natural language text and the physical characteristics of the recommended agricultural products. The actions include, but are not limited to: pointing to a specific product area, simulating tasting actions, and displaying product size or color comparisons.
4. The digital human live-streaming interactive system for personalized agricultural product recommendations based on a large model as described in claim 3, characterized in that, The system also includes a real-time decision-making module for: Based on the products recommended in the comprehensive interactive instructions, query the inventory database in real time; If the inventory is below a preset threshold, a low inventory warning will be added when generating the natural language text, or the multimodal analysis module will be triggered to start the alternative product recommendation process.
5. The digital human live-streaming interactive system for personalized agricultural product recommendations based on a large model as described in claim 1, characterized in that, The interactive acquisition module also includes a filtering submodule, used for: The obtained bullet screen text is filtered for invalid information. The filtering rules include removing sensitive words, duplicate content, and advertising information that is not related to the current recommended topic. Video chat requests are prioritized based on factors including user follower level, historical spending amount, and the relevance of the current interaction content to the live stream topic.
6. The digital human live-streaming interactive system for personalized agricultural product recommendations based on a large model as described in claim 1, characterized in that, The configuration module also provides a visual interface for operators to upload or link real-time agricultural product data sources, including product images, prices, origin traceability information, and test reports. It also allows setting recommended sales script templates and key display actions for digital human anchors for different products.
7. The digital human live-streaming interactive system for personalized agricultural product recommendations based on a large model, as described in any one of claims 1-6, is characterized in that... The synchronous streaming module aligns the audio and lip-sync animation sequences using a dynamic time warping algorithm and smooths the transitions between motion frames using an optical flow compensation algorithm to ensure audio-visual synchronization.
Citation Information
Patent Citations
Digital human live broadcast interaction method, system and device and storage medium
CN116996703A
Digital human live broadcast interaction method and system based on large model
CN119967197A