Live streaming scenario generation method and apparatus, electronic device, and storage medium

The live streaming scenario generation method addresses the limitations of digital human livestreaming by integrating voice, actions, and facial expressions, enhancing viewer engagement and adaptability.

JP2026016518APending Publication Date: 2026-02-03BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025177901
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-04-25
Filing Date
2025-10-22
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Existing digital human livestreaming technologies struggle with high interactivity and emotional rendering, lack of real-time adaptability, and poor multimodal integration, leading to unnatural appearances and reduced viewer engagement.

Method used

A live streaming scenario generation method that uses a large-scale language model to generate integrated, multimodal scenarios with flexible object description subsegments, enabling synchronized voice, actions, and facial expressions, and real-time interaction.

Benefits of technology

Enhances viewer engagement through natural and expressive digital human livestreaming, improving interaction and adaptability, and maintaining viewer immersion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026016518000001_ABST
    Figure 2026016518000001_ABST
Patent Text Reader

Abstract

A live streaming scenario generation method and apparatus, an electronic device, and a storage medium are provided.SOLUTION: The method includes generating at least one first scenario segment based on initial input information. The first scenario segment includes a speech content text and an object description sub-segment used for a live broadcast object, the object description sub-segment includes an object description text, and the object description text is used for describing at least one of the display action of the live broadcast object and a display mode used for the speech content text. The method also includes identifying a livecast scenario based on the at least one first scenario segment.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to the field of artificial intelligence, particularly to the fields of natural language processing, large-scale models, virtual digital humans, etc., which can be applied to virtual human live streaming scenes. More specifically, the present disclosure provides a live streaming scenario generation method, device, electronic device, storage medium, and computer program. [Background technology]

[0002] With the development of artificial intelligence technology, the application of virtual digital humans is increasing, which can replace human streamers and broadcast live 24 / 7 without interruption. Summary of the Invention

[0003] The present disclosure provides a live streaming scenario generation method, device, electronic device, storage medium, and computer program.

[0004] According to one aspect of the present disclosure, there is provided a live streaming scenario generation method including: generating at least one first scenario segment based on initial input information, the first scenario segment including utterance content text and an object description subsegment used for a live streaming object, the object description subsegment including object description text, the object description text being used to describe at least one of an exhibition operation of the live streaming object and an exhibition mode used for the utterance content text; and identifying a live streaming scenario based on the at least one first scenario segment.

[0005] According to another aspect of the present disclosure, there is provided a live streaming scenario generation device including: a first generation module that generates at least one first scenario segment based on initial input information, the first scenario segment including utterance content text and an object description subsegment used for a live streaming object, the object description subsegment including object description text, the object description text being used to describe at least one of an exhibition action of the live streaming object and an exhibition model used for the utterance content text; and a first identification module that identifies a live streaming scenario based on the at least one first scenario segment.

[0006] According to another aspect of the present disclosure, there is provided an electronic device including at least one processor and a memory communicatively connected to the at least one processor, the memory storing instructions executable by the at least one processor, the instructions being executed by the at least one processor such that the at least one processor performs a method according to the present disclosure.

[0007] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium having stored thereon computer instructions for causing a computer to perform a method according to the present disclosure.

[0008] According to another aspect of the present disclosure, there is provided a computer program product which, when executed by a processor, implements a method according to the present disclosure.

[0009] It should be noted that the content described in this section is not intended to identify key points or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will be readily understood from the following description.

[0010] The drawings are intended to provide a better understanding of the invention and are not intended to limit the invention. [Brief explanation of the drawings]

[0011] [Figure 1] FIG. 1 is a schematic diagram of an exemplary system architecture to which a live streaming scenario generation method and apparatus according to an embodiment of the present disclosure can be applied. [Figure 2] FIG. 2 is a flowchart of a live streaming scenario generation method according to an embodiment of the present disclosure. [Figure 3] FIG. 3 is a schematic diagram of a live streaming scenario generation method according to an embodiment of the present disclosure. [Figure 4] FIG. 4 is a flowchart of a live distribution scenario generation method according to another embodiment of the present disclosure. [Figure 5] FIG. 5 is a block diagram of a live streaming scenario generation device according to an embodiment of the present disclosure. [Figure 6] FIG. 6 is a block diagram of an electronic device to which a live streaming scenario generation method according to an embodiment of the present disclosure can be applied. DETAILED DESCRIPTION OF THE INVENTION

[0012]

[0023] The following describes exemplary embodiments of the present disclosure with reference to the drawings. Although various details of the embodiments of the present disclosure are included for ease of understanding, they should be considered merely as examples. Therefore, it should be understood that those skilled in the art can make various changes and modifications to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, the following description will omit descriptions of known functions and structures.

[0013] In the technical solution disclosed herein, the collection, storage, use, processing, transmission, provision, disclosure and other processes of relevant user personal information shall all comply with the provisions of relevant laws and regulations and shall not violate public order and morals.

[0014] With the rapid development of the e-commerce livestreaming industry, livestream sales have gradually become an important marketing tool. However, human streamers have limited resources and livestream operation costs are high. It is difficult for human streamers to maintain 24-hour continuous livestreaming. To reduce operation costs and achieve continuous livestreaming, artificial intelligence (AI)-based "digital human livestreaming" technology can be applied. "Digital human livestreaming" refers to a series of livestreaming tasks in a livestreaming room, such as product explanations, interactive Q&A, and user guidance, performed by virtual digital humans instead of real human streamers. Digital human livestreaming combines multiple front-end technologies: generating livestreaming scripts based on a large-scale language model (LLM), synthesizing natural-sounding voices using text-to-speech (TTS) technology, and combining it with a lip-motion synthesis system to achieve the "voice shape integration" driving effect. Each stage in this process is typically decoupled and has low interdependence, belonging to the "static generation + serial driving" technology path. Compared to human streamers, digital humans can achieve 24-hour continuous live streaming, with significant cost advantages, potential for improved efficiency, increased exposure, and greater control. Digital human live streaming technology has a relatively simple structure and is efficient at processing fixed-scene live streaming content. For example, digital humans can be used for low-interaction scenes, such as cyclical catchphrases or monologues. In these low-interaction scenes, the digital human's range of movement is small, facial expressions are limited, and real-time interaction with the viewer is essentially eliminated, while maintaining the digital human's expressiveness and stability.

[0015] However, as livestream sales gradually develop in line with the guidelines of high interactivity and emotional rendering, livestream scenes involve stronger expressive demands (e.g., emotional tension such as excitement, surprise, and emphasis), more complex action coordination (e.g., item interaction, dramatic actions, etc.), and more real-time interaction with users (e.g., barrage responses, lucky bag distribution, price fluctuations, etc.). The static and decoupled technology path described above has difficulty responding to scenes requiring such high interactivity and emotional rendering.

[0016] In low-interactivity scenarios, the scripts used for live streaming may be pre-generated, unable to respond to changes in the live streaming room in real time and lacking flexibility. Furthermore, the voices, actions, and facial expressions exhibited by digital humans during the live streaming process are generated by different, relatively independent modules, preventing advanced multimodal integration. This results in an unnatural appearance on the screen, an insufficient interactive experience, and a significant gap between the overall expressiveness of digital human live streaming and that of human streamers. Below, we will discuss issues such as live streaming scenarios, the behavior of digital human streamers, live streaming room tools, how to respond to changes in the live streaming room, and multi-person live streaming.

[0017] The content of live streaming scenarios used in digital human live streaming is unattractive. Most live streaming scenarios are generated in a templated format, resulting in the same content and lacking new meaning, making it difficult to effectively engage viewers. For example, excessive use of phrases such as "lowest on the entire network" and "if you miss today, you'll have to wait a year" can cause user fatigue and make it difficult to make a deep impression. Furthermore, scripted content is difficult to customize based on product characteristics, streamer character settings, and target audiences. This creates a flat atmosphere in the live streaming room, lacking mood swings and dramatic tension, making it difficult to create a "pop" or "crackling" impact on the content. Such scripts lack storytelling, interactivity, and emotional rendering, making it difficult for digital humans to truly engage viewers in live streaming.

[0018] Furthermore, digital human streamers have difficulty displaying expressive behaviors during live streaming. Their behavioral expressions remain monotonous. Digital human behaviors primarily involve synchronization between mouth shape and voice, a small number of basic facial expressions, and gesture-driven behaviors, making it difficult to display behaviors that closely match the semantics of the script. When expressing emotions such as excitement, surprise, and emphasis, digital humans often lack a wide range of movement and their facial expressions are not subtle, failing to truly convey emotional ups and downs and changes in tone. As a result, their overall expressiveness appears cognitively flat and "mechanical." Particularly in live streaming sales scenarios, such lackluster behavioral expressions make it difficult to lift viewers' spirits and create a viewing experience with a strong sense of immersion. Furthermore, delays, mismatches, and even incongruities can occur between the digital human's behavior, facial expressions, and voice and tone, further weakening the immersive and professional nature of the live stream. Compared with human streamers, digital humans are still not as natural and smooth, which puts them at a clear disadvantage in terms of strengthening brand impression and improving live streaming conversions.

[0019] Digital humans also have difficulty invoking various items in livestreaming rooms. Most digital human livestreaming systems have limited support for physical or virtual items in livestreaming scenes and lack deep scheduling capabilities for "interactive items" during livestreams. Digital human actions such as switching promotional materials, handing out lucky bags, presenting time limits, and barrage interactive prompts often fail to accurately synchronize with scripted content. Many actions performed by digital humans still rely on manual or background operation and maintenance intervention, resulting in poor automation and script-driven capabilities, disrupting the rhythm of livestreams and reducing the interactive effect. In scenes where two or more digital humans are livestreaming together, different characters repeatedly operate the same item, frequently resulting in information conflicts or logical confusion, affecting the overall quality of the livestream. Item invocation also lacks contextual design. When introducing skincare products, digital human streamers have difficulty "holding" the product naturally to showcase the packaging details.

[0020] Digital humans also struggle to dynamically adapt to changes in the livestreaming environment. Driven by static, preset scripts, digital humans struggle to detect and respond to real-time changes during livestreaming, preventing them from adapting to on-site conditions and lacking the ability to change. When product inventory temporarily fluctuates, users frequently ask about features, or interactive information pops up on the screen, digital human streamers are unable to instantly adjust the audio or change the focus of their commentary, instead relying solely on pre-defined lines, severely impacting the user experience. Digital humans also fundamentally lack the ability to detect and respond to real-time data, such as viewer mood, barrage feedback, product popularity, and viewer count fluctuations, preventing them from implementing a sales strategy of "advantaging with the situation." When a product significantly increases in popularity, digital humans are unable to quickly shift to key information, nor are they able to proactively tailor personalized responses or interactive Q&A to users' interests, resulting in the loss of a key element in guidance transformation.

[0021] Furthermore, multiple digital humans cannot achieve a natural multi-person live streaming effect. In complex live streaming room scenes, multiple digital humans with different characters work together. These characters include the streamer, coast streamer, host, on-site controller, and brand representative. Real-time cooperation, positioning, and interaction between these multiple characters is required to create a lively, smooth, and hierarchical live streaming atmosphere. However, multiple digital humans cannot mimic the natural interactive forms of "preemption," "breaks," "interjections," and "crosstalk" that occur between multiple anchors in real live streaming rooms.

[0022] To realize digital human live streaming with "strong expressive power, strong interactive opinions, and strong sales ability," it is necessary to create live streaming scenarios in a more integrated, intelligent, and real-time collaborative generation style. Therefore, the present invention provides a live streaming scenario generation method, which is described below.

[0023] 1 is a schematic diagram of an exemplary system architecture to which a live streaming scenario generation method and device according to an embodiment of the present disclosure can be applied. Note that FIG. 1 is merely an example of a system architecture to which the embodiments of the present disclosure can be applied, so that those skilled in the art can understand the technical content of the present disclosure, and does not mean that the embodiments of the present disclosure cannot be applied to other devices, systems, environments, or scenes.

[0024] 1, a system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 provides a medium for a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as, for example, wired and / or wireless communication links.

[0025] By using terminal devices 101, 102, and 103, users can interact with a server 105 via a network 104 to send and receive messages, etc. The terminal devices 101, 102, and 103 may be various electronic devices that have a display and support web page browsing, including, but not limited to, smartphones, tablet computers, laptop computers, and desktop computers.

[0026] Server 105 may be a server that provides various services, such as a background management server (by way of example only) that supports websites viewed by users using terminal devices 101, 102, and 103. The background management server performs processing such as analysis on data such as received user requests, and can feed back the processing results (e.g., web pages, information, or data obtained or generated based on the user requests) to the terminal devices.

[0027] Note that the live streaming scenario generation method according to the embodiments of the present disclosure may generally be executed by the server 105. Accordingly, the live streaming scenario generation device according to the embodiments of the present disclosure may generally be provided in the server 105. The live streaming scenario generation method according to the embodiments of the present disclosure may generally be executed by one or more of the terminal devices 101, 102, and 103. Accordingly, the live streaming scenario generation device according to the embodiments of the present disclosure may generally be provided in one or more of the terminal devices 101, 102, and 103. The live streaming scenario generation method according to the embodiments of the present disclosure may be executed by a server or server cluster that can communicate with the terminal devices 101, 102, and 103 and / or the server 105, separate from the server 105. Accordingly, the live streaming scenario generation device according to the embodiments of the present disclosure may be provided in a server or server cluster that can communicate with at least one of the terminal devices 101, 102, and 103 and the server 105, separate from the server 105.

[0028] Having described the system architecture of the present disclosure, the method of the present disclosure will now be described.

[0029] FIG. 2 is a flowchart of a live streaming scenario generation method according to an embodiment of the present disclosure.

[0030] As shown in FIG. 2, the method 200 may include operations S210 to S220.

[0031] In operation S210, at least one first scenario segment is generated based on the initial input information.

[0032] In an embodiment of the present disclosure, the initial input information may include information related to an object to be exhibited. For example, the object to be exhibited may be a product to be exhibited during the live broadcast. The initial input information may include a description of the product.

[0033] In an embodiment of the present disclosure, the first scenario segment includes a spoken content text and an object description sub-segment used for a live streaming object. For example, during the live streaming process, a digital human streamer may exhibit voice corresponding to the spoken content text in the form of conversation, speech, etc.

[0034] In an embodiment of the present disclosure, a large-scale language model can be used to generate the first scenario segment.

[0035] In an embodiment of the present disclosure, the object description subsegment includes object description text. The object description text is used for at least one of an action exhibited by the live-streaming object and an exhibition mode of the spoken content text. For example, the object description text may describe one or more actions exhibited by the live-streaming object. Also, for example, another object description text may describe an exhibition mode used for the spoken content text. The exhibition mode may include a tone of voice. The tone of voice may be calm, excited, etc.

[0036] In an embodiment of the present disclosure, the live broadcast object may be a digital human streamer.

[0037] In operation S220, a live broadcast scenario is identified based on the at least one first scenario segment.

[0038] In an embodiment of the present disclosure, one or more first scenario segments may be one or more live streaming scenario segments of a live streaming scenario. For example, the first scenario segment may be one live streaming scenario segment.

[0039] According to an embodiment of the present disclosure, when generating speech content text, object description subsegments are generated, which can better align the digital human's speech content with the behavior of live streaming objects and provide richer information for subsequent live streaming video generation. After providing the live streaming scenario to the large-scale model, the large-scale model can generate live streaming catchphrases that are fluent in content and logic and can also synchronously control the digital human's multimodal behavior. The live streaming scenario provided by the present disclosure can be realized as a true "executable scenario," transforming the digital human's live streaming from a single scene into a multi-dimensional, highly interactive, and highly expressive scene.

[0040] Having described the methods of the present disclosure, further discussion of live streaming objects and object description sub-segments follows.

[0041] In some embodiments, there is at least one live-streaming object, and the at least one live-streaming object includes at least one of a virtual character and an object to be displayed. For example, there may be one or more virtual characters. The one or more virtual characters may include characters played by various digital humans, such as a digital human streamer, a digital human coast streamer, a digital human host, a digital human on-site controller, or a digital human brand representative. For example, the object to be displayed may be a product or service to be displayed during the live broadcast. The service may be, for example, a travel service, a legal service, or the like. The product may be, for example, a skin care product, a bag, or the like.

[0042] In some implementations, the object description subsegments may be embedded in the utterance content text, for example, multiple object description subsegments may be embedded in multiple locations in the utterance content text.

[0043] In some embodiments, the object description subsegment may further include at least one of an exhibit identifier, a subsegment start identifier, a subsegment end identifier, and a delimiter identifier. The exhibit identifier may be the name of the exhibit mode, "action", or "expression". Based on the exhibit identifier, the large-scale model can generate audio, image, or text data corresponding to the exhibit identifier according to the object description text. The subsegment start identifier may be an identifier such as "(", "[", etc., and the subsegment end identifier may be an identifier such as ")", "]", etc. The delimiter identifier may be an identifier such as ":", "|", etc.

[0044] For example, a first scenario segment of the present disclosure is as follows:

[0045] "Streamer: (Tone: Calm) (Gestures: Holding two open boxes, putting down the box in his left hand, pointing to the contents in the box in his right hand and explaining how to use it) Look at this. It's easy to use (the coast streamers simultaneously say, "It's easy!"). Use it once a month for six consecutive months, with a six-month break, and you only need two boxes a year. Each time you use it, mix agent A and agent B and gently prick the skin with the roller provided. It's painless and the recovery period is short. Do not come into contact with water within 12 hours of finishing use. Just be careful with sunscreen. Streamer: (Tone: Calm) (Action: Pick up and display the wrapped roller on the table, while the streamer demonstrates the roller's movements in front of the face) This roller is a tool used in combination with other products, allowing nutrients to penetrate and be absorbed better. It softens the soil, allowing the skin to better absorb nutrients so that the seeds can take root better. Coast Streamer: (tone: calm) (action: take it from off-screen into the tool bag, open it and display it, while the streamer opens the packaging of the roll and takes out the roll and displays it) The tool bag is now complete (the streamer simultaneously says, "Very gentle!"), you can use it directly, which is very convenient. In addition, this product only lightly pricks the skin, is not painful, has a short recovery period, and will not affect your normal life. Coast Streamer: (Tone: Excited) (Movement: Take two bottles of liquid from the product box on the table, display them side by side, then put them back in the box) These two bottles are Agent A and Agent B, and when combined, they have a powerful effect. Like two superheroes working together, they can solve various skin problems and make the skin white and smooth. Streamer: (Tone: Excited) (Expression: Happy) Everyone, with such great products and so many benefits, what are you waiting for? Opportunities like this are rare and stock is limited, so if you miss this one, you may have to wait a long time for another discount like this. Coast Streamers: (Tone: Excited) (Action: Holding up a KT board displaying comparison photos of before and after use of the product, explaining the effects of use) (Expression: Surprised) Everyone, look at these comparison photos, the effects are really remarkable (Streamers simultaneously say, "This is amazing"). My skin had various problems before use, but my skin has improved after use. Wouldn't you like to have skin like this?

[0046] As shown in this first scenario segment, the multiple object description subsegments may include (tone of speech: calm), (action: holding two open-boxed products, putting down the box in the left hand, pointing to the contents in the box in the right hand, and explaining how to use them), (expression: happy), etc. Taking the object description subsegment (tone of speech: calm) as an example, "tone of speech" may be an exhibit identifier, "calm" may be the object description text, "(" may be the subsegment start identifier, and ")" may be the subsegment end identifier. ":" may be a separator identifier, located between the exhibit identifier "tone of speech" and the object description text "calm." Note that there are multiple virtual characters associated with this first scenario segment. The multiple virtual characters include a digital human streamer and a digital human coast streamer. It should be understood that the exhibit identifier, subsegment start identifier, subsegment end identifier, and separator identifier of this first scenario segment are merely examples. According to an embodiment of the present disclosure, an object description subsegment (exhibit identifier: object description text) is provided, and information such as tone of voice, actions, facial expressions, and item scheduling can be structured and embedded into a live streaming scenario, which enables a large-scale language model to generate logic-fluent live streaming catchphrases and effectively control the multimodal behaviors of digital humans, such as tone of voice, actions, and facial expressions.

[0047] In some embodiments, the object description subsegment may be one of an object exhibit mode description subsegment, an object action description subsegment, or an object expression description subsegment. The object exhibit mode description subsegment includes object exhibit mode description text. The object action description subsegment includes object action description text, and the object expression description subsegment includes object expression description text. For example, as shown in the partial first scenario segment above, the object description subsegment (tone: calm) may be an object exhibit mode description subsegment and may include the object exhibit mode description text "calm." The object description subsegment (action: pick up two open-boxed products, put down the box in the left hand, point to the contents in the box in the right hand, and explain how to use them) may be an object action description subsegment and may include the object action description text "pick up two open-boxed products, put down the box in the left hand, point to the contents in the box in the right hand, and explain how to use them." The object description subsegment (expression: happy) may be an object expression description subsegment and may include the object expression mode description text "happy." If the exhibition mode is tone, the object exhibition mode description sub-segment may be called an object tone description sub-segment, and the object exhibition mode description text may be called an object tone description text.

[0048] The object description sub-segments of the present disclosure have been described above, and the functions of the first scenario segment and the object description sub-segments will be further described below.

[0049] In some embodiments, a live-stream scenario may be used to generate a live-stream video. The live-stream video may include a first video segment corresponding to a first scenario segment. For example, the live-stream scenario may be input into a large-scale model to obtain the live-stream video. The first scenario segment may be one or more. The live-stream video may include one or more first video segments. Each first scenario segment corresponds to one first video segment.

[0050] In some embodiments, the audio data used for the spoken content text in the first video segment is generated according to the display mode described in the object display mode description text, for example, the audio data used for the spoken content text "See, it's very easy to use" is generated according to the "tone" described in the object display mode description text "calm."

[0051] In some embodiments, image data used for a virtual character in the first video segment is generated according to at least one of object action description text and object expression description text. The object action description text is used to describe at least one limb motion of the virtual character. The at least one limb motion may include a first limb motion used for an object to be exhibited, and may include a second limb motion of a live-streaming object. For example, the object action description text "Pick up two open-boxed products, put down the box in the left hand, point to the contents in the box in the right hand, and explain how to use them" can describe at least one first limb motion used for the object "products" to be exhibited. Also, for example, the object action description text "Go to the center of the screen" can describe a second limb motion of the virtual character itself.

[0052] In some embodiments, the object expression description text can describe at least one facial action of the virtual character. For example, the object expression description text "happy" can describe one or more facial actions of the virtual character. The one or more facial actions can represent a happy expression.

[0053] According to an embodiment of the present disclosure, the object description subsegments (display identifiers: object description text) can be flexibly adapted to different types of live streaming scenes, such as multiple vertical scenes, including makeup and skincare, health and nutrition, and digital home appliances. Taking a makeup scene as an example, a digital human streamer can explain the efficacy of a product. If the scenario includes object action description subsegments, the digital human streamer can automatically display actions such as "filling in the product," "tracing the facial area," and "displaying a fashion outfit," making the product information more visually appealing and convincing. Taking a health and nutrition scene as an example, if the scenario includes object action description subsegments, the digital human streamer can automatically display actions such as "opening the product packaging," "displaying dosage instructions," and "holding a portable sign to provide warnings," improving viewers' understanding and confidence in the usage instructions.

[0054] Furthermore, according to the embodiments of the present disclosure, the system can realize a three-dimensional collaborative expression of semantic, visual, and interactive content by controlling the scenario, digital human actions, and live streaming items based on the first limb movements used for the object to be displayed. This allows the digital human to not only "speak" but also "act," "move," and "exhibit" during live streaming. During the product introduction process, the digital human can naturally introduce the product, operation items, and display comparison diagrams, and further cooperate with the streamer to complete a "presentation-style explanation," significantly improving information transmission efficiency and viewing experience. This immersive presentation method significantly increases users' viewing concentration and content retention, providing a content basis for subsequent conversion actions.

[0055] In some embodiments, the first scenario segment includes a plurality of spoken content texts used for a plurality of virtual characters, respectively. An object description subsegment used for a live broadcast object is embedded in the spoken content text used for the virtual characters. As shown in some of the first scenario texts described above, the plurality of virtual characters include a digital human streamer and a digital human coast streamer. An object description subsegment for the digital human streamer (tone: calm) is embedded in the spoken content text for the digital human streamer. An object description subsegment for the digital human coast streamer (tone: excited) is embedded in the spoken content text for the digital human coast streamer.

[0056] According to an embodiment of the present disclosure, the object description subsegment (exhibit identifier: object description text) supports multi-character collaborative modeling. During the scenario generation phase, characters such as streamers, coast streamers, and field controllers can generate speech content and action positions, finely dividing angle positioning and functional division, improving the efficiency of multi-person collaboration. For example, coast streamers can focus on atmosphere, adding details, and promoting interactivity, while streamers can focus on explaining the core content of the product, leading to efficient collaboration in the scenario.

[0057] Having described the first scenario segment of the present disclosure, the method of the present disclosure will now be further described.

[0058] FIG. 3 is a schematic diagram of a live streaming scenario generation method according to an embodiment of the present disclosure.

[0059] In some embodiments, the initial input information includes at least one of character setting information used for the virtual character, initial object information used for the object to be exhibited, live streaming material information, and live streaming tool information. The initial object information may be multiple. For example, if the object to be exhibited is a product to be exhibited, the multiple initial object information may be information for multiple products. The live streaming material information may indicate multiple materials. The multiple materials may be used for one or more products. The live streaming tool information may indicate multiple live streaming tools available in the live streaming room. The multiple live streaming tools may include a barrage tool, a comment tool, etc. One or more pieces of character setting information may be used for each of the multiple virtual characters. For example, the character setting information may provide a style human setting for the virtual character. This allows the large-scale model, in combination with the style human setting request, to flexibly adjust the content style based on dimensions such as the live streaming target, live streaming viewer image, and product attributes, thereby ensuring a high level of matching between the scenario style and the scene. For example, in "knowledge dissemination" scenarios, content such as terminology, case analysis, and knowledge extension can be introduced based on the style corresponding to the scenario and the corresponding knowledge library, making live streaming more authoritative and educational. For another example, in "enlightenment sharing" scenarios, the scenarios generated by large-scale models adjust the speaking speed, emotional curve, and description style to emphasize semantic resonance and the viewer's sense of involvement. The introduction of character setting information allows for flexible and customized style control, ensuring that scenarios have good adaptability and infectiousness across different vertical categories and target audiences.

[0060] In some embodiments, in some embodiments of the above operation S210, generating at least one first scenario segment based on the initial input information includes identifying at least one target knowledge information based on character setting information, at least one initial object information, at least one of live streaming material information, and live streaming tool information, and a knowledge library. As shown in FIG. 3 , a deep thinking and knowledge strengthening operation S311 can be performed to identify the target knowledge information based on the knowledge library and one or more of the character setting information, at least one initial object information, live streaming material information, and live streaming tool information. Next, by performing operation S312, at least one first scenario segment can be generated based on the at least one target knowledge information. Operations S311 and S312 will be described in order below.

[0061] In some embodiments, identifying at least one target knowledge information based on the character setting information, at least one of the initial object information, the live streaming material information, and the live streaming tool information and the knowledge library includes identifying live streaming content planning information based on the character setting information, at least one of the initial object information, the live streaming material information, and the live streaming tool information. The live streaming content planning information may indicate that a first scenario segment includes at least one of introductory text for the object to be exhibited, historical case text for the object to be exhibited, and guide text for the object to be exhibited. For example, the live streaming content planning information may indicate an object to be exhibited for each of a plurality of first scenario segments. One or more first scenario segments may be generated for each object to be exhibited. The live streaming content planning information may also indicate that the first scenario segment used for the object to be exhibited includes introductory text, historical case text, and guide text. Alternatively, the live streaming content planning information may indicate that three first scenario segments used for the object to be exhibited each include introductory text, historical case text, and guide text. The introductory text may be a description of the object to be exhibited. The historical case text may be a description of the cases after different users have used the object to be displayed. The guide text may be used to guide users to perform actions on the object to be displayed. Taking the first scenario segment described above as an example, the speech content text of the first scenario segment, "Use it once a month for six consecutive months, with a six-month break, and you only need two boxes per year. Each time you use it, mix agent A and agent B and gently prick the skin with the included roller. It's painless and the recovery period is short," could be the introductory text. "Look at these comparison charts, the effect is truly remarkable," could be the historical case text. "Wouldn't you like skin like this?" could be the guide text.

[0062] In some embodiments, identifying at least one target knowledge information based on the character setting information, at least one initial object information, live streaming material information, and live streaming tool information and a knowledge library includes identifying a plurality of initial search terms based on live streaming content planning information. Identifying the target knowledge information based on the plurality of initial search terms and the knowledge library. As shown in FIG. 3 , a plurality of initial search term queries 30 can be identified based on the live streaming content planning information outline 30. Based on the plurality of initial search term queries 30 and the knowledge library, operation S3111 can be performed to identify knowledge information. In this way, the live streaming content planning information, combined with the in-depth search knowledge enrichment mechanism, interacts with a multi-source knowledge library during the scenario generation process. The multi-source knowledge library can include an object detail knowledge library, a case study knowledge library, and an industry encyclopedia knowledge library. The multi-source knowledge library can be used as an external knowledge system to realize "information-driven" scenario content generation. For example, if the object to be displayed is a product, content such as product effects, usage methods, core ingredients, and authoritative comments can be extracted as knowledge in the product detail knowledge library so that information about the object to be displayed can be naturally integrated into the comment content of the virtual character. For the actual case knowledge library, content such as user evaluations, usage feedback, and actual conversion results can be included as knowledge in the actual case knowledge library. For the industry encyclopedia library, external expertise can be adjusted and rationally expanded, for example, to introduce dermatological background into comment content for skin care ingredients. The knowledge library disclosed herein can realize the integration of multi-source knowledge, and the scenario can be logical, hierarchical, and specialized content rather than "empty likes" or "uniform promotions." The multi-source knowledge library may include a predetermined knowledge library or a knowledge library obtained through online data search.

[0063] In addition, multiple searches may be performed based on the initial search words to obtain knowledge more accurately and comprehensively. In some embodiments, the initial search words may be used as first-level search words to perform N searches in the knowledge library to obtain at least one target knowledge information. N may be an integer greater than 1. Based on the multiple n-th level search words, n-th level search results may be identified from a predetermined database. Based on the n-th level search results, n-th level knowledge information may be identified. Based on the n-th level search results and the multiple n-th level search words, operation S3112 may be performed to adjust the multiple search words to identify multiple n+1-th level search words. Then, the knowledge library may be searched based on the n+1-th level search words. At least one target knowledge information may be identified based on the N-th level search results. n may be an integer greater than or equal to 1 and less than N. If there are multiple objects to be exhibited, target knowledge information for each object to be exhibited may be identified from the N-th level search results. According to the embodiments of the present disclosure, when multiple searches are performed, by combining the object description subsegments, a fusion mechanism of structured information embedding, live streaming content planning, and in-depth search can be realized, significantly improving the quality of the scenario. The live streaming scenario can include multimodal collaborative content including tone of voice, actions, facial expressions, and item commands, giving the digital human true expressiveness. Furthermore, according to the embodiments of the present disclosure, the live streaming content can be more infectious, have a more rhythmic feel, have a more diverse language style, have clearer logic, and are richer in information, effectively avoiding the problems of traditional script template and uniformity. Users are more likely to be moved by the content while watching, creating emotional resonance.

[0064] Having described the deep thinking and knowledge enhancement operations of the present disclosure above, operation S312 will now be described.

[0065] 3, in operation S312, at least one first scenario segment is generated based on at least one target knowledge information. For example, a large-scale model can be used to generate the first scenario segment based on the target knowledge information used for one object to be exhibited.

[0066] Next, an operation S320 can be performed to identify a live broadcast scenario based on the at least one first scenario segment.

[0067] In some embodiments, it may be determined whether the difference between the first scenario segment and the character information is less than a predetermined difference threshold. For example, style features of a virtual character may be extracted from the first scenario segment. Differences between the style features and the character information features may be determined.

[0068] In some embodiments, the first scenario segment is identified as a live-stream scenario segment of the live-stream scenario in response to identifying a difference between the first scenario segment and the character character information as being less than a predetermined difference threshold. For example, if there is a small difference between the style characteristics of a virtual character in the first scenario segment and the character character information used for the virtual character, the first scenario segment can be identified as a live-stream scenario segment.

[0069] In some other embodiments, in response to identifying a difference between the first scenario segment and the characterization information as equal to or greater than a predetermined difference threshold, a first adjusted scenario segment is identified as a live-stream scenario segment of the live-stream scenario. The first adjusted scenario segment is obtained by adjusting the first scenario segment. For example, if there is a large difference between the virtual character style features of the first scenario segment and the characterization information used for the virtual character, the style of the first scenario segment can be adjusted using a large-scale model to reduce the difference between the style features and the characterization information.

[0070] As can be seen, the "persona profile" of a live streamer is key information for attracting users, forming memorization points, and building trust. To maintain consistency between a character's language style and behavioral logic and avoid "shifts" in the character profile during the long-term scenario generation process, the present disclosure introduces character profile information to label and construct the streamer character before scenario generation (e.g., by setting the virtual character's personality traits, language style, background, and expression habits). The differences between the scenario segments and the character profile information are then dynamically combined and limited to ensure that the virtual character's language style and behavioral logic are always consistent with the character profile. By identifying the differences between the first scenario segment and the character profile information, the consistency of the character profile expression can be continuously monitored during the scenario generation process, and potential character profile offsets can be automatically identified and corrected, avoiding issues such as inconsistencies in tone, value conflicts, and logic inversion. For example, for a "rational" virtual streamer character, the scenario segments used for that character should include speech content based on data, analysis, and logical derivation to support their perspective. For example, a "friendly" virtual streamer character is more likely to create a sense of familiarity by using familiar phrases such as "Babies" or "Look here, everyone." According to the embodiments of the present disclosure, full-cycle management and control of character settings, from "language style synthesis" to "behavior matching control," can be achieved, effectively resolving the problem of character setting fragmentation and collapse. To generate long scenarios, the present disclosure introduces a workflow for generating long scenarios in segments, which, combined with live streaming content planning information, generates one or more first scenario segments for each object to be exhibited. This allows scenarios of 5,000 to 10,000 characters to be generated at one time, meeting the demands of long-term live streaming of 15 to 30 minutes.

[0071] For ease of understanding, the above describes the method for generating a static live streaming scenario before live streaming, but the following describes the generation of a dynamic scenario during the live streaming process.

[0072] FIG. 4 is a flowchart of a live streaming scenario generation method according to another embodiment of the present disclosure.

[0073] Method 401 may be performed after method 200 above, and will be described below with reference to operations S431, S432, S441 and S442.

[0074] In operation S431, it is determined whether a task trigger signal has been received.

[0075] In some embodiments, in response to identifying the target action as not satisfying the trigger logic, operation S431 continues to be performed.

[0076] In some embodiments, operation S432 is performed in response to determining that a target behavior in a live streaming room satisfies the trigger logic. For example, after generating a live streaming scenario and generating a live streaming video using a large-scale model. A plurality of first video segments can be sequentially played to perform live streaming. During the playback of the first video segments, various behavior signals in the live streaming room can be monitored. The plurality of behaviors corresponding to the plurality of behavior signals can include user entry, exit, likes, comments, stay time, product clicks, etc. Questions or interactive requests submitted by users can also be acquired. If one or more of these behaviors satisfy a certain trigger logic, a task trigger signal can be generated. The trigger logic may be that the number of occurrences of one behavior is greater than or equal to a predetermined number. For example, in the case of product clicks, if the click volume on a certain product page rapidly increases, and the click volume on the product page is greater than or equal to a predetermined click volume threshold, a signal can be triggered as a task trigger signal for the first video segment. The large-scale model for generating the scenario can receive the signal and perform operation S432.

[0077] In operation S432, a task type to be used in the task trigger signal is identified based on the exhibition state information of the first video segment.

[0078] In some embodiments, the display status information may include at least one of playback progress information of the first video segment, task execution progress information used for the first video segment, and status information of a virtual character. The first video segment may be further divided to obtain multiple video subsegments. The playback progress information may indicate the playback status of each of the multiple video subsegments. The multiple video subsegments each correspond to multiple importance index values. The task execution progress information may indicate the execution progress of a preceding task used for the first video segment, indicating whether the preceding task has been completed. The status information may indicate whether the virtual character is present on the video screen. In response to identifying that playback of a target video subsegment has been completed and execution of the preceding task has been completed, a task type to be used for a task trigger signal is identified. The target video subsegment is a video subsegment whose importance index value is equal to or greater than a predetermined importance index threshold. For example, if playback of a target video subsegment related to important content has been completed, execution of the preceding task has been completed, and the action related to the preceding task is different from the action that generates the task trigger signal, a task type to be used for the task trigger signal may be identified. There may be multiple task types. The multiple task types include rating invitations, user questions and answers, and user behavior feedback. For example, a task queue can be set up, tasks to be performed can be added to the task queue, and the priorities of the tasks to be performed can be specified. The tasks in the task queue can be sorted based on priority, so that the live streaming rhythm and experience can still be consistent in the face of multi-task competition. The large-scale model can specify task priorities and pre-set priorities for different types of tasks.

[0079] In operation S441, a target add location is identified from at least one predetermined add location in the first video segment.

[0080] In some embodiments, the first video segment may include at least one predetermined adding position. The predetermined adding position may correspond to one possible time of adding the first video segment. Adding a new video segment based on the predetermined adding position has little impact on the consistency of the scenario content of the first video segment.

[0081] In some embodiments, a plurality of time offset values ​​are identified based on the plurality of possible addition times and the task trigger time at which the task trigger signal is received. A target addition position is identified from a plurality of predetermined addition positions based on the plurality of time offset values. A second scenario segment to be used for the task trigger signal is generated based on the task type and the target addition position. For example, the predetermined addition position corresponding to the smallest time offset value can be determined as the target addition position.

[0082] In operation S442, a second scenario segment to be used in the task trigger signal is generated based on the task type and the target add location of the at least one predetermined add location.

[0083] In some embodiments, the second scenario segment includes at least one of task content text and content linking text, and the content linking text is generated based on context data of the target addition location. The second scenario segment is used to generate a second video segment to be added to the target addition location. For example, the large-scale model can generate task content text based on the task type. For example, if the task type is "user question and answer," answer text can be generated as the task content text based on a question provided by the user. Also, for example, context text to be used at the target addition location can be obtained from the first scenario segment based on the target addition location. Then, the content linking text can be generated based on the context text. This can further match the style of the second scenario segment with the task setting information.

[0084] In some embodiments, the second scenario segment is used to generate a second video segment to be added to the target addition position. For example, in response to identifying a difference between the second scenario segment and the character setting information being less than a predetermined difference threshold, the second video segment is generated based on the second scenario segment. In response to identifying a difference between the second scenario segment and the character setting information being equal to or greater than the predetermined difference threshold, the second video segment is generated based on a second adjusted scenario segment. The second adjusted scenario segment is obtained by adjusting the second scenario segment. The above description based on the first scenario segment and the predetermined difference threshold also applies to the second scenario segment, and the present disclosure will not repeat the description here.

[0085] According to the embodiments of the present disclosure, the digital human can perform stable live streaming based on a predetermined first scenario segment, and can also realize intelligent responses based on signals during the live streaming process, thereby realizing a flexible, natural, and content-rich live streaming interactive experience and improving user stickiness and live streaming conversion rates.

[0086] Having described the method of the present disclosure, the apparatus of the present disclosure will now be described.

[0087] FIG. 5 is a block diagram of a live streaming scenario generation device according to an embodiment of the present disclosure.

[0088] As shown in FIG. 5, the apparatus 500 may include a first generating module 510 and a first identifying module 520 .

[0089] The first generation module 510 is used to generate at least one first scenario segment based on the initial input information, the first scenario segment including a speech content text and an object description subsegment used for the live distribution object, the object description subsegment including an object description text, and the object description text is used to describe at least one of an exhibition behavior of the live distribution object and an exhibition mode used for the speech content text.

[0090] The first identification module 520 is used to identify a live broadcast scenario based on at least one first scenario segment.

[0091] In some embodiments, the live-streaming object is at least one, and the at least one live-streaming object includes at least one of a virtual character and an object to be exhibited, and the initial input information includes at least one of character setting information used for the virtual character, initial object information used for the object to be exhibited, live-streaming material information, and live-streaming tool information. The live-streaming scenario is used to generate a live-streaming video including a first video segment corresponding to a first scenario segment.

[0092] In some embodiments, the object description subsegment is embedded in the utterance content text, and the object description subsegment further includes at least one of an exhibit identifier, a subsegment start identifier, a subsegment end identifier, and a delimiter identifier, the delimiter identifier being located between the exhibit identifier and the object description text.

[0093] In some embodiments, the object description subsegment is one of an object exhibition mode description subsegment, an object action description subsegment, and an object expression description subsegment, wherein the object exhibition mode description subsegment includes object exhibition mode description text, the object action description subsegment includes object action description text, and the object expression description subsegment includes object expression description text. The audio data used for the utterance content text in the first video segment is generated according to the exhibition mode described by the object exhibition mode description text. The image data used for the virtual character in the first video segment is generated according to at least one of the object action description text and the object expression description text, wherein the object action description text is used to describe at least one limb action of the virtual character, the at least one limb action including at least one first limb action used for the object to be exhibited, and the object expression description text is used to describe at least one facial action of the virtual character.

[0094] In some embodiments, the first generation module includes a first identification submodule for identifying at least one target knowledge information based on at least one of character setting information, at least one initial object information, live streaming material information, and live streaming tool information and a knowledge library, and a generation submodule for generating at least one first scenario segment based on the at least one target knowledge information.

[0095] In some embodiments, the first identification submodule includes: a first identification unit for identifying live streaming content planning information based on at least one of character setting information, initial object information, live streaming material information, and live streaming tool information, wherein the first identification unit is used to indicate that the live streaming content planning information includes at least one of introduction text of the object to be exhibited, history case text of the object to be exhibited, and guide text of the object to be exhibited; a second identification unit for identifying a plurality of initial search words based on the live streaming content planning information; and a third identification unit for identifying at least one target knowledge information based on the plurality of initial search words and the knowledge library.

[0096] In some embodiments, the third identification unit is further used to perform N searches in the knowledge library based on a plurality of first-level search words to obtain at least one target knowledge information, where N is an integer greater than or equal to 1, and the first-level search words are identified based on the initial search words.

[0097] In some embodiments, the third identification unit includes a first identification subunit for identifying n-th level search results in a predetermined database based on the plurality of n-th level search words, and a second identification subunit for identifying a plurality of n+1-th level search words based on the n-th level search results and the plurality of n-th level search words, wherein at least one target knowledge information is identified based on the N-th level search results, where n is an integer greater than or equal to 1 and less than N.

[0098] In some embodiments, the first identification module includes a second identification sub-module for identifying the first scenario segment as a live-stream scenario segment of the live-stream scenario in response to determining that a difference between the first scenario segment and the character setting information is less than a predetermined difference threshold, and a third identification sub-module for identifying the first adjusted scenario segment as a live-stream scenario segment of the live-stream scenario in response to determining that a difference between the first scenario segment and the character setting information is equal to or greater than the predetermined difference threshold, wherein the first adjusted scenario segment is obtained by adjusting the first scenario segment.

[0099] In some embodiments, the first video segment includes at least one predetermined add-on location. In response to receiving the task trigger signal for use with the first video segment, the device further includes: a second identification module for identifying a task type for use with the task trigger signal based on exhibition state information of the first video segment; and a second generation module for generating a second scenario segment for use with the task trigger signal based on the task type and a target add-on location among the at least one predetermined add-on location. The second scenario segment includes at least one of task content text and content linking text, where the content linking text is generated based on context data of the target add-on location, and the second scenario segment is used to generate a second video segment to be added to the target add-on location.

[0100] In some embodiments, the second scenario segment is used to instruct generating a second video segment to be added to the target addition location by the operation of: generating a second video segment based on the second scenario segment in response to determining that a difference between the second scenario segment and the character setting information is less than a predetermined difference threshold; generating a second video segment based on a second adjusted scenario segment in response to determining that a difference between the second scenario segment and the character setting information is equal to or greater than a predetermined difference threshold; the second adjusted scenario segment obtained by adjusting the second scenario segment.

[0101] In some embodiments, the exhibition status information includes at least one of playback progress information of a first video segment and task execution progress information used for the first video segment, where the first video segment includes a plurality of video subsegments corresponding to a plurality of importance index values, the playback progress information being used to indicate the playback status of each of the plurality of video subsegments, and the task execution progress information being used to indicate the execution progress of a preceding task used for the first video segment, the preceding task being a task prior to receiving the task trigger signal. The second identification module includes a fourth identification submodule for identifying a task type used for the task trigger signal in response to identifying that playback of the target video subsegment has been completed and that execution of the preceding task has been completed. The target video subsegment is a video subsegment whose importance index value is equal to or greater than a predetermined importance index threshold.

[0102] In some embodiments, the predetermined adding positions are multiple, and the multiple predetermined adding positions respectively correspond to multiple possible adding times of the first video segment. The second generating module includes a fifth identifying module for identifying multiple time offset values ​​based on the multiple possible adding times and the task trigger time at which the task trigger signal is received, and a sixth identifying sub-module for identifying a target adding position from the multiple predetermined adding positions based on the multiple time offset values. The second generating sub-module is used to generate a second scenario segment to be used for the task trigger signal based on the task type and the target adding position.

[0103] In the technical solution disclosed herein, the collection, storage, use, processing, transmission, provision, disclosure and other processes of relevant user personal information shall all comply with the provisions of relevant laws and regulations and shall not violate public order and morals.

[0104] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program.

[0105] 6 shows a schematic block diagram of an exemplary electronic device 600 for implementing embodiments of the present disclosure. The electronic device is intended to represent various types of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The electronic device may also represent various types of mobile devices, such as personal digital assistants, mobile phones, smartphones, wearable devices, and other similar computing devices. The components, their connections and relationships, and their functions shown herein are merely exemplary and do not limit the implementation of the present disclosure as described and / or claimed herein.

[0106] 6, the device 600 includes a computing unit 601, which may perform various appropriate operations and processes based on a computer program stored in a read-only memory (ROM) 602 or loaded from a storage unit 608 into a random access memory (RAM) 603. The RAM 603 may further store various programs and data necessary for the operation of the device 600. The computing unit 601, the ROM 602, and the RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0107] The components of device 600 are connected to an I / O interface 605, which includes an input unit 606 such as a keyboard, a mouse, etc., an output unit 607 such as various types of displays, speakers, etc., a storage unit 608 such as a magnetic disk, an optical disk, etc., and a communication unit 609 such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 enables device 600 to exchange information and data with other devices via a computer network such as the Internet and / or various electrical networks.

[0108] The computing unit 601 may be various general-purpose and / or specialized processing modules having processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graph processing unit (GPU), various specialized artificial intelligence (AI) computing chips, computing units running various machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 executes the methods and processes described above, such as the live streaming scenario generation method. For example, in some embodiments, the live streaming scenario video generation method may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 608. In some embodiments, some or all of the computer program may be loaded and / or installed into the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, it may perform one or more steps of the live streaming scenario generation method described above. Alternatively, in other embodiments, the computing unit 601 may be configured to perform the live streaming scenario generation method in any other suitable manner (e.g., via firmware).

[0109] Various embodiments of the systems and techniques described herein above may be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may be embodied in one or more computer programs that can be executed and / or interpreted by a programmable system that includes at least one programmable processor, which may be a special purpose or general purpose programmable processor, and that can receive data and instructions from, and transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0110] Program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, so that when the program code is executed by the processor or controller, the functions and operations specified in the flowcharts and / or block diagrams are performed. The program code may be executed entirely on a device, partially on a device, partially on a device as a separate software package, and partially on a remote device, or entirely on a remote device or server.

[0111] In the context of this disclosure, a machine-readable medium may be a tangible medium, and may contain or store a program for use in or in connection with an instruction execution system, device, or electronic device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or electronic device, or any suitable combination of the above. More specific examples of machine-readable storage media include an electrical connection of one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM) or flash memory, an optical fiber, a compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0112] To provide for user interaction, a computer may implement the systems and techniques described herein and include a display device (e.g., a cathode ray tube (CRT) display or a liquid crystal display (LCD)) for displaying information to a user, and a keyboard and pointing device (e.g., a mouse or trackball) through which a user can provide input to the computer. Other types of devices may also provide for user interaction; for example, the feedback provided to the user may be any form of sensing feedback (e.g., visual feedback, auditory feedback, or tactile feedback) and may receive input from the user in any form (including voice input, speech input, or tactile input).

[0113] The systems and techniques described herein can be implemented in a computing system including background components (e.g., a data server), or middleware components (e.g., an application server), or front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with embodiments of the systems and techniques described herein), or any combination of such background, middleware, or front-end components. The components of the system can be connected to each other by any form or medium of digital data communication (e.g., a communications network). Examples of communications networks include, by way of example, a local area network (LAN), a wide area network (WAN), and the Internet.

[0114] A computer system may include clients and servers. Clients and servers are generally remote and typically interact through a communication network. The relationship of client and server is created by computer programs running on the corresponding computers and having a client-server relationship.

[0115] It should be understood that various types of flows shown above may be used, and operations may be rearranged, added, or deleted. For example, the operations described in this disclosure may be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this specification is not limited thereto.

[0116] The above specific embodiments do not limit the scope of protection of the present disclosure. Those skilled in the art should understand that various modifications, combinations, subcombinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present disclosure should be included within the scope of protection of the present disclosure.

Claims

1. A live streaming scenario generation method, comprising: generating at least one first scenario segment based on initial input information, the first scenario segment including a speech content text and an object description sub-segment used for a live distribution object, the object description sub-segment including an object description text, the object description text describing at least one of an exhibition operation of the live distribution object and an exhibition mode used for the speech content text; and identifying a live broadcast scenario based on at least one of the first scenario segments. A live distribution scenario generation method comprising:

2. the live-streaming object is at least one, and the at least one live-streaming object includes at least one of a virtual character and an object to be exhibited; the initial input information includes at least one of character setting information used for the virtual character, initial object information used for the object to be exhibited, live-streaming material information, and live-streaming tool information; the live-streaming scenario is used to generate a live-streaming video including a first video segment corresponding to the first scenario segment; 2. The method of claim 1 .

3. the object description sub-segment is embedded in the utterance content text; the object description sub-segment further includes at least one of an exhibit identifier, a sub-segment start identifier, a sub-segment end identifier, and a delimiter identifier, the delimiter identifier being located between the exhibit identifier and the object description text; 2. The method of claim 1 .

4. the object description subsegment is one of an object exhibition mode description subsegment, an object action description subsegment, and an object expression description subsegment, the object exhibition mode description subsegment including object exhibition mode description text, the object action description subsegment including object action description text, and the object expression description subsegment including object expression description text; audio data used for the utterance content text in the first video segment is generated according to the exhibition mode described by the object exhibition mode description text; image data used for the virtual character in the first video segment is generated according to at least one of the object action description text and the object expression description text, the object action description text is used to describe at least one limb action of the virtual character, the at least one limb action includes at least one first limb action used for the object to be exhibited, and the object expression description text is used to describe at least one facial action of the virtual character; 4. The method of claim 3.

5. Generating at least one first scenario segment based on the initial input information includes: Identifying at least one piece of target knowledge information based on the character setting information, at least one piece of initial object information, the live streaming material information, and the live streaming tool information, and a knowledge library; generating at least one of the first scenario segments based on at least one of the target knowledge information; 3. The method of claim 2.

6. identifying at least one piece of target knowledge information based on at least one of the character setting information, at least one piece of initial object information, the live streaming material information, and the live streaming tool information, and a knowledge library; Identifying live streaming content planning information based on at least one of the character setting information, at least one of the initial object information, the live streaming material information, and the live streaming tool information, wherein the live streaming content planning information is used to indicate that the first scenario segment includes at least one of an introduction text of the object to be exhibited, a history case text of the object to be exhibited, and a guide text of the object to be exhibited; Identifying a plurality of initial search words based on the live streaming content planning information; and identifying at least one of the target knowledge information based on a plurality of the initial search words and the knowledge library.

6. The method of claim 5.

7. Identifying at least one of the target knowledge information based on a plurality of the initial search words and the knowledge library includes: performing N searches in the knowledge library based on a plurality of first-level search words to obtain at least one of the target knowledge information, where N is an integer greater than or equal to 1, and the first-level search words are identified based on the initial search words; 7. The method of claim 6.

8. conducting N searches in the knowledge library; identifying n-th level search results in the predetermined database based on a plurality of n-th level search words; Identifying a plurality of n+1-th level search words based on the n-th level search results and a plurality of the n-th level search words, wherein at least one of the target knowledge information is identified based on the N-th level search results, where n is an integer greater than or equal to 1 and less than N; 8. The method of claim 7.

9. Identifying a live streaming scenario based on at least one of the first scenario segments includes: In response to the difference between the first scenario segment and the character setting information being determined to be smaller than a predetermined difference threshold, identifying the first scenario segment as a live distribution scenario segment of the live distribution scenario; and identifying a first adjusted scenario segment obtained by adjusting the first scenario segment as a live distribution scenario segment of the live distribution scenario in response to the difference between the first scenario segment and the character setting information being identified as being equal to or greater than a predetermined difference threshold.

3. The method of claim 2.

10. the first video segment includes at least one predetermined additional position; In response to receiving a task trigger signal used for a first video segment, identifying a task type used for the task trigger signal based on display state information of the first video segment; generating a second plot segment to be used for the task trigger signal based on the task type and a target add location among at least one predetermined add locations, the second scenario segment including at least one of task content text and content link text, the content link text being generated based on context data of the target add location, and the second scenario segment being used to generate a second video segment to be added to the target add location.

3. The method of claim 2.

11. The second scenario segment comprises: generating the second video segment based on the second scenario segment in response to identifying a difference between the second scenario segment and the character setting information being less than a predetermined difference threshold; generating the second video segment based on a second adjusted scenario segment obtained by adjusting the second scenario segment in response to the difference between the second scenario segment and the character setting information being identified as being equal to or greater than a predetermined difference threshold; is used to instruct the generation of a second video segment to be added to the target addition position by the operation 11. The method of claim 10.

12. The display state information includes at least one of playback progress information of the first video segment and task execution progress information used for the first video segment, the first video segment including a plurality of video sub-segments, each of the plurality of video sub-segments corresponding to a plurality of importance index values, the playback progress information being used to indicate a playback state of each of the plurality of video sub-segments, the task execution progress information being used to indicate an execution progress of a preceding task used for the first video segment, the preceding task being a task before receiving the task trigger signal; Identifying a task type to be used in the task trigger signal based on exhibition state information of the first video segment includes: identifying a task type to be used in the task trigger signal in response to the target video sub-segment being identified as having completed playback and the preceding task being identified as having completed execution, the target video sub-segment being a video sub-segment having an importance index value equal to or greater than a predetermined importance index threshold; 11. The method of claim 10.

13. the predetermined adding positions are plural, and the plural predetermined adding positions correspond to plural addable times of the first video segment, respectively; generating a second scenario segment to be used in the task trigger signal based on the task type and a target additional location among at least one predetermined additional location; determining a plurality of time offset values ​​based on the plurality of addable times and the task trigger time at which the task trigger signal is received; identifying a target add location from the plurality of predetermined add locations based on the plurality of time offset values; generating a second scenario segment to be used in the task trigger signal based on the task type and the target addition position; 11. The method of claim 10.

14. A live streaming scenario generation device, comprising: a first generation module for generating at least one first scenario segment based on initial input information, the first scenario segment including a speech content text and an object description sub-segment used for a live distribution object, the object description sub-segment including an object description text, the object description text being used to describe at least one of an exhibition action of the live distribution object and an exhibition model used for the speech content text; a first identification module for identifying a live streaming scenario based on at least one of the first scenario segments; A live distribution scenario generation device characterized by:

15. An electronic device, at least one processor; a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor such that the at least one processor can perform the method of any one of claims 1 to 13; An electronic device characterized by:

16. A non-transitory computer-readable storage medium having computer instructions stored thereon, comprising: The computer instructions are used to cause the computer to carry out the method of any one of claims 1 to 13. A non-transitory computer-readable storage medium comprising:

17. A computer program which, when executed by a processor, implements the method according to any one of claims 1 to 13.