Live script generation method and device, electronic equipment and storage medium
Through a large language model, live scripts are generated and object description sub-snippets are combined, which solves the problem of insufficient expressiveness and interactive experience of digital human live broadcast in high interactive scenarios, realizes multimodal collaboration and real-time response, and improves the immersion and conversion rate of digital human live broadcast.
Patent Information
- Application Number
- CN202510536622.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-07-11
AI Technical Summary
The existing digital live broadcast technology has shortcomings in high interaction, emotional rendering and multimodal collaboration, resulting in a significant gap between expressiveness and live anchors, and it is impossible to flexibly respond to changes in the live broadcast room, insufficient interactive experience, unnatural movements and expressions, lack of emotional fluctuations and dramatic tension, making it difficult to impress the audience in the live broadcast.
Generate live scripts through a large language model, combine object description sub-fragments, synchronize the multi-modal behavior of digital people, realize the coordination of speech content and actions, embed object display identifiers, starting characters and separators, generate logically smooth live copywriting, support multi-character collaboration and real-time interaction, use the knowledge base to enhance the logic and professionalism of the content, and dynamically adjust the consistency of character expressions.
It improves the expressiveness and interactivity of digital people's live broadcasts, enhances the immersion and information transmission efficiency of live broadcasts, improves the audience's viewing focus and conversion rate, and achieves 24-hour uninterrupted and efficient live broadcasts.
Smart Images

Figure CN120302126A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to the fields of natural language processing, large models, virtual digital humans, etc., and can be applied to the virtual human live broadcast scenario. More specifically, the present disclosure provides a live script generation method, apparatus, electronic device, and storage medium. Background Art
[0002] With the development of artificial intelligence technology, the application of virtual digital humans is increasing continuously. Virtual digital humans can replace real human anchors to achieve all-weather and uninterrupted live broadcasts. Summary of the Invention
[0003] The present disclosure provides a live script generation method, apparatus, device, and storage medium.
[0004] According to one aspect of the present disclosure, there is provided a live script generation method, the method including: generating at least one first script segment according to initial input information, wherein the first script segment includes speech content text and an object description sub-segment for a live object, and the object description sub-segment includes object description text for describing at least one of the actions shown by the live object and the display mode for the speech content text; and determining a live script according to at least one first script segment.
[0005] According to another aspect of the present disclosure, there is provided a live script generation apparatus, the apparatus including: a first generation module for generating at least one first script segment according to initial input information, wherein the first script segment includes speech content text and an object description sub-segment for a live object, and the object description sub-segment includes object description text for describing at least one of the actions shown by the live object and the display mode for the speech content text; and a first determination module for determining a live script according to at least one first script segment.
[0006] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method provided by the present disclosure.
[0007] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method provided by the present disclosure.
[0008] According to another aspect of the present disclosure, there is provided a computer program product including a computer program which, when executed by a processor, implements the method provided by the present disclosure.
[0009] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0011] Figure 1 is a schematic diagram of an exemplary system architecture to which the live script generation method and apparatus according to an embodiment of the present disclosure can be applied;
[0012] Figure 2 is a flowchart of a live script generation method according to an embodiment of the present disclosure;
[0013] Figure 3 is a schematic diagram of a live script generation method according to an embodiment of the present disclosure;
[0014] Figure 4 is a flowchart of a live script generation method according to another embodiment of the present disclosure;
[0015] Figure 5 is a block diagram of a live script generation apparatus according to an embodiment of the present disclosure; and
[0016] Figure 6 is a block diagram of an electronic device to which the live script generation method according to an embodiment of the present disclosure can be applied. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] The following describes exemplary embodiments of the present disclosure with reference to the drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted for clarity and conciseness.
[0018] In the technical solution of the present disclosure, the processing of collection, storage, use, processing, transmission, provision, and disclosure of user personal information complies with the provisions of relevant laws and regulations and does not violate public order and good customs.
[0019] With the rapid development of the e-commerce live streaming industry, live commerce has gradually become an important marketing method. However, real human hosts have limited energy and high live streaming operation costs. It is difficult to achieve round-the-clock uninterrupted live streaming based on real human hosts. Therefore, in order to reduce operation costs and achieve uninterrupted live streaming, the "digital human live streaming" technology based on artificial intelligence can be applied. "Digital human live streaming" refers to replacing real human hosts with virtual digital humans to complete a series of live streaming tasks in the live streaming room, such as product explanation, interactive Q&A, and user guidance. Digital human live streaming integrates a number of cutting-edge technologies: relying on large language models (LLMs) to generate live streaming scripts, using text-to-speech (TTS) technology to synthesize natural voices, and combining lip movement synthesis systems to achieve the driving effect of "synchronization of voice and form". Each link in this process is usually decoupled and executed with relatively low mutual dependence, belonging to the technical path of "static generation + serial drive". Compared with real human hosts, digital humans can achieve 24-hour uninterrupted live streaming, with significant cost advantages, potential for efficiency improvement, ability to expand exposure, and stronger controllability. Digital human live streaming technology is relatively efficient in processing live streaming content with relatively simple structures and fixed scenarios. For example, digital humans can be used in low-interaction scenarios such as loop-played promotional video-style live streams or single-product displays. In these low-interaction scenarios, the movements of digital humans are small, the facial expressions change limitedly, and they basically do not interact with the audience in real time, and the expressiveness and stability of digital humans can be maintained.
[0020] However, as live commerce gradually develops towards a highly interactive and emotion-rendering direction, live streaming scenarios involve stronger performance requirements (such as emotional tensions like excitement, surprise, emphasis, etc.), more complex action coordination (such as item interaction, large-scale movements, etc.), and more real-time interactions with users (such as bullet chat responses, lucky bag distributions, price changes, etc.). The above-mentioned static and decoupled technical path is difficult to handle such high-interaction and emotion-rendering scenarios.
[0021] In low-interaction scenarios, the live streaming scripts used can be pre-generated, unable to respond to changes in the live streaming room in real time and lacking flexibility. In addition, during the live streaming process, the voices, movements, and facial expressions shown by digital humans are generated by different and relatively independent modules, unable to achieve high-level coordination of multiple modalities, resulting in unnatural visual effects and insufficient interactive experiences, and further leading to a significant gap in the overall expressiveness between digital human live streaming and real human hosts. The following will explain in combination with issues such as live streaming scripts, the actions of digital human hosts, tools in the live streaming room, response to changes in the live streaming room, and multi-person live streaming.
[0022] The content of the live broadcast scripts used by digital people for live broadcasts lacks appeal. Most live broadcast scripts have been generated in a templated form, with similar content and lack of novelty, which cannot effectively stimulate the interest of the audience. For example, expressions such as "the lowest price on the entire network" and "you will have to wait a year if you miss it today" are overused, and users are tired of these expressions and it is difficult to leave a deep impression. In addition, the script content is difficult to customize according to the characteristics of the product, the personality of the anchor and the target audience, resulting in a bland atmosphere in the live broadcast room, a lack of emotional fluctuations and dramatic tension, and it is difficult to form a "hot spot" or "break the circle" dissemination at the content level. This script lacks storytelling, interactivity and emotional rendering, making it difficult for digital people to truly impress the audience during live broadcasts.
[0023] In addition, during the live broadcast, it is difficult for digital human anchors to show highly expressive movements. Digital human anchors are still relatively simple in terms of movement performance. The movements of digital humans mainly involve lip shape and voice synchronization, a small number of basic expressions and gesture-driven movements, and it is difficult to show movements that are highly matched with the script semantics. When expressing emotions such as excitement, surprise, and emphasis, the amplitude of digital human movements is often not sufficient, and the changes in expression are not delicate enough, which cannot truly convey the ups and downs of emotions and changes in tone, resulting in the overall expression being dull and "mechanical". Especially in the live broadcast with goods scene, this lack of infectious movement expression is difficult to mobilize the emotions of the audience, and it is also difficult to create a "strong sense of substitution" viewing experience. In addition, there may be delays, dislocations, and even discord between the movements, expressions, voices, and intonations of digital humans, which further weakens the immersion and professionalism of the live broadcast. Compared with real anchors, digital humans are still not natural and smooth enough, and have obvious disadvantages in strengthening brand impressions and improving live broadcast conversions.
[0024] In addition, it is difficult for digital humans to call various props in the live broadcast room. Most digital human live broadcast systems have limited support for physical or virtual props in live broadcast scenes, and lack the ability to deeply schedule "interactive props" in live broadcasts. The actions implemented by digital humans, such as switching promotional materials, distributing lucky bags, limited-time prompts, and bullet screen interactive guidance, are often unable to be accurately synchronized with the script content. Many actions performed by digital humans still rely on manual or background operation and maintenance intervention, lacking automation and script-driven capabilities, resulting in a broken live broadcast rhythm and a discounted interactive effect. In scenes where two or more digital humans collaborate in live broadcasts, problems such as different characters repeatedly operating the same props, information conflicts, or logical confusion often occur, affecting the overall live broadcast quality. In addition, the calling of props also lacks contextual design. When introducing skin care products, it is difficult for digital human anchors to naturally "pick up" products and show packaging details.
[0025] In addition, it is difficult for digital humans to achieve dynamic explanations according to the changes in the live broadcast room. Based on the static preset script-driven method, digital humans are difficult to perceive and respond to the real-time changes during the live broadcast, resulting in the inability of the explanation content to flexibly adapt to the on-site situation and lacking the ability of "responding to changes on the spot". When the product inventory changes temporarily, users frequently ask questions about a certain function in the comments, or interactive information pops up on the screen, the digital human anchor often cannot immediately adjust the words or change the focus of the explanation, and can only mechanically broadcast according to the established lines, seriously affecting the user experience. In addition, digital humans basically lack the perception and response to real-time data such as the audience's emotions, bullet screen feedback, product popularity, and fluctuations in the number of viewers, and cannot implement a "tailor-made" sales strategy. When the attention of a certain product is significantly increased, the digital human cannot immediately turn to key promotion, nor can it actively generate personalized responses or interactive Q&A in combination with the user's interests, thus missing the key nodes for guiding conversion.
[0026] In addition, multiple digital humans cannot achieve a natural multi-person live broadcast effect. In a complex live broadcast room scenario, multiple digital humans with different roles need to work together. Multiple roles include the anchor, co-anchor, host, field controller, brand representative, etc. Multiple roles need to cooperate, fill in the gaps, and interact in real time to jointly create a lively, smooth, and layered live broadcast atmosphere. However, multiple digital humans cannot simulate the natural interaction forms of multiple anchors "grabbing the microphone", "interrupting", "echoing", and "cross-talking" in a real live broadcast room.
[0027] In order to realize digital human live broadcasts with "strong expressiveness, strong interactive opinions, and strong sales ability", a more integrated, intelligent, and real-time collaborative generation paradigm needs to be used to produce live broadcast scripts. Therefore, the present disclosure provides a method for generating a live broadcast script, which will be described below.
[0028] Figure 1 FIG. is a schematic diagram of an exemplary system architecture to which the live broadcast script generation method and apparatus according to an embodiment of the present disclosure can be applied. It should be noted that Figure 1 The figure shown is only an example of the system architecture to which the embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios.
[0029] As Figure 1 shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0030] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Terminal devices 101, 102, and 103 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablets, laptop computers, and desktop computers, etc.
[0031] Server 105 can be a server providing various services. For example, it can be a background management server (only for example) that supports the websites browsed by users using terminal devices 101, 102, and 103. The background management server can analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0032] It should be noted that the live script generation method provided by the embodiments of the present disclosure can generally be executed by server 105. Correspondingly, the live script generation device provided by the embodiments of the present disclosure can generally be set in server 105. The live script generation method provided by the embodiments of the present disclosure can generally also be executed by one or more of terminal devices 101, 102, and 103. Correspondingly, the live script generation device provided by the embodiments of the present disclosure can generally be set in one or more of terminal devices 101, 102, and 103. The live script generation method provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from server 105 and capable of communicating with terminal devices 101, 102, and 103 and / or server 105. Correspondingly, the live script generation device provided by the embodiments of the present disclosure can also be set in a server or a server cluster different from server 105 and capable of communicating with at least one of terminal devices 101, 102, and 103 and server 105.
[0033] It can be understood that the system architecture of the present disclosure has been described above, and the method of the present disclosure will be described below.
[0034] Figure 2 is a flowchart of a live script generation method according to an embodiment of the present disclosure.
[0035] As Figure 2 shown, the method 200 can include operation S210 to operation S220.
[0036] In operation S210, at least one first script segment is generated according to the initial input information.
[0037] In the embodiments of the present disclosure, the initial input information can include information related to the object to be displayed. For example, the object to be displayed can be a product shown during the live broadcast. The initial input information can include the description of the product.
[0038] In the embodiments of the present disclosure, the first script segment includes the speech content text and an object description sub-segment for the live object. For example, during the live broadcast, a digital human anchor can display the audio corresponding to the speech content text through forms such as dialogue and speech.
[0039] In the embodiments of the present disclosure, a large language model can be used to generate the first script segment.
[0040] In the embodiments of the present disclosure, the object description sub-segment includes object description text. The object description text is used to describe at least one of the actions shown by the live object and the display mode for the speech content text. For example, an object description text can describe one or more actions shown by the live object. Another example is that another object description text can describe the display mode for the speech content text. The display mode can include intonation. The intonation can be calm, enthusiastic, etc.
[0041] In the embodiments of the present disclosure, the live object can be a digital human anchor.
[0042] In operation S220, according to at least one first script segment, a live script is determined.
[0043] In the embodiments of the present disclosure, one or more first script segments can be used as one or more live script segments of the live script. For example, a first script segment can be used as a live script segment.
[0044] Through the embodiments of the present disclosure, when generating the speech content text, an object description sub-segment is generated, making the speech content of the digital human more coordinated with the actions of the live object, and providing richer information for the subsequent generation of live videos. After providing the live script to the large model, the large model can generate a live copy with smooth content logic and can also synchronously control the multi-modal behaviors of the digital human. The live script provided by the present disclosure can be implemented as a truly "executable script", which can enable the digital human live broadcast to transform from a single scene to a multi-dimensional, highly interactive, and high-performance scene.
[0045] It can be understood that the method of the present disclosure has been described above, and the live object and the object description sub-segment will be further described below.
[0046] In some embodiments, there is at least one live object, and the at least one live object includes at least one of a virtual character and an object to be displayed. For example, there can be one or more virtual characters. The one or more virtual characters include various roles played by digital humans such as digital human anchors, digital human assistants, digital hosts, digital human field controllers, and digital human brand representatives. Another example is that the object to be displayed can be a product or service shown during the live broadcast. Services can be, for example, travel services, legal services, etc. Products can be, for example, skin care products, luggage, etc.
[0047] In some embodiments, the object description sub-segment may be embedded in the speech content text. For example, multiple object description sub-segments may be embedded in multiple positions in the speech content text.
[0048] In some embodiments, the object description sub-segment may further include at least one of a display identifier, a sub-segment start symbol, a sub-segment end symbol, and a separator. The display identifier may be the name of the above-mentioned display mode, "action" or "expression". Based on the display identifier, the large model may generate audio, image or text data corresponding to the display identifier according to the object description text. The sub-segment start symbol may be symbols such as "(", "[", and the sub-segment end symbol may be symbols such as ")", "】". The separator may be symbols such as ":", "丨".
[0049] For example, part of the first script fragment of the present disclosure is:
[0050] "Host: (tone: calm) (action: pick up two opened boxes of products, put down the box in the left hand, point to the contents of the box in the right hand and explain how to use it) You see, the method of use is very simple (the assistant also said: very simple!), once a month, use it for 6 consecutive months, and take a break for 6 months, only 2 boxes are needed in a year. Each time you use it, mix agent A and agent B, and then use the matching roller to slightly prick the skin, which is painless and has a short recovery period. Do not touch water within 12 hours after use, and pay attention to sun protection.
[0051] Host: (Tone: calm) (Action: pick up the packaged roller on the table to show it, while the assistant uses his hands in front of his face to show how to use the roller) This roller is a tool that is used in combination. It can better help the nutrients penetrate and absorb. Just like loosening the soil, it allows the nutrients in the seeds to take root and sprout better, and allows the skin to absorb nutrients better.
[0052] Assistant: (tone: calm) (action: take the tool kit from the off-screen, open it and show it, while the host opens the roller packaging bag and takes out the roller to show it) My dears, the tool kit is fully equipped for you (host also says: very considerate!), you can use it directly after you get it, it is very convenient. Moreover, our product is a micro-puncture and micro-break of the skin, it is painless and the recovery period is short, and it will not affect your normal life.
[0053] Assistant: (Tone: enthusiastic) (Action: taking out two bottles of solution from the product box on the table, displaying them side by side, and then putting them back into the box) Look everyone, these two bottles are Agent A and Agent B. When they are combined together, they can exert a powerful effect. Just like two superheroes joining forces, they can defeat various skin problems and make the skin white, tender and smooth.
[0054] Host: (Tone: enthusiastic) (Expression: happy) All friends, with such a great product and so many benefits, what are you still waiting for? The opportunity is rare and the inventory is limited. Miss this time and it may be a long time before you can get such a discount again.
[0055] Assistant: (Tone: enthusiastic) (Action: Pick up the plastic board showing the before-and-after comparison photos of product use and introduce the usage effect) (Expression: surprised) Babies, take a look at these comparison pictures again. The effect is really immediate (Host says at the same time: This is exaggerated). There were various problems with the skin before using it, and it has improved a lot after using it. Do you all want to have such skin too?
[0056] As shown in the first script segment of this part, multiple object description sub-segments can include (Tone: calm), (Action: Pick up two opened products, put down the left one, and point to the contents of the right one to explain the usage method), (Expression: happy), etc. Taking the object description sub-segment (Tone: calm) as an example, "Tone" can be the display identifier, "calm" can be the object description text, "(" can be the sub-segment start symbol, and ")" can be the sub-segment end symbol. ":" can be the separator, located between the display identifier "Tone" and the object description text "calm". In addition, the virtual characters involved in the first script segment of this part are multiple. The multiple virtual characters include a digital human host and a digital human assistant. It can be understood that the display identifier, sub-segment start symbol, sub-segment end symbol, and separator in the first script segment of this part are only examples. Through the embodiments of the present disclosure, an object description sub-segment (display identifier: object description text) is provided, which can structurally embed information such as tone, action, expression, and prop scheduling into the live script, enabling the large language model to generate logically smooth live copywriting and effectively control multi-modal behaviors such as the tone, behavior, and expression of digital humans.
[0057] In some embodiments, the object descriptor fragment may be one of an object display mode descriptor fragment, an object action descriptor fragment, and an object expression descriptor fragment. The object display mode descriptor fragment includes object display mode description text. The object action descriptor fragment includes object action description text, and the object expression descriptor fragment includes object expression description text. For example, as shown in the first script fragment of the above part, the object descriptor fragment (intonation: calm) may be an object display mode descriptor fragment and may include the object display mode description text "calm". The object descriptor fragment (action: pick up two opened products, put down the left one, and point to the contents of the right one to explain the usage method) may be an object action descriptor fragment and may include the object action description text "pick up two opened products, put down the left one, and point to the contents of the right one to explain the usage method". The object descriptor fragment (expression: happy) may be an object expression descriptor fragment and may include the object expression mode description text "happy". It can be understood that when the display mode is intonation, the object display mode descriptor fragment may also be referred to as the object intonation descriptor fragment. The object display mode description text may also be referred to as the object intonation description text.
[0058] It can be understood that the object descriptor fragment of the present disclosure has been described above, and the functions of the first script fragment and the object descriptor fragment will be further described below.
[0059] In some embodiments, the live script can be used to generate a live video. The live video may include a first video fragment corresponding to the first script fragment. For example, inputting the live script into a large model can obtain the live video. The first script fragment may be one or more. The live video may include one or more first video fragments. Each first script fragment corresponds to a first video fragment.
[0060] In some embodiments, the audio data for the speech content text in the first video fragment is generated according to the display mode described by the object display mode description text. For example, the audio data for the speech content text "Look, the usage method is very simple" is generated according to the "intonation" described by the object display mode description text "calm".
[0061] In some embodiments, the image data for the virtual character in the first video clip is generated according to at least one of the object action description text and the object expression description text. The object action description text is used to describe at least one limb movement of the virtual character. The at least one limb movement may include a first limb movement for the object to be displayed, and may also include a second limb movement of the live object. For example, the object action description text "Pick up two boxes of opened products, put down the box in the left hand, and point to the contents of the box in the right hand to explain the usage method" can describe at least one first limb movement for the object to be displayed, namely "products". Another example, the object action description text "Walk to the center of the screen" can describe the second limb movement of the virtual character itself.
[0062] In some embodiments, the object expression description text can describe at least one facial movement of the virtual character. For example, the object expression description text "Happy" can describe one or more facial movements. The one or more facial movements can show a happy expression.
[0063] Through the embodiments of the present disclosure, the object description sub-fragment (display identifier: object description text) can be flexibly adapted to different types of live broadcast scenarios, such as multiple vertical scenarios like beauty and skincare, health and nutrition, home appliances and digital products, etc. Taking the beauty scenario as an example, while the digital human anchor is explaining the product efficacy, in the case of having an object action description sub-fragment in the script, actions such as "smearing action", "gesturing the facial area", "displaying the matching set" can be automatically shown, making the product information more vivid and persuasive. In addition, taking the health and nutrition scenario as an example, in the case where the script includes an object action description sub-fragment, the digital human co-anchor can automatically show actions such as "opening the product packaging", "showing the taking method", "holding a sign to prompt precautions", etc., improving the audience's understanding and trust in the usage method.
[0064] In addition, through the embodiments of the present disclosure, based on the first limb movement for the object to be displayed, the system can realize the linkage control of the script with the digital human actions and live broadcast props, achieving a three-dimensional collaborative expression of semantics-vision-interaction, enabling the digital human to not only "be able to speak" but also "be able to act", "be able to move", and "be able to display" during the live broadcast. During the product introduction process, the digital human can naturally pick up the product, operate the props, display comparison pictures, and even cooperate with the co-anchor to complete "demonstrative explanations", significantly enhancing the information transmission efficiency and viewing experience. This immersive expression method greatly improves the user's viewing concentration and content memory points, providing a content basis for subsequent conversion behaviors.
[0065] In some embodiments, the first script segment includes multiple speech content texts for multiple virtual characters respectively. The object description sub-segment for the live object is embedded in the speech text for the virtual characters. As shown in the above part of the first script text, the multiple virtual characters include a digital human host and a digital human co-host. The object description sub-segment (intonation: calm) for the digital human host is embedded in the speech content text for the digital human host. The object description sub-segment (intonation: enthusiastic) for the digital human co-host is embedded in the speech content text for the digital human co-host.
[0066] Through the embodiments of the present disclosure, the object description sub-segment (display identifier: object description text) supports multi-role collaborative modeling. In the script generation stage, the speech content and action positions can be generated respectively for roles such as the host, co-host, and field controller, finely dividing the role positioning and functional division of labor, and improving the multi-person collaboration efficiency. For example, the co-host can focus on creating an atmosphere, supplementing details, and promoting interaction, while the host can focus on explaining the core content of the product, forming an efficient cooperation in the script.
[0067] It can be understood that the above has described the first script segment of the present disclosure, and the method of the present disclosure will be further described below.
[0068] Figure 3 It is a schematic diagram of a live script generation method according to an embodiment of the present disclosure.
[0069] In some embodiments, the initial input information includes at least one of character setting information for a virtual character, initial object information for an object to be displayed, live broadcast material information, and live broadcast tool information. There can also be multiple pieces of initial object information. Taking the object to be displayed as a product to be displayed as an example, multiple pieces of initial object information can be information about multiple products. The live broadcast material information can indicate multiple materials. These multiple materials can be for one or more products. The live broadcast tool information can indicate multiple live broadcast tools available in the live broadcast room. The multiple live broadcast tools can include a bullet screen tool, a comment tool, etc. There can be one or more pieces of character setting information. Multiple pieces of character setting information are respectively used for multiple virtual characters. For example, the character setting information can provide a style persona for the virtual character, enabling the large model to flexibly adjust the content style according to dimensions such as the live broadcast goal, live broadcast audience portrait, and product attributes in combination with the style persona requirements, ensuring that the script style highly matches the scenario. In one example, in a "knowledge dissemination" scenario, based on the style corresponding to this scenario, content such as professional terms, case analyses, and knowledge extensions can be introduced based on the corresponding knowledge base, making the live broadcast more authoritative and educational. In another example, in a "sentiment sharing" scenario, the script generated by the large model will adjust the speech rate, emotional curve, and narrative method, and pay more attention to semantic resonance and audience immersion. By introducing character setting information, a flexible and customizable style control ability can be achieved, making the script have good adaptability and appeal in different vertical categories and for different target audiences.
[0070] In some embodiments, in some implementation manners of the above operation S210, generating at least one first script segment according to the initial input information includes: determining at least one target knowledge information according to at least one of the character setting information, at least one piece of initial object information, the live broadcast material information, the live broadcast tool information, and the knowledge base. As Figure 3 shown, one or more of the character setting information, at least one piece of initial object information, the live broadcast material information, and the live broadcast tool information can be used to perform in-depth thinking and knowledge enhancement operation S311 to determine the target knowledge information based on the knowledge base. Next, perform operation S312. According to at least one target knowledge information, at least one first script segment can be generated. Operations S311 and S312 will be described in sequence below.
[0071] In some embodiments, determining at least one target knowledge information based on at least one of the character setting information, at least one initial object information, live broadcast material information, and live broadcast tool information, and the knowledge base includes: determining live broadcast content planning information based on at least one of the character setting information, at least one initial object information, live broadcast material information, and live broadcast tool information. The live broadcast content planning information may indicate that the first script segment includes at least one of the introduction text of the object to be displayed, the historical case text of the object to be displayed, and the guiding text of the object to be displayed. For example, the live broadcast content planning information may indicate the object to be displayed for each of the multiple first script segments. One or more first script segments may be generated for each object to be displayed. The live broadcast content planning information may also indicate that the first script segment for the object to be displayed includes the introduction text, the historical case text, and the guiding text. Alternatively, the live broadcast content planning information may also indicate that the three first script segments for the object to be displayed respectively include the introduction text, the historical case text, and the guiding text. The introduction text may be an explanation of the object to be displayed. The historical case text may be a case description after different users use the object to be displayed. The guiding text may be used to guide the user to perform an action on the object to be displayed. Taking some of the above first script segments as examples, the speech content text of the first script segment, such as "Once a month, continuously use for 6 months, take a break for 6 months, only 2 boxes are needed in a year. Each time you use it, mix agent A and agent B, and then use the MTS roller for micro-needling, slightly break the skin, it doesn't hurt and the repair period is short", etc. may be the introduction text. "Let's take a look at these comparison pictures of usage again, the effect is really immediate" may be the historical case text. "Do you all want to have such skin too?" may be the guiding text.
[0072] In some embodiments, determining at least one target knowledge information based on at least one of the character setting information, at least one initial object information, live broadcast material information, and live broadcast tool information, and the knowledge base includes: determining a plurality of initial retrieval terms according to the live broadcast content planning information. Determining the target knowledge information according to the plurality of initial retrieval terms and the knowledge base. As Figure 3As shown, based on the live content planning information outline30, multiple initial retrieval terms query30 can be determined. According to the multiple initial retrieval terms query30 and the knowledge base, operation S3111 can be performed to determine knowledge information. Thus, through the live content planning information, combined with the knowledge enhancement mechanism of in-depth retrieval, it is linked with multi-source knowledge bases during the script generation process. The multi-source knowledge bases can include an object details knowledge base, a real case knowledge base, and an industry encyclopedia knowledge base. The multi-source knowledge bases can be used as an external knowledge system to achieve "information-driven" script content generation. For example, taking the object to be displayed as the product to be displayed as an example, content such as product efficacy, usage method, core ingredients, and authoritative endorsements can be extracted as knowledge in the product details knowledge base so that the information of the object to be displayed can be naturally incorporated into the speech content of the virtual character. For the real case knowledge base, content such as user evaluations, usage feedback, and real conversion results can be used as knowledge in the real case knowledge base. For the industry encyclopedia database, external professional knowledge can be mobilized and reasonably extended so that, for example, a dermatological background can be introduced in the speech content related to skincare ingredients. Through the knowledge base of the present disclosure, multi-source knowledge fusion can be achieved, making the script no longer "empty praise" or "stereotyped promotion", but content full of logic, hierarchy, and professionalism. It can be understood that the multi-source knowledge bases can include preset knowledge bases or knowledge bases obtained based on online data retrieval.
[0073] In addition, multiple rounds of retrieval can also be performed based on the initial retrieval terms to obtain knowledge more accurately and comprehensively. In some embodiments, the initial retrieval terms can be used as the first-level retrieval terms, and N rounds of retrieval can be performed in the knowledge base to obtain at least one target knowledge information. N can be an integer greater than 1. According to multiple n-level retrieval terms, the n-level retrieval results can be determined in the preset database. According to the n-level retrieval results, the n-level knowledge information can be determined. According to the n-level retrieval results and multiple n-level retrieval terms, operation S3112 can be executed to adjust the multiple retrieval terms to determine multiple (n + 1)-level retrieval terms. Next, the knowledge base can be retrieved based on the (n + 1)-level retrieval terms. At least one target knowledge information can be determined according to the N-level retrieval results. n can be an integer greater than or equal to 1 and less than N. In the case where there are multiple objects to be displayed, the target knowledge information of each object to be displayed can be determined from the N-level retrieval results. Through the embodiments of the present disclosure, in the case of performing multiple retrievals, combined with the above object description sub-fragments, a fusion mechanism of structured information embedding, live broadcast content planning, and in-depth retrieval is realized, which can significantly improve the quality of the script, and can make the live broadcast script contain multi-modal collaborative content including intonation, actions, expressions, and prop instructions, enabling the digital human to have real expressiveness. In addition, through the embodiments of the present disclosure, the live broadcast content can also be made more infectious, the narration more rhythmic, the language style diverse, logical, and information-rich, effectively avoiding the problems of previous script templatization and sameness. Users are more likely to be moved by the content during the viewing process, thus generating emotional resonance.
[0074] It can be understood that the above description of the in-depth thinking and knowledge enhancement operations of the present disclosure has been given, and the operation S312 will be described below.
[0075] As Figure 3 shown, in operation S312, at least one first script segment is generated according to at least one target knowledge information. For example, a large model can be used to generate a first script segment according to the target knowledge information for one object to be displayed.
[0076] Next, operation S320 can be executed to determine the live broadcast script according to at least one first script segment.
[0077] In some embodiments, it can be determined whether the difference between the first script segment and the character setting information is less than a preset difference threshold. For example, the style characteristics of the virtual character can be extracted from the first script segment. Determine the difference between the style characteristics and the characteristics of the character setting information.
[0078] In some embodiments, in response to determining that the difference between the first script segment and the character setting information is less than a preset difference threshold, the first script segment is determined as a live script segment of the live script. For example, if the difference between the virtual character style feature of the first script segment and the character setting information for the virtual character is small, the first script segment can be used as a live script segment.
[0079] In other embodiments, in response to determining that the difference between the first script segment and the character setting information is greater than or equal to a preset difference threshold, the first adjusted script segment is determined as a live script segment of the live script. The first adjusted script segment is obtained by adjusting the first script segment. For example, in the case where the difference between the style characteristics of the virtual character of the first script segment and the character setting information for the virtual character is large, the style of the first script segment can be adjusted using a large model to reduce the difference between the style characteristics and the character setting information.
[0080] It can be understood that the "personality" of the anchor and the assistant in the live broadcast is the key information to attract users, form memory points and build trust. Therefore, in order to maintain the consistency of the character's language style and behavior logic and avoid the problem of "drift" of the character in the long-term script generation process, the present disclosure introduces character setting information to realize the labeling construction of the anchor role before the script is generated (such as setting the personality characteristics, language style, identity background, expression habits, etc. of the virtual character), and then dynamically constrains the difference between the script fragment and the character setting information, so that the language style and behavior logic of the virtual character are always consistent with the character setting. By determining the difference between the first script fragment and the character setting information, the coherence of the character expression can be continuously monitored during the script generation process, and potential character deviations can be automatically identified and corrected to avoid problems such as inconsistent tone, conflict of values, and logical reversal. For example, for a "rational" virtual anchor role, the script fragment used for the role should include speech content based on data, analysis, and logical deduction to support the point of view. For another example, a "friendly" virtual anchor character is more likely to use intimate tones such as "babies" and "sisters, look here" to create a sense of closeness. Through the embodiments of the present disclosure, the full-cycle management of personality settings from "language style shaping" to "behavior consistency control" can be achieved, effectively solving the problem of personality breaks and collapse. In addition, in order to achieve the generation of longer scripts, the present disclosure also introduces a long script segmentation writing workflow, which combines the live content planning information to generate one or more first script segments for each object to be displayed, and can generate a script of 5,000 to 10,000 words at one time to meet the needs of long live broadcasts of 15 to 30 minutes.
[0081] It can be understood that the above describes the method of generating a static live broadcast script before the live broadcast, and the following will describe the generation of a dynamic script during the live broadcast.
[0082] Figure 4 It is a flowchart of a live script generation method according to another embodiment of the present disclosure.
[0083] Method 401 can be executed after the above method 200, and will be described below in combination with operation S431, operation S432, operation S441, and operation S442.
[0084] In operation S431, it is determined whether a task trigger signal is received.
[0085] In some embodiments, in response to determining that the target behavior does not meet the trigger logic, operation S431 is continued to be executed.
[0086] In some embodiments, in response to determining that the target behavior in the live broadcast room meets the trigger logic, operation S432 is executed. For example, after generating a live script and generating a live video using a large model, multiple first video segments can be played in sequence for live broadcast. During the playback of the first video segment, various behavior signals in the live broadcast room can be monitored. The various behaviors corresponding to the various behavior signals can include user entry, exit, like, comment, stay time, product click, etc. In addition, questions or interaction requests raised by users can also be obtained. When one or more of these behaviors meet a certain trigger logic, a task trigger signal can be generated. The trigger logic can be that the occurrence quantity of a behavior is greater than or equal to a preset quantity value. Taking the product click behavior as an example, in the case where the click volume of a certain product page rises rapidly, if the click volume of the product page is greater than or equal to the preset click volume threshold, a signal can be triggered as the task trigger signal for this first video segment. The large model used to generate the script can receive this signal and execute operation S432.
[0087] In operation S432, according to the display status information of the first video segment, the task type for the task trigger signal is determined.
[0088] In some embodiments, the display status information may include at least one of the playback progress information of the first video segment, the task execution progress information for the first video segment, and the status information of the virtual character. The first video segment can be further divided to obtain multiple video sub-segments. The playback progress information can indicate the playback status of each of the multiple video sub-segments. The multiple video sub-segments respectively correspond to multiple importance metric values. The task execution progress information can indicate the execution progress of the previous task for the first video segment to indicate whether the previous task has been executed and completed. The status information can indicate whether the virtual character is in the video frame. In response to determining that the target video sub-segment has been played and the previous task has been executed and completed, determine the task type for the task trigger logic. The target video sub-segment is a video sub-segment whose importance metric value is greater than or equal to a preset importance metric threshold. For example, if the target video sub-segment related to the key content has been played, the previous task has been executed and completed, and the behavior related to the previous task is different from the behavior for generating the task trigger signal, the task type for the task trigger signal can be determined. The task type can be multiple. The multiple task types include invitation for evaluation, user Q&A, user behavior feedback, etc. In addition, for another example, a task queue can be set up, the task to be executed is added to the task queue, and the priority of the task to be executed is determined. The tasks in the task queue can be sorted based on the priority so as to maintain the coherence of the live broadcast rhythm and experience under multi-task competition. It can be understood that the large model can determine the priority of the task, or the priorities of different types of tasks can be preset.
[0089] In operation S441, determine a target addition position among at least one preset addition position of the first video segment.
[0090] In some embodiments, the first video segment may include at least one preset addition position. The preset addition position can correspond to an addable moment of the first video segment. Adding a new video segment based on the preset addition position has less impact on the coherence of the script content of the first video segment.
[0091] In some embodiments, according to multiple addable moments and the task trigger moment when the task trigger signal is received, determine multiple time offset values. According to the multiple time offset values, determine the target addition position from the multiple preset addition positions. Generate a second script segment for the task trigger signal according to the task type and the target addition position. For example, the preset addition position corresponding to the smallest time offset value can be used as the target addition position.
[0092] In operation S442, generate a second script segment for the task trigger signal according to the task type and the target addition position among at least one preset addition position.
[0093] In some embodiments, the second script segment includes at least one of task content text and content connection text, and the content connection text is generated based on the context data of the target addition position. The second script segment is used to generate a second video segment to be added to the target addition position. For example, a large model can generate task content text according to the task type. Taking the task type being "user Q&A" as an example, answer text can be generated based on the questions provided by the user as the task content text. For another example, according to the target addition position, context text for the target addition position can be obtained from the first script segment. Next, based on the context text, content connection text can be generated. Thus, the style of the second script segment can be further made consistent with the task setting information.
[0094] In some embodiments, the second script segment is used to generate a second video segment to be added to the target addition position. For example, in response to determining that the difference between the second script segment and the character setting information is less than a preset difference threshold, a second video segment is generated according to the second script segment. In response to determining that the difference between the second script segment and the character setting information is greater than or equal to the preset difference threshold, a second adjusted script segment is generated according to the second adjusted script segment. The second adjusted script segment is obtained by adjusting the second script segment. It can be understood that the above description regarding the first script segment and the preset difference threshold also applies to the second script segment, and the present disclosure will not elaborate here.
[0095] Through the embodiments of the present disclosure, the digital human can conduct stable live broadcasts according to the established first script segment, and can also achieve intelligent responses based on the signals during the live broadcast, thereby realizing a flexible, natural, and content-rich live broadcast interaction experience, which can improve user stickiness and live broadcast conversion rate.
[0096] It can be understood that the method of the present disclosure has been described above, and the apparatus of the present disclosure will be described below.
[0097] Figure 5 It is a block diagram of a live script generation apparatus according to an embodiment of the present disclosure.
[0098] As Figure 5 shown, the apparatus 500 may include a first generation module 510 and a first determination module 520.
[0099] The first generation module 510 is configured to generate at least one first script segment according to the initial input information. The first script segment includes speech content text and an object description sub-segment for the live object, and the object description sub-segment includes object description text, which is used to describe at least one of the actions shown by the live object and the display mode for the speech content text.
[0100] The first determination module 520 is configured to determine a live script according to at least one first script segment.
[0101] In some embodiments, there is at least one live object, and the at least one live object includes at least one of a virtual character and an object to be displayed. The initial input information includes at least one of character setting information for the virtual character, initial object information for the object to be displayed, live material information, and live tool information. The live script is used to generate a live video, and the live video includes a first video segment corresponding to the first script segment.
[0102] In some embodiments, the object description sub-segment is embedded in the speech content text. The object description sub-segment further includes at least one of a display identifier, a sub-segment start symbol, a sub-segment end symbol, and a separator, and the separator is located between the display identifier and the object description text.
[0103] In some embodiments, the object description sub-segment is one of an object display mode description sub-segment, an object action description sub-segment, and an object expression description sub-segment. The object display mode description sub-segment includes object display mode description text, the object action description sub-segment includes object action description text, and the object expression description sub-segment includes object expression description text. The audio data for the speech content text in the first video segment is generated according to the display mode described in the object display mode description text. The image data for the virtual character in the first video segment is generated according to at least one of the object action description text and the object expression description text. The object action description text is used to describe at least one limb movement of the virtual character, and the at least one limb movement includes at least one first limb movement for the object to be displayed. The object expression description text is used to describe at least one facial movement of the virtual character.
[0104] In some embodiments, the first generation module includes: a first determination sub-module, configured to determine at least one target knowledge information according to at least one of the character setting information, at least one initial object information, live material information, and live tool information, and a knowledge base. A generation sub-module, configured to generate at least one first script segment according to the at least one target knowledge information.
[0105] In some embodiments, the first determination sub-module includes: a first determination unit, configured to determine live content planning information according to at least one of the character setting information, the initial object information, the live material information, and the live tool information. The live content planning information is used to indicate that the first script segment includes at least one of the introduction text of the object to be displayed, the historical case text of the object to be displayed, and the guiding text of the object to be displayed. A second determination unit, configured to determine a plurality of initial search terms according to the live content planning information. A third determination unit, configured to determine at least one target knowledge information according to the plurality of initial search terms and the knowledge base.
[0106] In some embodiments, the third determination unit is further configured to perform N rounds of retrieval in the knowledge base according to a plurality of first-level retrieval terms to obtain at least one piece of target knowledge information. N is an integer greater than or equal to 1, and the first-level retrieval terms are determined according to the initial retrieval terms.
[0107] In some embodiments, the third determination unit includes: a first determination subunit, configured to determine a first-level retrieval result in a preset database according to a plurality of n-level retrieval terms; a second determination subunit, configured to determine a plurality of (n + 1)-level retrieval terms according to the first-level retrieval result and the plurality of n-level retrieval terms. The at least one piece of target knowledge information is determined according to the N-level retrieval result, where n is an integer greater than or equal to 1 and less than N.
[0108] In some embodiments, the first determination module includes: a second determination sub-module, configured to determine the first script segment as the live script segment of the live script in response to determining that the difference between the first script segment and the character setting information is less than a preset difference threshold; a third determination sub-module, configured to determine the first adjusted script segment as the live script segment of the live script in response to determining that the difference between the first script segment and the character setting information is greater than or equal to the preset difference threshold. The first adjusted script segment is obtained by adjusting the first script segment.
[0109] In some embodiments, the first video segment includes at least one preset addition position. The apparatus further includes: a second determination module, configured to determine the task type for the task trigger signal according to the display status information of the first video segment in response to receiving the task trigger signal for the first video segment; a second generation module, configured to generate a second script segment for the task trigger signal according to the task type and the target addition position among the at least one preset addition position. The second script segment includes at least one of the task content text and the content connection text, and the content connection text is generated according to the context data of the target addition position. The second script segment is used to generate a second video segment added to the target addition position.
[0110] In some embodiments, the second script segment is used to indicate that the second video segment added to the target addition position is generated by: in response to determining that the difference between the second script segment and the character setting information is less than a preset difference threshold, generating a second video segment according to the second script segment; in response to determining that the difference between the second script segment and the character setting information is greater than or equal to the preset difference threshold, generating a second video segment according to the second adjusted script segment. The second adjusted script segment is obtained by adjusting the second script segment.
[0111] In some embodiments, the display status information includes at least one of the playback progress information of the first video segment and the task execution progress information for the first video segment. The first video segment includes a plurality of video sub-segments, and the plurality of video sub-segments respectively correspond to a plurality of importance metric values. The playback progress information is used to indicate the playback status of each of the plurality of video sub-segments, and the task execution progress information is used to indicate the execution progress of the previous task for the first video segment. The previous task is the task before receiving the task trigger signal. The second determination module includes: a fourth determination sub-module, configured to determine the task type for the task trigger logic in response to determining that the target video sub-segment has finished playing and the previous task has been executed. The target video sub-segment is a video sub-segment whose importance metric value is greater than or equal to a preset importance metric threshold.
[0112] In some embodiments, there are a plurality of preset addition positions, and the plurality of preset addition positions respectively correspond to a plurality of addable moments of the first video segment. The second generation module includes: a fifth determination module, configured to determine a plurality of time offset values according to the plurality of addable moments and the task trigger moment when receiving the task trigger signal. A sixth determination sub-module, configured to determine a target addition position from the plurality of preset addition positions according to the plurality of time offset values. A second generation sub-module, configured to generate a second script segment for the task trigger signal according to the task type and the target addition position.
[0113] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0114] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0115] Figure 6 FIG. shows a schematic block diagram of an exemplary electronic device 600 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0116] As Figure 6As shown, device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to computer programs stored in a read-only memory (ROM) 602 or computer programs loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0117] Multiple components in device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a disk, an optical disc, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0118] The computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 601 executes the various methods and processes described above, such as the live script generation method. For example, in some embodiments, the live script generation method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the live script generation method described above can be executed. Alternatively, in other embodiments, the computing unit 601 can be configured to execute the live script generation method by any other appropriate means (e.g., by means of firmware).
[0119] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), system on chip (SOC) systems, complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0120] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0121] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM) or flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0122] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) monitor or a liquid crystal display (LCD)) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0123] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0124] A computer system may include a client and a server. The client and the server are generally remote from each other and typically interact via a communication network. The relationship between the client and the server is generated by computer programs that run on respective computers and have a client-server relationship with each other.
[0125] It should be understood that the various forms of the processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.
[0126] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A method for generating a live script, comprising: Generating at least one first script segment according to initial input information, wherein the first script segment includes speech content text and an object description sub-segment for a live object, and the object description sub-segment includes object description text for describing at least one of the actions demonstrated by the live object and the display mode for the speech content text; and Determining a live script according to at least one of the first script segments.
2. The method according to claim 1, wherein, There is at least one live object, and at least one of the live objects includes at least one of a virtual character and an object to be demonstrated. The initial input information includes at least one of character setting information for the virtual character, initial object information for the object to be demonstrated, live material information, and live tool information. The live script is used to generate a live video, and the live video includes a first video segment corresponding to the first script segment.
3. The method according to claim 1, wherein, The object description sub-segment is embedded in the speech content text. The object description sub-segment further includes at least one of a display identifier, a sub-segment start symbol, a sub-segment end symbol, and a separator, and the separator is located between the display identifier and the object description text.
4. The method according to claim 3, wherein, The object description sub-segment is one of an object display mode description sub-segment, an object action description sub-segment, and an object expression description sub-segment. The object display mode description sub-segment includes object display mode description text, the object action description sub-segment includes object action description text, and the object expression description sub-segment includes object expression description text. The audio data for the speech content text in the first video segment is generated according to the display mode described by the object display mode description text. The image data for the virtual character in the first video segment is generated according to at least one of the object action description text and the object expression description text. The object action description text is used to describe at least one limb movement of the virtual character, and at least one of the limb movements includes at least one first limb movement for the object to be demonstrated. The object expression description text is used to describe at least one facial movement of the virtual character.
5. The method according to claim 2, wherein The generating at least one first script segment according to the initial input information includes: Determining at least one target knowledge information according to at least one of the character setting information, at least one of the initial object information, the live material information, the live tool information, and a knowledge base; and Generating at least one of the first script segments according to at least one of the target knowledge information.
6. The method according to claim 5, wherein, The determining at least one target knowledge information according to at least one of the character setting information, at least one of the initial object information, the live material information, the live tool information, and a knowledge base includes: Determine live broadcast content planning information according to at least one of the described character setting information, at least one of the described initial object information, the live broadcast material information, and the live broadcast tool information, wherein the live broadcast content planning information is used to indicate that the first script segment includes at least one of the introduction text of the object to be displayed, the historical case text of the object to be displayed, and the guiding text of the object to be displayed; Determine a plurality of initial search terms according to the live broadcast content planning information; and Determine at least one of the target knowledge information according to the plurality of initial search terms and the knowledge base.
7. The method according to claim 6, wherein The determining at least one of the target knowledge information according to the plurality of initial search terms and the knowledge base includes: Perform N rounds of searches in the knowledge base according to a plurality of first-level search terms to obtain at least one of the target knowledge information, where N is an integer greater than or equal to 1, and the first-level search terms are determined according to the initial search terms.
8. The method according to claim 7, wherein The performing N rounds of searches in the knowledge base includes: Determine the nth-level search results in the preset database according to a plurality of nth-level search terms; Determine a plurality of (n + 1)th-level search terms according to the nth-level search results and a plurality of the nth-level search terms, wherein at least one of the target knowledge information is determined according to the Nth-level search results, and n is an integer greater than or equal to 1 and less than N.
9. The method according to claim 2, wherein The determining the live broadcast script according to at least one of the first script segments includes: In response to determining that the difference between the first script segment and the character setting information is less than the preset difference threshold, determine the first script segment as the live broadcast script segment of the live broadcast script; In response to determining that the difference between the first script segment and the character setting information is greater than or equal to the preset difference threshold, determine the first adjusted script segment as the live broadcast script segment of the live broadcast script, wherein the first adjusted script segment is obtained by adjusting the first script segment.
10. The method according to claim 2, wherein The first video segment includes at least one preset addition position, Further includes: In response to receiving a task trigger signal for the first video segment, determine the task type for the task trigger signal according to the display status information of the first video segment; Generate a second script segment for the task trigger signal according to the task type and the target addition position among at least one preset addition position, the second script segment includes at least one of the task content text and the content connection text, the content connection text is generated according to the context data of the target addition position, and the second script segment is used to generate a second video segment added to the target addition position.
11. The method according to claim 10, wherein, The second script segment is used to indicate generating a second video segment added to the target addition position by the following operations: In response to determining that the difference between the second script segment and the character setting information is less than the preset difference threshold, generate the second video segment according to the second script segment; In response to determining that the difference between the second script segment and the character setting information is greater than or equal to a preset difference threshold, a second video segment is generated according to the second adjusted script segment, where the second adjusted script segment is obtained by adjusting the second script segment.
12. According to the method of claim 10, wherein The display status information includes at least one of the playback progress information of the first video segment and the task execution progress information for the first video segment. The first video segment includes multiple video sub-segments, and the multiple video sub-segments respectively correspond to multiple importance index values. The playback progress information is used to indicate the playback status of each of the multiple video sub-segments, and the task execution progress information is used to indicate the execution progress of the previous task for the first video segment. The previous task is the task before receiving the task trigger signal. Determining the task type for the task trigger signal according to the display status information of the first video segment includes: In response to determining that the target video sub-segment has been played and the previous task has been executed, determine the task type for the task trigger logic, where the target video sub-segment is a video sub-segment with an importance index value greater than or equal to a preset importance index threshold.
13. The method according to claim 10, wherein, There are multiple preset addition positions, and the multiple preset addition positions respectively correspond to multiple addable moments of the first video segment. Generating the second script segment for the task trigger signal according to the task type and the target addition position in at least one preset addition position includes: Determine multiple time offset values according to the multiple addable moments and the task trigger moment when receiving the task trigger signal; Determine the target addition position from the multiple preset addition positions according to the multiple time offset values; and Generate a second script segment for the task trigger signal according to the task type and the target addition position.
14. A live script generation device, comprising: A first generation module, configured to generate at least one first script segment according to initial input information, where the first script segment includes a speech content text and an object description sub-segment for a live object, and the object description sub-segment includes an object description text, and the object description text is used to describe at least one of the actions shown by the live object and the display mode for the speech content text; A first determination module, configured to determine a live script according to at least one of the first script segments.
15. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to any one of claims 1 to 13.
16. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 13.
17. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 13.