Generation method of digital human broadcast video, electronic equipment, medium and program product
By performing semantic analysis and constructing logical relationships on the content information of digital human broadcast videos, and combining them with a multimodal large language model, digital human broadcast videos are automatically generated, solving the problem of low efficiency in existing technologies and achieving efficient and accurate video creation.
Patent Information
- Application Number
- CN202511085219.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2025-11-11
AI Technical Summary
Currently, generating digital human broadcast videos requires users to manually input content and adjust the video scene, which is inefficient and difficult to accurately adapt, resulting in insufficient creation efficiency and practicality.
By acquiring the content information of the target digital human to be broadcast, semantic analysis is performed to determine the timeline and logical relationships, a storyboard content description is constructed, and a video is generated based on user instructions. Key information and shot intent are extracted using a multimodal large language model to achieve automated video generation.
It enables the rapid and accurate generation of digital human narration videos, improving creation efficiency and practicality, and meeting users' needs for intelligent generation of storyboards.
Smart Images

Figure CN120935429A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, electronic device, medium, and program product for generating digital human broadcast videos. Background Technology
[0002] With the rapid development of computer and artificial intelligence technologies, virtual characters such as digital humans have been widely used in various fields. Currently, when generating videos for digital humans to broadcast, users need to manually input the video content and adjust the video scene, which is inefficient and difficult to accurately adapt to the broadcast content.
[0003] How to quickly and accurately generate digital human broadcasting videos, and improve the efficiency and practicality of digital human video creation, is a key research issue in the industry. Summary of the Invention
[0004] This invention provides a method, electronic device, medium, and program product for generating digital human broadcast videos, so as to generate digital human broadcast videos quickly and accurately, thereby improving the efficiency and practicality of digital human video creation.
[0005] According to one aspect of the present invention, a method for generating digital human broadcasting video is provided, the method comprising:
[0006] Obtain the content information of the target digital human's to be broadcast;
[0007] Semantic analysis is performed on each of the content information, and the timeline and logical relationships of the target video to be generated are determined based on the semantic analysis results.
[0008] Determine the content description corresponding to each storyboard, and display the content description of each storyboard based on the timeline and logical relationship; the content description includes at least one of the following: image description, text content, and shot type;
[0009] In response to the target user's processing instructions for each display result, a target video containing the target digital human is generated.
[0010] In an optional implementation of this embodiment, obtaining the content information of the target digital human to be broadcast includes:
[0011] Obtain the target text information of the content to be broadcast, and perform semantic understanding on the target text information;
[0012] The content information of the content to be broadcast is determined based on the semantic understanding results;
[0013] Alternatively, it can receive content information entered by the target user on the target interface;
[0014] The content information includes at least one of the following: topic type, target audience, content requirements, and applicable scenarios.
[0015] In an optional implementation of this embodiment, the step of performing semantic analysis on each of the content information and determining the timeline and logical relationships of the target video to be generated based on the semantic analysis results includes:
[0016] Each of the aforementioned content information is input into a pre-tuned multimodal large language model, and key information of each of the aforementioned content information is extracted based on the multimodal large language model; wherein, each of the aforementioned key information is structured information;
[0017] Based on the key information, the intent of each shot is inferred, and the timeline and logical relationship are constructed according to the intent of each shot.
[0018] In an optional implementation of this embodiment, the step of inferring the intent of each shot based on the key information and constructing the timeline and the logical relationship according to each shot intent includes:
[0019] Based on the intent of each shot, the text information of the content to be broadcast by the target digital human is mapped into narrative logic; the narrative logic includes multiple paragraphs, and each paragraph corresponds to a scene.
[0020] The timeline is obtained by assigning each scene to a corresponding time period based on the total video duration.
[0021] In an optional implementation of this embodiment, determining the content description corresponding to each scene and displaying the content description of each scene based on the timeline and logical relationships includes:
[0022] Determine the key information and time period of the target corresponding to the target storyboard;
[0023] The key information of the target and the target time period are converted into content descriptions corresponding to the target storyboard.
[0024] The scenes are sorted according to the timeline, and the content descriptions of each scene are rendered and displayed based on the sorting results.
[0025] In an optional implementation of this embodiment, after displaying the content description of each scene based on the timeline and logical relationships, the method further includes:
[0026] Receive the target user's processing instructions for the displayed results, and process each of the storyboards based on the processing instructions;
[0027] The processing instructions include: editing a storyboard, deleting a storyboard, or adding a storyboard.
[0028] In an optional implementation of this embodiment, generating a target video containing the target digital human in response to the target user's processing instructions for each display result includes:
[0029] Once it is determined that the target user has completed processing of each of the aforementioned scenes, the processed target scenes are received, and the target scenes are spliced together according to the time period of each target scene to obtain the target video.
[0030] According to another aspect of the present invention, an apparatus for generating digital human broadcast videos is provided, the apparatus comprising:
[0031] The content information acquisition module is used to acquire the content information of the target digital human to be broadcast;
[0032] The first determining module is used to perform semantic analysis on each of the content information and determine the timeline and logical relationship of the target video to be generated based on the semantic analysis results.
[0033] The second determining module is used to determine the content description corresponding to each scene, and to display the content description of each scene based on the timeline and logical relationship; the content description includes at least one of the following: scene description, text content, and shot type;
[0034] The target video generation module is used to generate a target video containing the target digital human in response to the target user's processing instructions for each display result.
[0035] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0036] At least one processor; and
[0037] A memory communicatively connected to the at least one processor; wherein,
[0038] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the digital human broadcasting video generation method according to any embodiment of the present invention.
[0039] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the digital human broadcasting video generation method according to any embodiment of the present invention.
[0040] According to another aspect of the present invention, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method for generating digital human broadcast video according to any embodiment of the present invention.
[0041] The solution of this invention involves obtaining content information of the target digital human to be broadcast; performing semantic analysis on each piece of content information; determining the timeline and logical relationship of the target video to be generated based on the semantic analysis results; determining the content description corresponding to each scene; and displaying the content description of each scene based on the timeline and logical relationship; the content description includes at least one of the following: image description, text content, and shot type; and generating a target video containing the target digital human in response to the target user's processing instructions for each display result. This method can quickly and accurately generate digital human broadcast videos, improving the efficiency and practicality of digital human video creation.
[0042] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0043] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0044] Figure 1 This is a flowchart of a method for generating a digital human broadcasting video according to Embodiment 1 of the present invention;
[0045] Figure 2 This is a flowchart of a method for generating a digital human broadcasting video according to Embodiment 2 of the present invention;
[0046] Figure 3 This is a schematic diagram of the structure of a digital human broadcasting video generation device according to Embodiment 3 of the present invention;
[0047] Figure 4 This is a schematic diagram of the structure of an electronic device that implements the digital human broadcasting video generation method of this invention. Detailed Implementation
[0048] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0049] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0050] Example 1
[0051] Figure 1 This is a flowchart of a method for generating a digital human broadcasting video according to Embodiment 1 of the present invention. This embodiment is applicable to the generation of digital human broadcasting videos. The method can be executed by a digital human broadcasting video generation device, which can be implemented in hardware and / or software and can be configured in electronic devices such as computers, servers, or tablet computers. Figure 1 As shown, the method includes:
[0052] Step 110: Obtain the content information of the target digital human to be broadcast.
[0053] The target digital person can be any virtual character, such as a cartoon character, a female host character, a male host character, or a child character, etc., and this embodiment does not limit it.
[0054] In this embodiment, the content information of the content to be broadcast may include at least one of the following: topic type, target audience, content requirements, and applicable scenarios.
[0055] Optionally, in this embodiment, the content to be broadcast by the target digital human can be news, articles, or product promotions, etc., and is not limited thereto in this embodiment; in this embodiment, when it is determined that the target digital human needs to broadcast relevant content, the content information of the content can be further obtained, for example, the topic type, target audience, content requirements and applicable scenarios of the content can be obtained at the same time.
[0056] Optionally, in this embodiment, obtaining the content information of the target digital human to be broadcast may include: obtaining the target text information of the content to be broadcast, performing semantic understanding on the target text information; determining the content information of the content to be broadcast based on the semantic understanding result; or, receiving the content information input by the target user on the target interface;
[0057] In an optional implementation of this embodiment, after determining that the target digital human will broadcast the content to be broadcast, the text information of the content to be broadcast can be further obtained, which is referred to as target text information in this embodiment; for example, if the content to be broadcast is news, then the target text information can be a press release written by a relevant reporter.
[0058] Furthermore, semantic understanding can be performed on the acquired text information to determine the topic type, target audience, content requirements, and applicable scenarios of the content to be broadcast based on the semantic understanding results. For example, the target text information can be input into a large semantic understanding model, and the content information of the content to be broadcast can be output through the large semantic understanding model.
[0059] In another optional implementation of this embodiment, the user can directly input relevant information such as the topic type, target audience, content requirements, and applicable scenarios of the content to be broadcast on the target interface (e.g., display interface or management interface).
[0060] Step 120: Perform semantic analysis on each of the content information, and determine the timeline and logical relationships of the target video to be generated based on the semantic analysis results.
[0061] Optionally, in this embodiment, after obtaining the content information of the content to be broadcast, semantic analysis can be further performed on the obtained content information, and the timeline and logical relationship of the target video to be generated can be determined based on the semantic analysis results; wherein, the target video to be generated is the digital human broadcast video that matches the content to be broadcast.
[0062] It is understood that in this embodiment, the timeline of the target video is the time arrangement of each shot (storyboard) or the total duration of the target video; for example, the time period of the first shot is (0:00-0:08), the time period of the second shot is (0:09-0:15), and the time period of the third shot is (0:16-0:22), where the total duration of the target video is 22 seconds.
[0063] The logical relationship of the target video, that is, the relationship between different scenes, and how to connect them to form a coherent storyline.
[0064] Optionally, in this embodiment, the step of performing semantic analysis on each of the content information and determining the timeline and logical relationship of the target video to be generated based on the semantic analysis results may include: inputting each of the content information into a pre-tuned multimodal large language model, extracting key information of each of the content information based on the multimodal large language model; wherein, each of the key information is structured information; inferring the intent of each shot based on the key information, and constructing the timeline and logical relationship based on the intent of each shot.
[0065] In one optional implementation of this embodiment, after obtaining the topic type, target audience, content requirements, and applicable scenarios of the content to be broadcast, the obtained topic type, target audience, content requirements, and applicable scenarios can be input into a pre-tuned multimodal large language model. The input information is processed based on the multimodal large language model to obtain various structured key information. Furthermore, the intent of each shot can be determined based on the structured key information, and the timeline and logical relationships of the target video can be constructed.
[0066] It is understood that shot intent refers to the purpose and function of each shot in the video narrative, which determines how each shot serves the overall storytelling. In this embodiment, shot intent not only involves the design of visual content, but also includes emotional communication, information delivery, and interaction with the audience. For example, shot intent can be divided into displaying background or environment, introducing characters or objects, highlighting key actions or events, expressing emotions or feelings, driving the plot forward, or providing information or explanations, etc., which are not limited in this embodiment.
[0067] In an optional implementation of this embodiment, the function and purpose of each shot can be determined through structured key information; for example, if the following structured information is obtained through the above steps: {
[0068] Subject Type: "Product Promotion";
[0069] Core Functions: ["Customer Data Integration", "Automated Follow-up", "Sales Funnel Analysis"];
[0070] User pain points: ["Information fragmentation", "Low follow-up efficiency", "Data untraceability"];
[0071] "Solutions": ["Unified Customer Database", "AI-Powered Automatic Reminders", "Visualized Reports"];
[0072] "Role":[
[0073] {"Name":"Marketing Manager","Description":"Responsible for managing the team and developing strategies"};
[0074] {"Name":"Salesperson","Description":"Uses a CRM system for daily customer management"} ];
[0076] "dialogue":[
[0077] {"Character":"Marketing Manager","Dialogue":"The marketing department faces daily challenges of information chaos and inefficient follow-up."};
[0078] {"Role":"Salesperson","Dialogue":"With the intelligent CRM system, everything becomes simple."} ];
[0080] "Scene":[
[0081] {"Name":"Office","Description":"Modern office environment with multi-window switching"};
[0082] {"Name":"CRM System Interface","Description":"A concise and technologically advanced user interface"};
[0083] {"Name":"Data Dashboard","Description":"Visualized customer conversion rate increased by 47%"} ]
[0085] It can be determined that the intention of this shot is to "show the marketing manager facing multiple messy spreadsheets, emphasizing the problems of scattered information and low follow-up efficiency; show the sales staff clicking the screen to open the CRM system interface, emphasizing the simplicity and ease of use of the system; and show the data dashboard showing a data growth animation of a 47% increase in customer conversion rate, emphasizing the significant effect after using the system."
[0086] Optionally, in this embodiment, inferring the intent of each shot based on the key information and constructing the timeline and logical relationship according to the intent of each shot may include: mapping the text information of the content to be broadcast by the target digital human to narrative logic based on the intent of each shot; the narrative logic includes multiple paragraphs, and each paragraph corresponds to a scene; allocating each scene to a corresponding time period based on the total video duration to obtain the timeline.
[0087] In one optional implementation of this embodiment, after determining the intent of each shot, the text information of the content to be broadcast by the target digital human, i.e., the target text information, can be further mapped to narrative logic, that is, the target text information is mapped to multiple paragraphs, each paragraph corresponding to a storyboard.
[0088] In the example above, it can be seen that three shot intentions are determined. Based on these three shot intentions, the target text information can be mapped into three paragraphs, and each paragraph, i.e., each shot intention, corresponds to one storyboard.
[0089] Furthermore, the time for each scene can be allocated based on the total video duration. For example, if the total video duration is 22 seconds, the time periods for these three scenes can be (0:00-0:08), (0:09-0:15), and (0:16-0:22), respectively.
[0090] Step 130: Determine the content description corresponding to each scene, and display the content description of each scene based on the timeline and logical relationship.
[0091] The content description includes at least one of the following: image description, text content, and shot type.
[0092] In one optional implementation of this embodiment, after determining each scene, the scene description, text content, and shot type of each scene can be further determined, and the scene description, text content, and shot type of each scene can be displayed based on the timeline and logical relationship.
[0093] Optionally, in this embodiment, determining the content description corresponding to each scene and displaying the content description of each scene based on the timeline and logical relationship may include: determining the target key information and target time period corresponding to the target scene; converting the target key information and target time period into the content description corresponding to the target scene; sorting each scene according to the chronological order of the timeline, and rendering and displaying the content description of each scene based on the sorting result.
[0094] The target scene can be any scene from the target video, and this embodiment does not limit it.
[0095] In an optional implementation of this embodiment, after determining each scene, the target key information and target time period corresponding to the target scene can be further determined. The target key information can be the shot intention or specific content of the target scene, which is not limited in this embodiment. The target time period is the time period of the target scene in the target video. Furthermore, the target key information and target time period can be converted into a content description of the target decomposition, such as a picture description, text content and shot type of the target scene.
[0096] Furthermore, the scenes can be sorted according to the playback time sequence of each scene in the timeline of the target video, and the content description of each scene can be rendered and displayed based on the sorting results, thus obtaining the display results of different scenes based on the target digital human's broadcast.
[0097] Step 140: In response to the target user's processing instructions for each display result, generate a target video containing the target digital human.
[0098] Among them, the target user can be any user who has the right to modify the target video.
[0099] Optionally, in this embodiment, after obtaining the rendered display of each scene, the target user can process each displayed scene, such as modifying the content of a scene, adding a scene, or deleting a scene, etc. This embodiment does not limit this.
[0100] After the target user has finished processing, the final target video broadcast by the target digital human can be obtained based on the target user's modification results.
[0101] The solution in this embodiment obtains the content information of the target digital human to be broadcast; performs semantic analysis on each piece of content information, determines the timeline and logical relationship of the target video to be generated based on the semantic analysis results; determines the content description corresponding to each scene, and displays the content description of each scene based on the timeline and logical relationship; the content description includes at least one of the following: image description, text content, and shot type; responds to the target user's processing instructions for each display result, generates a target video containing the target digital human, which can quickly and accurately generate digital human broadcast videos, improving the efficiency and practicality of digital human video creation.
[0102] Example 2
[0103] Figure 2This is a flowchart of a method for generating a digital human broadcast video according to Embodiment 2 of the present invention. This embodiment is a further refinement of the above technical solution, and the technical solution in this embodiment can be combined with various optional solutions in one or more of the above embodiments. Figure 2 As shown, the method includes:
[0104] Step 210: Obtain the content information of the target digital human to be broadcast.
[0105] Step 220: Perform semantic analysis on each of the content information, and determine the timeline and logical relationships of the target video to be generated based on the semantic analysis results.
[0106] Step 230: Determine the content description corresponding to each scene, and display the content description of each scene based on the timeline and logical relationship.
[0107] Step 240: Receive the target user's processing instructions for the display results, and process each of the storyboards based on the processing instructions.
[0108] The processing instructions include: editing a storyboard, deleting a storyboard, or adding a storyboard.
[0109] Optionally, in this embodiment, after obtaining the rendered display of each scene, the system can further receive processing instructions from the target user for each displayed scene. For example, the instructions can be for editing the first scene, modifying its layout, changing the foreground and background, etc.; or for deleting the second scene; or for adding the fourth scene, etc. Furthermore, after receiving the above processing instructions, the system can further process each scene accordingly based on the processing instructions.
[0110] Step 250: After determining that the target user has completed the processing of each of the target scenes, receive the processed target scenes and splice them together according to the time period of each target scene to obtain the target video.
[0111] Optionally, in this embodiment, after determining that the target user has completed processing each segment, the processed target segments can be further received, and the target segments can be spliced together according to the time period of each target segment to obtain the target video.
[0112] In this embodiment, after obtaining the rendered display of each scene, the system can further receive processing instructions from the target user for each displayed scene, process each scene based on the processing instructions, and splice the processed scenes together to obtain the final target video, thus providing a basis for improving the accuracy and efficiency of digital human video broadcasting.
[0113] To better understand the method for generating digital human broadcast videos involved in this embodiment, a specific example is used below for illustration:
[0114] In this embodiment, the user inputs four detailed parameters: theme, target audience, content requirements, and applicable scenarios. The large model generates customized storyboards based on these parameters and matching them with corresponding materials. Specifically, the narrative structure can be determined based on the theme: a multimodal large language model is used to break down core elements, and a theme classifier is used to identify the theme type, such as product promotion, educational science popularization, or short drama films; after determining the theme type, relevant entities are expanded using a knowledge graph, for example, for the promotion of an enterprise intelligent CRM system, materials with cluttered spreadsheets or computer interfaces are used; thus, the emotional tone is determined to be technological; further... It can match audience profiles based on input parameters, such as matching Gen Z with fast-paced editing, internet memes, or vertical screens; matching enterprise clients with data visualization / technical terminology; and matching senior citizens with enlarged fonts / slow-paced narration. Furthermore, it can constrain the visuals through a multimodal rule engine. For example, if the input requirement is no real people appearing on screen, it will automatically trigger a 3D animation material library. Furthermore, it can constrain the scene format based on input parameters. For example, in a news feed scene, the pace is fast, and there must be a conflict point in the first 3 seconds. For example, in a corporate press conference scene, brand color cards can be added as constraints, and 10 seconds can be reserved for product logo finalization, etc.
[0115] In one specific embodiment, a digital human is pre-generated to promote a company's products. The user's input parameters are "Topic: Intelligent Customer Relationship Management System; Audience: Marketing Department of SMEs; Scene: Company Page". The corresponding video generation process can be as follows: Strategy layer: adopt the narrative of "data chaos → system introduction → efficiency improvement"; emphasize the visualization of return on investment data; Structure layer: [0:00-0:08] chaotic scene of the marketing department (close-up of multi-window switching); [0:09-0:15] the interface of the customer relationship management system covers the cluttered screen; [0:16-0:22] the data dashboard shows a 47% increase in customer conversion rate; Shot layer: match the "spreadsheet chaos screen" material; generate a dynamic chart of "data flow dashboard"; Validation layer: automatically add the company's logo watermark; avoid competitor software interfaces; Output: can be output in tabular form according to the sequence of scenes, including the scene, text structure, text content, shot size, shooting scene, duration, etc.
[0116] The solution of this invention can meet users' needs for intelligent generation of storyboards when generating digital human videos; provide a more convenient and intelligent way to generate video storyboards; and improve the controllability of digital human video creation.
[0117] Example 3
[0118] Figure 3This is a schematic diagram of the structure of a digital human broadcasting video generation device according to Embodiment 3 of the present invention. Figure 3 As shown, the device includes: a content information acquisition module 310, a first determination module 320, a second determination module 330, and a target video generation module 340.
[0119] Content information acquisition module 310 is used to acquire content information of the content to be broadcast by the target digital human;
[0120] The first determining module 320 is used to perform semantic analysis on each of the content information and determine the timeline and logical relationship of the target video to be generated based on the semantic analysis results.
[0121] The second determining module 330 is used to determine the content description corresponding to each scene, and to display the content description of each scene based on the timeline and logical relationship; the content description includes at least one of the following: scene description, text content, and shot type;
[0122] The target video generation module 340 is used to generate a target video containing the target digital human in response to the processing instructions of the target user for each display result.
[0123] In this embodiment, the solution involves: acquiring content information of the target digital human to be broadcast through a content information acquisition module; performing semantic analysis on each piece of content information through a first determination module, and determining the timeline and logical relationship of the target video to be generated based on the semantic analysis results; determining the content description corresponding to each scene through a second determination module, and displaying the content description of each scene based on the timeline and logical relationship; the content description includes at least one of the following: scene description, text content, and shot type; and generating a target video containing the target digital human through a target video generation module in response to the target user's processing instructions for each display result. This method can quickly and accurately generate digital human broadcast videos, improving the efficiency and practicality of digital human video creation.
[0124] In an optional implementation of this embodiment, the content information acquisition module 310 is specifically used for:
[0125] Obtain the target text information of the content to be broadcast, and perform semantic understanding on the target text information;
[0126] The content information of the content to be broadcast is determined based on the semantic understanding results;
[0127] Alternatively, it can receive content information entered by the target user on the target interface;
[0128] The content information includes at least one of the following: topic type, target audience, content requirements, and applicable scenarios.
[0129] In an optional implementation of this embodiment, the first determining module 320 is specifically used for:
[0130] Each of the aforementioned content information is input into a pre-tuned multimodal large language model, and key information of each of the aforementioned content information is extracted based on the multimodal large language model; wherein, each of the aforementioned key information is structured information;
[0131] Based on the key information, the intent of each shot is inferred, and the timeline and logical relationship are constructed according to the intent of each shot.
[0132] In an optional implementation of this embodiment, the first determining module 320 is further specifically used for:
[0133] Based on the intent of each shot, the text information of the content to be broadcast by the target digital human is mapped into narrative logic; the narrative logic includes multiple paragraphs, and each paragraph corresponds to a scene.
[0134] The timeline is obtained by assigning each scene to a corresponding time period based on the total video duration.
[0135] In an optional implementation of this embodiment, the second determining module 330 is specifically used for:
[0136] Determine the key information and time period of the target corresponding to the target storyboard;
[0137] The key information of the target and the target time period are converted into content descriptions corresponding to the target storyboard.
[0138] The scenes are sorted according to the timeline, and the content descriptions of each scene are rendered and displayed based on the sorting results.
[0139] In an optional implementation of this embodiment, the digital human broadcasting video generation device further includes: a storyboard processing module, used for:
[0140] Receive the target user's processing instructions for the displayed results, and process each of the storyboards based on the processing instructions;
[0141] The processing instructions include: editing a storyboard, deleting a storyboard, or adding a storyboard.
[0142] In an optional implementation of this embodiment, the target video generation module 340 is specifically used for:
[0143] Once it is determined that the target user has completed processing of each of the aforementioned scenes, the processed target scenes are received, and the target scenes are spliced together according to the time period of each target scene to obtain the target video.
[0144] The digital human broadcasting video generation device provided in this embodiment of the invention can execute the digital human broadcasting video generation method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0145] In the technical solutions of this invention, the collection, storage, use, processing, transmission, provision and disclosure of the content to be broadcast (such as content information) all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0146] Example 4
[0147] Figure 4 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0148] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0149] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0150] Processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, central processing unit (CPU), graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as a method for generating a digital human broadcasting video, which may include: acquiring content information of the content to be broadcast by the target digital human; performing semantic analysis on each of the content information, and determining the timeline and logical relationship of the target video to be generated based on the semantic analysis results; determining the content description corresponding to each scene, and displaying the content description of each scene based on the timeline and logical relationship; the content description includes at least one of the following: image description, text content, and shot type; and generating a target video containing the target digital human in response to the processing instructions of the target user for each display result.
[0151] In some embodiments, the method for generating digital human-broadcast videos can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the method for generating digital human-broadcast videos described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the method for generating digital human-broadcast videos by any other suitable means (e.g., by means of firmware).
[0152] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0153] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0154] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0155] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0156] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0157] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0158] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0159] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
[0160] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements a database detection method as provided in any embodiment of this application.
[0161] In implementing the computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0162] It should be noted that in the embodiments of this application, certain software, components, models and other existing solutions in the industry may be mentioned. These should be regarded as exemplary and are only intended to illustrate the feasibility of implementing the technical solution of this application. However, it does not mean that the applicant has used or necessarily used the solution.
[0163] Note that the above description is merely a preferred embodiment of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of the present invention, the scope of which is determined by the scope of the appended claims.
Claims
1. A method for generating digital human broadcasting videos, characterized in that, include: Obtain the content information of the target digital human's to be broadcast; Semantic analysis is performed on each of the content information, and the timeline and logical relationships of the target video to be generated are determined based on the semantic analysis results. Determine the content description corresponding to each storyboard, and display the content description of each storyboard based on the timeline and logical relationship; the content description includes at least one of the following: image description, text content, and shot type; In response to the target user's processing instructions for each display result, a target video containing the target digital human is generated.
2. The method for generating digital human broadcasting video according to claim 1, characterized in that, The acquisition of content information for the target digital human to be broadcast includes: Obtain the target text information of the content to be broadcast, and perform semantic understanding on the target text information; The content information of the content to be broadcast is determined based on the semantic understanding results; Alternatively, it can receive content information entered by the target user on the target interface; The content information includes at least one of the following: topic type, target audience, content requirements, and applicable scenarios.
3. The method for generating digital human broadcasting video according to claim 1, characterized in that, The step of performing semantic analysis on each of the content information, and determining the timeline and logical relationships of the target video to be generated based on the semantic analysis results, includes: Each of the aforementioned content information is input into a pre-tuned multimodal large language model, and key information of each of the aforementioned content information is extracted based on the multimodal large language model; wherein, each of the aforementioned key information is structured information; Based on the key information, the intent of each shot is inferred, and the timeline and logical relationship are constructed according to the intent of each shot.
4. The method for generating digital human broadcasting video according to claim 3, characterized in that, The step of inferring the intent of each shot based on the key information, and constructing the timeline and logical relationships based on the intent of each shot, includes: Based on the intent of each shot, the text information of the content to be broadcast by the target digital human is mapped into narrative logic; the narrative logic includes multiple paragraphs, and each paragraph corresponds to a scene. The timeline is obtained by assigning each scene to a corresponding time period based on the total video duration.
5. The method for generating digital human broadcasting video according to claim 4, characterized in that, The process of determining the content description corresponding to each scene and displaying the content description of each scene based on the timeline and logical relationships includes: Determine the key information and time period of the target corresponding to the target storyboard; The key information of the target and the target time period are converted into content descriptions corresponding to the target storyboard. The scenes are sorted according to the timeline, and the content descriptions of each scene are rendered and displayed based on the sorting results.
6. The method for generating digital human broadcast video according to claim 1, characterized in that, After presenting the content descriptions of each scene based on the timeline and logical relationships, the following is also included: Receive the target user's processing instructions for the displayed results, and process each of the storyboards based on the processing instructions; The processing instructions include: editing a storyboard, deleting a storyboard, or adding a storyboard.
7. The method for generating digital human broadcasting video according to claim 6, characterized in that, The step of generating a target video containing the target digital human in response to the target user's processing instructions for each display result includes: Once it is determined that the target user has completed processing of each of the aforementioned scenes, the processed target scenes are received, and the target scenes are spliced together according to the time period of each target scene to obtain the target video.
8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the method for generating digital human broadcast video according to any one of claims 1-7.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the method for generating a digital human broadcasting video according to any one of claims 1-7.
10. A computer program product comprising a computer program that, when executed by a processor, implements a method for generating a digital human broadcasting video according to any one of claims 1-7.
Citation Information
Cited By
Digital human audio and video processing method and device, storage medium and program product
CN122205198A