Video generation method and apparatus, storage medium, and electronic device
By acquiring a set of target video materials and combining them with video generation methods, the problem of low quality of video scripts generated in existing technologies has been solved. This enables semantic-driven video generation and precise content arrangement, improving the logic and quality of video generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TAOBAO CHINA SOFTWARE
- Filing Date
- 2026-03-13
- Publication Date
- 2026-07-21
AI Technical Summary
Existing technologies cannot effectively handle complex inputs with high noise, long context, and multimodal collaboration in real workflows, resulting in problems such as logical breaks, misuse of materials, audiovisual disconnect, and insufficient narrative appeal in the generated video scripts.
By acquiring a set of target video materials, including a first video clip related to the target object and a second video clip unrelated to it, and combining them with video generation instructions, filtering and semantic alignment are performed to generate a target video script containing visual and auditory descriptive information. This clearly distinguishes between relevant and unrelated materials, constructs a dual-channel audiovisual description, and ensures that the video generation process is driven by semantic instructions.
It improves the logic, consistency, and guidance of video scripts, and solves problems such as messy video content, deviation from the theme, and audiovisual disconnect caused by low script quality in related technologies, thereby improving the accuracy, thematic focus, and overall quality of video generation.
Smart Images

Figure CN122437981A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and more specifically, to a video generation method and apparatus, a storage medium, and an electronic device. Background Technology
[0002] With the rapid development of short video content, video creation has become a core aspect of social media operations. Currently, the mainstream video generation methods fall into two main categories: one is end-to-end video generation technology, which directly generates complete video clips based on text prompts; the other is text-based scriptwriting methods, which utilize large language models to generate storyboards or narration based on user descriptions.
[0003] End-to-end video generation models lack the ability to perceive and reuse real-world footage, failing to distinguish between usable and distracting material. They rely solely on text prompts to generate content, resulting in a disconnect between the generated results and the actual footage, lacking realism and controllability. While text-based scriptwriting methods can output structured scripts, they completely ignore the existence of visual material, failing to align the script with actual usable video clips. Therefore, existing technologies cannot effectively handle the complex inputs of high-noise, long-context, and multimodal collaboration in real-world workflows, leading to logical breaks, misuse of footage, audiovisual disconnect, and insufficient narrative appeal in the generated video scripts. Ultimately, this severely impacts the quality and application value of the generated videos.
[0004] There is currently no effective solution to the above problems. Summary of the Invention
[0005] This application provides a video generation method and apparatus, storage medium and electronic device to at least solve the technical problem in the related art where the script quality of the generated video is low, resulting in poor video generation quality.
[0006] According to one aspect of the embodiments of this application, a video generation method is provided, comprising: acquiring a target video material set and a video generation instruction, wherein the target video material set includes at least: a first video segment related to a target object, a second video segment unrelated to the target object, and text description information corresponding to the target object in the first video segment; generating a target video script based on the video generation instruction and the target video material set, wherein the target video script includes: visual description information of multiple video segments and auditory description information corresponding to the multiple video segments respectively; and generating a target video corresponding to the target object based on the target video script.
[0007] Furthermore, generating a target video script based on the video generation instructions and the target video material set includes: filtering the target video material set based on the video generation instructions and the target object to obtain multiple target video segments; generating a video storyline based on the video generation instructions and the multiple target video segments; and generating a target video script based on the video storyline and the multiple target video segments.
[0008] Furthermore, generating a target video script based on the video storyline and multiple target video segments includes: sorting the multiple target video segments according to the video storyline to obtain sorted target video segments; performing connection recognition on the sorted target video segments to obtain the position information of the video segment to be supplemented; generating visual description information of the supplementary video segment at the position information of the video segment to be supplemented based on the video storyline and multiple target video segments; and generating a target video script based on the sorted target video segments and the visual description information of the supplementary video segments.
[0009] Furthermore, generating the target video script based on the visual description information of the sorted target video segments and the supplementary video segments includes: generating corresponding auditory description information for the sorted target video segments and the supplementary video segments respectively, based on the video storyline, wherein the auditory description information includes at least: voice text information and tone style information; obtaining the identification information corresponding to the sorted target video segments; and obtaining the target video script based on the identification information, the visual description information of the supplementary video segments, and the auditory description information.
[0010] Furthermore, obtaining the target video material set includes: obtaining multiple first historical videos of the target object, and segmenting them based on the multiple first historical videos to obtain first video segments; obtaining multiple second historical videos of objects other than the target object, and segmenting them based on the multiple second historical videos to obtain second video segments; and obtaining the target video material set based on the first video segments and the second video segments.
[0011] Furthermore, the process of segmenting multiple first historical videos to obtain first video segments includes: analyzing any first historical video to obtain multiple initial video segments; scoring the multiple initial video segments to obtain scores for each initial video segment; and filtering the multiple initial video segments based on their scores to obtain the first video segment.
[0012] Furthermore, after obtaining multiple initial video segments, the method further includes: performing context-aware semantic induction on the multiple initial video segments to obtain visual description information corresponding to each of the multiple initial video segments; and obtaining video generation instructions based on the visual description information corresponding to each of the multiple initial video segments and the multiple initial video segments, wherein the video generation instructions include at least: style constraint information, target total duration, and description information of the target audience.
[0013] Furthermore, the first historical video is analyzed to obtain multiple initial video segments, including: extracting image frames from the first historical video based on a fixed frame rate to obtain multiple image frames; calculating the visual features of the multiple image frames to obtain visual feature values of objects in the multiple image frames; calculating the visual difference values between adjacent image frames in the multiple image frames based on the visual feature values; determining the scene switching points in the first historical video based on the visual difference values, and segmenting the first historical video based on the scene switching points to obtain multiple initial video segments.
[0014] Furthermore, after generating the target video script based on the video generation instructions and the target video material set, the method also includes: quantifying the target video script based on preset rule indicators to obtain a first score value; quantifying the target video script through a natural language model to obtain a second score value; and optimizing the target video script based on the first score value and the second score value.
[0015] Furthermore, the target video script is quantified based on preset rule indicators to obtain a first score value, including: calculating the number of second video segments referenced from the target video material set in the target video script to obtain the proportion of incorrectly referenced video segments; calculating the number of video segments reused in the target video script to obtain the richness of video segments; calculating the planned total duration of the target video script and the target total duration in the video generation instruction to obtain the duration deviation value; and calculating the first score value based on the proportion value, richness, and duration deviation value.
[0016] According to another aspect of the embodiments of this application, a video generation method is also provided, comprising: acquiring a target video material set and a video generation instruction uploaded by a client, wherein the target video material set includes at least: a first video segment related to a target object, a second video segment unrelated to the target object, and text description information corresponding to the target object in the first video segment; generating a target video script in a cloud server based on the video generation instruction and the target video material set, wherein the target video script includes: visual description information of multiple video segments and auditory description information corresponding to each of the multiple video segments; generating a target video corresponding to the target object based on the target video script; and returning the target video to the client.
[0017] According to another aspect of the embodiments of this application, a video generation apparatus is also provided, comprising: an acquisition unit, configured to acquire a target video material set and a video generation instruction, wherein the target video material set includes at least: a first video segment related to a target object, a second video segment unrelated to the target object, and text description information corresponding to the target object in the first video segment; a first generation unit, configured to generate a target video script based on the video generation instruction and the target video material set, wherein the target video script includes: visual description information of multiple video segments and auditory description information corresponding to the multiple video segments respectively; and a second generation unit, configured to generate a target video corresponding to the target object based on the target video script.
[0018] Furthermore, the first generation unit includes: a filtering subunit, used to filter from a set of target video materials based on video generation instructions and target objects to obtain multiple target video segments; a first generation subunit, used to generate a video storyline based on video generation instructions and the multiple target video segments; and a second generation subunit, used to generate a target video script based on the video storyline and the multiple target video segments.
[0019] Furthermore, the second generation subunit includes: a sorting module, used to sort multiple target video segments according to the video storyline to obtain sorted target video segments; a recognition module, used to perform connection recognition on the sorted target video segments to obtain the position information of the video segment to be supplemented; a first generation module, used to generate visual description information of the supplementary video segment at the position information of the video segment to be supplemented based on the video storyline and multiple target video segments; and a second generation module, used to generate a target video script based on the sorted target video segments and the visual description information of the supplementary video segments.
[0020] Furthermore, the second generation module includes: a generation submodule, used to generate corresponding auditory description information for the sorted target video segments and the supplementary video segments respectively based on the video storyline, wherein the auditory description information includes at least: voice text information and tone style information; an acquisition submodule, used to acquire the identification information corresponding to the sorted target video segments; and a first determination submodule, used to obtain the target video script based on the identification information, the visual description information of the supplementary video segments, and the auditory description information.
[0021] Furthermore, the acquisition unit includes: a first acquisition subunit, used to acquire multiple first historical videos of the target object, and segment the multiple first historical videos to obtain a first video segment; a second acquisition subunit, used to acquire multiple second historical videos of objects other than the target object, and segment the multiple second historical videos to obtain a second video segment; and a determination subunit, used to obtain a target video material set based on the first video segment and the second video segment.
[0022] Furthermore, the first acquisition subunit includes: an analysis module, used to analyze any first historical video to obtain multiple initial video segments; a scoring module, used to score the multiple initial video segments to obtain the scores corresponding to the multiple initial video segments; and a filtering module, used to filter the multiple initial video segments based on the scores to obtain the first video segment.
[0023] Furthermore, the device also includes: an induction unit, used to perform context-aware semantic induction on the multiple initial video segments after obtaining multiple initial video segments, to obtain visual description information corresponding to each of the multiple initial video segments; and a determination unit, used to obtain video generation instructions based on the visual description information corresponding to each of the multiple initial video segments and the multiple initial video segments, wherein the video generation instructions include at least: style constraint information, target total duration, and description information of the target audience.
[0024] Furthermore, the analysis module includes: an extraction submodule, used to extract image frames from the first historical video based on a fixed frame rate to obtain multiple image frames; a first calculation submodule, used to calculate the visual features of the multiple image frames to obtain the visual feature values of the objects in the multiple image frames; a second calculation submodule, used to calculate the visual difference values between adjacent image frames in the multiple image frames based on the visual feature values; and a second determination submodule, used to determine the scene switching points in the first historical video based on the visual difference values, and to segment the first historical video based on the scene switching points to obtain multiple initial video segments.
[0025] Furthermore, the device also includes: a first quantization unit, used to quantize the target video script based on preset rule indicators after generating the target video script according to the video generation instructions and the target video material set, to obtain a first score value; a second quantization unit, used to quantize the target video script through a natural language model to obtain a second score value; and an optimization unit, used to optimize the target video script based on the first score value and the second score value.
[0026] Furthermore, the first quantification unit includes: a first calculation subunit, used to calculate the number of second video segments referenced from the target video material set in the target video script, to obtain the proportion of incorrectly referenced video segments; a second calculation subunit, used to calculate the number of video segments reused in the target video script, to obtain the richness of video segments; a third calculation subunit, used to calculate the planned total duration of the target video script and the target total duration in the video generation instruction, to obtain the duration deviation value; and a fourth calculation subunit, used to calculate based on the proportion value, richness, and duration deviation value, to obtain a first score value.
[0027] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored program, wherein, when the program is running, it controls the device where the storage medium is located to execute the above-described video generation method.
[0028] According to another aspect of the embodiments of this application, an electronic device is also provided, including: a memory storing an executable program; and a processor for running the program, wherein the program executes the video generation method described above when it runs.
[0029] According to another aspect of the embodiments of this application, a computer program product is also provided, including a computer program or instructions that, when executed by a processor, implement the video generation method described above.
[0030] In this embodiment, the following steps are employed: obtaining a target video material set and a video generation instruction, wherein the target video material set includes at least: a first video segment related to the target object, a second video segment unrelated to the target object, and textual description information corresponding to the target object in the first video segment; generating a target video script based on the video generation instruction and the target video material set, wherein the target video script includes: visual description information of multiple video segments and auditory description information corresponding to each of the multiple video segments; and generating a target video corresponding to the target object based on the target video script, thereby solving the technical problem in related technologies where the quality of the generated video script is low, resulting in poor video generation quality.
[0031] In this application, a target video material set is obtained, comprising a first video segment related to the target object, a second video segment unrelated to the target object, and textual description information corresponding to the target object in the first video segment. Combined with video generation instructions, multimodal input materials are structurally filtered and semantically aligned to generate a target video script containing visual description information of multiple video segments and their corresponding auditory description information. This achieves proactive differentiation of redundant and interfering materials and precise arrangement of audiovisual content. Based on this, a target video corresponding to the target object is generated based on the structured script, ensuring that the video generation process is driven by explicit semantic instructions, rather than relying on unconstrained end-to-end splicing. By clearly distinguishing between relevant and irrelevant materials, introducing visual description information as semantic anchors, and simultaneously constructing audiovisual dual-channel description information, the logic, consistency, and guidance of the video script are improved. This effectively solves the problems of messy generated video content, thematic deviation, and audiovisual disconnect caused by low script quality in related technologies, achieving the technical effect of improving video generation accuracy, thematic focus, and overall quality. Attached Figure Description
[0032] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0033] Figure 1 This is a hardware structure block diagram of a computer terminal provided according to Embodiment 1 of this application;
[0034] Figure 2 This is a flowchart of the video generation method provided according to Embodiment 1 of this application;
[0035] Figure 3 This is a schematic diagram of the video generation method provided in Embodiment 1 of this application. Figure 1 ;
[0036] Figure 4 This is a schematic diagram of the video generation method provided in Embodiment 1 of this application. Figure 2 ;
[0037] Figure 5 This is a flowchart of the video generation method provided according to Embodiment 2 of this application;
[0038] Figure 6 This is a schematic diagram of a video generation apparatus according to Embodiment 3 of this application;
[0039] Figure 7 This is a structural block diagram of an electronic device provided according to Embodiment 4 of this application. Detailed Implementation
[0040] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0041] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0042] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0043] Example 1
[0044] According to an embodiment of this application, a video generation method is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0045] The method embodiment provided in Embodiment 1 of this application can be executed on a mobile terminal, computer terminal, or similar computing device. Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing a video generation method is shown. Figure 1As shown, the computer terminal (or mobile device) 10 may include a processor set 102 (the processor set 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA, and the processor set 102 may include a processor set, Figure 1 The data is illustrated using 102a, 102b, ..., 102n. A memory 104 is used for storing data, and a transmission module 106 is used for communication functions. In addition, it may include: a display, an input / output interface (I / O interface), a Universal Serial Bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0046] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10 (or mobile device). As involved in the embodiments of this application, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).
[0047] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the video generation method in this embodiment. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby realizing the aforementioned video generation method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0048] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.
[0049] The display may be, for example, a touchscreen LCD display that allows the user to interact with the user interface of the computer terminal 10 (or mobile device).
[0050] Under the aforementioned operating environment, this application provides the following: Figure 2 The video generation method shown. Figure 2 This is a flowchart of a video generation method according to Embodiment 1 of this application. The video generation method includes:
[0051] Step S201: Obtain a target video material set and a video generation instruction. The target video material set includes at least: a first video segment related to the target object, a second video segment unrelated to the target object, and text description information corresponding to the target object in the first video segment.
[0052] Optionally, raw videos can be collected from external data sources or local media libraries, and a target video material set can be constructed based on these raw videos. The target video material set contains two core types of materials: one type consists of first-level video clips highly relevant to the target object (such as a product, brand, or person), for example, video clips showcasing product usage scenarios, user feedback, or brand logos; the other type consists of second-level video clips not directly related to the target object. These second-level video clips are mixed in as distractors to simulate redundant or irrelevant content common in real-world creation, such as irrelevant scenes, background vignettes, or non-thematic video clips. It should be noted that video clips can also be called shot materials. Shot material refers to an indivisible semantic unit constituting video content, an independent video clip with complete visual action or state changes.
[0053] The first video clip, as reusable and valid content, has textual descriptions that are clearly associated with the material, ensuring that subsequent processing can accurately identify its semantic content; the second video clip, as noise interference, is used to simulate the situation of redundant material and irrelevant content mixed in the actual creative environment.
[0054] It should be noted that the text description information can be characteristic descriptions and product introductions of the target object in the first video clip (e.g., toiletries). Video generation instructions refer to the constraints input by the user, which may include information such as target duration, video style, and target audience.
[0055] In an optional embodiment, visual description information can be obtained through a natural language model. For example, a pre-trained multimodal large language model can be invoked, and the first video segment can be input into the model in the form of a frame sequence or keyframe sampling. Combined with temporal context information, the model can automatically generate a natural language visual description that conforms to human expression habits. This process does not rely on manual annotation, but is based on the model's semantic understanding of the content of the scene, outputting structured descriptive text including elements such as the main object, actions, scene environment, color atmosphere, and camera movement.
[0056] For example, "A woman smiles as she pours product into a glass in a sunlit kitchen, against a backdrop of a minimalist, modern countertop; the camera slowly zooms in on the product." To ensure accuracy and consistency in the description, prompt templates can be used to guide the model to focus on visually observable information, avoiding subjective speculation or the introduction of non-visual content such as audio or text. The generated descriptive text will be bound to the corresponding video clip, forming "visual-text" paired data, which will serve as the semantic basis for subsequent filtering and inference.
[0057] In an optional embodiment, an image frame is converted into a high-dimensional visual feature vector by an encoder, and the temporal relationship between frames is fused through a spatiotemporal attention mechanism to form a coherent visual context representation. Then, this visual representation is fed into the multimodal fusion layer of the model and cross-attention aligned with the instruction semantics in the text prompt (i.e., "Please describe the content of the image in natural language, based only on visible elements, without inferring sound or text") to achieve deep interaction between visual and language modalities. In the language decoding layer, the model generates a natural language description that conforms to grammatical structure, is semantically complete, and rich in detail, word by word, based on the fused context. The output is strictly limited to include only observable visual elements, such as human posture, object position, ambient lighting, and camera movement. Finally, the generated description text is standardized and cleaned by a post-processing module, including removing redundant modifiers, unifying terminology, and correcting punctuation, to ensure that the output format is uniform, the semantics are clear, and it can be directly parsed and used by downstream modules.
[0058] Step S202: Based on the video generation instructions and the target video material set, generate a target video script, wherein the target video script includes: visual description information of multiple video segments and auditory description information corresponding to each of the multiple video segments.
[0059] Optionally, a target video script is generated based on the video generation instructions and the target video material set. The target video script consists of visual descriptions of multiple video clips and corresponding auditory descriptions of those clips. It should be noted that the visual descriptions can be detailed textual descriptions of the visual content of a shot, including elements such as characters, scenes, actions, and camera movements. The video generation instructions provide target constraints for script generation, including elements such as duration, style, and target audience, while the target video material set serves as the input multimodal content foundation, containing available video clips and interfering material. Through processing the instructions and material set, the visual content description of the shot is extracted and organized, clarifying its visual elements such as composition, actions, and camera movements. Simultaneously, corresponding auditory content, including voice-over text and tone of voice, is matched to the shot, ultimately forming a structured audiovisual aligned script, achieving a direct transformation from instructions and materials to an executable creative plan.
[0060] It should be noted that the multiple video clips in the target video script include existing footage directly referenced from the target video footage set (i.e., the first video clip or the second video clip), as well as additional video clips that need to be generated (or generative footage). The additional video clips that need to be generated are represented in the target video script through visual and auditory descriptive information.
[0061] In an optional embodiment, the target video script described above can be obtained through the following steps: Relevance assessment of video clips in the target video material set is performed to identify their semantic matching degree with the instruction theme, filtering out low-relevance or interfering materials. Then, the selected effective materials are arranged chronologically to determine the connection relationship between shots and identify narrative breaks or missing links. For the identified narrative gaps, it is determined whether generative shots need to be introduced to complete the plot. Simultaneously, for retained or newly added shot nodes, corresponding visual description information (such as "the camera slowly pulls away from the side of the product, the background gradually changes to a warm-toned city night scene, and the character raises a glass and smiles") and auditory description information (such as "the narration states the core functions of the product in a light and trusting tone, with a moderate speaking speed, and a slight pause at the end to enhance memorability") are generated, ultimately resulting in the structured target video script described above.
[0062] In an optional embodiment, the multiple video segments in the target video script include valid footage and generative shot footage selected from the target video footage set. For valid footage, the ID information of the video segments in the target video footage set can be directly used as the visual description information of the valid footage. For generative shot footage, corresponding visual description information is generated based on the identified narrative gaps, such as elements including characters, scenes, actions, and camera movement.
[0063] Step S203: Based on the target video script, generate the target video corresponding to the target object.
[0064] Optionally, the shot entries in the target video script are parsed and their types are distinguished: if they are "selected material" (i.e., existing shot materials directly referenced from the target video material set), the corresponding original shot (i.e., the aforementioned video clip) is extracted from the target video material set; if they are "to be generated" (i.e., generative shot materials), the visual description information (including subject, action, environment, camera movement, etc.) and auditory description information (such as narration text, tone, and rhythm) of the shot are used as input, and a high-quality, duration-matched dynamic visual content is generated through a text-to-video generation model, ensuring that the generated image is accurately aligned with the description and synchronized with the narration in semantics and emotion. It should be noted that the auditory description information (such as narration text, tone, and rhythm) corresponding to directly referenced existing shot materials needs to be regenerated and the original auditory description information cannot be used.
[0065] Then, all processed footage and generated footage are seamlessly spliced together in the time sequence in the script, with corresponding audio overlays added synchronously. Audio-visual synchronization calibration, transition effects (such as fade-in / fade-out, cross-dissolve), and global volume balancing are then completed.
[0066] The final generated video passes through a quality verification module to check narrative coherence, audiovisual consistency, and compliance with instructions. After confirming that there are no logical breaks or style conflicts, it is output as a finished video in a standard format, realizing end-to-end automated production from structured scripts to high-fidelity, publishable complete videos.
[0067] In an optional embodiment, generating high-quality, duration-matched dynamic visual content via a text-to-video generation model can be achieved through the following steps:
[0068] Visual descriptive information is input into the text-to-video generation model. This descriptive information can include: the subject (e.g., "a young woman"), actions (e.g., "holding a product, smiling, and turning around"), the scene environment (e.g., "a modern minimalist living room with natural light streaming through floor-to-ceiling windows"), camera movement (e.g., "slowly zooming in from an over-the-shoulder shot to a close-up of the product"), lighting and atmosphere (e.g., "warm orange tones, soft shadows"), and duration constraints (e.g., "lasting 5 seconds"). These elements are formatted into structured prompts using a predefined template to ensure the model accurately understands each dimension's requirements and avoids ambiguity.
[0069] Based on the duration of the shot (e.g., 3s, 5s, 8s) and the complexity of the content, the required number of frames (e.g., 24fps corresponds to 72 frames) is automatically calculated and linked to the timeline control module of the generated model. If the description contains rhythmic changes (e.g., "slowly zooming in for the first 2 seconds, then accelerating and rotating for the next 3 seconds"), the action stages are divided using time-segmented prompts, guiding the model to generate rhythmic motion trajectories within different time intervals, ensuring that the visual rhythm is precisely aligned with the semantic pauses and emphasis points of the narration.
[0070] During the generation process, the model simultaneously receives "auditory description information" as auxiliary input (such as the narration text "This product awakens skin radiance in M days" and its tone markers "warm, confident, and slightly surprised"). Through the cross-modal alignment module, the video generation model jointly optimizes speech semantics and visual content: when the narration mentions "M days," a calendar page-turning animation appears simultaneously in the video; when the tone changes to surprise, the camera slightly tilts up and is accompanied by soft lighting enhancement.
[0071] In an optional embodiment, for generative shots, corresponding video segments can also be obtained by shooting on location based on the visual description information of the shot (including subject, action, environment, camera movement, etc.) and the auditory description information (such as narration text, tone, rhythm).
[0072] In one optional embodiment, the e-commerce platform's operators upload a set of target video materials, including product demonstration videos, user usage clips, and irrelevant scenic footage. They then input a video generation command to "generate a 30-second, lively promotional video for a beauty product aimed at young women." Based on this command and the set of materials, the system automatically filters out footage related to the beauty product, excludes irrelevant footage, and generates a target video script that includes the sequence of shots, visual descriptions of the shots (such as "close-up of a woman applying lipstick, with the background blurred into a pink halo"), and corresponding auditory descriptions (such as "lighthearted narration: a touch of brightness, awakening confidence all day long"). Finally, based on this script, the system calls a video generation model to synthesize a complete target video.
[0073] In summary, by acquiring a set of target video materials containing a first video segment related to the target object, a second video segment unrelated to the target object, and textual descriptions of the target object within the first video segment, and combining this with video generation instructions, the multimodal input materials are structurally filtered and semantically aligned. This generates a target video script containing visual descriptions of multiple video segments and their corresponding auditory descriptions, enabling proactive differentiation of redundant and interfering materials and precise arrangement of audiovisual content. Based on this, a target video corresponding to the target object is generated using the structured script, ensuring that the video generation process is driven by explicit semantic instructions, rather than relying on unconstrained end-to-end splicing. By clearly distinguishing between relevant and irrelevant materials, introducing visual descriptions as semantic anchors, and simultaneously constructing audiovisual dual-channel descriptions, the logic, consistency, and guidance of the video script are improved. This effectively solves problems in related technologies where low script quality leads to messy generated video content, thematic deviation, and audiovisual disconnect, achieving the technical effect of improving video generation accuracy, thematic focus, and overall quality.
[0074] To improve the logical consistency of the video script, in the video generation method provided in Embodiment 1 of this application, generating a target video script based on the video generation instructions and the target video material set includes: filtering the target video material set based on the video generation instructions and the target object to obtain multiple target video segments; generating a video storyline based on the video generation instructions and the multiple target video segments; and generating a target video script based on the video storyline and the multiple target video segments.
[0075] Optionally, the video generation instructions (such as "target duration 30 seconds, targeting Gen Z women, fresh and healing style, highlighting the natural ingredients of the product") can be semantically decomposed to extract key constraints, including: target audience (Gen Z women, who prefer natural, authentic, and light luxury feel), visual style (fresh and healing, soft light, low saturation, slow motion), and product features (natural ingredients, which need to reflect elements such as plants, water droplets, and organic textures).
[0076] For a set of target video materials, the visual embedding vector of each frame of the target video clip is extracted; the keywords in the video generation instructions (such as "plant", "water droplet" and "natural light") are encoded into text embeddings, and cosine similarity is calculated with the video frame embeddings to filter out multiple target video clips.
[0077] Then, based on the video generation instructions and multiple target video clips, a video storyline is constructed. For example, a classic three-act narrative structure is constructed: Beginning: Establishing characters and scenes (e.g., "The Connection Between Women and Nature"); Development: Introducing problems or contrasts (e.g., "The Stress of Modern Life vs. Natural Nourishment"); Climax and Resolution: The product appears as a solution (e.g., "The Product Becomes an Extension of the Power of Nature").
[0078] Finally, based on the video storyline and multiple target video clips, a target video script is generated.
[0079] In an alternative embodiment, a video storyline can be constructed using a natural language model based on video generation instructions and multiple target video segments. For example, the model first performs semantic fusion and context encoding on the input video generation instructions and target video segments, mapping them to a unified semantic representation space, eliminating modal differences, and establishing potential associations between instruction intent and visual elements.
[0080] Then, the model analyzes the target video clips, extracts the subject of the action, spatial relationships, dynamic features and emotional tendencies, forms a deep understanding of the semantic connotation of the shot, and identifies semantic repetition, conflict or gap between shots, and constructs a semantic association map across shots.
[0081] Building upon this foundation, the model, based on narrative logic deduction mechanisms, reorganizes the shot sequence according to its inherent structure of beginning, development, turning point, and conclusion. Combining the emotional guidance and rhythmic requirements of the instructions, it automatically plans the narrative arc. The model further performs consistency checks on the generated narrative sequence, examining semantic coherence, emotional plausibility, and compliance with instruction constraints. It corrects any breaks, redundancies, or deviations from the theme, ensuring a rigorous structure without logical jumps. Finally, the model transforms the validated narrative structure into natural language text, outputting the video storyline.
[0082] First, the source material is selectively filtered to remove irrelevant and distracting material, retaining only semantically relevant target video clips to ensure the accuracy of the subsequent generation. Then, using the filtered target video clips as input, a video storyline with chronological order, logical coherence, and thematic consistency is constructed in conjunction with video generation instructions. Based on this, the structured storyline and corresponding target video clips are further combined to generate a complete target video script, including shot sequences and aligned visual and auditory descriptions. This ensures the script not only conforms to the instructions but also possesses inherent narrative rationality and audiovisual synergy. Finally, the target video is generated based on this high-quality script. This effectively solves the problems of loose script structure, disjointed content, and instruction deviation caused by directly mixing redundant material and lacking narrative guidance in traditional solutions, significantly improving the coherence, relevance, and professionalism of the generated video.
[0083] To ensure the coherence of the target video script, the video generation method provided in Embodiment 1 of this application generates a target video script based on a video storyline and multiple target video segments, including: sorting the multiple target video segments according to the video storyline to obtain sorted target video segments; performing connection recognition on the sorted target video segments to obtain position information of the video segment to be supplemented; generating visual description information of the supplementary video segment at the position information of the video segment to be supplemented based on the video storyline and multiple target video segments; and generating the target video script based on the sorted target video segments and the visual description information of the supplementary video segments.
[0084] Optionally, to ensure the narrative coherence and structural integrity of the target video script, the final target video script is generated based on the video storyline and multiple existing target video segments.
[0085] First, based on the overall narrative logic and emotional flow constructed by the video storyline, semantic matching and sequence adjustment are performed on multiple target video clips. The video storyline implicitly contains a rhythm and emotional trajectory of introduction, development, climax, and resolution. Based on this, the functional positioning of each shot in the narrative can be determined, such as opening setup, emotional turning point, or climax conclusion. Their order is then rearranged to ensure that the shot sequence is consistent with the internal logic of the storyline, forming a coherent shot arrangement sequence.
[0086] Then, a connection analysis is performed on the sorted sequence of shots. For example, based on the semantic transition and visual fluency between adjacent shots, it is identified whether there are logical gaps, emotional jumps, or incoherent actions. For instance, if the previous shot shows a character sitting quietly indoors, and the next shot suddenly switches to running outdoors without any transitional elements, this is identified as a weak point in the narrative connection and marked as the location information for shots to be added.
[0087] After identifying the locations to be supplemented, visual descriptions of supplementary video clips are generated by combining the overall intent of the video storyline with existing shots. The generation process needs to remain consistent with the main narrative. For example, inserting a scene of "curtains being blown by the wind, sunlight slowly streaming into the room" between indoor and outdoor shots maintains stylistic consistency while naturally guiding the viewer's emotional shift.
[0088] Finally, the visual description information of the sorted original target video clips and the newly generated supplementary video clips is integrated, organized in a unified time sequence, and the corresponding auditory elements (such as narration and sound effects) and time lengths are matched for each shot, forming a target video script that is structurally complete, logically rigorous, and rhythmically smooth.
[0089] Based on the video storyline, the target video segments are first sequentially ordered to form a logically coherent sequence of shots. By analyzing the narrative flow and visual semantic consistency between the ordered shots, positions with content breaks or missing transitions are identified. Based on the overall context of the video storyline and existing shots, visual descriptions of the supplementary video segments needed to fill these positions are generated, effectively bridging the narrative gaps between shots. Finally, the visual descriptions of the ordered target shots and supplementary shots are merged to generate a target video script that is structurally complete, logically sound, and narratively coherent. Therefore, this method can solve the problems of script breaks, abrupt transitions, and decreased video quality caused by narrative breaks between target video segments in related technologies, thereby improving the coherence and completeness of the video script and enhancing the narrative expressiveness and audience immersion of the generated video.
[0090] To further improve the accuracy of the target video script, in the video generation method provided in Embodiment 1 of this application, generating the target video script based on the visual description information of the sorted target video segments and supplementary video segments includes: generating corresponding auditory description information for the sorted target video segments and supplementary video segments according to the video storyline, wherein the auditory description information includes at least: voice text information and tone style information; obtaining the identification information corresponding to the sorted target video segments; and obtaining the target video script based on the identification information, the visual description information of the supplementary video segments, and the auditory description information.
[0091] Optionally, based on the established visual sequence structure, a collaborative construction of the auditory dimension can be further introduced to organically unify the visuals and sound. First, for the ordered target video segments, corresponding auditory descriptive information is generated based on their visual content and functional role in the narrative. This information includes voice-over text that highly matches the mood, rhythm, and semantics of the visuals, as well as an appropriate tone of voice. For example, a gentle, slow narration is generated for a calm morning scene, while a slightly expectant or insightful tone is used at turning points, ensuring that the audio expression does not deviate from the visual context and enhancing the narrative's impact.
[0092] For the visual description information of supplementary video clips, corresponding auditory description information is also generated to ensure that the new shots seamlessly connect with the overall rhythm in terms of auditory performance. Although supplementary shots are transitional or restorative content, their accompanying voice-over text and tone must still conform to the overall narrative tone to avoid abruptness or stylistic deviation. A standardized JSON script is output, containing N shots arranged in chronological order. Each shot identifies its type (a source shot referencing an existing video clip or a generative shot) and includes: Visual Track information: For source-based shots, referencing the existing video clip ID; for generative shots, a detailed description of the characters, scenes, actions, and camera movements in the shot. Auditory Track information: Dubbing content and tone.
[0093] By generating auditory descriptions containing voice-over text and tone information from sorted target video segments, and generating auditory descriptions from visual descriptions of supplementary video segments, differentiated content is constructed for the two types of shots in the auditory dimension. Simultaneously, by acquiring the corresponding identifiers of the target video segments, a mapping relationship is established between them and the generated auditory descriptions, ensuring that voice-over text and tone are accurately bound to the corresponding shots. Finally, a structurally complete target video script with highly aligned audiovisual elements is constructed, effectively solving the problem of poor script executability caused by the lack of accurate correspondence and structured expression of audiovisual elements in traditional solutions. This significantly improves the logic, rhythm, and content consistency of video generation, achieving automated generation of high-quality video scripts.
[0094] In an alternative embodiment, it can be achieved through, as follows: Figure 3 The diagram shown illustrates the target video script described above, which includes:
[0095] Input: Video Footage Collection: Contains two types of footage: first, usable footage related to the current theme (i.e., the first video clip); second, interfering footage (i.e., the second video clip), used to simulate redundant data in a real environment. Text Footage: Contains unstructured text (i.e., text description information) such as product features and background introductions. User Instructions (i.e., video generation instructions): Contains constraints such as target duration, video style, and target audience.
[0096] Model reasoning: (1) Material selection: Identify and extract relevant segments from the video material collection; (2) Narrative planning: Construct a narrative arc; (3) Completion and generation: Plan generative shots to satisfy narrative coherence. Finally, generate audiovisual aligned narration and visual descriptions for all shots.
[0097] Output: A standardized JSON script containing N shots arranged chronologically. Each shot clearly identifies its type (a source shot referencing an existing video clip or a generative shot) and includes: Visual Track: For source-based shots, a reference to the existing video clip ID; for generative shots, a detailed description of the characters, scenes, actions, and camera movements in the shot. Auditory Track: Dubbing content and tone. (e.g., ...) Figure 3 The structured script shown includes a generative shot (Shot 1), comprising a visual track and an auditory track. The visual track includes: "Background; Characters; Camera Movement; Action." The auditory track includes: "Tone; Dubbing Content." Shot 2 is a stock shot, also comprising a visual track and an auditory track. The visual track includes: "Stock ID:1.mp4." The auditory track includes: "Tone; Dubbing Content."
[0098] Based on the formatted script, the final complete video can be obtained directly using video generation models or manual shooting methods, which has practical application value and high scalability.
[0099] In an optional embodiment, the model inference can be guided by prompts as shown below:
[0100] "You are an excellent screenwriting master with precise visual understanding and narrative planning skills. Your core task is to create a screenplay based on provided multimodal materials (videos, texts) and user instructions."
[0101] Input content:
[0102] You need to parse the following structured information:
[0103] List of video footage (given in video frame format):<video_materials>
[0104] Text material:<text_material>
[0105] User commands: <instruction>
[0106] Workflow (Please think it through internally first, then output the final JSON):
[0107] 1. Analyze the footage: Understand the visual content, emotional tone, and potential use of each video clip. Combine this with textual materials and user instructions to select the appropriate footage.
[0108] 2. Narrative Conception: Combining video footage, text materials, and user instructions, devise a storyline that connects the available materials (e.g., posing a problem -> demonstrating a solution -> emphasizing the effect -> elevating the emotional impact), and arrange the video footage accordingly.
[0109] 3. Shot Creation: Considering the continuity and appeal of the footage, new shots that need to be filmed are added between the sorted video clips to connect or enhance the overall effect.
[0110] 4. Script generation: Create scripts and voice-overs for each piece of footage. For new shots, create elements such as scenes, characters, and visuals so that the filmmaker can directly shoot the final video based on the generated script.
[0111] Output format (strictly adhered to):
[0112] 1. Existing video footage:
[0113] {
[0114] "shot_id": "A unique lens identifier, an integer, such as 1",
[0115] "duration": "Shot duration (seconds), e.g., 4.0",
[0116] "material_usage":{
[0117] "video_id":"The filename of the referenced video clip, for example, '1.mp4'",
[0118] },
[0119] "dub":{
[0120] "voice": "Voice characteristics of the voice actor, for example: a young female voice",
[0121] "style": "Dubbing tone style, for example: warm and friendly",
[0122] "content": "Speaker and voice-over content, such as: voice-over: This is a great product or young woman in the video: Choosing safety protection for your family is the most thoughtful decision your loved ones can make."
[0123] }
[0124] }
[0125] 2. New generative lens:
[0126] {
[0127] "shot_id": "A unique lens identifier, an integer, such as 1",
[0128] "duration": "Shot duration (seconds), e.g., 2.5",
[0129] "visual":{
[0130] "setting": "A detailed description of the scene's environment (e.g., a city rooftop at dusk, a dimly lit study)",
[0131] "character": "A description of a character's clothing, appearance, and emotional state".
[0132] "shot_type":"Camera type (e.g., close-up, medium shot, full shot, bird's-eye view)",
[0133] "camera_movement": "Camera movement (e.g., fixed camera, handheld shaky camera, slow zoom-in, fast zoom-out, tracking shot)"
[0134] },
[0135] "action": "Describe the specific actions of the people in the shot in sequence, for example: a family member interacts with a child with a smile, and then picks up the child."
[0136] "dub":{
[0137] "voice": "Voice characteristics of the voice actor, for example: a young female voice",
[0138] "style": "Dubbing tone style, for example: warm and friendly",
[0139] "content": "Speaker and voice-over content, such as: voice-over: This is a great product or young woman in the video: Choosing safety protection for your family is the most thoughtful decision your loved ones can make."
[0140] }
[0141] }".
[0142] To improve the usability of target video materials, in the video generation method provided in Embodiment 1 of this application, obtaining the target video material set includes: obtaining multiple first historical videos of the target object, and segmenting them according to the multiple first historical videos to obtain first video segments; obtaining multiple second historical videos of objects other than the target object, and segmenting them according to the multiple second historical videos to obtain second video segments; and obtaining the target video material set according to the first video segments and the second video segments.
[0143] Optionally, in order to improve the usability of the target video material, in the video generation method provided in Embodiment 1 of this application, the construction of the target video material set is achieved by collecting and structurally segmenting multi-source historical videos, ensuring that the material library has both relevance and diversity, thereby providing high-quality and reusable visual resources for subsequent script creation.
[0144] First, acquire multiple first-historical videos of the target object. These videos can originate from official content, events, or user-generated content from the target brand, product, or individual. Then, automatically segment these first-historical videos using scene changes, camera movement, or semantic boundary recognition technology, breaking down the continuous video into multiple independent, semantically complete first-historical video segments.
[0145] Next, acquire multiple second historical videos of objects other than the target object. These videos can be sourced from excellent examples of similar scenarios or video content on related topics. Similarly, segment the videos into shots to extract the second video segments.
[0146] Finally, the first and second video clips are merged to form the target video material set. During the merging process, low-quality, repetitive, or semantically ambiguous video clips can be removed, and the remaining materials undergo standardized preprocessing, such as removing the original audio, subtitles, and watermarks, to ensure that subsequent inference relies solely on visual content.
[0147] In an optional embodiment, when making a video for a certain brand of sports shoes, firstly, multiple first historical videos previously released by the brand are obtained, and then they are segmented into several first video segments using a scene detection algorithm; then, multiple second historical videos of other brands or unrelated products are obtained, and similarly segmented into second video segments; finally, the first video segments and the second video segments are mixed to form a target video material set, which serves as a multimodal input source for subsequent script creation.
[0148] By acquiring and segmenting multiple first historical videos of the target object, first video segments directly related to the target object are extracted. Simultaneously, multiple second historical videos other than the target object are acquired and segmented to construct second video segments unrelated to the target object. These two types of video segments are then merged to form a mixed video material set containing target relevance and external interference. This allows for accurate matching and semantic filtering of visual and auditory descriptive information based on structured video segments collected from real-world scenes when generating target video scripts according to video generation instructions. This effectively suppresses script misassociations, narrative breaks, and instruction deviations caused by ambiguous material sources or unstructured interference information, thereby improving the accuracy, coherence, and instruction compliance of the generated script, ultimately achieving high-quality and highly reliable target video generation.
[0149] To improve the accuracy of the first video segment, the video generation method provided in Embodiment 1 of this application, which divides multiple first historical videos to obtain the first video segment, includes: analyzing any one first historical video to obtain multiple initial video segments; scoring the multiple initial video segments to obtain scores corresponding to each initial video segment; and filtering the multiple initial video segments based on the scores to obtain the first video segment.
[0150] Optionally, frame-by-frame analysis is performed on any first historical video, and motion detection, scene boundary recognition, and semantic segmentation techniques are combined to decompose the continuous video into multiple initial video segments. An initial shot (i.e., an initial video segment) is identified as a segment with relatively independent visual semantics, such as a complete product demonstration, a character's appearance, or an emotional expression unit, ensuring that the segmentation results have the characteristics of the smallest semantic unit in terms of time and content.
[0151] Then, the initial video clips are scored across multiple dimensions. Scoring criteria may include: visual clarity (such as resolution, focus accuracy, and lighting appropriateness), content completeness (whether it contains a complete action or event), information density (whether it conveys clear semantics rather than redundant segments), stylistic consistency (whether it matches the visual tone of the target audience's past content), and potential reusability (whether it has adaptability to general scenarios, such as no lip-sync conflicts or sensitive information). The scoring can be quantified by a pre-trained evaluation module, ultimately yielding a comprehensive score that reflects the shot's usability potential in actual creative work.
[0152] Finally, all initial video clips are filtered based on the scoring results, retaining only those with scores above a preset threshold as the first video clip. Low-scoring clips, such as blurry, excessively short, semantically fragmented, or stylistically off-target shots, are removed. This filtering process not only eliminates noise interference but also enhances the professionalism and consistency of the material library, ensuring that the visual foundation upon which subsequent script generation relies is accurate, reliable, and scalable, thereby significantly improving the overall quality and efficiency of script generation.
[0153] After segmenting multiple first-historical videos to obtain initial video clips, the initial video clips in the videos are further comprehensively analyzed and quantitatively scored in terms of semantics, duration, composition, rhythm, and reuse frequency. Based on the scoring results, high-quality shot clips with high relevance, high reuse value, and low interference are selected to form a set of precisely optimized first video clips. This avoids the problem of low-quality, redundant, or irrelevant clips mixed in due to simply relying on segmentation, and improves the quality and consistency of the material base on which the subsequent video script generation depends. Ultimately, this achieves a synergistic improvement in the accuracy of the video generation script and the overall quality of the finished product.
[0154] In the video generation method provided in Embodiment 1 of this application, after obtaining multiple initial video segments, the method further includes: performing context-aware semantic induction on the multiple initial video segments to obtain visual description information corresponding to each of the multiple initial video segments; and obtaining video generation instructions based on the visual description information corresponding to each of the multiple initial video segments and the multiple initial video segments, wherein the video generation instructions include at least: style constraint information, target total duration, and description information of the target audience.
[0155] Optionally, context-aware semantic summarization is performed on the initial video clips, combining their temporal relationship, emotional trajectory, and overall narrative intent within the original video, and a multimodal large language model is used to deeply understand the content of the shots. For example, by comprehensively analyzing the characters' actions, scene setting, camera movement, color atmosphere, and temporal rhythm in the shots, visual descriptive information can be generated. For instance, a shot of "a product slowly rotating in the morning light" could have a visual description that includes "the product is a metallic water cup, the background is a light-colored wooden tabletop, and the light shines obliquely from the left," and could also include, by incorporating contextual semantics, adding "the camera slowly zooms in, accompanied by soft ambient sounds, creating a high-end and tranquil user experience," thus elevating the description beyond pixel-level observation to an understandable creative semantic.
[0156] Then, a global summary can be made based on all initial video clips and their corresponding visual descriptions. By analyzing the tone, emotional inclination, and content structure commonly conveyed by these shots, three key video generation instruction elements can be extracted: First, style constraint information, such as "light luxury and simplicity," "youthful vitality," and "professional authority," describing the unified paradigm of the overall visual and auditory style; second, the target total duration, based on the total duration and narrative density of the current set of shots, to infer the appropriate finished video duration, such as "30 seconds" or "45 seconds"; third, the target audience description information, by analyzing historical video interaction data, audience feedback, and visual symbols, to infer the core characteristics served by the content.
[0157] Finally, video generation instructions are formed based on the aforementioned style constraints, target duration, and target audience description.
[0158] By performing context-aware semantic induction on multiple initial video clips, the visual semantic features contained in each video clip are extracted to form structured visual description information. Combined with the original video clips themselves, the style constraints, target total duration, and target audience description information required for video creation are dynamically inferred. This generates video generation instructions with clear creative guidance, solving the problem of vague generation instructions and lack of creative constraints caused by the lack of semantic understanding of the initial materials in traditional solutions. This enables the subsequent video generation process to automatically adapt based on precise style, duration, and audience guidance, significantly improving the content consistency, logical integrity, and user adaptability of the generated video. Ultimately, it achieves end-to-end intelligent guidance from original materials to high-quality video output.
[0159] To improve the accuracy of the initial video segments, the video generation method provided in Embodiment 1 of this application analyzes the first historical video to obtain multiple initial video segments, including: extracting image frames from the first historical video based on a fixed frame rate to obtain multiple image frames; calculating the visual features of the multiple image frames to obtain visual feature values of objects in the multiple image frames; calculating the visual difference values between adjacent image frames in the multiple image frames based on the visual feature values; determining scene switching points in the first historical video based on the visual difference values, and segmenting the first historical video based on the scene switching points to obtain multiple initial video segments.
[0160] Optionally, image frame sequences can be uniformly extracted from the first historical video at a fixed frame rate (e.g., 5 frames per second) to obtain multiple image frames. This sampling strategy ensures computational efficiency while avoiding redundancy introduced by excessively high frame rates or the omission of key visual changes due to excessively low frame rates.
[0161] Visual feature extraction is performed on each frame of the image. A pre-trained deep convolutional neural network can be used to map each frame to a high-dimensional semantic feature vector, i.e., the visual feature value of that frame. The visual difference between adjacent image frames is calculated. This difference can be obtained by calculating the Euclidean distance or cosine similarity between the visual feature vectors of two frames, used to quantify the intensity of visual content change in consecutive frames. When the difference between adjacent frames is significantly higher than the background fluctuation threshold, a significant visual transition can be identified, i.e., a potential scene switching point. To avoid false positives, a time sliding window and trend smoothing processing can be further combined to filter out transient fluctuations caused by camera shake, slight changes in lighting, or rapid movement, ensuring that only switching points with semantic discontinuity are retained.
[0162] Finally, based on the identified scene transition points, the original first historical video is precisely segmented on the timeline. The video segment between each adjacent transition point is defined as an independent initial video clip, ensuring that the shots have relative visual integrity and internal consistency, such as a complete product feature demonstration, a continuous dialogue, or a complete environmental transition.
[0163] By extracting image frames from the first historical video at a fixed frame rate, discrete but temporally continuous visual samples are obtained. Then, visual feature values of the corresponding objects in the image frames are extracted to quantify the visual representation of key content within the frame. Next, by calculating the differences in visual feature values between adjacent image frames, a continuous difference sequence reflecting the dynamic changes in the scene content is formed. Based on significant abrupt changes in this sequence, semantically coherent scene transition boundaries are automatically identified. Finally, the original video is precisely segmented based on the transition points to generate multiple semantically complete and clearly defined initial video segments. This solves the problem that traditional methods rely on manual annotation or coarse-grained temporal segmentation, resulting in inaccurate shot units and an inability to meet the needs of structured script generation. It achieves automated and high-precision analysis of shot structures in historical videos, providing reliable and semantically consistent visual basic units for subsequent script generation.
[0164] In an alternative embodiment, it can be achieved through, as follows: Figure 4 The diagram illustrates the construction of the target video material set, including: acquiring multiple high-quality videos. Phase 1: Using scene detection algorithms, the original high-quality videos are segmented into independent shots. A multimodal large model, combining the relative position and contextual information of the current shot, generates detailed visual descriptions and narration for each shot, forming the original script. Phase 2: Based on the reconstructed complete script, the large model summarizes the core features and product information corresponding to the video (as text material input), as well as style and audience constraints (as user command input). Reusability scoring: The model scores each shot, selecting high-quality shots with universality and no lip-sync limitations as usable material. Phase 3: Constructing a hybrid material pool, containing highly reusable shots selected from relevant videos and random distracting shots drawn from unrelated videos. All materials are preprocessed, removing original audio and subtitles to ensure the model relies solely on visual information for inference, such as... Figure 4 As shown, the mixed media pool includes video footage (available footage + distraction footage), text footage, and user commands.
[0165] To further improve the quality of the target video script, in the video generation method provided in Embodiment 1 of this application, after generating the target video script based on the video generation instructions and the target video material set, the method further includes: quantifying the target video script based on preset rule indicators to obtain a first score value; quantifying the target video script through a natural language model to obtain a second score value; and optimizing the target video script based on the first score value and the second score value.
[0166] Optionally, the target video script can be automatically quantitatively analyzed based on preset hard rule indicators to generate a first score. For example, the quantification may include whether interfering material was misused, whether the same shot was excessively repeated, whether the total duration deviates from the user-specified target duration, whether the JSON format is complete and standardized, and whether the shot type labeling is accurate.
[0167] It can also utilize natural language models to perform in-depth semantic understanding and creative quality assessment of the script content, generating a second score. For example, it can comprehensively evaluate whether the script's narrative logic is coherent, whether the narration is engaging, whether the visual descriptions are highly consistent with the auditory content, whether the shot sequence matches the emotional rhythm, and whether the generated shots truly fill narrative gaps rather than being redundant additions.
[0168] After obtaining the first and second scores, the two scores are weighted and combined to form a comprehensive judgment on the script quality. If the first score is low, structural repairs can be prioritized, such as automatically replacing incorrectly referenced materials, cropping out time-out shots, and correcting formatting errors to ensure the script can be stably called downstream. If the second score is low, semantic-level optimization is initiated, such as rewriting awkward narration, adjusting shot descriptions to enhance visual expressiveness, deleting unnecessary generated paragraphs, or adding key turning points to make the content more attractive and persuasive.
[0169] After generating the initial target video script, a dual quantitative evaluation mechanism is introduced: on the one hand, the script's structural integrity, shot sequence logic, and audiovisual element matching are quantitatively scored according to preset rule indicators to obtain a first score value; on the other hand, a natural language model is used to conduct deep semantic analysis on the script's semantic consistency, instruction fit, descriptive accuracy, and expression fluency to obtain a second score value. Then, the combined scores from both are used as feedback signals to make targeted corrections and optimizations to problems such as structural defects, semantic deviations, or audiovisual inconsistencies in the script. This achieves closed-loop control from script generation to quality improvement, effectively solving the problem of low script quality and poor final video effects caused by the lack of multi-dimensional evaluation and intelligent feedback mechanisms in traditional technologies, and improving the accuracy, rationality, and usability of video scripts.
[0170] To improve the accuracy of scoring the target video script, the video generation method provided in Embodiment 1 of this application quantifies the target video script based on preset rule indicators to obtain a first score value, including: calculating the number of second video segments referenced from the target video material set in the target video script to obtain the proportion of incorrectly referenced video segments; calculating the number of video segments reused in the target video script to obtain the richness of video segments; calculating the planned total duration of the target video script and the target total duration in the video generation instruction to obtain a duration deviation value; and calculating the first score value based on the proportion value, richness, and duration deviation value.
[0171] Optionally, in order to improve the accuracy of the target video script scoring, the video generation method provided in Embodiment 1 of this application uses three key and quantifiable engineering indicators to evaluate the structural rationality and material utilization efficiency of the target video script, and calculates the first score accordingly.
[0172] First, a source analysis is performed on all video clips referenced in the script to count the number of interfering clips (i.e., non-preset usable clips or low-quality redundant clips) from outside the target video clip set, and the proportion of these clips to the total number of referenced shots is calculated to obtain the proportion of incorrectly referenced video clips.
[0173] Secondly, the frequency of repeatedly used video clips in the script is analyzed to identify instances where a single shot is called multiple times, thereby calculating the richness of the video clips. This metric measures the breadth of the script's use of the media library by calculating the coverage ratio of effective footage.
[0174] Then, the total planned duration of all shots in the script (including source footage and generated footage) is compared with the target total duration specified in the video generation instructions. The deviation between the two is calculated and normalized to a percentage to form a duration deviation value. This value is used to evaluate the ability to control global time constraints. Excessive deviation (such as exceeding ±10%) will lead to loss of control over the pacing of the final cut, seriously affecting the user experience.
[0175] Finally, the three indicators mentioned above—error citation ratio, material richness, and duration deviation—are weighted and integrated according to their importance in actual creation to construct a comprehensive scoring function, outputting a first score. For example, the error citation ratio has a high weight because it directly relates to the script's credibility and compliance; richness is secondary, affecting creative diversity; duration deviation is a binding indicator, and excessive deviation directly results in disqualification. Through this weighted calculation, the first score not only objectively reflects the accuracy of the script's technical implementation but also provides a clear and traceable direction for subsequent optimization, ensuring that the evaluation process is rigorous, reliable, and closely reflects real creative needs.
[0176] By statistically analyzing the number of incorrectly referenced second video segments from the target video material set in the target video script, the proportion of incorrectly referenced video segments is calculated to assess the accuracy of the script's material selection. Simultaneously, the number of repeatedly used video segments in the script is analyzed to calculate the richness of video segments, measuring the diversity and creativity of material usage. Furthermore, the total planned duration of the script is compared with the target total duration set in the video generation instructions to obtain a duration deviation value, reflecting the compliance and consistency of the script's time planning. The proportion, richness, and duration deviation values are then comprehensively calculated to generate a first score. This achieves an objective and structured evaluation of the video script in three key dimensions: accuracy of material selection, content diversity, and instruction compliance. This effectively solves the problem of previous methods that only performed general quantification and could not accurately identify defects in the core creative process of the script, achieving the technical effect of improving the quality of video generation scripts and enhancing the matching degree between the generated results and user intent.
[0177] In the video generation method provided in Embodiment 1 of this application, a target video material set and a video generation instruction are obtained. The target video material set includes at least: a first video segment related to the target object, a second video segment unrelated to the target object, and text description information corresponding to the target object in the first video segment. Based on the video generation instruction and the target video material set, a target video script is generated. The target video script includes: visual description information of multiple video segments and auditory description information corresponding to each of the multiple video segments. Based on the target video script, a target video corresponding to the target object is generated. This solves the technical problem in related technologies where the quality of the generated video script is low, resulting in poor video generation quality.
[0178] In this application, a target video material set is obtained, comprising a first video segment related to the target object, a second video segment unrelated to the target object, and textual description information corresponding to the target object in the first video segment. Combined with video generation instructions, multimodal input materials are structurally filtered and semantically aligned to generate a target video script containing visual description information of multiple video segments and their corresponding auditory description information. This achieves proactive differentiation of redundant and interfering materials and precise arrangement of audiovisual content. Based on this, a target video corresponding to the target object is generated based on the structured script, ensuring that the video generation process is driven by explicit semantic instructions, rather than relying on unconstrained end-to-end splicing. By clearly distinguishing between relevant and irrelevant materials, introducing visual description information as semantic anchors, and simultaneously constructing audiovisual dual-channel description information, the logic, consistency, and guidance of the video script are improved. This effectively solves the problems of messy generated video content, thematic deviation, and audiovisual disconnect caused by low script quality in related technologies, achieving the technical effect of improving video generation accuracy, thematic focus, and overall quality.
[0179] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0180] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.
[0181] Example 2
[0182] According to embodiments of this application, a video generation method is also provided, such as... Figure 5 As shown, the video generation method includes:
[0183] Step S501: Obtain the target video material set and video generation instructions uploaded by the client. The target video material set includes at least: a first video segment related to the target object, a second video segment unrelated to the target object, and text description information corresponding to the target object in the first video segment.
[0184] Step S502: Generate a target video script in the cloud server based on the video generation instructions and the target video material set. The target video script includes: visual description information of multiple video segments and auditory description information corresponding to each of the multiple video segments. Based on the target video script, generate a target video corresponding to the target object.
[0185] Step S503: Return the target video to the client.
[0186] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0187] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.
[0188] Example 3
[0189] According to embodiments of this application, a video generation apparatus for implementing the above-described video generation method is also provided, such as... Figure 6 As shown, the device includes: an acquisition unit 601, a first generation unit 602, and a second generation unit 603.
[0190] The acquisition unit 601 is used to acquire a target video material set and a video generation instruction, wherein the target video material set includes at least: a first video segment related to the target object, a second video segment unrelated to the target object, and text description information corresponding to the target object in the first video segment;
[0191] The first generation unit 602 is used to generate a target video script based on the video generation instructions and the target video material set. The target video script includes: visual description information of multiple video segments and auditory description information corresponding to each of the multiple video segments.
[0192] The second generation unit 603 is used to generate the target video corresponding to the target object based on the target video script.
[0193] In the video generation apparatus provided in Embodiment 3 of this application, the acquisition unit 601 acquires a target video material set and a video generation instruction. The target video material set includes at least: a first video segment related to the target object, a second video segment unrelated to the target object, and text description information corresponding to the target object in the first video segment. The first generation unit 602 generates a target video script based on the video generation instruction and the target video material set. The target video script includes: visual description information of multiple video segments and auditory description information corresponding to each of the multiple video segments. The second generation unit 603 generates a target video corresponding to the target object based on the target video script, thus solving the technical problem in related technologies where the quality of the generated video script is low, resulting in poor video generation quality.
[0194] In this application, a target video material set is obtained, comprising a first video segment related to the target object, a second video segment unrelated to the target object, and textual description information corresponding to the target object in the first video segment. Combined with video generation instructions, multimodal input materials are structurally filtered and semantically aligned to generate a target video script containing visual description information of multiple video segments and their corresponding auditory description information. This achieves proactive differentiation of redundant and interfering materials and precise arrangement of audiovisual content. Based on this, a target video corresponding to the target object is generated based on the structured script, ensuring that the video generation process is driven by explicit semantic instructions, rather than relying on unconstrained end-to-end splicing. By clearly distinguishing between relevant and irrelevant materials, introducing visual description information as semantic anchors, and simultaneously constructing audiovisual dual-channel description information, the logic, consistency, and guidance of the video script are improved. This effectively solves the problems of messy generated video content, thematic deviation, and audiovisual disconnect caused by low script quality in related technologies, achieving the technical effect of improving video generation accuracy, thematic focus, and overall quality.
[0195] Optionally, in the video generation apparatus provided in Embodiment 3 of this application, the first generation unit includes: a filtering subunit, used to filter in a target video material set based on video generation instructions and target objects to obtain multiple target video segments; a first generation subunit, used to generate a video storyline based on video generation instructions and multiple target video segments; and a second generation subunit, used to generate a target video script based on the video storyline and multiple target video segments.
[0196] Optionally, in the video generation apparatus provided in Embodiment 3 of this application, the second generation subunit includes: a sorting module, used to sort multiple target video segments according to the video storyline to obtain sorted target video segments; an identification module, used to perform connection identification on the sorted target video segments to obtain position information of the video segment to be supplemented; a first generation module, used to generate visual description information of the supplementary video segment at the position information of the video segment to be supplemented based on the video storyline and multiple target video segments; and a second generation module, used to generate a target video script based on the sorted target video segments and the visual description information of the supplementary video segments.
[0197] Optionally, in the video generation apparatus provided in Embodiment 3 of this application, the second generation module includes: a generation submodule, used to generate corresponding auditory description information for the sorted target video segments and supplementary video segments respectively according to the video storyline, wherein the auditory description information includes at least: voice text information and tone style information; an acquisition submodule, used to acquire the identification information corresponding to the sorted target video segments; and a first determination submodule, used to obtain the target video script based on the identification information, the visual description information of the supplementary video segments, and the auditory description information.
[0198] Optionally, in the video generation apparatus provided in Embodiment 3 of this application, the acquisition unit includes: a first acquisition subunit, used to acquire multiple first historical videos of a target object, and segment the multiple first historical videos to obtain a first video segment; a second acquisition subunit, used to acquire multiple second historical videos of objects other than the target object, and segment the multiple second historical videos to obtain a second video segment; and a determination subunit, used to obtain a target video material set based on the first video segment and the second video segment.
[0199] Optionally, in the video generation apparatus provided in Embodiment 3 of this application, the first acquisition subunit includes: an analysis module, used to analyze any first historical video to obtain multiple initial video segments; a scoring module, used to score the multiple initial video segments to obtain the score values corresponding to the multiple initial video segments respectively; and a filtering module, used to filter the multiple initial video segments according to the score values to obtain a first video segment.
[0200] Optionally, in the video generation apparatus provided in Embodiment 3 of this application, the apparatus further includes: an induction unit, configured to perform context-aware semantic induction on the multiple initial video segments after obtaining multiple initial video segments, to obtain visual description information corresponding to each of the multiple initial video segments; and a determination unit, configured to obtain a video generation instruction based on the visual description information corresponding to each of the multiple initial video segments and the multiple initial video segments, wherein the video generation instruction includes at least: style constraint information, target total duration, and description information of the target audience.
[0201] Optionally, in the video generation apparatus provided in Embodiment 3 of this application, the analysis module includes: an extraction submodule, used to extract image frames from a first historical video based on a fixed frame rate to obtain multiple image frames; a first calculation submodule, used to calculate the visual features of the multiple image frames to obtain visual feature values of the objects in the multiple image frames; a second calculation submodule, used to calculate the visual difference value between adjacent image frames in the multiple image frames based on the visual feature value; and a second determination submodule, used to determine the scene switching point in the first historical video based on the visual difference value, and to segment the first historical video based on the scene switching point to obtain multiple initial video segments.
[0202] Optionally, in the video generation apparatus provided in Embodiment 3 of this application, the apparatus further includes: a first quantization unit, used to quantize the target video script based on preset rule indicators after generating the target video script according to the video generation instructions and the target video material set, to obtain a first score value; a second quantization unit, used to quantize the target video script through a natural language model to obtain a second score value; and an optimization unit, used to optimize the target video script based on the first score value and the second score value.
[0203] Optionally, in the video generation apparatus provided in Embodiment 3 of this application, the first quantization unit includes: a first calculation subunit, used to calculate the number of second video segments referenced from the target video material set in the target video script to obtain a proportion value of incorrectly referenced video segments; a second calculation subunit, used to calculate the number of video segments reused in the target video script to obtain the richness of video segments; a third calculation subunit, used to calculate the planned total duration of the target video script and the target total duration in the video generation instruction to obtain a duration deviation value; and a fourth calculation subunit, used to calculate based on the proportion value, richness, and duration deviation value to obtain a first score value.
[0204] It should be noted that the aforementioned acquisition unit 601, first generation unit 602, and second generation unit 603 correspond to steps S201 to S203 in Embodiment 1. The three units and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should also be noted that the aforementioned units, as part of the device, can operate in the computer terminal 10 provided in Embodiment 1.
[0205] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.
[0206] Example 4
[0207] Embodiments of this application may provide an electronic device, which may be any one of a group of electronic device terminals. Optionally, the aforementioned electronic device may also be replaced by a terminal device such as a mobile terminal.
[0208] Optionally, the aforementioned electronic device may be located in at least one of a plurality of network devices in a computer network.
[0209] The aforementioned electronic device can execute the program code for the following steps in the video generation method: obtaining a target video material set and a video generation instruction, wherein the target video material set includes at least: a first video segment related to the target object, a second video segment unrelated to the target object, and text description information corresponding to the target object in the first video segment; generating a target video script based on the video generation instruction and the target video material set, wherein the target video script includes: visual description information of multiple video segments and auditory description information corresponding to each of the multiple video segments; and generating a target video corresponding to the target object based on the target video script.
[0210] The aforementioned electronic device can execute the program code for the following steps in the video generation method: generating a target video script based on video generation instructions and a target video material set, including: filtering the target video material set based on video generation instructions and target objects to obtain multiple target video segments; generating a video storyline based on video generation instructions and multiple target video segments; and generating a target video script based on the video storyline and multiple target video segments.
[0211] The aforementioned electronic device can execute program code for the following steps in the video generation method: generating a target video script based on a video storyline and multiple target video segments, including: sorting the multiple target video segments according to the video storyline to obtain sorted target video segments; performing connection recognition on the sorted target video segments to obtain position information of the video segment to be supplemented; generating visual description information of the supplementary video segment at the position information of the video segment to be supplemented based on the video storyline and multiple target video segments; and generating a target video script based on the sorted target video segments and the visual description information of the supplementary video segments.
[0212] The aforementioned electronic device can execute the program code for the following steps in the video generation method: generating a target video script based on the visual description information of the sorted target video segments and supplementary video segments, including: generating corresponding auditory description information for the sorted target video segments and supplementary video segments respectively based on the video storyline, wherein the auditory description information includes at least: voice text information and tone style information; obtaining the identification information corresponding to the sorted target video segments; and obtaining the target video script based on the identification information, the visual description information of the supplementary video segments, and the auditory description information.
[0213] The aforementioned electronic device can execute the program code for the following steps in the video generation method: obtaining a target video material set includes: obtaining multiple first historical videos of the target object, and segmenting them based on the multiple first historical videos to obtain a first video segment; obtaining multiple second historical videos of objects other than the target object, and segmenting them based on the multiple second historical videos to obtain a second video segment; and obtaining a target video material set based on the first video segment and the second video segment.
[0214] The aforementioned electronic device can execute the program code for the following steps in the video generation method: segmenting multiple first historical videos to obtain a first video segment, including: analyzing any one first historical video to obtain multiple initial video segments; scoring the multiple initial video segments to obtain the scores corresponding to each initial video segment; and filtering the multiple initial video segments based on the scores to obtain the first video segment.
[0215] The aforementioned electronic device can execute the program code for the following steps in the video generation method: after obtaining multiple initial video segments, the method further includes: performing context-aware semantic induction on the multiple initial video segments to obtain visual description information corresponding to each of the multiple initial video segments; and obtaining a video generation instruction based on the visual description information corresponding to each of the multiple initial video segments and the multiple initial video segments, wherein the video generation instruction includes at least: style constraint information, target total duration, and description information of the target audience.
[0216] The aforementioned electronic device can execute the program code for the following steps in the video generation method: analyzing the first historical video to obtain multiple initial video segments, including: extracting image frames from the first historical video based on a fixed frame rate to obtain multiple image frames; calculating the visual features of the multiple image frames to obtain visual feature values for objects in the multiple image frames; calculating the visual difference values between adjacent image frames in the multiple image frames based on the visual feature values; determining scene switching points in the first historical video based on the visual difference values, and segmenting the first historical video based on the scene switching points to obtain multiple initial video segments.
[0217] The aforementioned electronic device can execute the program code for the following steps in the video generation method: after generating a target video script based on the video generation instructions and the target video material set, the method further includes: quantifying the target video script based on preset rule indicators to obtain a first score value; quantifying the target video script through a natural language model to obtain a second score value; and optimizing the target video script based on the first score value and the second score value.
[0218] The aforementioned electronic device can execute program code for the following steps in the video generation method: quantifying the target video script based on preset rule indicators to obtain a first score value, including: calculating the number of second video segments referenced from the target video material set in the target video script to obtain the proportion of incorrectly referenced video segments; calculating the number of video segments reused in the target video script to obtain the richness of video segments; calculating the planned total duration of the target video script and the target total duration in the video generation instruction to obtain a duration deviation value; and calculating the first score value based on the proportion value, richness, and duration deviation value.
[0219] Optionally, Figure 7 This is a structural block diagram of an electronic device according to an embodiment of this application. Figure 7 As shown, the electronic device 70 may include one or more (only one is shown in the figure) processors 702 and memory 704. The electronic device 70 may also include a memory controller to control and manage the memory 704; the electronic device 70 may also include a peripheral interface to connect to a radio frequency module, an audio module, and a display screen, etc.
[0220] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the video generation method and apparatus in this embodiment. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby realizing the aforementioned video generation method. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to the electronic device 70 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0221] The processor can invoke information and application programs stored in memory via a transmission device to perform the following steps: acquiring a target video material set and video generation instructions, wherein the target video material set includes at least: a first video segment related to the target object, a second video segment unrelated to the target object, and text description information corresponding to the target object in the first video segment; generating a target video script based on the video generation instructions and the target video material set, wherein the target video script includes: visual description information of multiple video segments and auditory description information corresponding to each of the multiple video segments; and generating a target video corresponding to the target object based on the target video script.
[0222] The processor can access information and applications stored in memory via a transmission device to perform the following steps: generating a target video script based on video generation instructions and a target video material set, including: filtering multiple target video segments from the target video material set based on video generation instructions and target objects; generating a video storyline based on video generation instructions and multiple target video segments; and generating a target video script based on the video storyline and multiple target video segments.
[0223] The processor can invoke information and application programs stored in the memory through the transmission device to perform the following steps: generating a target video script based on a video storyline and multiple target video segments, including: sorting the multiple target video segments according to the video storyline to obtain sorted target video segments; performing connection recognition on the sorted target video segments to obtain the position information of the video segment to be supplemented; generating visual description information of the supplementary video segment at the position information of the video segment to be supplemented based on the video storyline and multiple target video segments; and generating a target video script based on the sorted target video segments and the visual description information of the supplementary video segment.
[0224] The processor can invoke information and applications stored in the memory via a transmission device to perform the following steps: generating a target video script based on the visual description information of the sorted target video segments and supplementary video segments, including: generating corresponding auditory description information for the sorted target video segments and supplementary video segments according to the video storyline, wherein the auditory description information includes at least: voice text information and tone style information; obtaining the identification information corresponding to the sorted target video segments; and obtaining the target video script based on the identification information, the visual description information of the supplementary video segments, and the auditory description information.
[0225] The processor can invoke information and applications stored in the memory through the transmission device to perform the following steps: obtaining a target video material set includes: obtaining multiple first historical videos of the target object, and segmenting them according to the multiple first historical videos to obtain first video segments; obtaining multiple second historical videos of objects other than the target object, and segmenting them according to the multiple second historical videos to obtain second video segments; and obtaining a target video material set based on the first video segments and the second video segments.
[0226] The processor can access information and applications stored in the memory via a transmission device to perform the following steps: segmenting multiple first historical videos to obtain a first video segment, including: analyzing any one first historical video to obtain multiple initial video segments; scoring the multiple initial video segments to obtain scores corresponding to each initial video segment; and filtering the multiple initial video segments based on the scores to obtain a first video segment.
[0227] The processor can invoke information and application programs stored in memory through a transmission device to perform the following steps: After obtaining multiple initial video segments, the method further includes: performing context-aware semantic induction on the multiple initial video segments to obtain visual description information corresponding to each of the multiple initial video segments; and obtaining video generation instructions based on the visual description information corresponding to each of the multiple initial video segments and the multiple initial video segments, wherein the video generation instructions include at least: style constraint information, target total duration, and description information of the target audience.
[0228] The processor can invoke information and application programs stored in the memory through the transmission device to perform the following steps: analyzing the first historical video to obtain multiple initial video segments, including: extracting image frames from the first historical video based on a fixed frame rate to obtain multiple image frames; calculating the visual features of the multiple image frames to obtain visual feature values of objects in the multiple image frames; calculating the visual difference values between adjacent image frames in the multiple image frames based on the visual feature values; determining scene switching points in the first historical video based on the visual difference values, and segmenting the first historical video based on the scene switching points to obtain multiple initial video segments.
[0229] The processor can access information and applications stored in the memory via a transmission device to perform the following steps: after generating a target video script based on video generation instructions and a target video material set, the method further includes: quantifying the target video script based on preset rule indicators to obtain a first score value; quantizing the target video script using a natural language model to obtain a second score value; and optimizing the target video script based on the first and second score values.
[0230] The processor can access information and applications stored in the memory via a transmission device to perform the following steps: quantifying the target video script based on preset rule indicators to obtain a first score value includes: calculating the number of second video segments referenced from the target video material set in the target video script to obtain the proportion of incorrectly referenced video segments; calculating the number of video segments reused in the target video script to obtain the richness of video segments; calculating the planned total duration of the target video script and the target total duration in the video generation instruction to obtain a duration deviation value; and calculating the first score value based on the proportion value, richness, and duration deviation value.
[0231] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0232] Example 5
[0233] Embodiments of this application also provide a computer program product. Optionally, the computer program product can be used to store the program code executed by the video generation method provided in Embodiment 1.
[0234] Optionally, the aforementioned computer program product may be located in any computer terminal in a group of computer terminals in a computer network, or in any mobile terminal in a group of mobile terminals.
[0235] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0236] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0237] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some interfaces; indirect couplings or communication connections between units or modules may be electrical or other forms.
[0238] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0239] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0240] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0241] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.< / instruction>
Claims
1. A video generation method, characterized in that, include: Obtain a target video material set and video generation instructions, wherein the target video material set includes at least: a first video segment related to the target object, a second video segment unrelated to the target object, and text description information corresponding to the target object in the first video segment; Based on the video generation instructions and the target video material set, a target video script is generated, wherein the target video script includes: visual description information of multiple video segments and auditory description information corresponding to the multiple video segments respectively; Based on the target video script, generate the target video corresponding to the target object.
2. The method according to claim 1, characterized in that, Based on the video generation instructions and the target video material set, the target video script is generated as follows: Based on the video generation instructions and the target object, multiple target video clips are obtained by filtering the target video material set; Based on the video generation instructions and the multiple target video segments, a video storyline is generated; Based on the video storyline and the multiple target video segments, the target video script is generated.
3. The method according to claim 2, characterized in that, Generating the target video script based on the video storyline and the multiple target video segments includes: Based on the video storyline, the multiple target video segments are sorted to obtain sorted target video segments; The sorted target video segments are spliced together to obtain the location information of the video segments to be supplemented. Based on the video storyline and the multiple target video segments, visual description information of the supplementary video segment is generated at the location information of the video segment to be supplemented; The target video script is generated based on the visual description information of the sorted target video segments and the supplementary video segments.
4. The method according to claim 3, characterized in that, Generating the target video script based on the visual description information of the sorted target video segments and the supplementary video segments includes: Based on the video storyline, corresponding auditory description information is generated for the sorted target video segments and the supplementary video segments, wherein the auditory description information includes at least: voice text information and tone style information; Obtain the identification information corresponding to the sorted target video segments; The target video script is obtained based on the identification information, the visual description information of the supplementary video segment, and the auditory description information.
5. The method according to claim 1, characterized in that, The collection of target video footage includes: Multiple first historical videos of the target object are obtained, and the video segments are obtained by segmenting the multiple first historical videos. Obtain multiple second historical videos of objects other than the target object, and segment them according to the multiple second historical videos to obtain the second video segment; The target video material set is obtained based on the first video segment and the second video segment.
6. The method according to claim 5, characterized in that, The first video segment is obtained by segmenting the multiple first historical videos, including: For any given first historical video, analyze the first historical video to obtain multiple initial video segments; The multiple initial video segments are scored to obtain the scores corresponding to each of the multiple initial video segments; The first video segment is obtained by filtering the multiple initial video segments based on the scores.
7. The method according to claim 6, characterized in that, After obtaining the plurality of initial video segments, the method further includes: Context-aware semantic induction is performed on the multiple initial video segments to obtain visual description information corresponding to each of the multiple initial video segments; Based on the visual description information corresponding to the plurality of initial video segments and the plurality of initial video segments, the video generation instruction is obtained, wherein the video generation instruction includes at least: style constraint information, target total duration, and description information of the target audience.
8. The method according to claim 6, characterized in that, Analysis of this first historical video yielded several initial video clips, including: Multiple image frames are obtained by extracting image frames from the first historical video based on a fixed frame rate; The visual features of the plurality of image frames are calculated to obtain the visual feature values of the objects in the plurality of image frames respectively; Based on the visual feature values, calculate the visual difference values between adjacent image frames in the plurality of image frames; The scene switching points in the first historical video are determined based on the visual difference values, and the first historical video is segmented based on the scene switching points to obtain the plurality of initial video segments.
9. The method according to claim 1, characterized in that, After generating the target video script based on the video generation instructions and the target video material set, the method further includes: The target video script is quantified based on preset rule indicators to obtain a first score value; The target video script is quantified using a natural language model to obtain a second score. Based on the first score and the second score, the target video script is optimized.
10. The method according to claim 9, characterized in that, The target video script is quantified based on preset rule indicators to obtain a first score value, including: The proportion of incorrectly referenced video segments is calculated by determining the number of second video segments referenced from the target video material set in the target video script; The number of video segments that are repeatedly used in the target video script is calculated to obtain the richness of the video segments; The total planned duration of the target video script is calculated against the total target duration in the video generation instruction to obtain the duration deviation value; The first score is obtained by calculating based on the ratio, the richness, and the duration deviation.
11. A video generation method, characterized in that, include: The system obtains a set of target video materials uploaded by the client and a video generation instruction. The set of target video materials includes at least: a first video segment related to the target object, a second video segment unrelated to the target object, and text description information of the target object in the first video segment. In a cloud server, a target video script is generated based on the video generation instructions and the target video material set. The target video script includes visual description information of multiple video segments and auditory description information corresponding to each of the multiple video segments. Based on the target video script, a target video corresponding to the target object is generated. The target video is returned to the client.
12. A video generation apparatus, characterized in that, include: The acquisition unit is used to acquire a target video material set and a video generation instruction, wherein the target video material set includes at least: a first video segment related to the target object, a second video segment unrelated to the target object, and text description information corresponding to the target object in the first video segment; The first generation unit is used to generate a target video script based on the video generation instructions and the target video material set, wherein the target video script includes: visual description information of multiple video segments and auditory description information corresponding to the multiple video segments respectively; The second generation unit is used to generate the target video corresponding to the target object based on the target video script.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein, when the program is executed, it controls the device on which the storage medium is located to perform the video generation method according to any one of claims 1 to 11.
14. An electronic device, characterized in that, include: Memory, which stores executable programs; A processor for running the program, wherein the program, when running, performs the video generation method according to any one of claims 1 to 11.
15. A computer program product, characterized in that, Includes a computer program or instructions that, when executed by a processor, implement the video generation method according to any one of claims 1 to 11.