Video production method and device, electronic equipment, storage medium and program product

The method enables user interaction with virtual objects to generate and edit video content, addressing the lack of flexibility in existing video production methods by allowing direct expression and personalization, thereby improving user experience and efficiency.

CN120321450APending Publication Date: 2025-07-15BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510229632.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing video production methods lack flexibility and users are unable to participate in the creative process, resulting in low flexibility in video creation.

Method used

By obtaining the session between the user and the virtual object, the initial video material is generated, and the user is allowed to perform editing operations, including scene adjustment, material addition and deletion, audio adjustment and visual effect adjustment, combined with semantic analysis and AI generation technology, video material that meets user needs is automatically generated.

Benefits of technology

It improves the flexibility and autonomy of video creation, reduces users' time and energy in material creation, ensures the matching of video materials with user needs, and improves user experience and creative efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120321450A_ABST
    Figure CN120321450A_ABST
Patent Text Reader

Abstract

The invention relates to a video production method and device, electronic equipment, a storage medium and a program product, and the method comprises the steps: obtaining a session between a user and a virtual object, and generating a corresponding initial video material based on the session; in response to an editing operation of the user on the initial video material, obtaining an edited video material; and generating a target video file based on the edited video material. The method can improve the flexibility of video production.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a video production method, apparatus, electronic device, storage medium, and program product. Background Art

[0002] With the rapid development of artificial intelligence and multimedia technologies, video creation has gradually moved from the professional field to the public. In related technologies, there are video production methods based on video templates and those based on AI automatic editing. However, the video production method based on video templates is limited by the templates, and in the video production method based on AI automatic editing, users cannot participate in the creation process. Therefore, the existing video production methods lack flexibility. Summary of the Invention

[0003] The present disclosure provides a video production method, apparatus, electronic device, storage medium, and program product to at least solve the problem of low flexibility in video creation in the related video production methods. The technical solutions of the present disclosure are as follows:

[0004] According to a first aspect of the embodiments of the present disclosure, a video production method is provided, including:

[0005] Obtaining a conversation between a user and a virtual object, and generating corresponding initial video materials based on the conversation;

[0006] Responding to an editing operation of the user on the initial video materials to obtain edited video materials;

[0007] Generating a target video file based on the edited video materials.

[0008] In an exemplary embodiment, the method further includes:

[0009] Receiving role information set by the user and scene description information of the video to be produced; the role information includes role information corresponding to the user and role information corresponding to the virtual object;

[0010] Interacting with the user based on the role information and the scene description information to obtain the conversation.

[0011] In an exemplary embodiment, the generating corresponding initial video materials based on the conversation includes:

[0012] Performing semantic analysis on the conversation to obtain an initial video script;

[0013] Generating a plurality of sub-shot scripts according to the initial video script;

[0014] Generate video materials corresponding to each of the sub-shot scripts, and use the video materials corresponding to each of the sub-shot scripts as the initial video materials corresponding to the session.

[0015] In an exemplary embodiment, after performing semantic analysis on the session to obtain an initial video script, the method further includes:

[0016] Performing refinement processing on at least one of the plot, character actions, and scene details of the initial video script to obtain an adjusted video script;

[0017] Generate a plurality of sub-shot scripts according to the adjusted video script.

[0018] In an exemplary embodiment, the generating video materials corresponding to each of the sub-shot scripts includes:

[0019] For each of the sub-shot scripts, generate a character image and / or a scene background that matches the sub-shot script;

[0020] According to the character information and plot information in the sub-shot script, perform dynamic processing on the character corresponding to the character image; and perform visual effect addition operations on the scene background and / or the character to form the video materials corresponding to the sub-shot script.

[0021] In an exemplary embodiment, the method further includes:

[0022] In response to the user's editing operation on the initial video materials, display the corresponding editing effect; the editing operation includes at least one of scene adjustment, addition and deletion of materials, audio adjustment, and visual effect adjustment.

[0023] In an exemplary embodiment, the obtaining the edited video materials in response to the user's editing operation on the initial video materials includes:

[0024] In response to the user's editing operation on the initial video materials, display editing suggestion information and / or display optimization function controls; the editing suggestion information is used to provide operation guidance for the user's current editing operation;

[0025] Based on the editing operation adjusted by the user according to the editing suggestion information, and / or the optimization operation performed on the selection operation of any optimization function in the optimization function controls, obtain the edited video materials.

[0026] In an exemplary embodiment, the method further includes:

[0027] In the case where incoherent video frames are detected in the initial video material, generate transition frames for the incoherent video frames and insert the transition frames between the incoherent video frames; and / or,

[0028] Match corresponding sound effects and background music according to the picture content and plot information of the initial video material to obtain the edited video material.

[0029] According to a second aspect of the embodiments of the present disclosure, there is provided a video production device, including:

[0030] A material generation unit configured to execute obtaining a session between a user and a virtual object and generating corresponding initial video material based on the session;

[0031] A video editing unit configured to execute obtaining the edited video material in response to an editing operation of the user on the initial video material;

[0032] A video generation module configured to execute generating a target video file based on the edited video material.

[0033] In an exemplary embodiment, the device further includes a session acquisition unit configured to execute receiving role information set by the user and scene description information of a video to be produced; the role information includes role information corresponding to the user and role information corresponding to the virtual object; and interact with the user based on the role information and the scene description information to obtain the session.

[0034] In an exemplary embodiment, the material generation unit is further configured to execute performing semantic analysis on the session to obtain an initial video script; generating a plurality of sub-shot scripts according to the initial video script; generating video materials corresponding to each of the sub-shot scripts, and using the video materials corresponding to each of the sub-shot scripts as the initial video material corresponding to the session.

[0035] In an exemplary embodiment, the material generation unit is further configured to execute performing refinement processing on at least one of the plot, character actions, and scene details of the initial video script to obtain an adjusted video script; and generating a plurality of sub-shot scripts according to the adjusted video script.

[0036] In an exemplary embodiment, the material generation unit is further configured to execute, for each of the sub-shot scripts, generating a character image and / or a scene background that matches the sub-shot script; performing dynamic processing on the character corresponding to the character image according to the character information and plot information in the sub-shot script; and performing a visual effect adding operation on the scene background and / or the character to form the video material corresponding to the sub-shot script.

[0037] In an exemplary embodiment, the video editing unit is further configured to perform, in response to an editing operation of the user on the initial video material, presenting a corresponding editing effect; the editing operation includes at least one of scene adjustment, addition or deletion of materials, audio adjustment, and visual effect adjustment.

[0038] In an exemplary embodiment, the video editing unit is further configured to perform, in response to an editing operation of the user on the initial video material, presenting editing suggestion information and / or presenting optimization function controls; the editing suggestion information is used to provide operation guidance for the user's current editing operation; based on the editing operation adjusted by the user according to the editing suggestion information, and / or an optimization operation performed on a selection operation of any one of the optimization functions in the optimization function controls, the edited video material is obtained.

[0039] In an exemplary embodiment, the video editing unit is further configured to perform, when it is detected that there are incoherent video frames in the initial video material, generating transition frames for the incoherent video frames and inserting the transition frames between the incoherent video frames; and / or, matching corresponding sound effects and background music according to the picture content and plot information of the initial video material to obtain the edited video material.

[0040] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including:

[0041] A processor;

[0042] A memory for storing instructions executable by the processor;

[0043] Wherein, the processor is configured to execute the instructions to implement the method described in any one of the above.

[0044] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the method described in any one of the above.

[0045] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product, which includes instructions that, when executed by a processor of an electronic device, enable the electronic device to execute the method described in any one of the above.

[0046] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:

[0047] By obtaining the conversation between the user and the virtual object, the system can understand the user's intentions and needs, breaking the previous fixed and single creation mode. This enables the user to more intuitively express their ideas during the creation process. Automatically generating corresponding initial video materials based on the user conversation reduces the time and effort required by the user for material creation and ensures the matching degree between the video materials and the user's needs. At the same time, the user is allowed to edit the initial video materials to ensure that the user can make sufficient personalized adjustments according to their own needs during the creation process. This flexible editing process makes video creation more in line with the user's style and expression intention. Therefore, through intelligent generation, flexible user editing, and efficient result output, this video production method effectively solves the technical problem of low flexibility in traditional video production, improves the user experience and creation efficiency, and thus makes video creation more flexible and autonomous.

[0048] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.

[0050] Figure 1 is a schematic flowchart of a video production method shown according to an exemplary embodiment.

[0051] Figure 2 is a schematic flowchart of the steps of generating initial video materials based on a conversation shown according to an exemplary embodiment.

[0052] Figure 3 is a schematic flowchart of a video production method shown according to another exemplary embodiment.

[0053] Figure 4 is a schematic flowchart of a video production method shown according to still another exemplary embodiment.

[0054] Figure 5 is a structural block diagram of a video production device shown according to an exemplary embodiment.

[0055] Figure 6 is a block diagram of an electronic device shown according to an exemplary embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0057] It should be noted that the embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims. It should also be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in the present disclosure are all information and data authorized by the user or fully authorized by all parties.

[0058] In an exemplary embodiment, as Figure 1 shown, a video production method is provided. In this embodiment, it is exemplified that the method is applied to a terminal. It can be understood that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. Among them, the terminal can be but is not limited to various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server can be implemented by an independent server or a server cluster composed of multiple servers. In this embodiment, the method includes the following steps:

[0059] In step S110, obtain the session between the user and the virtual object, and generate corresponding initial video materials based on the session.

[0060] In a specific implementation, an interaction system is constructed. The interaction system includes an interaction interface. The user can select one role from multiple roles in the interaction interface as their own role identity, and at the same time select a role identity for the AI (Artificial Intelligence) as the virtual object. Or, the user only selects a role identity for the AI as the virtual object and does not need to select a role for themselves. The user conducts a conversation with other roles played by the AI in the role identity they selected or from the perspective of a player. The interaction system can record the entire conversation process in real time, including information such as conversation content, role lines, and emotional expressions. Thus, a complete conversation can be obtained as the basis for subsequent creation. Among them, the roles provided in the interaction interface can be preset for the user to select, or the user can customize the roles. The role setting can include setting the appearance, personality, and clothing of the role.

[0061] Further, after obtaining the conversation between the user and other roles played by the AI, semantic analysis can be performed on the conversation, and a corresponding video script can be generated based on the analysis results, and then the initial video material corresponding to the video script can be generated. Specifically, to generate the initial video material corresponding to the video script, the character image and / or scene background matching the video script can be generated first, and then, according to the character information and plot information in the video script, the character corresponding to the character image can be dynamicized. For example, the actions and expression changes of the character can be generated to ensure the smoothness and authenticity of the animation; and, visual effects can be added to the scene background and / or the character, thereby obtaining the initial video material.

[0062] In some embodiments, during the process of the user having a conversation with other roles played by the AI, the conversation interface of both parties can be displayed, and the conversation content, character avatars, and emotional states of both parties can be displayed in the conversation interface.

[0063] In some embodiments, the initial video material can be added to the material library, and the generated initial video material can be displayed in the interaction interface, and in response to the user's preview instruction for the initial video material, the preview page of the initial video material can be displayed.

[0064] In step S120, in response to the user's editing operation on the initial video material, the edited video material is obtained.

[0065] In specific implementation, after the initial video material is generated, the user can also edit and adjust it. For example, the user's editing operation on the initial video material can include the adjustment of the scene; adding or deleting materials; the adjustment of the audio; the adjustment of the visual effects, etc.

[0066] In step S130, based on the edited video material, a target video file is generated.

[0067] In specific implementation, after the user completes the editing of the initial video material to obtain the edited video material, the generation and output of the target video file can be performed. Specifically, the output parameters selected by the user are received, including video format, resolution, frame rate, bit rate, etc., and the target video file is rendered and exported according to the output parameters. In this process, acceleration processing can be performed through efficient rendering technology. After the target video file is obtained, it can be saved locally or stored in the cloud for the user to manage and share. For example, the target video file can be shared on social media, video platforms, or sent to others for viewing.

[0068] In some embodiments, during the process of generating the target video file from the edited video material, the interaction interface can also display an export progress bar, and after the export is completed, the user can be prompted that the export is completed and sharing options can be displayed for the user to perform sharing operations.

[0069] In the above video production method, by obtaining the conversation between the user and the virtual object, the system can understand the user's intentions and needs, breaking the previous fixed and single creation mode of templates, enabling the user to more intuitively express their ideas during the creation process. Automatically generating corresponding initial video materials according to the user's conversation reduces the time and effort required by the user for material creation and can ensure the matching degree between the video materials and the user's needs. At the same time, the user is allowed to edit the initial video materials to ensure that the user can make full personalized adjustments according to their own needs during the creation process. This flexible editing process makes video creation more in line with the user's style and expression intention. Therefore, through intelligent generation, flexible user editing, and efficient result output, this video production method effectively solves the technical problem of low flexibility in traditional video production, improves the user experience and creation efficiency, making video creation more flexible and autonomous.

[0070] In an exemplary embodiment, the method further includes: receiving role information set by the user and scene description information of the video to be produced; the role information includes the role information corresponding to the user and the role information corresponding to the virtual object; based on the role information and the scene description information, interacting with the user to obtain a conversation.

[0071] Among them, the scene description information is used to describe the video background and may include location, time, environment, etc.

[0072] In specific implementation, after the interaction system receives that the user selects or customizes the role played by themselves and the role played by the virtual object through the interaction interface, it can construct a personality characteristic model of the role, including appearance, language style, emotional tendency, etc., to ensure that during the conversation between the virtual object and the user, the role's behavior and language conform to the role setting. At the same time, the user can also input the scene description information of the video to be produced in text or voice mode. When the virtual object has a conversation with the user, it can also, based on the scene description information on the basis of understanding the conversation content, make responses and action descriptions that conform to the situation to improve the coherence and naturalness of the conversation and enhance the user's sense of immersion.

[0073] In some embodiments, after the user selects the role information and inputs the scene description information, a scene preview interface corresponding to the role image selected by the user and the scene description information can be displayed on the interaction interface.

[0074] In some embodiments, the interaction system can also support the conversation interaction of multiple roles, and the virtual object can respectively simulate the languages and behaviors of different roles, so as to enrich the plot content and support complex role relationships and interactions.

[0075] In this embodiment, the virtual object can have a conversation with the user based on the role information set by the user and the scene description information of the video to be produced. This can not only make the role behavior and language of the virtual object conform to the role setting, but also ensure that the responses and actions it makes conform to the context through the context understanding of the conversation and the scene description information, improve the coherence and naturalness of the conversation, and enhance the user's sense of immersion.

[0076] In one exemplary embodiment, as Figure 2 shown, generating the corresponding initial video material based on the conversation in step S110 includes:

[0077] Step S111, performing semantic analysis on the conversation to obtain an initial video script;

[0078] Step S112, generating multiple sub-shot scripts according to the initial video script;

[0079] Step S113, generating the video material corresponding to each sub-shot script, and using the video material corresponding to each sub-shot script as the initial video material corresponding to the conversation.

[0080] In specific implementation, natural language processing technology can be used to perform semantic analysis on the conversation between the user and the virtual object, and generate an initial video script based on the semantic analysis result. Automatically design sub-shots according to the initial video script, determine the shot type, angle, scene size, movement, etc., and thus obtain multiple sub-shot scripts. Use each sub-shot script as the video script corresponding to the conversation between the user and the virtual object. Further, for each sub-shot script, generate the video material corresponding to the sub-shot script. Specifically, to generate the video material corresponding to the sub-shot script, first generate a character image and / or scene background that matches the sub-shot script, and then, according to the character information and plot information in the sub-shot script, perform dynamic processing on the character corresponding to the character image. For example, generate the action and expression changes of the character to ensure the fluency and authenticity of the animation; and perform visual effect addition operations on the scene background and / or the character, thereby obtaining the video material corresponding to the sub-shot script. Combine the video materials corresponding to each sub-shot script to obtain the initial video material.

[0081] In some embodiments, after the design of the sub-shot script is completed, a sub-shot preview interface can be displayed on the interaction interface to show the thumbnails and descriptions of each sub-shot.

[0082] In this embodiment, a video script is automatically generated according to the user's conversation, which reduces the time and effort required by the user in script creation and ensures the matching degree between the script content and the user's needs. This intelligent generation enables the creator to quickly obtain a suitable script, thus focusing on other creative aspects of the video. Further, by automatically performing shot design and generating a shot script, the difficulty of manual design by the user can be reduced, and the design efficiency of the shots can be improved.

[0083] In an exemplary embodiment, after performing semantic analysis on the conversation in step S111 to obtain an initial video script, the method further includes: improving at least one of the plot, character actions, and scene details of the initial video script to obtain an adjusted video script.

[0084] Correspondingly, step S112 further includes: generating a plurality of shot scripts according to the adjusted video script.

[0085] In a specific implementation, after using natural language processing technology to perform semantic analysis on the conversation content to obtain an initial video script, the initial video script can also be adjusted based on the semantic analysis results. For example, improving the plot description, supplementing action and scene details, etc., to generate a more complete and expressive script.

[0086] In this embodiment, by adjusting the initial video script to generate a more complete and expressive script, a good foundation is provided for subsequent production.

[0087] In an exemplary embodiment, generating the video materials corresponding to each shot script in the above step S113 includes: for each shot script, generating a character image and / or a scene background that matches the shot script; dynamically processing the character corresponding to the character image according to the character information and plot information in the shot script; and performing an operation of adding visual effects to the scene background and / or the character to form the video materials corresponding to the shot script.

[0088] Among them, the dynamic processing means generating changes in the actions and expressions of the character to make the actions and expressions of the character coherent and smooth.

[0089] Among them, the visual effects may include special effects, filters, subtitles, stickers, etc.

[0090] In specific implementation, deep generation technologies such as diffusion models can be utilized to generate corresponding character images and scene backgrounds respectively according to each sub-shot script, and further combine character information and plot information to generate the action and expression changes of the characters, so as to ensure the smoothness and authenticity of the animation. Also, necessary visual effects or filters can be added according to the plot requirements. For example, the parts that need to add special effects are identified, and large model technologies are used to generate smoke, fire, water flow, etc.; or, special effects and filters can be recommended according to the plot information for the user to choose, so as to enrich the visual effects and meet personalized needs. For subtitles, the interaction system can automatically generate subtitle content, add stickers, etc., to assist in information transmission, improve the readability and viewing experience of the video, increase the interest, and form video scripts corresponding to each sub-shot script with the animated characters and the scene backgrounds and characters added with visual effects.

[0091] In this embodiment, by automatically generating high-quality image and video materials such as character images and scene backgrounds that meet the requirements, the dependence on external materials and manpower can be reduced; by combining character settings and plot requirements, the natural action and expression changes of the characters are generated to ensure the smoothness and authenticity of the animation, create vivid character images, and enhance the viewing experience of the video; by using AI to identify the parts that need to add special effects and using large model technologies to generate complex special effects such as smoke, fire, and water flow, and being able to automatically generate subtitles, add filters and stickers, the visual effects of the video can be enriched and the plot expressiveness can be enhanced.

[0092] In an exemplary embodiment, the method further includes: in response to a user's editing operation on the initial video material, displaying the corresponding editing effect; the editing operation includes at least one of scene adjustment, addition and deletion of materials, audio adjustment, and visual effect adjustment.

[0093] Specifically, the user can perform multi-track timeline editing. Specifically, in the video editor, the video, audio, visual effects, etc. can be edited. Among them, the editor has function modules such as multi-tracks, timeline, and material library. The multi-track timeline displays each media track, improving the flexibility of editing. For scene adjustment, for example, dragging to adjust the order of video clips, modifying the duration of the scene, and adjusting the shot transition. Adding / deleting materials, that is, new video or picture materials can be inserted, and the unnecessary parts can be deleted, and AI is supported to generate new materials. Audio processing, for example, adjusting the speaking speed and volume of the character lines, adding background music and sound effects. Visual effect processing, for example, applying transition effects, filters, subtitles, etc. to enhance the visual effects. After the editing operation, the edited video material is obtained. At the same time, during the user's editing process, the video preview window can be updated according to the editing operation to display the editing effect.

[0094] In this embodiment, during the user's editing process, the video preview window can be updated according to the editing operation to display the editing effect, so that the user can timely see the result of the editing operation and improve the editing efficiency.

[0095] In an exemplary embodiment, step S120 above responds to the user's editing operation on the initial video material to obtain the edited video material, including: responding to the user's editing operation on the initial video material, displaying editing suggestion information and / or displaying optimization function controls; the editing suggestion information is used to provide operation guidance for the user's current editing operation; based on the editing operation adjusted by the user according to the editing suggestion information, and / or the optimization operation executed for the selection operation of any optimization function in the optimization function controls, the edited video material is obtained.

[0096] In specific implementation, during the user's editing of the initial video material, the interaction system can also display editing suggestion information according to the user's editing content to provide editing suggestions, such as adjusting the rhythm, adding close-ups, optimizing transitions, color correction, etc. The user can adopt the editing suggestion information to make corresponding adjustments to optimize the editing effect. On the other hand, during the user's editing of the initial video material, the interaction system can also display at least one optimization function control, where the optimization function control can be an optimization function control for parameters such as the color, brightness, contrast, or stability of the video. After receiving the user's selection operation on any optimization function control, the optimization operation corresponding to the selected optimization function control is automatically executed. For example, if the user's selection operation on the color optimization control is received, the intelligent execution of the color optimization process of the video is performed to improve the picture quality.

[0097] In this embodiment, on the one hand, by displaying editing suggestion information based on the user's editing operation on the initial video material for the user to adjust the video editing to assist the user in improving the video quality; on the other hand, by displaying optimization function controls with one-key optimization functions for the user to select, the user can more quickly realize the editing adjustment of the video, save editing time and effort, and improve the editing efficiency.

[0098] In an exemplary embodiment, the method further includes: in the case of detecting that there are incoherent video frames in the initial video material, generating transition frames for the incoherent video frames and inserting the transition frames between the incoherent video frames; and / or, matching corresponding sound effects and background music according to the picture content and plot information of the initial video material to obtain the edited video material.

[0099] In specific implementation, when the user performs editing operations on the initial video material, the interaction system can also perform intelligent frame interpolation processing on the initial video material. Specifically, it can detect the coherence of video frames in the initial video material. If incoherent video frames are detected, for example, missing frames or inconsistent character movements, it uses generation techniques such as diffusion models to generate at least one transitional frame between the incoherent video frames, and inserts the at least one transitional frame between the incoherent video frames to improve the smoothness of the picture. In addition, the interaction system can also automatically match and adjust the sound effects and background music according to the picture content and plot information of the initial video material.

[0100] In some embodiments, before automatically matching and adjusting the sound effects and background music, an inquiry window can also be displayed. The inquiry window is used to inquire whether the user accepts the automatic matching and adjustment operations of the music and background music. If a confirmation instruction from the user is received, the sound effects and background music can be automatically matched and adjusted according to the picture content and plot information of the initial video material.

[0101] In this embodiment, through the coherence detection operation of the initial video material and performing intelligent frame interpolation operations in the case of incoherent video frames, the smoothness of the video picture can be improved; through the picture content and plot information of the initial video material, intelligent matching of sound effects and background music is performed to enhance the auditory effect and improve the overall quality of the video, and the editing efficiency of the initial video material can also be improved.

[0102] In another exemplary embodiment, as Figure 3 shown, is a flowchart of another video production method shown according to an exemplary embodiment. In this embodiment, the method includes the following steps:

[0103] Step S310, receiving the role information set by the user and the scene description information of the video to be produced; the role information includes the role information corresponding to the user and the role information corresponding to the virtual object;

[0104] Step S320, interacting with the user based on the role information and the scene description information to obtain a conversation;

[0105] Step S330, performing semantic analysis on the conversation to obtain an initial video script;

[0106] Step S340, performing adjustment processing on the initial video script to obtain an adjusted video script; the adjustment processing is used to perfect the video script;

[0107] Step S350, generating multiple sub-shot scripts according to the adjusted video script;

[0108] Step S360, generating a character image and / or a scene background that matches each sub-shot script;

[0109] Step S370: For each sub-shot script, perform dynamic processing on the characters corresponding to the character images according to the character information and plot information in the sub-shot script; and perform visual effect addition operations on the scene background and / or characters to form the initial video material for each sub-shot script.

[0110] Step S380: In response to the user's editing operation on the initial video material, obtain the edited video material and display the corresponding editing effect in real time; the editing operation includes at least one of scene adjustment, addition and deletion of materials, audio adjustment, and visual effect adjustment.

[0111] Step S390: Generate a target video file based on the edited video material.

[0112] The video production method provided in this embodiment allows the user to participate more actively in video creation by receiving the character information and scene description set by the user and interacting with the user to obtain a conversation. This interaction not only enhances the user experience but also makes the final video more in line with the user's expectations. The video script generated through semantic analysis of the conversation is adjusted to ensure that it is more logical and meets the narrative requirements. This intelligent processing reduces the burden on the user in script creation while improving the quality of the script. Dividing the video script into multiple sub-shot scripts makes the video creation process more structured, which also lays a foundation for subsequent dynamic processing of characters and scenes and addition of visual effects; and performing dynamic processing and adding visual effects to the character images and scene backgrounds makes the video of higher quality and more visually appealing. Responding to the user's editing operation on the initial video material and displaying the editing effect in real time enables the user to immediately see the effect during the creation process, making it easier to make adjustments. Such immediate feedback can promote the user's better participation in creation and improve the creation efficiency. During the editing process, the user can perform scene adjustment, addition and deletion of materials, audio and visual effect adjustment, ensuring the personalization and flexibility of the video, so that each user can create a video work that conforms to their own style.

[0113] In another exemplary embodiment, an interactive system for implementing the video production method is further provided. The system may include a character and scene management module, a conversation interaction module, an AI content generation module, a video editing module, an AI-assisted optimization module, and a rendering and publishing module. Among them,

[0114] The character and scene management module includes a character setting and management function (including handling the creation, editing, and attribute setting of characters), a scene description processing function (including managing the input and processing of scene information), and a preview generator function (for real-time display of character and scene effects).

[0115] The session interaction module includes a real-time session system (for supporting real-time sessions between users and AI characters), a session recorder (for recording and managing session content), and an emotion recognizer (for analyzing emotional features in the session).

[0116] The AI content generation module includes a script generator (for optimizing session content and generating a complete script), a storyboard planner (for automatically designing storyboard details), a material generation engine (for creating visual content using deep generation models), an action and expression generator (for processing character actions and expressions), and a special effect generator (for creating video special effects and animations).

[0117] The video editing module includes a multi-track editor (for managing video, audio, and special effect tracks), a timeline controller (for handling the temporal relationships of video clips), a material manager (for organizing and managing media resources), and a real-time preview system (for providing instant previews of editing effects).

[0118] The AI-assisted optimization module includes an intelligent suggestion system (for providing editing suggestions), a video optimizer (for processing picture quality optimization), a frame interpolation engine (for generating transitional frames), and an intelligent audio effect system (for handling audio matching and optimization).

[0119] The rendering and publishing module includes a rendering engine (for processing the rendering of the final video), a format converter (for handling different format outputs), and a publishing manager (for handling the storage and sharing of works).

[0120] The process of implementing the video production method through the above-mentioned various modules is as Figure 4 shown, and includes the following steps:

[0121] (1) Character setting and scene description

[0122] Input: Character information selected or customized by the user, scene setting.

[0123] Processing process:

[0124] 1. Character setting: The user selects a preset character or customizes a character through the interface, including the appearance, personality, clothing, etc. of the character.

[0125] 2. Scene description: The user describes the scene in text or voice, including the location, time, environment, etc.

[0126] Output result: Complete character information and scene description, serving as the basis for script and material generation.

[0127] Page display change: The interface displays the character image and scene preview selected by the user.

[0128] (2) The user conducts a session with the AI character

[0129] Input: Character setting and scene description.

[0130] Processing process:

[0131] 1. Conversation interaction: The user plays one of the characters and has a conversation with other characters played by the AI.

[0132] 2. Conversation record: The system records the conversation content in real time, including information such as character lines and emotional expressions.

[0133] Output result: A complete conversation script, including character lines and interaction plots.

[0134] Page display change: The conversation interface displays the conversation content of both sides, character avatars and emotional states.

[0135] (3)AI generates scripts and storyboards

[0136] Input: Conversation script and scene description.

[0137] Processing process:

[0138] 1. Script optimization: The AI performs semantic analysis on the conversation content, improves the plot description, and supplements necessary action and scene details.

[0139] 2. Storyboard generation: The AI automatically generates storyboards based on the script content, determining camera angles, shot sizes, movements, etc.

[0140] Output result: A refined script and the corresponding storyboard script.

[0141] Page display change: The storyboard preview interface shows thumbnails and descriptions of each storyboard.

[0142] (4)AI generates video materials

[0143] Input: Storyboard script.

[0144] Processing process:

[0145] 1. Deep generation models generate characters and scenes: Using deep generation technologies such as diffusion models, generate character images, actions, and scene backgrounds that meet the requirements of the storyboard.

[0146] 2. Action and expression generation: Combine character settings and plot requirements to generate changes in characters' actions and expressions.

[0147] 3. Special effects and animation addition: Generate necessary special effects and animation elements according to the plot needs.

[0148] Output result: Initial video materials corresponding to each storyboard.

[0149] Page display change: The generated materials are shown in the material library, and users can preview them.

[0150] (5) Users edit and adjust in the video editor.

[0151] Input: Initial video materials generated by AI.

[0152] Processing process:

[0153] 1. Multi-track timeline editing: Users edit videos, audios, special effects, etc. in a professional video editor. The editor has function modules such as multi-tracks, timeline, and material library.

[0154] 2. Scene adjustment: Drag and drop to adjust the order of video clips, modify the duration of the scene, and adjust the lens switching.

[0155] 3. Adding / deleting materials: Insert new video or picture materials, delete unnecessary parts, and support AI to generate new materials.

[0156] 4. Audio processing: Adjust the speech rate and volume of the character lines, and add background music and sound effects.

[0157] 5. Special effect processing: Apply transition effects, filters, subtitles, etc. to enhance the visual effect.

[0158] Output result: The video project file perfected by the user.

[0159] Page display change: The video preview window is updated in real time to show the editing effect; the multi-track timeline shows each media track, and users can operate flexibly.

[0160] (6) AI-assisted optimization

[0161] Input: The video project file being edited by the user.

[0162] Processing process:

[0163] 1. Intelligent suggestions: AI provides editing suggestions according to the editing content, such as adjusting the rhythm, adding close-ups, optimizing transitions, etc.

[0164] 2. Automatic optimization: One-key optimize parameters such as the color, brightness, contrast, and stability of the video to improve the picture quality.

[0165] 3. Intelligent frame interpolation: For parts with discontinuous actions or missing frames, use technologies such as diffusion models to generate transition frames to improve the picture fluency.

[0166] 4. Intelligent sound effect matching: AI automatically matches and adjusts sound effects and background music according to the picture content and plot.

[0167] Output result: Optimized video project file.

[0168] Page display changes: The positions where AI suggestions are marked in the editor, and users can choose to accept or ignore them; the preview window shows the optimization effect.

[0169] (7) Work output and publication

[0170] Input: The finally edited video project file.

[0171] Processing process:

[0172] 1. Output settings: Users select parameters such as video format, resolution, frame rate, bit rate, etc.

[0173] 2. Video rendering: The system starts to render and export the video file, and uses efficient rendering technology to accelerate the processing.

[0174] 3. Work saving: Support local saving or cloud storage, which is convenient for users to manage and share.

[0175] 4. Sharing and publication: Users can choose to share the work on social media, video platforms or send it to others for viewing.

[0176] Output result: The final video file for users to watch, share and spread.

[0177] Page display changes: The work export progress bar, and after the export is completed, it prompts the user and provides sharing options.

[0178] Through the interactive content generation method provided in this embodiment, the following effects can be achieved: 1. Lower the creation threshold: Through human-computer interaction and AI assistance, users can create high-quality video content without complex professional skills. 2. Improve user participation and creation experience: Users deeply participate in the creation process, cooperate with AI to complete the script and video production, and increase the fun of creation. 3. Provide powerful AI assistance functions: AI provides all-round support in aspects such as script generation, material production, and editing optimization, improving the creation efficiency. 4. Implement professional-level editing functions: Have professional functions such as multi-track timeline, real-time preview, and AI intelligent suggestions to meet the advanced needs of users. 5. Support multi-modal content: Through large model technology, support the generation and editing of various media forms such as text, image, audio, and video, enriching the creation content. 6. Meet personalized needs: Users can freely edit and customize the work according to their personal preferences to create unique videos.

[0179] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the indications of the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0180] It can be understood that the same / similar parts among the various embodiments of the above method in this specification can be referred to each other. Each embodiment focuses on the differences from other embodiments. For the relevant parts, refer to the descriptions of other method embodiments.

[0181] Based on the same inventive concept, the embodiments of the present disclosure also provide a video production device for implementing the above-mentioned video production method.

[0182] Figure 5 It is a structural block diagram of a video production device shown according to an exemplary embodiment. Referring to Figure 5 , the device includes:

[0183] A material generation unit 510, configured to execute obtaining a session between a user and a virtual object, and generating corresponding initial video materials based on the session;

[0184] A video editing unit 520, configured to execute obtaining edited video materials in response to a user's editing operation on the initial video materials;

[0185] A video generation module 530, configured to execute generating a target video file based on the edited video materials.

[0186] In an exemplary embodiment, the device further includes a session acquisition unit, configured to execute receiving role information set by the user and scene description information of the video to be produced; the role information includes role information corresponding to the user and role information corresponding to the virtual object; interacting with the user based on the role information and the scene description information to obtain a session.

[0187] In an exemplary embodiment, the material generation unit 510 is further configured to perform semantic analysis on the session to obtain an initial video script; generate a plurality of sub-shot scripts according to the initial video script; generate video materials corresponding to each of the sub-shot scripts, and use the video materials corresponding to each of the sub-shot scripts as the initial video materials corresponding to the session.

[0188] In an exemplary embodiment, the material generation unit 510 is further configured to perform refinement processing on at least one of the plot, character actions, and scene details of the initial video script to obtain an adjusted video script; generate a plurality of sub-shot scripts according to the adjusted video script.

[0189] In an exemplary embodiment, the material generation unit 510 is further configured to, for each of the sub-shot scripts, generate a character image and / or a scene background that matches the sub-shot script; perform dynamic processing on the character corresponding to the character image according to the character information and plot information in the sub-shot script; and perform a visual effect addition operation on the scene background and / or the character to form the video material corresponding to the sub-shot script.

[0190] In an exemplary embodiment, the video editing unit 520 is further configured to perform, in response to a user's editing operation on the initial video material, display the corresponding editing effect; the editing operation includes at least one of scene adjustment, addition and deletion of materials, audio adjustment, and visual effect adjustment.

[0191] In an exemplary embodiment, the video editing unit 520 is further configured to perform, in response to a user's editing operation on the initial video material, display editing suggestion information and / or display optimization function controls; the editing suggestion information is used to provide operation guidance for the user's current editing operation; perform an optimization operation based on the editing operation adjusted by the user according to the editing suggestion information and / or the selection operation of any one of the optimization functions in the optimization function controls to obtain the edited video material.

[0192] In an exemplary embodiment, the video editing unit 520 is further configured to, when it is detected that there are incoherent video frames in the initial video material, generate transition frames for the incoherent video frames and insert the transition frames between the incoherent video frames; and / or match corresponding sound effects and background music according to the picture content and plot information of the initial video material to obtain the edited video material.

[0193] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0194] Figure 6FIG. 0 is a block diagram of an electronic device 600 for implementing a video production method according to an exemplary embodiment. For example, the electronic device 600 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0195] Referring to Figure 6 , the electronic device 600 may include one or more of the following components: a processing component 602, a memory 604, a power component 606, a multimedia component 608, an audio component 610, an input / output (I / O) interface 612, a sensor component 614, and a communication component 616.

[0196] The processing component 602 generally controls the overall operation of the electronic device 600, such as operations associated with display, telephone calls, data communications, camera operations, and recording operations. The processing component 602 may include one or more processors 620 to execute instructions to complete all or part of the steps of the above-described method. In addition, the processing component 602 may include one or more modules to facilitate the interaction between the processing component 602 and other components. For example, the processing component 602 may include a multimedia module to facilitate the interaction between the multimedia component 608 and the processing component 602.

[0197] The memory 604 is configured to store various types of data to support the operation of the electronic device 600. Examples of such data include instructions for any application or method operating on the electronic device 600, contact data, phone book data, messages, pictures, videos, etc. The memory 604 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, optical disk, or graphene memory.

[0198] The power component 606 provides power to the various components of the electronic device 600. The power component 606 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 600.

[0199] The multimedia component 608 includes a screen that provides an output interface between the electronic device 600 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of the touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 608 includes a front camera and / or a rear camera. When the electronic device 600 is in an operation mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.

[0200] The audio component 610 is configured to output and / or input audio signals. For example, the audio component 610 includes a microphone (MIC), which is configured to receive external audio signals when the electronic device 600 is in an operation mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 604 or transmitted via the communication component 616. In some embodiments, the audio component 610 further includes a speaker for outputting audio signals.

[0201] The I / O interface 612 provides an interface between the processing component 602 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power button, and a lock button.

[0202] The sensor component 614 includes one or more sensors for providing status assessments of various aspects of the electronic device 600. For example, the sensor component 614 can detect the on / off state of the electronic device 600, the relative positioning of components, such as the display and keypad of the electronic device 600. The sensor component 614 can also detect a change in the position of the electronic device 600 or an electronic device 600 component, the presence or absence of user contact with the electronic device 600, the orientation or acceleration / deceleration of the device 600, and a change in the temperature of the electronic device 600. The sensor component 614 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 614 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 614 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0203] The communication component 616 is configured to facilitate communication between the electronic device 600 and other devices in a wired or wireless manner. The electronic device 600 can access a communication standard-based wireless network, such as WiFi, a carrier network (such as 2G, 3G, 4G, or 5G), or a combination thereof. In an exemplary embodiment, the communication component 616 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 616 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0204] In an exemplary embodiment, the electronic device 600 can be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.

[0205] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory 604 including instructions, and the above instructions can be executed by a processor 620 of the electronic device 600 to complete the above method. For example, the computer-readable storage medium can be a ROM, Random Access Memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0206] In an exemplary embodiment, a computer program product is also provided, and the computer program product includes instructions that can be executed by a processor 620 of the electronic device 600 to complete the above method.

[0207] It should be noted that the above-mentioned device, electronic device, computer-readable storage medium, computer program product, etc. may also include other implementation manners according to the description of the method embodiments. The specific implementation manners can refer to the description of the relevant method embodiments and will not be elaborated herein one by one.

[0208] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and embodiments are only considered exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.

[0209] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A video production method, characterized in that, including: Obtain a session between a user and a virtual object, and generate corresponding initial video material based on the session; In response to the user's editing operation on the initial video material, obtain the edited video material; Generate a target video file based on the edited video material.

2. The method according to claim 1, wherein The method further includes: Receive the role information set by the user and the scene description information of the video to be produced; the role information includes the role information corresponding to the user and the role information corresponding to the virtual object; Interact with the user based on the role information and the scene description information to obtain the session.

3. The method according to claim 1, wherein The generating corresponding initial video material based on the session includes: Perform semantic analysis on the session to obtain an initial video script; Generate multiple sub-shot scripts according to the initial video script; Generate video material corresponding to each sub-shot script, and use the video material corresponding to each sub-shot script as the initial video material corresponding to the session.

4. The method according to claim 3, wherein After performing semantic analysis on the session to obtain an initial video script, it further includes: Perform improvement processing on at least one of the plot, character actions, and scene details of the initial video script to obtain an adjusted video script; Generate multiple sub-shot scripts according to the adjusted video script.

5. The method according to claim 3, wherein The generating video material corresponding to each sub-shot script includes: For each sub-shot script, generate a character image and / or a scene background that matches the sub-shot script; Dynamically process the character corresponding to the character image according to the character information and the plot information in the sub-shot script; and perform a visual effect addition operation on the scene background and / or the character to form the video material corresponding to the sub-shot script.

6. The method according to claim 1, wherein The method further includes: In response to the user's editing operation on the initial video material, display the corresponding editing effect; the editing operation includes at least one of scene adjustment, addition and deletion of materials, audio adjustment, and visual effect adjustment.

7. The method according to claim 1, wherein The obtaining the edited video material in response to the user's editing operation on the initial video material includes: In response to the user's editing operation on the initial video material, display editing suggestion information and / or display optimization function controls; the editing suggestion information is used to provide operation guidance for the user's current editing operation; Obtain the edited video material based on the editing operation adjusted by the user according to the editing suggestion information and / or the optimization operation performed on the selection operation of any optimization function in the optimization function controls.

8. The method according to claim 7, wherein The method further includes: In the case of detecting incoherent video frames in the initial video material, generate transition frames for the incoherent video frames and insert the transition frames between the incoherent video frames; and / or Match corresponding sound effects and background music according to the picture content and plot information of the initial video material to obtain the edited video material.

9. A video production device, characterized in that, including: A material generation unit configured to execute obtaining a session between a user and a virtual object and generating corresponding initial video material based on the session; A video editing unit, configured to perform an editing operation on the initial video material in response to the user, and obtain the edited video material; A video generation module, configured to perform generating a target video file based on the edited video material.

10. An electronic device, characterized in that, Comprising: A processor; A memory for storing executable instructions of the processor; Wherein, the processor is configured to execute the instructions to implement the video production method according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is enabled to execute the video production method according to any one of claims 1 to 8.

12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the video production method according to any one of claims 1 to 8.

Citation Information

Cited By

  • Video generation method and device, equipment and medium

    CN120769107A

  • Digital human interaction method and system based on multi-mode sensing intelligent action switching

    CN121050590A