Video processing method and device, readable medium, electronic equipment and program product

By displaying the editing area in the video editing interface and receiving editing operation information to generate new video clips, the inefficiency problem in generative AI video tools is solved, achieving efficient and accurate video editing and a user-friendly operating experience.

CN121908086APending Publication Date: 2026-04-21BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2026-01-19
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing generative AI video tools are inefficient when modifying videos and struggle to make precise changes. Users need to manually regenerate or trim and splice video clips, which is cumbersome.

Method used

A video processing method is provided that displays an editing area in a video editing interface, receives editing operation information, and generates new video clips based on this information. It supports flexible editing at the clip level and avoids regenerating the entire video.

Benefits of technology

It enables efficient and accurate video editing, simplifies the operation process, improves the user interaction experience, and reduces the waste of computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121908086A_ABST
    Figure CN121908086A_ABST
Patent Text Reader

Abstract

A video processing method and apparatus, a readable medium, an electronic device and a program product, relating to the technical field of computers, in response to an editing operation for a video, displaying the video in a video editing interface, the video comprising at least one video clip, in response to a trigger operation for a first video clip in the at least one video clip, and in response to a trigger operation for a second video clip in the at least one video clip, displaying the first video clip in the video editing interface. A first editing area is displayed, editing operation in the first editing area is received, the editing operation comprises first information, a second video clip is displayed, and the second video clip is generated based on the first information. The video can be flexibly edited and adjusted according to the fragment dimension, the whole video does not need to be regenerated, clipping and splicing do not need to be carried out, operation is simple, efficient and accurate video editing can be achieved, and the interaction experience of a user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technical solution relates to the field of computer technology, specifically to a video processing method, apparatus, readable medium, electronic device, and program product. Background Technology

[0002] Generative AI (Artificial Intelligence) video tools typically employ a "one-shot generation" model, meaning they generate videos based on user input.

[0003] If a user is not satisfied with a segment of the generated video, they usually need to modify the input and regenerate the entire video, or generate a new segment independently and then cut and splice the new segment with the original video. This process is cumbersome, inefficient, and makes it difficult to make precise modifications. Summary of the Invention

[0004] This section is provided to provide a brief overview of the concepts, which will be described in detail in the subsequent Detailed Description section. This section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0005] Firstly, a video processing method is provided, the method comprising: In response to an editing operation on a video, the video is displayed in a video editing interface, the video including at least one video segment; In response to a trigger operation on a first video segment of the at least one video segment, a first editing area is displayed; Receive an editing operation in the first editing area, the editing operation including first information; The second video clip is displayed, which was generated based on the first information.

[0006] Secondly, a video processing apparatus is provided, the apparatus comprising: A first display module is configured to display the video in a video editing interface in response to an editing operation on the video, the video including at least one video segment; The second display module is configured to display a first editing area in response to a trigger operation on the first video segment of the at least one video segment; A receiving module is configured to receive editing operations in the first editing area, wherein the editing operations include first information; The third display module is used to display a second video clip, which is generated based on the first information.

[0007] Thirdly, a computer-readable medium is provided having a computer program stored thereon, which, when executed by a processing device, implements the steps of the method described in the first aspect.

[0008] Fourthly, an electronic device is provided, comprising: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the method described in the first aspect.

[0009] Fifthly, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the steps of the method described in the first aspect.

[0010] The above technical solution, in response to video editing operations, displays a video including at least one video segment in the video editing interface. In response to a trigger operation on a first video segment within the at least one video segment, a first editing area is displayed, and first information corresponding to the editing operation in the first editing area is received. Then, a second video segment generated based on the first information is displayed. Using this method, videos can be flexibly edited and adjusted along the segment dimension without regenerating the entire video or performing cropping and splicing. The operation is simple, enabling efficient and accurate video editing and improving the user's interactive experience.

[0011] Other features and advantages of the above technical solution will be described in detail in the following detailed implementation section. Attached Figure Description

[0012] The above and other features, advantages, and aspects of the technical solution will become more apparent when taken in conjunction with the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale. In the drawings: Figure 1 This is a schematic diagram illustrating the implementation environment based on certain scenarios.

[0013] Figure 2 This is a flowchart illustrating video processing methods under certain circumstances.

[0014] Figure 3 These are schematic diagrams illustrating the video generation and editing interfaces under certain circumstances.

[0015] Figure 4 This is a schematic diagram of a video editing interface shown in other cases.

[0016] Figure 5 This is a schematic diagram of a video editing interface shown under certain circumstances.

[0017] Figure 6 This is a schematic diagram of a video editing interface shown under certain circumstances.

[0018] Figure 7 This is a schematic diagram of the module connection of a video processing device provided according to certain circumstances.

[0019] Figure 8 This is a schematic diagram of the module connections of an electronic device provided under certain circumstances. Detailed Implementation

[0020] The technical solution will now be described in more detail with reference to the accompanying drawings. Although certain scenarios are shown in the drawings, it should be understood that the technical solution can be implemented in various forms and should not be construed as limited to the scenarios described herein. Rather, these scenarios are provided to provide a more thorough and complete understanding of the technical solution. It should be understood that the accompanying drawings and the scenarios described are for illustrative purposes only and are not intended to limit the scope of protection of the technical solution.

[0021] It should be understood that the steps described in the method implementation may be performed in different orders and / or in parallel. Furthermore, the method implementation may include additional steps and / or omit the steps shown. The scope of the technical solution is not limited in this respect.

[0022] The term "comprising" and its variations as used herein can be open-ended, meaning "including but not limited to". The term "based on" can mean "at least partially based on". The term "one case" means "at least one case"; the term "another case" means "at least one additional case"; the term "some cases" means "at least some cases". Definitions of other terms will be given in the following description.

[0023] It should be noted that the concepts of "first" and "second" mentioned here are only used to distinguish different devices, modules or units, and are not used to limit the order of the functions performed by these devices, modules or units or their interdependencies.

[0024] It should be noted that the terms "one" and "more" used here are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0025] The names of messages or information exchanged between the multiple devices in the implementation are for illustrative purposes only and are not intended to limit the scope of these messages or information.

[0026] It is understandable that before using the technical solutions provided here, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in accordance with relevant laws and regulations, and their authorization should be obtained through appropriate means.

[0027] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose, based on the prompt message, whether to provide personal information to the software or hardware such as electronic devices, applications, servers, or storage media performing the operations described herein.

[0028] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose whether to "agree" or "disagree" to provide personal information to the electronic device.

[0029] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of the technical solution. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the technical solution.

[0030] At the same time, it is understood that the data involved in the technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of relevant laws, regulations and related provisions.

[0031] Generative AI video tools typically employ a "one-shot generation" model, meaning they generate a video based on user input. If the video length needs to be changed, the input content, duration, and other information must be manually modified, and the entire video must be regenerated. If the user modifies a narration (causing the audio to lengthen) or extends the video length, the synchronization relationship between subsequent footage and the current footage (such as audio-visual alignment) is highly susceptible to disruption. Modifying a scene within the video (e.g., changing a character or action) requires manually deleting the old segment, generating a new segment, and manually aligning the timeline (video editing track). This process is cumbersome, inefficient, and makes precise modifications difficult.

[0032] In view of this, the technical solution provides a video processing method, apparatus, readable medium, electronic device, and program product to solve the above-mentioned technical problems.

[0033] The video processing method provided by the technical solution can be executed by an electronic device, which can be provided as at least one of a terminal and a server. Figure 1 This is an exemplary schematic diagram illustrating an implementation environment; see [link / reference]. Figure 1The implementation environment includes: terminal 101 and server 102.

[0034] For example, a target application is installed on terminal 101, which is used for video editing. Server 102 is the backend server for the target application, used to provide backend services for the target application.

[0035] For example, after a video segment is extended on the application interface of terminal 101, the extended video segment can be displayed on terminal 101.

[0036] Terminal 101 can be at least one of the following devices: smartphone, smartwatch, desktop computer, laptop, virtual reality terminal, augmented reality terminal, wireless terminal, and laptop computer. Terminal 101 has communication capabilities and can access wired or wireless networks. Terminal 101 can refer to one of multiple terminals, and those skilled in the art will understand that the number of such terminals can be more or less. Server 102 can be an independent physical server, a server cluster composed of multiple physical servers, or a distributed file system. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0037] For example, server 102 and terminal 101 are connected directly or indirectly via wired or wireless communication, without limitation.

[0038] Optionally, the number of servers 102 can be more or less, and there is no limitation thereto. Of course, servers 102 may also include other functional servers to provide more comprehensive and diversified services. Server 102 undertakes the main computing work, and terminal 101 undertakes the secondary computing work; or, server 102 undertakes the secondary computing work, and terminal 101 undertakes the main computing work; or, server 102 or terminal 101 can each undertake computing work independently, and there is no limitation thereto.

[0039] Figure 2 This is a flowchart illustrating a video processing method as shown in the example. Figure 2 As shown, the technical solution provides a video processing method, which can be executed by a video processing device, which can be implemented by software and / or hardware. Figure 2 As shown, the method may include the following steps.

[0040] S201: In response to an editing operation on a video, display the video in the video editing interface, the video including at least one video segment.

[0041] In one scenario, the method further includes: generating at least one video segment in response to fourth information configured in the video generation interface to obtain a video; or, in response to a video import operation, obtaining the video corresponding to the video import operation, segmenting the video, and obtaining at least one video segment.

[0042] For example, such as Figure 3 As shown, a video can be generated in response to input information at the video generation interface. The video can include at least one video segment, such as segment 1, segment 2, etc. Then, in response to editing operations on the video, a video editing interface is displayed. On the video editing track in the video editing interface, multiple video segments are displayed in chronological order. Each video segment can be associated with corresponding video generation information, such as video content description information, subtitle information, duration, etc., without limitation. The video generation information is determined based on the input information; for example, the video generation model determines the video segments to be generated based on the input information, as well as the corresponding video generation information for each video segment.

[0043] Alternatively, in response to a video import operation in the video editing interface, if the video includes at least one video segment generated based on video generation information, each video segment can be associated with the corresponding video generation information. Alternatively, the video processing model can segment the video to determine each video segment and its corresponding description information, allowing for individual processing of each segment later. The specific settings can be configured according to requirements and are not limited in this regard.

[0044] This enables video processing at the segment level for videos obtained through different methods, effectively improving the flexibility and accuracy of video processing.

[0045] In one scenario, a dialog generation area (video generation interface) and a project editing area (video editing interface) can be set up. When a user opens a video file, the dialog generation area can automatically collapse or be converted into a parameter configuration sidebar to avoid information interference. Additionally, every editing operation can be recorded, including track changes, regeneration, moving segments, or changing settings, facilitating subsequent undoing of edits without restriction.

[0046] It should be noted that this technical solution can generate video clips that include audio, video, and subtitles simultaneously, or it can generate audio clips, video clips, and subtitle clips separately and display them on the audio track, video track, and subtitle track respectively. The specific settings can be configured according to requirements, and there are no restrictions on this.

[0047] S202: In response to a trigger operation on a first video segment in at least one video segment, a first editing area is displayed.

[0048] For example, such as Figure 3 As shown, you can process the segments that need adjustment as required without adjusting the entire video.

[0049] S203: Receive an editing operation in the first editing area, the editing operation including first information.

[0050] S204: Display the second video clip, which was generated based on the first information.

[0051] For example, video clips can be extended or regenerated, and video editing operations such as cropping and adding audio are also supported. The video, as well as the extended or regenerated video clips, can be generated by connecting to a video generation agent with a video generation model. The specific video generation model can be selected according to needs, and there are no restrictions on this.

[0052] Using the above method, videos can be flexibly edited and adjusted at the segment level without regenerating the entire video or performing cropping and splicing. The operation is simple, thus enabling efficient and accurate video editing and improving the user's interactive experience.

[0053] In one scenario, in response to a triggering operation on a first video segment in at least one video segment, displaying a first editing area includes: in response to a selection operation on the first video segment, displaying a second editing area corresponding to the first video segment, the second editing area displaying video generation information corresponding to the first video segment, the first editing area including the second editing area; receiving an editing operation in the first editing area includes: receiving a modification operation on the video generation information in the second editing area, the first information including the modified video generation information.

[0054] For example, taking video generation based on user input as an example, multiple video scripts corresponding to multiple video segments can be generated first. Each video script can include the video type, video style, duration, included characters, scenes, storyboard images, and storyboard descriptions for the segment to be generated, etc., for user confirmation. After user confirmation, the corresponding content segment is generated based on each video script. Furthermore, parameters such as the storyboard images, prompts (e.g., storyboard descriptions), and segment duration of the corresponding video segment can be filled into a file such as... Figure 4 The editing area shown.

[0055] For example, such as Figure 4As shown, if segment 2 is selected, the corresponding editing area for segment 2 will be displayed. The parameters within this editing area can then be edited, and segment 2 can be regenerated based on the edited information. This allows users to modify and regenerate unsatisfactory segments according to their needs, without having to regenerate the entire video, thus reducing the waste of computing resources and improving video editing efficiency.

[0056] In one scenario, the second editing area includes a first duration configuration item and / or a first content configuration item, and the modified video generation information includes the duration configured in the first duration configuration item and / or the information configured in the first content configuration item.

[0057] For example, such as Figure 4 As shown, the editing area corresponding to a video clip can include duration configuration item 43 and content configuration items (such as...). Figure 4 The storyboard image configuration item 41 and / or input box 42 shown can be used to modify the video duration and video content of video clips to meet different video modification needs.

[0058] In one scenario, the second editing area includes a first content configuration item, which includes a first image configuration item and a first input box. The first input box displays first descriptive information in natural language, and the first image configuration item displays a first storyboard image corresponding to the first video segment. The first descriptive information and the first storyboard image can generate the first video segment. Receiving modification operations on video generation information in the second editing area includes receiving modification operations on the first storyboard image and / or the first descriptive information, where the first information includes the modified first storyboard image and / or the modified first descriptive information.

[0059] For example, the first content configuration item may include storyboard image configuration item 41 and / or input box 42. Storyboard image configuration item 41 displays the storyboard image of generated segment 2, and input box 42 displays video description information in natural language form for generated segment 2 (e.g., prompts for generating the video segment), which may include subtitles, video action descriptions, etc., and can be set according to needs without limitation. After submitting modifications, the generation progress is displayed through animation at the current position on the timeline. After a new video segment is generated, it automatically replaces the original video segment while maintaining its position. This allows modification of the storyboard image and video description information of video segments to generate new video segments, meeting different video modification needs and improving the accuracy of video modification.

[0060] For example, such as Figure 4As shown, the duration of a video clip can be adjusted by regenerating it. For example, the duration configuration item can be reset, including shortening or lengthening the clip duration. Then, based on the description information in input box 42 and the storyboard image in storyboard image configuration item 41, a video clip with the modified duration can be regenerated and replaced with the original video clip. Other clips after this clip are adaptively moved. For example, if a video clip is extended by 2 seconds, all other clips after this clip will be moved back by 2 seconds.

[0061] Of course, you can also replace the video clip with a video clip of the same length by modifying the description information in input box 42 and / or the storyboard image configuration item 41 without modifying the video clip duration. There are no restrictions on this.

[0062] This enables flexible and efficient video editing of video clips in terms of duration, content, and other dimensions, effectively improving video editing efficiency and enhancing the user experience.

[0063] In addition, input box 42 can be an input box that supports inputting video action descriptions and subtitle content, or it can include an input box for inputting video action descriptions and an input box for inputting subtitle content; there are no restrictions on this.

[0064] In one scenario, receiving a modification operation on the first storyboard image includes: in response to the modification operation on the first storyboard image, displaying a third editing region corresponding to the first storyboard image; in response to a first editing operation in the third editing region, displaying at least one candidate storyboard image, the at least one candidate storyboard image being generated based on image description information corresponding to the first editing operation; and in response to a selection operation on the at least one candidate storyboard image, using the candidate storyboard image corresponding to the selection operation as the modified first storyboard image.

[0065] For example, such as Figure 4 As shown, users can trigger the storyboard image configuration item 41 to modify the storyboard image. They can manually import storyboard images or generate them using AI, and can also generate multiple storyboard images for the user to choose from, etc., without any restrictions. This improves the flexibility and accuracy of the storyboard images, thereby generating video clips that meet the user's needs and improving the efficiency of video clip generation.

[0066] Of course, other parameters can also be set, such as mode, generation model, and segment duration. The mode can include at least one of narration mode, audio-visual synchronization mode, and image-dubbing mode. In narration mode, a narration audio segment is generated based on the subtitles in input box 42, and a video segment is generated based on the storyboard image; the duration of the video segment is controlled by the duration of the narration audio segment. In audio-visual synchronization mode, an audio segment is generated based on the descriptive information in input box 42, and a video segment is generated based on the audio segment and the storyboard image. In image-dubbing mode, a dubbing audio segment is generated based on the descriptive information in input box 42, multiple still images are generated based on the storyboard image, and a video segment is generated based on the still images and the corresponding camera movement. The specific mode can be selected according to requirements and is not limited.

[0067] Different underlying models can be selected according to needs. In addition, for newly created blank segments, the model can be locked once the content is generated to prevent style discrepancies between cross-model generation. Of course, this can also be left unrestricted.

[0068] In addition, the narration tone and other settings corresponding to the video clip can be configured. For example, the video clip's visual content, narration audio, and subtitle files can be generated separately and then merged into a single video clip. Alternatively, the visual content, narration audio, and subtitle files can be displayed independently on the timeline, without any limitations.

[0069] In one scenario, displaying a first editing area in response to a trigger operation on a first video segment in at least one video segment includes: displaying a first blank segment in the first video segment in response to an extension operation on the first video segment in at least one video segment, and displaying a fourth editing area corresponding to the first blank segment on a video editing interface, the first editing area including the fourth editing area; receiving an editing operation in the first editing area includes: receiving a second editing operation in the fourth editing area, the second editing operation including second information, the video content of the second video segment including the video content of the first video segment and the first blank segment, the video content of the first blank segment being generated based on the second information.

[0070] For example, in addition to being able to regenerate fragments, such as Figure 5 As shown, video clips can also be extended. For example... Figure 5 As shown, in response to the extension operation of segment 2, a blank segment 51 and its corresponding editing area are displayed. This allows for the acquisition of corresponding segment generation information based on the editing operations within this area, and then the generation of the blank segment's video content. This eliminates the need to regenerate the entire video or video segment, or to generate separate video segments for trimming and splicing, effectively improving the efficiency and flexibility of video extension.

[0071] In one scenario, the fourth editing area includes a second content configuration item, and the method further includes: in response to a first configuration operation on the second content configuration item, obtaining first configuration information corresponding to the first configuration operation, wherein the first configuration information is video description information in natural language form, and the second information includes the first configuration information.

[0072] For example, the second content configuration item may include, Figure 5 The input box 52 shown allows users to input video description information in natural language, which is then used to generate video content for blank segments, effectively improving the efficiency and flexibility of video extension.

[0073] In one scenario, displaying a first blank segment in a first video segment includes: if the duration of the video segment after the extension operation is greater than the original duration of the first video segment, displaying a first blank segment in the first video segment; the method further includes: if the duration of the video segment after the extension operation is less than or equal to the original duration of the first video segment, displaying the extended first video segment.

[0074] For example, it can support fine-grained editing capabilities for editing and generating extended videos, improving the efficiency and controllability of AI video creation. Editing extended videos: When the current display length of a video clip is less than the display length of its original clip (i.e., it was previously cropped), if the user extends the video clip, for example by dragging the edge of the clip, the hidden original content can be restored without consuming generation computing power. Generating extended videos: For example... Figure 5 As shown, when a video clip has already displayed its original length, and the user wants to extend it, the portion exceeding the original length is displayed as a blank segment 51 to be filled. Then, video content corresponding to the blank segment 51 is generated based on the user's input in the input box 52. A prompt message, such as "What will happen next?", can also be displayed in the blank segment 51 to inform the user that no video content has been generated for that part. This allows for meeting the user's need to extend the video while reducing the waste of computing resources.

[0075] In one scenario, the method further includes: obtaining first video information corresponding to the first video segment; the first video information includes the video content of the first video segment and / or the video description information of the first video segment; and the video content of the first blank segment generated based on the first video information and the second information.

[0076] For example, the first video information may include information from the video script of the aforementioned video segment, the generated video content, etc., so that based on the second information, the video content of the blank segment is generated in combination with the video segment, so that the video content of the blank segment can be better connected with the video content of the original video segment, effectively improving the video quality of the generated video extension and enhancing the user's visual experience.

[0077] In one scenario, in response to an operation to extend the tail of the first video segment, the video content of the first video segment includes at least the last video frame of the first video segment; or, in response to an operation to extend the head of the first video segment, the video content of the first video segment includes at least the first video frame of the first video segment.

[0078] For example, for a video clip, you can extend the video content after the clip or extend the video content before the clip. When extending the video content after the clip, such as... Figure 5 As shown, the video content of the blank segment 51 of segment 2 can be generated based on the description information of segment 2, the last video frame of segment 2, and the information entered in input box 52, ensuring at least a better connection between the newly generated segment content and the last video frame. Alternatively, when extending the video content before the video segment, the description information of the video segment, the first video frame, and the information entered in input box 52 can be used to generate corresponding video content, ensuring at least a better connection between the newly generated segment content and the first video frame.

[0079] For example, after generation, a confirmation dialog box can pop up for the user to confirm whether it is overwrite mode or continuation mode. Overwrite mode: Replace the old video clip (segment 2) with the newly generated long video clip (e.g., clip 2 + the video content of the blank clip). Continuation mode: Keep the old video clip and splice the newly generated part (the video content of the blank clip) as an independent clip after the old video clip. The specific settings can be configured according to needs and there are no restrictions.

[0080] In one scenario, the extension operation includes at least one of the following: dragging the edge of a first video segment to lengthen the segment; or triggering an entry point to a fourth editing area in the video editing interface.

[0081] For example, such as Figure 5 As shown, the length of a video clip can be extended directly on the video editing track, thereby triggering events such as... Figure 5 The editing interface for "Video Extension" is shown. Alternatively, you can also access it from... Figure 4 The "Regenerate" editing interface switches to the "Video Extension" editing interface. After switching, a default extension duration is set, which can be adjusted later as needed. There are no restrictions on the specific adjustments. This effectively improves the flexibility of video editing, thereby enhancing the user's interactive experience.

[0082] In one scenario, the method further includes: moving a third video segment backward based on the distance corresponding to the first blank segment, the third video segment comprising at least one video segment that is later in time sequence than the first video segment.

[0083] For example, after extending a video segment, the video segment following it can be moved forward to the extended length, including video extensions resulting from regeneration or extended generation. Conversely, shortening a video segment can also move the video segment following it forward. This ensures a consistently tight, gap-free, and overlap-free timeline, eliminating the need for manual "gap-filling" or "repositioning" by the user, effectively improving video editing efficiency and accuracy.

[0084] In one scenario, the fourth editing area includes a second duration configuration item, and the method further includes: in response to an operation of dragging the edge of a first video segment to lengthen the segment, displaying the duration corresponding to the operation in the second duration configuration item; and / or, in response to a second configuration operation on the second duration configuration item, displaying a blank segment in the first video segment with a duration corresponding to the second configuration operation.

[0085] For example, by establishing a two-way linkage mechanism between the parameter editing area and the visual timeline, real-time synchronization between changes in video clip duration and generated parameters can be achieved.

[0086] For example, such as Figure 5 As shown, in response to the length extension operation of segment 2, the corresponding duration (which can be retained to two decimal places) is displayed in real time in the duration configuration item. Correspondingly, modifying the duration configuration item to display the corresponding duration will also synchronously adjust the length of the blank segment, realizing a strong correlation between the timeline and the parameter configuration items. Both modification methods will trigger the above-mentioned movement logic, thereby always keeping the timeline compact, without gaps or overlaps, eliminating the need for users to manually "fill in the gaps" or "shift positions," effectively improving video editing efficiency and accuracy.

[0087] In one embodiment, the method further includes: in response to a movement operation on a first content segment, displaying the moved first content segment and simultaneously displaying the moved second content segment, the second content segment being a content segment that is later in time sequence than the first content segment.

[0088] For example, content segments can be video segments, audio segments, or subtitle segments.

[0089] For example, such as Figure 6As shown, taking moving video clips as an example, when a user drags a video clip on the video editing track and it comes into contact with other video clips on the same track, the dragged clip appears as a semi-transparent "ghost" that follows the mouse movement. The overlap ratio between the dragged clip and the target clip can be calculated in real time. If the mouse position at the starting position of the target clip is less than or equal to half the length of the target clip, the dragged clip is inserted before the target clip; otherwise, it is inserted after the target clip. After insertion, the existing subsequent clips on the track automatically shift backward to make room for the new clip, without requiring manual movement from the user. Of course, subsequent clips in the original position of the dragged clip can be moved forward. The specific actions are based on the actual situation, with automatic sorting and adaptive movement to ensure that the clips do not overlap or have gaps; there are no restrictions on this.

[0090] Of course, one or more video clips can be selected to move. If multiple video clips are moved, the relative time difference between the multiple video clips is locked and they are treated as a whole block to be moved. The interaction process is the same as the movement process described above.

[0091] Besides dragging and dropping video clips, the drag-and-drop logic for subtitle and audio clips is similar and will not be elaborated further. Additionally, the specific interactive style of the interface can be customized as needed, without any restrictions. This ensures that when the position of a content clip is changed, other content clips on the same track can adjust accordingly, avoiding overlap or gaps between content clips, reducing user operations, and improving video editing efficiency.

[0092] In one scenario, the video editing page includes a video editing track, which includes a video track, an audio track, and a subtitle track. The method further includes: in response to a movement operation on a third content segment on a first track, synchronously moving a content segment associated with the third content segment on a second track, wherein the first track and the second track are different tracks on the video editing track.

[0093] It's important to note that the video track, audio track, and subtitle track describe how different types of media elements (video, audio, and subtitles) are organized, overlaid, and played synchronously on the timeline of video content. The video track is the timeline channel that carries the video frames, containing all the video frames that change over time. The audio track is the timeline channel that carries the sound content, and the subtitle track is the timeline channel that carries text information, used to display text content related to the audio or video frames, such as dialogue subtitles, explanatory text, titles, keywords, etc.

[0094] For example, taking video clips, audio clips, and subtitle clips as generated and displayed on the video track, audio track, and subtitle track respectively, assuming that video clip 1 is associated with audio clip 1 and subtitle clip 1, then when moving video clip 1, it is also necessary to move the associated audio clip 1 and subtitle clip 1 synchronously to avoid audio-visual asynchrony, reduce user operations, and improve video editing efficiency.

[0095] For example, if a user modifies the narration, causing the TTS (Text-to-Speech) audio to become longer, or extends the duration of a video segment, and the original length of the current segment is detected to be less than the new length, the time difference between the two can be calculated. All content on the same track and related tracks after the current segment can be uniformly moved backward by the distance corresponding to the time difference, preventing content on the same track from overlapping or overlapping, as well as inconsistencies between content on different tracks.

[0096] In other words, ripple editing of the generated video—that is, when shortening, lengthening, deleting, or inserting a segment of content on the timeline—automatically shifts all subsequent content on all tracks forward or backward, thus maintaining a compact, gap-free, and overlap-free timeline. This eliminates the need for manual "filling in the gaps" or "repositioning" by the user, effectively improving video editing efficiency and accuracy. Specifically, it utilizes an intelligent track conflict detection algorithm to achieve automatic track alignment and compression when generated content changes.

[0097] In one embodiment, the method further includes: in response to an operation on adding a new segment to the video, displaying a second blank segment in the video and displaying a fifth editing area corresponding to the second blank segment; receiving a third editing operation in the fifth editing area, the third editing operation including third information; displaying the video content of the second blank segment, the video content of the second blank segment being generated based on the third information; wherein the fifth editing area includes a third duration configuration item, a second input box, a second image configuration item, a model configuration item, and a video type configuration item, wherein the second input box is used to input descriptive information of the first video content, the second image configuration item is used to configure the storyboard image of the first video content, the model configuration item is used to configure the generation model of the first video content, and the video type configuration item is used to configure the video type of the first video content.

[0098] For example, new video clips can also be generated, such as Figure 3 As shown, control 31 can be triggered to add a new video clip at the end of the video, and then display it as shown. Figure 4The editing area for the shown video clip differs in that its input boxes and storyboard images are empty, while other parameters can display default values ​​or remain empty; there are no restrictions on this. Similar to the logic of regeneration, users perform editing operations within the editing area as needed, such as setting video duration, video description information, video generation model, and video type (i.e., the aforementioned video mode), and then generate a new video clip based on the information corresponding to the editing operations. Furthermore, new video clips can also be generated before or after existing clips, without restriction. This allows for the generation of new video clips as needed, improving the efficiency and flexibility of video editing.

[0099] In the aforementioned technical solution, the strong linkage between generated parameters and the timeline allows users to intuitively see the impact of modifying prompts (descriptive information) or duration on the entire AI video stream, achieving a video editing effect that generates content while editing. Through intelligent track compression and conflict handling mechanisms, the timeline structure of the entire video remains intact even when frequently modifying generated parameters, eliminating the need for users to repeatedly manually align tracks. Ghost drag and automatic sorting logic reduce the difficulty of complex narrative editing for users. By distinguishing between editing extended videos and generating extended videos, the flexibility of the original editing is preserved while seamlessly integrating the infinite extensibility of AI, avoiding unnecessary computational consumption due to user errors. Furthermore, when generating extended videos, the last frame and prompts of the previous segment are automatically included, ensuring the visual and logical coherence of the newly generated segment and resolving the issue of discontinuity caused by modifications to AI videos.

[0100] It should be noted that the editing logic for audio clips and subtitle clips is similar to that for video clips, and will not be elaborated upon here.

[0101] Figure 7 This is a schematic diagram of the module connections of a video processing device, provided according to certain situations. For example... Figure 7 As shown, the video processing apparatus 700 may include: A first display module 701 is configured to display the video in a video editing interface in response to an editing operation on the video, the video including at least one video segment; The second display module 702 is configured to display a first editing area in response to a trigger operation on the first video segment of the at least one video segment; The receiving module 703 is configured to receive an editing operation in the first editing area, wherein the editing operation includes first information; The third display module 704 is used to display a second video clip, which is generated based on the first information.

[0102] Optionally, the second display module 702 is used for: In response to the selection operation of the first video segment, a second editing area corresponding to the first video segment is displayed. The second editing area displays the video generation information corresponding to the first video segment. The first editing area includes the second editing area. The receiving module 703 is used for: The system receives modification operations on the video generation information in the second editing area, wherein the first information includes the modified video generation information.

[0103] Optionally, the second editing area includes a first duration configuration item and / or a first content configuration item, and the modified video generation information includes the duration configured in the first duration configuration item and / or the information configured in the first content configuration item.

[0104] Optionally, the second editing area includes a first content configuration item, which includes a first image configuration item and a first input box. The first input box displays first descriptive information in natural language, and the first image configuration item displays a first storyboard image corresponding to the first video clip. The first descriptive information and the first storyboard image can generate the first video clip. The receiving module 703 is used for: The system receives modification operations on the first storyboard image and / or the first description information, wherein the first information includes the modified first storyboard image and / or the modified first description information.

[0105] Optionally, the receiving module 703 is used for: In response to a modification operation on the first storyboard image, a third editing area corresponding to the first storyboard image is displayed; In response to a first editing operation in the third editing area, at least one candidate storyboard image is displayed, the at least one candidate storyboard image being generated based on image description information corresponding to the first editing operation; In response to the selection operation of the at least one candidate storyboard image, the candidate storyboard image corresponding to the selection operation is used as the modified first storyboard image.

[0106] Optionally, the second display module 702 is used for: In response to the extension operation of the first video segment in the at least one video segment, a first blank segment is displayed in the first video segment, and a fourth editing area corresponding to the first blank segment is displayed in the video editing interface, wherein the first editing area includes the fourth editing area; The receiving module 703 is used for: The system receives a second editing operation in the fourth editing area. The second editing operation includes second information. The video content of the second video segment includes the video content of the first video segment and the video content of the first blank segment. The video content of the first blank segment is generated based on the second information.

[0107] Optionally, the fourth editing area includes a second content configuration item, and the video processing device 700 further includes a configuration module for: In response to a first configuration operation on the second content configuration item, first configuration information corresponding to the first configuration operation is obtained. The first configuration information is video description information in natural language form, and the second information includes the first configuration information.

[0108] Optionally, the fourth editing area includes a second content configuration item, and the video processing device 700 further includes an acquisition module for: Obtain the first video information corresponding to the first video segment; the first video information includes the video content of the first video segment and / or the video description information of the first video segment. The video content of the first blank segment generated based on the first video information and the second information.

[0109] Optionally, in response to the tail extension operation of the first video segment, the video content of the first video segment includes at least the last video frame of the first video segment; or, in response to the head extension operation of the first video segment, the video content of the first video segment includes at least the first video frame of the first video segment.

[0110] Optionally, the extension operation includes at least one of the following: The operation of dragging the edge of the first video clip to lengthen the clip; The triggering operation of the fourth editing area entrance in the video editing interface.

[0111] Optionally, the fourth editing area includes a second duration configuration item, and the video processing device 700 further includes a fourth display module for: In response to an operation of dragging the edge of the first video clip to lengthen the clip, the duration corresponding to the operation is displayed in the second duration configuration item; and / or, In response to a second configuration operation on the second duration configuration item, a blank segment corresponding to the duration of the second configuration operation is displayed in the first video segment.

[0112] Optionally, the second display module 702 is used for: If the length of the video after the extension operation is greater than the original length of the first video segment, the first blank segment is displayed in the first video segment; The video processing device 700 further includes a fifth display module, used for: If the length of the video after the extension operation is less than or equal to the original length of the first video segment, the extended first video segment is displayed.

[0113] Optionally, the video processing device 700 further includes a moving module for: Based on the distance corresponding to the first blank segment, a third video segment is moved backward, the third video segment including a video segment that is later in time sequence than the first video segment among the at least one video segment.

[0114] Optionally, the video processing device 700 further includes a new module for: In response to the operation of adding a new segment to the video, a second blank segment is displayed in the video, and a fifth editing area corresponding to the second blank segment is also displayed; Receive a third editing operation in the fifth editing area, the third editing operation including third information; The video content of the second blank segment is displayed, and the video content of the second blank segment is generated based on the third information; The fifth editing area includes a third duration configuration item, a second input box, a second image configuration item, a model configuration item, and a video type configuration item. The second input box is used to input the description information of the first video content, the second image configuration item is used to configure the storyboard images of the first video content, the model configuration item is used to configure the generation model of the first video content, and the video type configuration item is used to configure the video type of the first video content.

[0115] Optionally, the video processing apparatus 700 further includes: The generation module is configured to generate the at least one video segment in response to the fourth information configured in the video generation interface, thereby obtaining the video; or, The import module is used to respond to a video import operation, obtain the video corresponding to the video import operation, segment the video, and obtain the at least one video segment.

[0116] The method logic of each functional module of the aforementioned video processing device 700 has been described in detail in the section on methods, and will not be repeated here.

[0117] Based on the same concept, this technical solution also provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of any of the above-described video processing methods.

[0118] Based on the same concept, this technical solution also provides an electronic device, which may include: A storage device on which computer programs are stored; A processing device for executing a computer program stored in a storage device to implement the steps of any of the above video processing methods.

[0119] Based on the same concept, this technical solution also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described video processing methods.

[0120] The following is for reference. Figure 8 The diagram illustrates a structural schematic of an electronic device 800 suitable for implementing the above-described technical solution. The terminal device may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Personal Computers), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs (Televisions), desktop computers, etc. Figure 8 The electronic device shown is merely an example and should not be construed as limiting its functionality or scope of use.

[0121] like Figure 8 As shown, the electronic device 800 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 into a random access memory (RAM) 803. The RAM 803 also stores various programs and data required for the operation of the electronic device 800. The processing unit 801, the read-only memory 802, and the RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0122] Typically, the following devices can be connected to the input / output interface 805: input devices 806 including, for example, a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; output devices 807 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 808 including, for example, magnetic tape, hard disk, etc.; and communication devices 809. Communication device 809 allows electronic device 800 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 8 An electronic device 800 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0123] In particular, depending on certain circumstances, the process described in the above-referenced flowchart can be implemented as a computer software program. For example, a computer program product is provided, comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowchart. This computer program can be downloaded and installed from a network via a communication device 809, or installed from a storage device 808, or installed from a read-only memory 802. When the computer program is executed by the processing device 801, it performs the functions defined in the above-described method.

[0124] It should be noted that the aforementioned computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM, or flash memory), optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In one case, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In another case, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. The transmitted data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (Radio Frequency), etc., or any suitable combination thereof.

[0125] In some implementations, communication can be conducted using any currently known or future-developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can be interconnected with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (LANs), wide area networks (WANs), the internet (e.g., the Internet), and end-to-end networks (e.g., ad-hoc end-to-end networks), as well as any currently known or future-developed networks.

[0126] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0127] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: display a video in a video editing interface in response to an editing operation on the video, the video including at least one video segment; display a first editing area in response to a triggering operation on a first video segment of the at least one video segment; receive an editing operation in the first editing area, the editing operation including first information; and display a second video segment generated based on the first information.

[0128] Computer program code for performing the above operations can be written in one or more programming languages ​​or a combination thereof. These programming languages ​​include, but are not limited to, object-oriented programming languages, as well as conventional procedural programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0129] The flowcharts and block diagrams in the accompanying figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products under various scenarios. In this respect, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the figures. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0130] The modules mentioned above can be implemented in software or hardware. In some cases, the name of a module does not necessarily limit the functionality of that module.

[0131] The functions described above can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field-Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application-Specific Standard Parts (ASSPs), Systems on Chips (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0132] In this context, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0133] The above description is merely illustrative and explains the technical principles employed. Those skilled in the art should understand that the scope of the technical solution is not limited to specific combinations of the above-described technical features, but also includes other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features provided herein that have similar functions.

[0134] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. Multitasking and parallel processing may be advantageous in certain environments. Similarly, although some specific implementation details are included in the above discussion, these should not be interpreted as limitations on the scope of the technical solution. Certain features described in the context of a single example can also be implemented in combination in a single example. Conversely, various features described in the context of a single example can also be implemented individually or in any suitable sub-combination in multiple examples.

[0135] Although the technical solution has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims. Regarding the aforementioned apparatus, the specific manner in which each module performs its operation has already been described in detail in the section concerning the method, and will not be elaborated upon here.

Claims

1. A video processing method, the method comprising: In response to an editing operation on the video, the video, which includes at least one video segment, is displayed in a video editing interface; In response to a trigger operation on a first video segment of the at least one video segment, a first editing area is displayed; Receive an editing operation in the first editing area, the editing operation including first information; The second video clip is displayed, which is generated based on the first information.

2. The method according to claim 1, wherein displaying the first editing area in response to a trigger operation on the first video segment of the at least one video segment comprises: In response to the selection operation of the first video segment, a second editing area corresponding to the first video segment is displayed. The second editing area displays the video generation information corresponding to the first video segment. The first editing area includes the second editing area. The receiving of editing operations in the first editing area includes: The system receives modification operations on the video generation information in the second editing area, wherein the first information includes the modified video generation information.

3. The method according to claim 2, wherein the second editing area includes a first duration configuration item and / or a first content configuration item, and the modified video generation information includes the duration configured in the first duration configuration item and / or the information configured in the first content configuration item.

4. The method according to claim 2, wherein the second editing area includes a first content configuration item, the first content configuration item includes a first image configuration item and a first input box, the first input box displays first descriptive information in natural language, and the first image configuration item displays a first storyboard image corresponding to the first video segment, wherein, The first description information and the first storyboard image can generate the first video segment; The receiving of modification operations on the video generation information in the second editing area includes: The system receives modification operations on the first storyboard image and / or the first description information, wherein the first information includes the modified first storyboard image and / or the modified first description information.

5. The method according to claim 4, wherein receiving the modification operation on the first storyboard image includes: In response to a modification operation on the first storyboard image, a third editing area corresponding to the first storyboard image is displayed; In response to a first editing operation in the third editing area, at least one candidate storyboard image is displayed, the at least one candidate storyboard image being generated based on image description information corresponding to the first editing operation; In response to the selection operation of the at least one candidate storyboard image, the candidate storyboard image corresponding to the selection operation is used as the modified first storyboard image.

6. The method according to claim 1, wherein displaying the first editing area in response to a trigger operation on the first video segment of the at least one video segment comprises: In response to the extension operation of the first video segment in the at least one video segment, a first blank segment is displayed in the first video segment, and a fourth editing area corresponding to the first blank segment is displayed in the video editing interface, wherein the first editing area includes the fourth editing area; The receiving of editing operations in the first editing area includes: The system receives a second editing operation in the fourth editing area. The second editing operation includes second information. The video content of the second video segment includes the video content of the first video segment and the video content of the first blank segment. The video content of the first blank segment is generated based on the second information.

7. The method according to claim 6, wherein the fourth editing area includes a second content configuration item, and the method further includes: In response to a first configuration operation on the second content configuration item, first configuration information corresponding to the first configuration operation is obtained. The first configuration information is video description information in natural language form, and the second information includes the first configuration information.

8. The method according to claim 6, further comprising: Obtain the first video information corresponding to the first video segment; The first video information includes the video content of the first video segment and / or the video description information of the first video segment; The video content of the first blank segment generated based on the first video information and the second information.

9. The method according to claim 7, In response to the tail-extending operation of the first video segment, the video content of the first video segment includes at least the last video frame of the first video segment; or, In response to the header extension operation of the first video segment, the video content of the first video segment includes at least the first video frame of the first video segment.

10. The method according to any one of claims 6-9, wherein the extension operation comprises at least one of the following: The operation of dragging the edge of the first video clip to lengthen the clip; The triggering operation of the fourth editing area entrance in the video editing interface.

11. The method according to any one of claims 6-9, wherein the fourth editing area includes a second duration configuration item, and the method further includes: In response to an operation of dragging the edge of the first video clip to lengthen the clip, the duration corresponding to the operation is displayed in the second duration configuration item; And / or, In response to a second configuration operation on the second duration configuration item, a blank segment corresponding to the duration of the second configuration operation is displayed in the first video segment.

12. The method according to any one of claims 6-9, wherein displaying a first blank segment in the first video segment comprises: If the length of the video after the extension operation is greater than the original length of the first video segment, the first blank segment is displayed in the first video segment; The method further includes: If the length of the video after the extension operation is less than or equal to the original length of the first video segment, the extended first video segment is displayed.

13. The method according to any one of claims 6-9, further comprising: Based on the distance corresponding to the first blank segment, a third video segment is moved backward, the third video segment including a video segment that is later in time sequence than the first video segment among the at least one video segment.

14. The method according to any one of claims 1-9, further comprising: In response to the operation of adding a new segment to the video, a second blank segment is displayed in the video, and a fifth editing area corresponding to the second blank segment is also displayed; Receive a third editing operation in the fifth editing area, the third editing operation including third information; The video content of the second blank segment is displayed, and the video content of the second blank segment is generated based on the third information; The fifth editing area includes a third duration configuration item, a second input box, a second image configuration item, a model configuration item, and a video type configuration item. The second input box is used to input the description information of the first video content, the second image configuration item is used to configure the storyboard images of the first video content, the model configuration item is used to configure the generation model of the first video content, and the video type configuration item is used to configure the video type of the first video content.

15. The method according to any one of claims 1-9, further comprising: In response to the fourth information configured in the video generation interface, the at least one video segment is generated to obtain the video; or, In response to a video import operation, the video corresponding to the video import operation is obtained, and the video is segmented to obtain at least one video segment.

16. A video processing apparatus, the apparatus comprising: A first display module is configured to display the video in a video editing interface in response to an editing operation on the video, the video including at least one video segment; The second display module is configured to display a first editing area in response to a trigger operation on the first video segment of the at least one video segment; A receiving module is configured to receive editing operations in the first editing area, wherein the editing operations include first information; The third display module is used to display a second video clip, which is generated based on the first information.

17. A computer-readable medium having a computer program stored thereon, characterized in that, When executed by a processing device, the computer program performs the steps of the method described in any one of claims 1-15.

18. An electronic device, characterized in that, include: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the method according to any one of claims 1-15.

19. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-15.