A video processing method, apparatus, device, and storage medium

The video processing method organizes editing interfaces by script nodes to align multimedia content with script structures, addressing the complexity of script-based video editing and enhancing editing efficiency and accuracy.

CN115811632BActive Publication Date: 2025-07-15BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111081785.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-15
Publication Date
2025-07-15
Estimated Expiration
2041-09-15

AI Technical Summary

Technical Problem

The prior art is difficult to meet people's demand for diverse video editing, especially when the editing process related to language content is complicated and cannot achieve fast and convenient editing operations.

Method used

By displaying the material editing area of the video clip according to the first script structure, the material editing area is divided into multiple sub-regions, each sub-region corresponds to the script node, and the target multimedia material is displayed according to the time axis track in the target sub-region, and the target video is finally generated.

Benefits of technology

It realizes the convenience and efficiency of video editing, reduces the complexity of language content-related editing processing, improves the efficiency and accuracy of video processing, and meets the diverse video editing needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115811632B_ABST
    Figure CN115811632B_ABST
Patent Text Reader

Abstract

The present disclosure provides a video processing method, apparatus, device, and storage medium. The method includes: displaying a material editing area for video editing according to a first script structure; wherein, the material editing area is divided into multiple sub-areas, one sub-area corresponds to one script node in the first script structure, the first script structure indicates the content paragraph structure of the target video, and one script node indicates one content paragraph of the target video; in the target sub-area among the multiple sub-areas, displaying the target multimedia material according to the time axis track; generating the target video according to the multimedia material displayed in the material editing area; wherein, the target multimedia material is filled in the target content paragraph of the target video. It can be seen that the embodiments of the present disclosure can implement video editing based on the material editing area including multiple sub-areas corresponding to script nodes, enrich the video processing methods, and further meet the diverse video editing needs of people.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a video processing method, apparatus, device, and storage medium. Background Art

[0002] With the development of computer technologies, videos are increasingly widely used in work and life, and people's demands for video editing are also becoming more diverse.

[0003] Therefore, how to meet people's diverse demands for video editing is a technical problem that urgently needs to be solved at present. Summary of the Invention

[0004] To solve the above technical problem or at least partially solve the above technical problem, embodiments of the present disclosure provide a video processing method, apparatus, device, and storage medium.

[0005] In a first aspect, the present disclosure provides a video processing method, and the method includes:

[0006] Display a material editing area for video editing according to a first script structure; wherein, the material editing area is divided into multiple sub-areas, one of the sub-areas corresponds to one script node in the first script structure, the first script structure is used to indicate the content paragraph structure of a target video, and one script node is used to indicate one content paragraph of the target video;

[0007] In a target sub-area among the multiple sub-areas, display target multimedia materials according to a time-axis track; wherein, the target multimedia materials are multimedia materials selected for a target script node, and the target script node is the script node in the first script structure corresponding to the target sub-area;

[0008] Generate the target video according to the multimedia materials displayed in the material editing area; wherein, the target multimedia materials are filled in a target content paragraph of the target video, and the target content paragraph corresponds to the target script node.

[0009] In an optional implementation manner, the interface layout manner of the multiple sub-areas in the material editing area is vertical alignment arrangement.

[0010] In an optional implementation manner, before generating the target video according to the multimedia materials displayed in the material editing area, it further includes:

[0011] In response to an adjustment operation on the target text content of the first script node in the first script structure, determine the multimedia material corresponding to the first script node in the material editing area, and determine the multimedia segment in the multimedia material corresponding to the target text content;

[0012] According to the adjustment operation, clip the multimedia segment in the multimedia material.

[0013] In an optional implementation manner, before generating the target video according to the multimedia material displayed in the material editing area, it further includes:

[0014] In response to an operation of adding text content to the target text position of the second script node in the first script structure, determine the multimedia material corresponding to the second script node in the material editing area, and determine the timeline position in the multimedia material corresponding to the target text position;

[0015] According to the operation of adding text content, add a multimedia segment corresponding to the text content at the timeline position in the multimedia material.

[0016] In an optional implementation manner, before generating the target video according to the multimedia material displayed in the material editing area, it further includes:

[0017] In response to a clipping operation on the target multimedia segment of the first multimedia material in the material editing area, determine the script node corresponding to the first multimedia material, and determine the text content in the script node corresponding to the target multimedia segment;

[0018] According to the clipping operation, adjust the text content in the script node.

[0019] In an optional implementation manner, before generating the target video according to the multimedia material displayed in the material editing area, it further includes:

[0020] In response to an operation of adjusting the order between the second script node and the third script node in the first script structure, determine the sub-areas respectively corresponding to the second script node and the third script node in the material editing area;

[0021] According to the order adjustment operation, adjust the order of the multimedia materials in the sub-areas respectively corresponding to the second script node and the third script node in the material editing area.

[0022] In an alternative embodiment, the target multimedia material has alternative multimedia materials. Before generating the target video according to the multimedia materials displayed in the material editing area, it further includes:

[0023] In response to a switching operation on the target multimedia material and the alternative multimedia material in the target sub-region, switch the target multimedia material displayed in the target sub-region to the alternative multimedia material.

[0024] In a second aspect, the present disclosure also provides a video processing device, the device includes:

[0025] A first display module, configured to display a material editing area of a video clip according to a first script structure; wherein, the material editing area is divided into multiple sub-regions, one of the sub-regions corresponds to one script node in the first script structure, the first script structure is used to indicate the content paragraph structure of the target video, and one script node is used to indicate one content paragraph of the target video;

[0026] A second display module, configured to display a target multimedia material according to a timeline track in a target sub-region among the multiple sub-regions; wherein, the target multimedia material is a multimedia material selected for a target script node, and the target script node is the script node corresponding to the target sub-region in the first script structure;

[0027] A generation module, configured to generate the target video according to the multimedia materials displayed in the material editing area; wherein, the target multimedia material is filled in a target content paragraph of the target video, and the target content paragraph corresponds to the target script node.

[0028] In a third aspect, the present disclosure provides a computer-readable storage medium, in which instructions are stored. When the instructions run on a terminal device, the terminal device is enabled to implement the above method.

[0029] In a fourth aspect, the present disclosure provides a device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the above method is implemented.

[0030] In a fifth aspect, the present disclosure provides a computer program product, the computer program product includes computer programs / instructions, and when the computer programs / instructions are executed by a processor, the above method is implemented.

[0031] The technical solutions provided by the embodiments of the present disclosure have at least the following advantages compared with the prior art:

[0032] An embodiment of the present disclosure provides a video processing method, which displays a material editing area for video clips according to a first script structure, so that sub-areas in the material editing area correspond to script nodes in the first script structure. Additionally, in a target sub-area of the material editing area, multimedia materials selected for a target script node corresponding to the target sub-area are displayed according to a timeline track. Furthermore, a target video is generated according to the multimedia materials displayed in the material editing area. The embodiments of the present disclosure can implement video editing based on a material editing area including multiple sub-areas corresponding to script nodes, enrich the video processing methods, and further meet the diverse video editing needs of people. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure and, together with the specification, are used to explain the principles of the present disclosure.

[0034] To more clearly illustrate the technical solutions in the embodiments of the present disclosure or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0035] Figure 1 It is a schematic flowchart of a video processing method provided by an embodiment of the present disclosure;

[0036] Figure 2 It is a schematic diagram of the relationship among a script node, a content paragraph, and a sub-area provided by an embodiment of the present disclosure;

[0037] Figure 3a It is a schematic diagram of the alignment method of a material editing area provided by an embodiment of the present disclosure;

[0038] Figure 3b It is another schematic diagram of the alignment method of a material editing area provided by an embodiment of the present disclosure;

[0039] Figure 4a It is a schematic diagram of a material editing area and a first script structure provided by an embodiment of the present disclosure;

[0040] Figure 4b It is another schematic diagram of a material editing area and a first script structure provided by an embodiment of the present disclosure;

[0041] Figure 4c It is yet another schematic diagram of a material editing area and a first script structure provided by an embodiment of the present disclosure;

[0042] Figure 5Schematic diagram of the relationship among a target script node, a target sub-region, a target content paragraph, and a target multimedia material provided by an embodiment of the present disclosure;

[0043] Figure 6a Schematic diagram of the display of a target multimedia material provided by an embodiment of the present disclosure;

[0044] Figure 6b Another schematic diagram of the display of a target multimedia material provided by an embodiment of the present disclosure;

[0045] Figure 7 Schematic diagram of generating a target video provided by an embodiment of the present disclosure;

[0046] Figure 8 Schematic diagram of switching between a target multimedia material and an alternative multimedia material provided by an embodiment of the present disclosure;

[0047] Figure 9 Schematic diagram of the structure of a video processing device provided by an embodiment of the present disclosure;

[0048] Figure 10 Schematic diagram of the structure of a video processing device provided by an embodiment of the present disclosure. Detailed implementation manners

[0049] In order to more clearly understand the above objects, features, and advantages of the present disclosure, the solutions of the present disclosure will be further described below. It should be noted that, without conflict, the embodiments of the present disclosure and the features in the embodiments may be combined with each other.

[0050] Many specific details are set forth in the following description in order to provide a thorough understanding of the present disclosure, but the present disclosure may be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of the present disclosure, rather than all the embodiments.

[0051] To meet the diverse needs of users for video editing and enrich the video processing methods, embodiments of the present disclosure propose a video processing method. First, according to the first script structure, a material editing area for video editing is displayed. Among them, the material editing area is divided into multiple sub-areas, and one sub-area corresponds to one script node in the first script structure. The first script structure indicates the content paragraph structure of the target video, and one script node indicates one content paragraph of the target video. Then, in the target sub-area among the multiple sub-areas, the target multimedia material is displayed according to the time axis track. Among them, the target multimedia material is the multimedia material selected for the target script node, and the target script node is the script node in the first script structure corresponding to the target sub-area. Furthermore, according to the multimedia material displayed in the material editing area, the target video is generated. Among them, the target multimedia material is filled in the target content paragraph of the target video, and the target content paragraph corresponds to the target script node.

[0052] It can be seen that embodiments of the present disclosure display the material editing area for video editing according to the first script structure, so that the sub-areas in the material editing area correspond to the script nodes in the first script structure. In addition, in the target sub-area of the material editing area, the multimedia material selected for the target script node corresponding to the target sub-area is displayed according to the time axis track. Furthermore, according to the multimedia material displayed in the material editing area, the target video is generated. Embodiments of the present disclosure can realize video editing based on the material editing area including multiple sub-areas corresponding to script nodes, enrich the video processing methods, and further meet the diverse video editing needs of people.

[0053] Based on this, embodiments of the present disclosure provide a video processing method, refer to Figure 1 , which is a schematic flowchart of a video processing method provided by embodiments of the present disclosure. This method can be executed by a video processing device, where the device can be implemented by software and / or hardware and is generally integrated in an electronic device.

[0054] As Figure 1 shown, this method may include:

[0055] Step 101, display the material editing area for video editing according to the first script structure.

[0056] Among them, the material editing area is divided into multiple sub-areas, one sub-area corresponds to one script node in the first script structure, the first script structure is used to indicate the content paragraph structure of the target video, and one script node is used to indicate one content paragraph of the target video.

[0057] A script is the draft during the process of film and television creation. Usually, the script includes descriptions of multiple scenes to guide the shooter to shoot and generate corresponding film and television works. For example, the script includes the description of scene a for indicating the shooting of the first shot, and also includes the description of scene b for indicating the shooting of the second shot, etc. When conducting film and television creation, the shooter can shoot the first shot containing video clip A according to the description of scene a, and shoot the second shot containing video clip B according to the description of scene b, and then splice the second shot after the first shot to obtain the film and television work corresponding to the script.

[0058] In this embodiment, the first script structure can refer to the structure of the above script, such as the description content paragraph structure, etc. It can be understood that the description contents corresponding to the first shot and the second shot included in the above example script respectively correspond to the script nodes in the first script structure. For example, the description content of the first shot corresponds to the first script node in the first script structure, and the description content of the second shot corresponds to the second script node in the first script structure.

[0059] In this embodiment, after determining the first script structure, the material editing area for video editing is displayed according to the first script structure, and the multimedia materials to be edited can be displayed in this material editing area. Specifically, the material editing area can be divided according to the script nodes in the first script structure to obtain sub-areas respectively corresponding to each script node. Among them, each sub-area corresponds to a script node in the first script structure.

[0060] To more vividly understand the relationship between the sub-areas and the script nodes in the embodiments of the present disclosure, it can be described in combination with Figure 2 the content shown. Among them, Figure 2 the first script structure in includes Q script nodes, where Q is a positive integer. Based on the first script structure including Q script nodes, the material editing area can be divided into Q sub-areas, and each script node corresponds to a sub-area in the material editing area.

[0061] It should be noted that there are various interface layout methods for the multiple sub-areas in this material editing area, which can be selected according to needs, and this embodiment does not make any restrictions. For example, as Figure 3a shown in the vertical alignment arrangement, that is, different rows are vertically aligned. Figure 3a In, each sub-area is left-aligned, or each sub-area can also be right-aligned. In addition, as Figure 3b shown in the horizontal alignment arrangement, that is, different columns are horizontally aligned, Figure 3b in, each sub-area is upward-aligned, or each sub-area can also be downward-aligned.

[0062] It should be noted that in an optional implementation, the script nodes in the first script structure may include: script annotations and / or script paragraphs in the script, that is, there is a corresponding relationship between the script nodes and the script annotations and / or script paragraphs in the script. Among them, the script annotation is used to generally represent the content of the multimedia material corresponding to the script node, and the script paragraph includes the detailed text content corresponding to the script node. In an optional implementation, the detailed text content included in the script paragraph may be the text information obtained by performing speech recognition on the video.

[0063] Specifically, an example of the material editing area for displaying video clips according to the first script structure is described as follows:

[0064] Example 1: The script nodes in the first script structure include script annotations. As Figure 4a shown, assume that the first script structure includes a first script annotation and a second script annotation. Among them, the first script annotation is " / / Opening remarks", and the second script annotation is " / / Environment introduction". In the material editing area displayed according to this first script structure, the first sub-area of the material editing area corresponds horizontally to " / / Opening remarks", and the second sub-area of the material editing area corresponds horizontally to " / / Environment introduction".

[0065] Example 2: The script nodes in the first script structure include script paragraphs. As Figure 4b shown, assume that the first script structure includes a first script paragraph and a second script paragraph. Among them, the first script paragraph is the opening remarks text obtained by speech recognition, and the second script paragraph is the environment introduction text obtained by speech recognition. In the material editing area displayed according to the first script structure, the first sub-area of the material editing area corresponds horizontally to the opening remarks text script paragraph, and the second sub-area of the material editing area corresponds horizontally to the environment introduction text script paragraph.

[0066] Example 3: The script nodes in the first script structure include script annotations and script paragraphs. As Figure 4c shown, assume that the first script structure includes a first script node and a second script node. Among them, the first script node includes a first script annotation and a first script paragraph. The first script annotation is " / / Opening remarks", and the first script paragraph is the opening remarks text obtained by speech recognition. The second script node includes a second script annotation and a second script paragraph. The second script annotation is " / / Environment introduction", and the second script paragraph is the environment introduction text obtained by speech recognition. The first sub-area in the material editing area displayed according to this first script structure corresponds horizontally to both the opening remarks text script paragraph and the first script annotation in the first script structure, and the second sub-area in the material editing area corresponds horizontally to both the environment introduction text script paragraph and the second script annotation in the first script structure.

[0067] After presenting the material editing area according to the first script structure, continue to execute the following step 102.

[0068] Step 102, in the target sub-region among multiple sub-regions, present the target multimedia material according to the timeline track.

[0069] Among them, the target multimedia material is the multimedia material selected for the target script node, and the target script node is the script node in the first script structure corresponding to the target sub-region.

[0070] In this embodiment, the target sub-region can be any one of the multiple sub-regions in the material editing area. There is a corresponding target script node for this target sub-region in the first script structure, and the corresponding target multimedia material can be selected based on this target script node.

[0071] In an alternative implementation, after obtaining the target multimedia material imported by the user, voice recognition can be performed on the target multimedia material, and the voice recognition result can be textually matched with each script node in the first script structure to determine the target script node corresponding to the target multimedia material. Then, in the target sub-region corresponding to the target script node, the target multimedia material is presented according to the timeline track. The target multimedia material in the embodiments of the present disclosure can be a whole video obtained by shooting, or a clip in a whole video obtained by shooting. This embodiment does not make any restrictions.

[0072] For ease of understanding, refer to Figure 5 , select the target multimedia material according to the target script node. This target node also corresponds to the target sub-region, and thus the target multimedia material presented in the target sub-region can be determined.

[0073] In an alternative implementation, as Figure 6a shown, assume that the first sub-region in the material editing area is the target sub-region. First, determine that the target script node corresponding to the first sub-region in the first script structure is " / / Opening remarks", then select the multimedia material according to the target script node as the target multimedia material, and then present the target multimedia material in the first sub-region.

[0074] Among them, the method of selecting the multimedia material according to the target script node can include determining the multimedia material with the highest matching degree with the " / / Opening remarks" script node as the target multimedia material by performing image recognition and / or voice recognition on the candidate multimedia materials, etc., and presenting the target multimedia material in the first sub-region.

[0075] In another alternative implementation, as Figure 6bAs shown in the figure, assume that the first sub-region in the material editing area is the target sub-region. The multimedia material includes an opening video material. Perform speech recognition on the opening material to obtain the corresponding opening text, and use this opening text as the target script node in the first script structure. Select the target multimedia material with the highest matching degree in the multimedia material according to this opening text. For example, select the opening video material as the target multimedia material, and then display the opening video material in the first sub-region according to the time axis track.

[0076] Step 103: Generate a target video according to the multimedia material displayed in the material editing area.

[0077] Among them, the target multimedia material is filled in the target content paragraph of the target video, and the target content paragraph corresponds to the target script node.

[0078] In this embodiment, after the multimedia material is displayed in each sub-region, a target video can be generated according to the multimedia material displayed in the material editing area.

[0079] For the sake of easy understanding, refer to Figure 5 , the target video includes Q content paragraphs, where Q is a positive integer. Each content paragraph has a corresponding relationship with a script node in the first script structure. Each content paragraph can be filled with the multimedia material selected by the script node corresponding to this content paragraph. The multimedia material includes, but is not limited to, any one or more of video and audio.

[0080] In this embodiment, the first script structure is used to indicate the content paragraph structure of the target video. Specifically, a script node in the first script structure is used to indicate a content paragraph of the target video, that is, the content paragraph corresponding to the script node meets the requirements of this script node. Therefore, the content paragraph can be adjusted according to the script node in the first script structure, so as to generate a target video that meets the first script structure.

[0081] Continuing with Figure 5 as an example, the first script structure includes Q script nodes, where Q is a positive integer, and each script node has a corresponding content paragraph. According to the target script node, the corresponding relationship among the target sub-region, the target content paragraph, and the target multimedia material can be determined. Furthermore, fill the corresponding target multimedia material in each content paragraph, and splice each content paragraph according to the first script structure, so as to obtain the corresponding target video.

[0082] In an alternative implementation manner, as Figure 7As shown, the first sub-region of the material editing area displays the opening video material, which includes n frames. The second sub-region of the material editing area displays the environment introduction material, which includes m frames. Here, n and m are positive integers. According to the first script structure, the first sub-region is determined to correspond to the first content paragraph of the target video, and the second sub-region corresponds to the second content paragraph of the target video. Thus, the n-frame opening video material is used to fill the first content paragraph, and the m-frame environment introduction video material is used to fill the second content paragraph, thereby generating the target video.

[0083] In summary, the video processing method of the present disclosure embodiment displays the material editing area for video editing according to the first script structure, making the sub-regions in the material editing area correspond to the script nodes in the first script structure. Additionally, in the target sub-region of the material editing area, the multimedia materials selected for the target script node corresponding to the target sub-region are displayed according to the time axis track. Furthermore, according to the multimedia materials displayed in the material editing area, the target video is generated. The present disclosure embodiment can realize the editing of the video based on the material editing area including multiple sub-regions corresponding to the script nodes, enriching the video processing methods and further meeting the diverse video editing needs of people.

[0084] Generally, a video work is generated by editing multiple sub-videos. During the editing process, it is necessary to edit the sub-videos according to the time axis corresponding to the sub-videos and splice the sub-videos according to the time axis corresponding to the total video. However, this time-axis-based editing method is complex in operation when performing editing processing related to language content, and it is necessary to repeatedly compare the content of each frame in the time axis of the sub-videos. Therefore, this technical solution cannot achieve fast and convenient editing operations on the video. Thus, the video editing operation can be realized based on the above embodiments. Specifically, before generating the target video according to the multimedia materials displayed in the material editing area, corresponding operation steps can be added according to requirements. The example is as follows:

[0085] In an optional implementation manner, since there are slips of the tongue or other situations in the multimedia materials and corresponding segments in the multimedia materials need to be cut off, the steps that need to be added before step 103 of the above embodiment include:

[0086] First, in response to the adjustment operation for the target text content of the first script node in the first script structure, determine the multimedia material corresponding to the first script node in the material editing area, and determine the multimedia segment corresponding to the target text content in the multimedia material.

[0087] In this example, the first script node in the first script structure has a corresponding multimedia material, and the first script node is the text content corresponding to the multimedia material, and the text content is obtained by: obtaining text information according to speech recognition technology, manually configured subtitles, etc. The user can adjust the target text content in the first script node according to the needs, and in response to the adjustment, the multimedia material corresponding to the first script node is determined in the material editing area, and in order to determine the content that needs to be adjusted, it is also necessary to determine the multimedia segment corresponding to the target text content in the multimedia material.

[0088] Furthermore, according to the adjustment operation, the multimedia segments in the multimedia material are edited, wherein the editing includes but is not limited to: deletion, shifting, etc.

[0089] For example, the target text content is "Good morning and good afternoon", and the multimedia material is a greeting video. The correspondence between the target text content and the multimedia material is: "早" corresponds to the 1st frame of the greeting video, "上" corresponds to the 2nd frame of the greeting video, "中" corresponds to the 3rd frame of the greeting video, "午" corresponds to the 4th frame of the greeting video, and "好" corresponds to the 5th frame of the greeting video. In this example, "早上" is a slip of the tongue, and the corresponding segment needs to be deleted in the target video. Therefore, the target text content can be operated on to delete the "早上" in "Good morning and good afternoon", and the corresponding 1st and 2nd frames of the greeting video will also be deleted.

[0090] In another example, the target text content can establish a correspondence with the multimedia material through a timestamp. The timestamp can establish a connection between the text content and the timeline of the multimedia material. Specifically, assuming that the text content is "Good morning and good afternoon", the target text content is "Morning", and the multimedia material is a greeting video, the correspondence between the text content and the multimedia material is: "Morning" corresponds to the 0th to 1.5th seconds of the multimedia material, "Noon" corresponds to the 1.5th to 3rd seconds of the multimedia material, and "Good" corresponds to the 3rd to 4th seconds of the multimedia material. If the target text content is operated and "Morning" is deleted from "Good morning and good afternoon", the corresponding greeting video from the 0th to 1.5th seconds will also be deleted.

[0091] In this implementation, by operating the first script structure, the troublesome operation of manually positioning the text to be processed on the timeline of the multimedia material is avoided, thereby improving the efficiency and accuracy of video processing.

[0092] In another optional implementation, if it is necessary to add a multimedia segment to the target video, the steps that need to be added before step 103 of the above embodiment include:

[0093] First, in response to the operation of adding text content to the target text position of the second script node in the first script structure, the multimedia material corresponding to the second script node is determined in the material editing area, and the timeline position corresponding to the target text position in the multimedia material is determined.

[0094] In this example, the second script node in the first script structure has a corresponding multimedia material, and the second script node is the text content corresponding to the multimedia material. There are multiple ways to obtain the text content, including: text information obtained by recognizing speech recognition technology, manually configured subtitles, etc. The user can add text content to the target text position of the second script node as required, and in response to the adjustment, the multimedia material corresponding to the second script node is determined in the material editing area, and in order to determine the position where the multimedia segment needs to be added, the timeline position corresponding to the target text position in the multimedia material needs to be determined.

[0095] Furthermore, according to the operation of adding text content, a multimedia segment corresponding to the text content is added to the timeline position in the multimedia material.

[0096] In an optional implementation, the multimedia material before the timeline position can be determined as the front multimedia material, and the multimedia material after the timeline position can be determined as the rear multimedia material, and the adding operation can be to connect the multimedia segment after the front multimedia material, and to connect the rear multimedia material after the multimedia segment. By operating the first script structure, the troublesome operation of manually positioning the text to be added on the timeline of the multimedia material is avoided, and the efficiency and accuracy of video processing are improved.

[0097] For example, the second script node is "Hello everyone", the multimedia material is a greeting video, and the correspondence between the target text content and the multimedia material is: "Big" corresponds to the first frame of the greeting video, "Home" corresponds to the second frame of the greeting video, and "Good" corresponds to the third frame of the greeting video. In this example, it is necessary to add "Noon" between "Home" and "Good". In response to the operation of adding "Noon" in the second script node, the video clip corresponding to "Noon" is obtained, including the first frame of the noon video and the second frame of the noon video. Therefore, the first frame and the second frame of the noon video are connected after the second frame of the greeting video, and the third frame of the greeting video is connected after the second frame of the noon video.

[0098] In another optional implementation, when a multimedia material is edited, the script node corresponding to the multimedia material will also change accordingly. In this application scenario, the steps that need to be added before step 103 of the above embodiment include:

[0099] First, in response to a clipping operation on a target multimedia segment of a first multimedia material in a material editing area, determine a script node corresponding to the first multimedia material, and determine the text content in the script node corresponding to the target multimedia segment. Further, adjust the text content in the script node according to the clipping operation.

[0100] In this example, if the user performs a clipping operation on a target multimedia segment in the first multimedia material in the material editing area, in response to this operation, corresponding operations need to be performed on the first script structure. Therefore, it is necessary to determine the script node corresponding to the first multimedia material, and determine the text content in the script node corresponding to the target multimedia segment. Then, adjust the text content in the script node accordingly according to the clipping operation on the first multimedia material.

[0101] For example, the first multimedia material is a greeting video, and the correspondence between the text content in the script node and the greeting video is as follows: "Morning" corresponds to the first frame of the greeting video, "Up" corresponds to the second frame of the greeting video, "Middle" corresponds to the third frame of the greeting video, "Noon" corresponds to the fourth frame of the greeting video, and "Good" corresponds to the fifth frame of the greeting video. In this example, the third and fourth frames of the greeting video are deleted. According to the deletion operations on the third and fourth frames of the greeting video, "Middle" and "Noon" are deleted from the text content of the script node accordingly. The processed script node is "Good morning". Thus, the changes in the multimedia material and the corresponding script node are unified, maintaining the consistency between the multimedia material and the script node.

[0102] In another alternative implementation manner, if it is possible to adjust the order of multimedia materials based on the first script structure, the steps that need to be added before step 103 in the above embodiment include:

[0103] First, in response to an order adjustment operation between a second script node and a third script node in the first script structure, determine sub-regions corresponding to the second script node and the third script node respectively in the material editing area. Further, adjust the order of the multimedia materials in the sub-regions corresponding to the second script node and the third script node respectively in the material editing area according to the order adjustment operation.

[0104] When the user needs to adjust the order of multimedia materials, the second script node and the third script node in the first script structure can be adjusted. Moreover, there is a corresponding second sub-region for the second script node, and a corresponding third sub-region for the third script node. In response to this adjustment, the second sub-region and the third sub-region are determined in the material editing area, and the second sub-region and the third sub-region are adjusted according to the user's adjustment of the script structure. In this example, by adjusting the first script structure, the order of multimedia materials can be adjusted, which improves the efficiency of video processing. At the same time, it also eliminates the step of manually viewing multimedia materials to determine the content of multimedia materials, making video processing more intuitive.

[0105] For example, in this example, the second script node in the first script structure is " / / Opening remarks", and the third script node is " / / Video of environment introduction". And " / / Opening remarks" is located after " / / Video of environment introduction". In the corresponding material editing area, the opening remarks material is located after the environment introduction material. If the user needs to move the opening remarks video before the environment introduction video, the " / / Opening remarks" in the first script structure can be moved before " / / Video of environment introduction". In response to this operation of the user, the opening remarks material in the material editing area is moved before the environment introduction material.

[0106] In another alternative implementation, in order to improve the quality of the generated target video during shooting, multiple videos of similar types are shot. Therefore, the target multimedia material has alternative multimedia materials. Then, the one with the best effect is selected from the alternative multimedia materials. The steps that need to be added before step 103 in the above embodiment include:

[0107] In response to the switching operation of the target multimedia material in the target sub-region and the alternative multimedia materials, the target multimedia material displayed in the target sub-region is switched to the alternative multimedia material.

[0108] In this example, the alternative multimedia materials can be set by the user, or obtained by comparing the similarity between the target multimedia material and the alternative multimedia materials through image recognition and speech recognition technologies. The user can switch the target multimedia material to the alternative multimedia material. In response to this switching operation, the target multimedia material displayed in the material editing area is switched to the alternative multimedia material. It should be noted that the target script node corresponding to the target sub-region in the first script structure can be adjusted to the text information corresponding to the alternative multimedia material according to the alternative multimedia material. Through the alternative operation, the most suitable one can be conveniently and quickly selected from multiple multimedia materials, which improves the efficiency of video processing.

[0109] For example, in this example, as Figure 8As shown, the target sub-region is the first sub-region, the target multimedia material in the first sub-region is the target opening material, the alternative multimedia materials are the first alternative opening material and the second alternative opening material, and an alternative display control is further included in the material alternative region. The alternative display control will display the alternative multimedia materials in the first sub-region in response to the user's touch operation. In this example, the user touches the alternative display control and selects the second alternative opening material to switch with the target multimedia material, and then the second alternative opening material is displayed in the first sub-region.

[0110] In summary, the video processing method of the embodiments of the present disclosure, based on the correspondence relationship between the sub-regions, content paragraphs, and multimedia materials established by the first script structure, can intuitively and conveniently adjust the target video and / or the first script structure, while reducing the complexity of editing and processing videos with language content or plot as the core, and improving the video processing efficiency.

[0111] Based on the above method embodiments, the present disclosure further provides a video processing device. Refer to Figure 9 , which is a schematic structural diagram of a video processing device provided by the embodiments of the present disclosure. The device includes:

[0112] A first display module 901, configured to display a material editing area for video editing according to the first script structure; wherein, the material editing area is divided into multiple sub-regions, and one of the sub-regions corresponds to one script node in the first script structure. The first script structure is used to indicate the content paragraph structure of the target video, and one script node is used to indicate one content paragraph of the target video;

[0113] A second display module 902, configured to display target multimedia materials in a target sub-region among the multiple sub-regions according to a time axis track; wherein, the target multimedia materials are multimedia materials selected for a target script node, and the target script node is the script node in the first script structure corresponding to the target sub-region;

[0114] A generation module 903, configured to generate the target video according to the multimedia materials displayed in the material editing area; wherein, the target multimedia materials are filled in a target content paragraph of the target video, and the target content paragraph corresponds to the target script node.

[0115] In an optional implementation manner, the interface layout manner of the multiple sub-regions in the material editing area is vertically aligned arrangement.

[0116] In an optional implementation manner, the device further includes:

[0117] A first determination module, configured to, in response to an adjustment operation on the target text content of a first script node in the first script structure, determine the multimedia material corresponding to the first script node in the material editing area, and determine the multimedia segment corresponding to the target text content in the multimedia material;

[0118] A clipping module, configured to clip the multimedia segment in the multimedia material according to the adjustment operation.

[0119] In an optional implementation manner, the apparatus further includes:

[0120] A second determination module, configured to, in response to an operation of adding text content at a target text position of a second script node in the first script structure, determine the multimedia material corresponding to the second script node in the material editing area, and determine the timeline position corresponding to the target text position in the multimedia material;

[0121] An adding module, configured to add a multimedia segment corresponding to the text content at the timeline position in the multimedia material according to the operation of adding text content.

[0122] In an optional implementation manner, the apparatus further includes:

[0123] A third determination module, configured to, in response to a clipping operation on a target multimedia segment of a first multimedia material in the material editing area, determine the script node corresponding to the first multimedia material, and determine the text content corresponding to the target multimedia segment in the script node;

[0124] A first adjustment module, configured to adjust the text content in the script node according to the clipping operation.

[0125] In an optional implementation manner, the apparatus further includes:

[0126] A fourth determination module, configured to, in response to an order adjustment operation between a second script node and a third script node in the first script structure, determine the sub-areas corresponding to the second script node and the third script node respectively in the material editing area;

[0127] A second adjustment module, configured to adjust the order of the multimedia materials in the sub-areas corresponding to the second script node and the third script node respectively in the material editing area according to the order adjustment operation.

[0128] In an optional implementation manner, the target multimedia material has alternative multimedia materials, and the apparatus further includes:

[0129] A switching module, configured to switch the target multimedia material displayed in the target sub-region to the alternative multimedia material in response to a switching operation on the target multimedia material and the alternative multimedia material in the target sub-region.

[0130] In the video processing apparatus provided by the embodiments of the present disclosure, a material editing area for displaying video clips is presented according to a first script structure, such that the sub-regions in the material editing area correspond to the script nodes in the first script structure. Additionally, in the target sub-region of the material editing area, the multimedia materials selected for the target script node corresponding to the target sub-region are presented according to a time-axis track. Furthermore, a target video is generated according to the multimedia materials presented in the material editing area. The embodiments of the present disclosure can implement video editing based on a material editing area including multiple sub-regions corresponding to script nodes, enriching the video processing methods and further meeting the diverse video editing needs of people.

[0131] In addition to the above methods and apparatuses, the embodiments of the present disclosure further provide a computer-readable storage medium, in which instructions are stored. When the instructions are run on a terminal device, the terminal device implements the video processing method described in the embodiments of the present disclosure.

[0132] The embodiments of the present disclosure also provide a computer program product, where the computer program product includes computer programs / instructions, and when the computer programs / instructions are executed by a processor, the video processing method described in the embodiments of the present disclosure is implemented.

[0133] In addition, the embodiments of the present disclosure further provide a video processing device, as shown in Figure 10 It may include:

[0134] A processor 1001, a memory 1002, an input device 1003, and an output device 1004. The number of processors 1001 in the video processing device may be one or more, Figure 10 Taking one processor as an example. In some embodiments of the present disclosure, the processor 1001, the memory 1002, the input device 1003, and the output device 1004 may be connected by a bus or other means. Among them, Figure 10 Taking the connection by a bus as an example.

[0135] The memory 1002 can be used to store software programs and modules. The processor 1001 executes various functional applications and data processing of the video processing device by running the software programs and modules stored in the memory 1002. The memory 1002 mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function, etc. In addition, the memory 1002 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices. The input device 1003 can be used to receive input digital or character information, and generate signal inputs related to the user settings and function controls of the video processing device.

[0136] Specifically in this embodiment, the processor 1001 loads the executable files corresponding to the processes of one or more application programs into the memory 1002 according to the following instructions, and the processor 1001 runs the application programs stored in the memory 1002, so as to implement various functions of the above video processing device.

[0137] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0138] The above are only specific implementation manners of the present disclosure, enabling those skilled in the art to understand or implement the present disclosure. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to these embodiments described herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A video processing method, characterized in that, The method includes: Displaying a material editing area of a video clip according to a first script structure; wherein, the material editing area is divided into multiple sub-areas, one of the sub-areas corresponds to one script node in the first script structure, the first script structure is used to indicate the content paragraph structure of a target video, and one script node is used to indicate one content paragraph of the target video; In a target sub-area among the multiple sub-areas, displaying target multimedia materials according to a time axis track; wherein, the target multimedia materials are multimedia materials selected for a target script node, and the target script node is the script node in the first script structure corresponding to the target sub-area; According to the multimedia materials displayed in the material editing area, splicing each content paragraph according to the first script structure to generate the target video; wherein, the target multimedia materials are filled in a target content paragraph of the target video, and the target content paragraph corresponds to the target script node.

2. The method according to claim 1, wherein The interface layout mode of the multiple sub-areas in the material editing area is vertical alignment arrangement.

3. The method according to claim 1, wherein Before generating the target video according to the multimedia materials displayed in the material editing area, it further includes: In response to an adjustment operation on target text content of a first script node in the first script structure, determining the multimedia materials corresponding to the first script node in the material editing area, and determining the multimedia segments corresponding to the target text content in the multimedia materials; According to the adjustment operation, clipping the multimedia segments in the multimedia materials.

4. The method according to claim 1, wherein Before generating the target video according to the multimedia materials displayed in the material editing area, it further includes: In response to an operation of adding text content at a target text position of a second script node in the first script structure, determining the multimedia materials corresponding to the second script node in the material editing area, and determining the time axis position corresponding to the target text position in the multimedia materials; According to the operation of adding text content, adding multimedia segments corresponding to the text content at the time axis position in the multimedia materials.

5. The method according to claim 1, wherein Before generating the target video according to the multimedia materials displayed in the material editing area, it further includes: In response to a clipping operation on a target multimedia segment of a first multimedia material in the material editing area, determining the script node corresponding to the first multimedia material, and determining the text content corresponding to the target multimedia segment in the script node; According to the clipping operation, adjusting the text content in the script node.

6. The method according to claim 1, wherein Before generating the target video according to the multimedia materials displayed in the material editing area, it further includes: In response to an order adjustment operation between a second script node and a third script node in the first script structure, determining the sub-areas respectively corresponding to the second script node and the third script node in the material editing area; Adjust the order of the multimedia materials in the sub-regions corresponding to the second script node and the third script node in the material editing area according to the said order adjustment operation.

7. The method according to claim 1, characterized in that, The target multimedia material has alternative multimedia materials. Before generating the target video according to the multimedia materials displayed in the material editing area, it further includes: In response to the switching operation of the target multimedia material and the alternative multimedia material in the target sub-region, switch the target multimedia material displayed in the target sub-region to the alternative multimedia material.

8. A video processing device, characterized in that, The device includes: A first display module, configured to display the material editing area of the video clip according to the first script structure; wherein, the material editing area is divided into multiple sub-regions, one sub-region corresponds to one script node in the first script structure, the first script structure is used to indicate the content paragraph structure of the target video, and one script node is used to indicate one content paragraph of the target video; A second display module, configured to display the target multimedia material according to the time axis track in the target sub-region among the multiple sub-regions; wherein, the target multimedia material is the multimedia material selected for the target script node, and the target script node is the script node corresponding to the target sub-region in the first script structure; A generation module, configured to splice each content paragraph according to the first script structure according to the multimedia materials displayed in the material editing area to generate the target video; wherein, the target multimedia material is filled in the target content paragraph of the target video, and the target content paragraph corresponds to the target script node.

9. A computer-readable storage medium, characterized in that, Instructions are stored in the computer-readable storage medium, and when the instructions are run on the terminal device, the terminal device realizes the method described in any one of claims 1-7.

10. An electronic device, characterized in that, It includes: A memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it realizes the method described in any one of claims 1-7.

11. A computer program product, characterized in that, The computer program product includes computer programs / instructions, and when the computer programs / instructions are executed by the processor, they realize the method described in any one of claims 1-7.

Citation Information

Patent Citations

  • Video display and processing method, device and system, equipment and medium

    CN112579826A

  • Video generation method and device, electronic equipment and storage medium

    CN113364999A