Image processing method, image processing device, device, storage medium, and computer program
The video processing method and device address the challenge of diverse user needs by displaying and adjusting multimedia materials in sub-areas corresponding to script nodes, enhancing video clipping efficiency and accuracy through intuitive alignment and selection.
Patent Information
- Application Number
- JP2023577720
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-09-15
- Filing Date
- 2022-09-08
- Publication Date
- 2026-03-02
- Estimated Expiration
- 2042-09-08
AI Technical Summary
Existing video processing methods struggle to efficiently meet diverse user needs for video clips, particularly in terms of quick and easy clipping and adjustment of multimedia materials based on script structures.
A video processing method and device that displays a material editing area divided into sub-areas corresponding to script nodes, allowing for the display and adjustment of multimedia materials according to a timeline track, and generates videos based on these materials, enabling intuitive and efficient clipping, addition, and replacement of multimedia segments.
Enriches video processing methods by allowing quick and easy adjustment of video clips, improving efficiency and accuracy by aligning multimedia materials with script structures, reducing manual operations, and enabling selection of optimal multimedia content.
Smart Images

Figure 0007822405000001 
Figure 0007822405000002 
Figure 0007822405000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to the field of computer technology, and in particular to a video processing method, device, apparatus, and storage medium. [Background technology]
[0002] With the development of computer technology, video is widely used in work and life, and the needs for video clips are becoming more diverse.
[0003] Therefore, how to meet the various demands for video clips is a technical problem that needs to be solved immediately. Summary of the Invention [Problem to be solved by the invention]
[0004] SUMMARY OF THE INVENTION To solve the above technical problems or to solve at least part of the above technical problems, embodiments of the present disclosure provide a video processing method, a device, an apparatus, and a storage medium. [Means for solving the problem]
[0005] In a first aspect, the present disclosure provides a video processing method, the method comprising: Displaying a material editing area of a video clip according to a first script structure, the material editing area being divided into a plurality of sub-areas, each sub-area corresponding to a script node in the first script structure, the first script structure being for indicating a content paragraph structure of a target video, and each script node being for indicating a content paragraph of the target video; displaying target multimedia material in a target sub-region of the plurality of sub-regions according to a timeline track, the target multimedia material being multimedia material selected for a target script node, the target script node being a script node in the first script structure corresponding to the target sub-region; The target video is generated according to the multimedia material displayed in the material editing area, and includes filling the target content paragraph of the target video with the target multimedia material and corresponding the target content paragraph to the target script node.
[0006] In an alternative embodiment, the sub-areas within the material editing area are arranged in vertical alignment in the interface layout.
[0007] In an alternative embodiment, before generating the target video in accordance with the multimedia material displayed in the material editing area, According to an adjustment operation on a target text content of a first script node in the first script structure, determine a multimedia material corresponding to the first script node in the material editing area, and determine a multimedia segment in the multimedia material corresponding to the target text content; The method further includes clipping the multimedia segments within the multimedia material according to the adjustment operation.
[0008] In an alternative embodiment, before generating the target video in accordance with the multimedia material displayed in the material editing area, In response to an operation of adding text content to a target text position of a second script node in the first script structure, determining a multimedia material corresponding to the second script node in the material editing area, and determining a timeline position in the multimedia material corresponding to the target text content; The method further includes adding a multimedia segment corresponding to the text content at the timeline position within the multimedia material in accordance with the operation of adding the text content.
[0009] In an alternative embodiment, before generating the target video in accordance with the multimedia material displayed in the material editing area, determining a script node corresponding to the multimedia material in accordance with a clip operation on a target multimedia segment of the multimedia material in the material editing area, and determining text content in the script node corresponding to the target multimedia segment; and adjusting the text content in the script node according to the clip operation.
[0010] In an alternative embodiment, before generating the target video in accordance with the multimedia material displayed in the material editing area, determining sub-areas corresponding to the second script node and the third script node, respectively, within the material editing area in accordance with an adjustment operation on the order of the second script node and the third script node in the first script structure; The method further includes adjusting the order of the multimedia materials in the sub-areas corresponding to the second script node and the third script node in the material editing area according to the adjustment operation on the order.
[0011] In an alternative embodiment, the target multimedia material comprises candidate multimedia material, and prior to generating the target video in accordance with the multimedia material displayed in the material editing area: The method further includes replacing the target multimedia material displayed in the target sub-region with the candidate multimedia material in response to a replacement operation on the target multimedia material and the candidate multimedia material in the target sub-region.
[0012] In a second aspect, the present disclosure further provides a video processing device, the device comprising: a first display module, a second display module, and a generating module.
[0013] The first display module is for displaying a material editing area of a video clip according to a first script structure, the material editing area being divided into a plurality of sub-areas, one of the sub-areas corresponding to one script node in the first script structure, the first script structure being for indicating a content paragraph structure of a target video, and one of the script nodes being for indicating one content paragraph of the target video. The second display module is for displaying target multimedia material for each timeline track in a target sub-area within the plurality of sub-areas, the target multimedia material being multimedia material selected for a target script node, and the target script node being a script node corresponding to the target sub-area in the first script structure. The generation module generates the target video according to the multimedia material displayed in the material editing area, and fills the target content paragraph of the target video with the target multimedia material, and the target content paragraph corresponds to the target script node.
[0014] In a third aspect, the present disclosure provides a computer-readable storage medium having stored thereon instructions that, when executed on a terminal device, cause the terminal device to implement the above-described method.
[0015] In a fourth aspect, the present disclosure provides an apparatus comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, the processor executing the computer program to realize the above method.
[0016] In a fifth aspect, the present disclosure provides a computer program product, the computer program product including computer programs / instructions that, when executed by a processor, cause the computer programs / instructions to implement the above-described methods.
[0017] The technical solutions according to the embodiments of the present disclosure have at least the following advantages over the prior art:
[0018] A video processing method according to an embodiment of the present disclosure displays a material editing area for a video clip according to a first script structure, whereby sub-areas in the material editing area correspond to script nodes in the first script structure. Furthermore, in a target sub-area of the material editing area, a multimedia material selected for a target script node corresponding to the target sub-area is displayed according to a timeline track. Furthermore, a target video is generated according to the multimedia material displayed in the material editing area. The embodiment of the present disclosure can realize a clip for a video based on a material editing area including multiple sub-areas corresponding to script nodes, thereby enriching video processing methods and further meeting people's diverse needs for video clips. [Brief explanation of the drawings]
[0019] The figures herein illustrate examples consistent with the present disclosure, are incorporated herein as part of the specification, and together with the specification serve to explain the principles of the disclosure.
[0020] In order to more clearly explain the technical solutions in the embodiments or prior art of the present disclosure, the figures necessary for explaining the embodiments or prior art will be briefly described below, and it will be obvious that a person skilled in the art can derive other figures based on these figures without expending creative labor.
[0021] [Figure 1] 1 is a flowchart illustrating a video processing method according to an embodiment of the present disclosure. [Figure 2] FIG. 2 is a schematic diagram illustrating the relationship between script nodes, content paragraphs, and sub-regions according to an embodiment of the present disclosure. [Figure 3a] 10A and 10B are schematic diagrams illustrating a method for aligning material editing regions according to an embodiment of the present disclosure. [Figure 3b]10A and 10B are schematic diagrams illustrating another method of aligning material editing regions according to an embodiment of the present disclosure. [Figure 4a] 1 is a schematic diagram illustrating a material editing area and a first script structure according to an embodiment of the present disclosure. [Figure 4b] FIG. 10 is a schematic diagram illustrating another material editing area and a first script structure according to an embodiment of the present disclosure. [Figure 4c] FIG. 10 is a schematic diagram illustrating yet another material editing area and a first script structure according to an embodiment of the present disclosure. [Figure 5] FIG. 2 is a schematic diagram illustrating the relationship between a target script node, a target sub-region, a target content paragraph, and a target multimedia material according to an embodiment of the present disclosure. [Figure 6a] 1 is a schematic diagram illustrating a display of target multimedia material according to an embodiment of the present disclosure. [Figure 6b] FIG. 10 is a schematic diagram illustrating a display of another target multimedia material according to an embodiment of the present disclosure. [Figure 7] FIG. 1 is a schematic diagram illustrating generation of a target image according to an embodiment of the present disclosure. [Figure 8] FIG. 2 is a schematic diagram illustrating the replacement of target multimedia material with candidate multimedia material according to an embodiment of the present disclosure. [Figure 9] 1 is a schematic diagram illustrating a structure of a video processing device according to an embodiment of the present disclosure. [Figure 10] 1 is a schematic diagram illustrating a structure of a video processing device according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0022] In order to more clearly understand the above-mentioned objectives, features and advantages of the present disclosure, the technical solution of the present disclosure will be further described below. It should be noted that, where not inconsistent, the embodiments and features in the embodiments of the present disclosure can be combined with each other.
[0023] In the following description, numerous details are set forth to provide a thorough understanding of the present disclosure, but the present disclosure may be implemented in a manner different from that described herein, and it is clear that the examples in the specification are only a portion and not all of the examples.
[0024] In order to meet various user needs for video clips and enrich video processing methods, a video processing method according to an embodiment of the present disclosure first displays a material editing area of a video clip according to a first script structure. The material editing area is divided into a plurality of sub-areas, each sub-area corresponding to a script node in the first script structure. The first script structure is for indicating a content paragraph structure of the target video, and each script node is for indicating a content paragraph of the target video. Next, target multimedia material is displayed in a target sub-area of the plurality of sub-areas according to a timeline track. The target multimedia material is multimedia material selected for the target script node, and the target script node is a script node in the first script structure corresponding to the target sub-area. Furthermore, the target video is generated according to the multimedia material displayed in the material editing area. The target content paragraph of the target video is filled with the target multimedia material, and the target content paragraph corresponds to the target script node.
[0025] As can be seen from the above, an embodiment of the present disclosure displays a material editing area for a video clip according to a first script structure, whereby sub-areas in the material editing area correspond to script nodes in the first script structure. Furthermore, in a target sub-area of the material editing area, multimedia material selected for a target script node corresponding to the target sub-area is displayed according to a timeline track, and a target video is generated according to the multimedia material displayed in the material editing area. An embodiment of the present disclosure can realize a clip for a video based on a material editing area including multiple sub-areas corresponding to script nodes, thereby enriching video processing methods and further meeting people's diverse needs for video clips.
[0026] Based on this, an embodiment of the present disclosure provides an image processing method. Figure 1 is a flowchart illustrating the image processing method according to an embodiment of the present disclosure. Referring to Figure 1, the method can be performed by an image processing device, where the device can be realized in software and / or hardware and can generally be incorporated into electronic equipment.
[0027] As shown in FIG. 1, the method includes the following steps:
[0028] Step 101: display a material editing area of a video clip according to a first script structure, where the material editing area is divided into multiple sub-areas, each sub-area corresponds to a script node in the first script structure, the first script structure is for indicating a content paragraph structure of the target video, and each script node is for indicating a content paragraph of the target video.
[0029] A script is a manuscript used in the production process of a film or television program. A script typically includes multiple segment descriptions to guide a cinematographer in the filming and production of a corresponding film or television program. For example, a script may include explanatory content for screen a to instruct the filming of a first shot, and explanatory content for screen b to instruct the filming of a second shot. When producing a film or television program, the cinematographer follows the explanatory content of screen a to capture a first shot containing video segment A, and follows the explanatory content of screen b to capture a second shot containing video segment B, and then splices the second shot after the first shot to obtain a film or television program corresponding to the script.
[0030] In this embodiment, the first script structure may be a script structure such as the described content paragraph structure. Note that in the above example, explanatory content corresponding to each of the first shot and second shot included in the script corresponds to a script node in the first script structure. For example, the explanatory content of the first shot corresponds to the first script node in the first script structure, and the explanatory content of the second shot corresponds to the second script node in the first script structure.
[0031] In this embodiment, once the first script structure is determined, a material editing area for the video clip is displayed according to the first script structure. The multimedia material to be clipped may also be displayed in the material editing area. Specifically, the material editing area is divided into sub-areas corresponding to each script node according to the script nodes in the first script structure. Each sub-area corresponds to one script node in the first script structure.
[0032] To more intuitively understand the relationship between sub-areas and script nodes according to an embodiment of the present disclosure, the following description will be given with reference to the content shown in Fig. 2. The first script structure in Fig. 2 includes Q script nodes, where Q is a positive integer. Based on the first script structure including the Q script nodes, the material editing area can be divided into Q sub-areas, and each script node corresponds to one sub-area in the material editing area.
[0033] The interface layout of the multiple sub-areas in the material editing area can be selected as needed, and this embodiment is not limited thereto. For example, as shown in Figure 3a, the sub-areas are aligned vertically, i.e., each row is aligned vertically. In Figure 3a, each sub-area is aligned left. Alternatively, each sub-area may be aligned right. Also, as shown in Figure 3b, the sub-areas are aligned horizontally, i.e., each column is aligned horizontally. In Figure 3b, each sub-area is aligned top. Alternatively, each sub-area may be aligned bottom.
[0034] In an alternative embodiment, a script node in the first script structure may include a script annotation and / or a script paragraph in the script. That is, the script node is associated with a script annotation and / or a script paragraph in the script. Here, the script annotation provides an outline of the multimedia material content corresponding to the script node, and the script paragraph includes detailed text content corresponding to the script node. In an alternative embodiment, the detailed text content included in the script paragraph may be character information obtained by performing speech recognition on a video.
[0035] Specifically, an example of the material editing area of a video clip according to the first script structure is shown below.
[0036] Example 1: A script node in a first script structure includes a script annotation. As shown in Figure 4a, the first script structure includes a first script annotation and a second script annotation. Here, the first script annotation is " / / Introduction" and the second script annotation is " / / Environment Introduction". In a material editing area displayed according to the first script structure, a first sub-area of the material editing area corresponds horizontally to " / / Introduction", and a second sub-area of the material editing area corresponds horizontally to " / / Environment Introduction".
[0037] Example 2: A script node in a first script structure includes a script paragraph. As shown in Figure 4b, the first script structure includes a first script paragraph and a second script paragraph. Here, the first script paragraph is an introductory character obtained by speech recognition, and the second script paragraph is an environment introduction character obtained by speech recognition. In a material editing area displayed according to the first script structure, a first sub-area of the material editing area corresponds horizontally to the introductory character script paragraph, and a second sub-area of the material editing area corresponds horizontally to the environment introduction character script paragraph.
[0038] Example 3: A script node in a first script structure includes a script annotation and a script paragraph. As shown in FIG. 4c, the first script structure includes a first script node and a second script node. Here, the first script node includes a first script annotation and a first script paragraph, where the first script annotation is " / / introduction," and the first script paragraph is an introductory character obtained by speech recognition. The second script node includes a second script annotation and a second script paragraph, where the second script annotation is " / / environment introduction," and the second script paragraph is an environment introduction character obtained by speech recognition. A first sub-area in a material editing area displayed according to the first script structure corresponds horizontally to both the introductory character script paragraph and the first script annotation in the first script structure, and a second sub-area in the material editing area corresponds horizontally to both the environment introduction character script paragraph and the second script annotation in the first script structure.
[0039] After displaying the material editing area in accordance with the first script structure, the process proceeds to the next step 102 .
[0040] Step 102: Displaying target multimedia material for each timeline track in a target sub-region within the plurality of sub-regions.
[0041] Here, the target multimedia material is the multimedia material selected for the target script node, and the target script node is the script node corresponding to the target sub-region in the first script structure.
[0042] In this embodiment, the target sub-region may be any one of a plurality of sub-regions in the material editing region, and the first script structure has a target script node corresponding to the target sub-region, and the target sub-region can select corresponding multimedia material based on the target script node.
[0043] In an alternative embodiment, after receiving target multimedia material imported by a user, speech recognition is performed on the target multimedia material, and the speech recognition result is subjected to text matching with each script node in the first script structure to determine a target script node corresponding to the target multimedia material. Then, the target multimedia material is displayed in a target sub-area corresponding to the target script node according to a timeline track. The target multimedia material in the embodiment of the present disclosure may be the entire video captured by filming, or a segment of the entire video captured by filming. This embodiment is not limited thereto.
[0044] For easier understanding, refer to FIG. 5, the target multimedia material is selected based on the target script node, and the target node further corresponds to the target sub-area, so that the target multimedia material to be displayed in the target sub-area can be determined.
[0045] In an alternative embodiment, as shown in Fig. 6a, if the first sub-area of the material editing area is the target sub-area, first determine the target script node corresponding to the first sub-area in the first script structure as " / / preface", then select the multimedia material as the target multimedia material based on the target script node, and display the target multimedia material in the first sub-area.
[0046] Here, as a method of selecting multimedia material based on the target script node, image recognition and / or voice recognition may be performed on the selected multimedia material, and the multimedia material that has the highest degree of match with the script node `` / / preface'' may be determined as the target multimedia material, and the target multimedia material may be displayed in the first sub-area.
[0047] In another alternative embodiment, as shown in Figure 6b, if a first sub-area in the material editing area is a target sub-area and the multimedia material includes introductory video material, speech recognition is performed on the introductory material to obtain corresponding introductory text. The introductory text is set as a target script node in the first script structure, and a target multimedia material with the highest matching score is selected from the multimedia material based on the introductory text. For example, the introductory video material is selected as the target multimedia material, and the introductory video material is displayed in the first sub-area according to the timeline track.
[0048] Step 103: Generate a target video according to the multimedia material displayed in the material editing area.
[0049] Here, the target content paragraph of the target video is filled with the target multimedia material, and the target content paragraph corresponds to the target script node.
[0050] In this embodiment, after the multimedia material is displayed in each sub-area, the target video can be generated according to the multimedia material displayed in the material editing area.
[0051] For easier understanding, referring to Figure 5, the target video includes Q content paragraphs, where Q is a positive integer. Each content paragraph is associated with one script node in the first script structure, and each content paragraph is filled with multimedia material selected for the script node corresponding to the content paragraph. The multimedia material includes, but is not limited to, one or more of video and audio.
[0052] In this embodiment, the first script structure is used to represent the content paragraph structure of the target video. Specifically, one script node in the first script structure is used to represent one content paragraph of the target video. That is, the content paragraph corresponding to the script node satisfies the requirements of the script node, so that the content paragraph can be adjusted based on the script node in the first script structure to generate a target video that complies with the first script structure.
[0053] Continuing with reference to FIG. 5, the first script structure includes Q script nodes, where Q is a positive integer. Each script node has a corresponding content paragraph. From the target script node, a three-way correspondence between the target sub-region, the target content paragraph, and the target multimedia material can be determined. Thus, by filling each content paragraph with the corresponding target multimedia material and stitching the content paragraphs together according to the first script structure, a corresponding target video can be obtained.
[0054] In an alternative embodiment, as shown in Figure 7, a first sub-area of the material editing area displays n frames of introductory video material, and a second sub-area of the material editing area displays m frames of environment introduction video material, where n and m are positive integers. If it is determined based on the first script structure that the first sub-area corresponds to a first content paragraph of the target video and the second sub-area corresponds to a second content paragraph of the target video, the n frames of introductory video material are filled into the first content paragraph, and the m frames of environment introduction video material are filled into the second content paragraph to generate the target video.
[0055] As described above, a video processing method according to an embodiment of the present disclosure displays a material editing area for a video clip according to a first script structure, thereby corresponding sub-areas in the material editing area to script nodes in the first script structure. Furthermore, in a target sub-area of the material editing area, selected multimedia material is displayed for the target script node corresponding to the target sub-area according to a timeline track, and a target video is generated according to the multimedia material displayed in the material editing area. An embodiment of the present disclosure realizes a video clip based on a material editing area including multiple sub-areas corresponding to script nodes. This enriches video processing methods and enables the provision of video clips that meet diverse user needs.
[0056] Typically, a video work is generated by clipping multiple child videos. In the clipping process, the child videos must be clipped according to the timeline corresponding to each child video and then spliced together according to the timeline corresponding to the overall video. However, this timeline-based clipping method requires repeated comparison of the content of each screen frame on the child video's timeline when clipping language content. Therefore, this technical solution does not allow for quick and easy clipping of video. Therefore, based on the above embodiment, quick and easy clipping of video can be achieved. Specifically, the following operation steps may be added as needed before generating a target video according to the multimedia material displayed in the material editing area. An example of this is described below.
[0057] In an alternative embodiment, if a situation such as a slip of the tongue in the multimedia material requires clipping of the corresponding segment in the multimedia material, the following step needs to be added before step 103 in the above embodiment.
[0058] First, according to an adjustment operation on a target text content of a first script node in a first script structure, a multimedia material corresponding to the first script node is determined in a material editing area, and a multimedia segment corresponding to the target text content is determined in the multimedia material.
[0059] In this example, a first script node in a first script structure has corresponding multimedia material, and the first script node is text content corresponding to the multimedia material. The text content can be obtained by using voice recognition technology to recognize character information, manually set subtitles, etc. A user can adjust the target text content in the first script node as needed. Based on the adjustments, the multimedia material corresponding to the first script node is determined in the material editing area. Furthermore, to determine the content requiring adjustment, it is necessary to further determine the multimedia segment in the multimedia material that corresponds to the target text content.
[0060] Furthermore, clipping the multimedia segments within the multimedia material according to the adjustment operation, which may include, but is not limited to, deleting, shifting, etc.
[0061] For example, if the target text content is "Good morning, good afternoon" (Good morning) and the multimedia material is a greeting video, the correspondence between the target text content and the multimedia material is such that "early" corresponds to the first frame of the greeting video, "up" corresponds to the second frame of the greeting video, "mid" corresponds to the third frame of the greeting video, "go" corresponds to the fourth frame of the greeting video, and "good" corresponds to the fifth frame of the greeting video. In this example, "early" is a misspelling, and the segment corresponding to the target video needs to be deleted. Therefore, when an operation is performed to delete "early" from "good morning, good afternoon" in the target text content, the first and second frames of the corresponding greeting video are also deleted.
[0062] In another example, the target text content may be associated with the multimedia material by a timestamp. The timestamp can associate the text content with a timeline of the multimedia material. Specifically, if the text content is "Early in the morning, good in the afternoon," the target text content is "Early in the morning," and the multimedia material is a greeting video, the correspondence between the text content and the multimedia material is such that "Early" corresponds to seconds 0 to 1.5 of the multimedia material, "Mid-afternoon" corresponds to seconds 1.5 to 3 of the multimedia material, and "Good" corresponds to seconds 3 to 4 of the multimedia material. When an operation is performed to delete "Early in the morning" from "Early in the morning, good in the afternoon" in the target text content, seconds 0 to 1.5 of the corresponding greeting video are also deleted.
[0063] In this embodiment, by operating on the first script structure, the tedious manual placement of the text to be processed on the timeline of the multimedia material is avoided, thereby improving the efficiency and accuracy of video processing.
[0064] In another alternative embodiment, if it is necessary to add multimedia segments to the target video, the following steps should be added before step 103 in the above embodiment.
[0065] First, in response to an operation of adding text content to a target text position of a second script node in a first script structure, a multimedia material corresponding to the second script node is determined in a material editing area, and a timeline position in the multimedia material corresponding to the target text content is determined.
[0066] In this example, a second script node in the first script structure has corresponding multimedia material, and the second script node is text content corresponding to the multimedia material. The text content can be obtained by using voice recognition technology to recognize character information, manually set subtitles, etc. The user can add text content to the target text position of the second script node as needed. Based on this adjustment, the multimedia material corresponding to the second script node is determined in the material editing area. Furthermore, to determine the position of the added multimedia segment, it is necessary to further determine the timeline position within the multimedia material that corresponds to the target text content.
[0067] Furthermore, in accordance with the operation of adding the text content, a multimedia segment corresponding to the text content is added to the timeline position within the multimedia material.
[0068] In an alternative embodiment, the multimedia material before the timeline position may be determined as the leading multimedia material, and the multimedia material after the timeline position may be determined as the trailing multimedia material, so that the appending operation can splice the multimedia segment after the leading multimedia material and the trailing multimedia material after the multimedia segment. By operating on the first script structure, the tedious operation of manually arranging the appended text on the timeline of the multimedia material can be avoided, thereby improving the efficiency and accuracy of video processing.
[0069] For example, if the second script node is "Hello everyone" and the multimedia material is a greeting video, as the correspondence between the target text content and the multimedia material, "big" corresponds to the first frame of the greeting video, "home" corresponds to the second frame of the greeting video, and "good" corresponds to the third frame of the greeting video. In this example, it is necessary to add "noon" between "home" and "good". Therefore, in response to the operation of adding "noon" to the second script node, obtain a video segment corresponding to "noon" to include the first frame of the "noon" video and the second frame of the "noon" video. Therefore, after the second frame of the greeting video, connect the first frame and the second frame of the "noon" video, and after the second frame of the "noon" video, connect the third frame of the greeting video.
[0070] In another alternative embodiment, when a clip operation is performed on the multimedia material, the script node corresponding to the multimedia material also changes. In the case of such an application segment, the following steps need to be added before step 103 of the above embodiment.
[0071] First, in response to a clip operation on the target multimedia segment of the multimedia material in the material editing area, determine the script node corresponding to the multimedia material, and determine the text content corresponding to the target multimedia segment within the script node. Further, perform adjustments on the text content within the script node according to the clip operation.
[0072] In this example, when the user performs a clip operation on the target multimedia segment within the multimedia material in the material editing area, since it is necessary to perform a corresponding operation on the first script structure in response to the operation, it is necessary to determine the script node corresponding to the multimedia material and determine the text content corresponding to the target multimedia material within the script node. Then, appropriately adjust the text content within the script node based on the clip operation of the multimedia material.
[0073] For example, if the multimedia material is a greeting video, the correspondence between the text content of the script node and the greeting video is such that "early" corresponds to the first frame of the greeting video, "up" corresponds to the second frame of the greeting video, "mid" corresponds to the third frame of the greeting video, "horse" corresponds to the fourth frame of the greeting video, and "good" corresponds to the fifth frame of the greeting video. In this example, the third and fourth frames of the greeting video are deleted. In accordance with the operation of deleting the third and fourth frames of the greeting video, "mid" and "horse" are deleted from the text content of the script node, and the script node after processing becomes "Early up, good." This unifies the changes between the multimedia material and the corresponding script node, maintaining consistency between the multimedia material and the script node.
[0074] In another alternative embodiment, the order of the multimedia materials can be adjusted based on the first script structure, so the following step needs to be added before step 103 in the above embodiment:
[0075] First, in accordance with an adjustment operation on the order of the second script node and the third script node in the first script structure, sub-areas corresponding to the second script node and the third script node are determined in the material editing area, and then, in accordance with the adjustment operation on the order, an order adjustment is performed on the multimedia materials in the sub-areas corresponding to the second script node and the third script node in the material editing area.
[0076] If the user needs to make a ranking adjustment to the multimedia material, the adjustment can be made to the second script node and the third script node in the first script structure, and a second sub-area corresponding to the second script node and a third sub-area corresponding to the third script node exist. The second sub-area and the third sub-area are determined in the material editing area according to the adjustment, and the second sub-area and the third sub-area are adjusted according to the user's adjustment to the script structure. In this example, adjusting the first script structure adjusts the ranking of the multimedia material, improving the efficiency of video processing and eliminating the need to view the multimedia material to determine its content, making video processing more intuitive.
[0077] For example, in this example, the second script node in the first script structure is " / / Introduction", the third script node is " / / Environment Introduction Video", and " / / Introduction" is located after " / / Environment Introduction Video", and in the corresponding material editing area, the introduction material is located after the environment introduction material. If the user needs to move the introduction video before the environment introduction video, he or she can move " / / Introduction" in the first script structure before " / / Environment Introduction Video". In response to this user operation, the introduction material in the material editing area is moved before the environment introduction material.
[0078] In another alternative embodiment, in order to obtain a target video of better quality at the time of shooting, multiple similar videos may be shot as candidates for the target multimedia material, and the best one may be selected from the candidate multimedia material. Therefore, the following step needs to be added before step 103 in the above embodiment:
[0079] In response to a replacement operation on the target multimedia material and the candidate multimedia material in the target sub-region, the target multimedia material displayed in the target sub-region is replaced with the candidate multimedia material.
[0080] In this example, the candidate multimedia materials may be set by the user or may be obtained by comparing the similarity with the target multimedia material using image or voice recognition technology. When the user replaces the target multimedia material with the candidate multimedia material, the target multimedia material displayed in the material editing area is replaced with the candidate multimedia material in accordance with the replacement operation. In addition, the target script node corresponding to the target sub-area in the first script structure may be adjusted to include text information corresponding to the candidate multimedia material based on the candidate multimedia material. The candidate operation allows the user to easily and quickly select the multimedia material that best suits their needs from multiple multimedia materials, thereby improving the efficiency of video processing.
[0081] 8, the target sub-region is a first sub-region, the target multimedia material in the first sub-region is a target introductory material, the candidate multimedia material is a first candidate introductory material and a second candidate introductory material, and the material candidate region further includes a candidate display control. The candidate display control displays the candidate multimedia material in the first sub-region in response to a user's touch operation. In this example, when the user touches the candidate display control to select and replace the second candidate introductory material with the target multimedia material, the second candidate introductory material is displayed in the first sub-region.
[0082] As described above, the video processing method according to the embodiment of the present disclosure can intuitively and easily make adjustments to the target video and / or the first script structure based on the correspondences between sub-regions, content paragraphs, and multimedia materials established by the first script structure, thereby reducing the complexity of clip processing for language content or story-centric video and improving the efficiency of video processing.
[0083] Based on the above method embodiment, the present disclosure further provides an image processing device. Figure 9 is a schematic diagram showing the structure of an image processing device according to an embodiment of the present disclosure. Referring to Figure 9, the device includes a first display module 901, a second display module 902, and a generating module 903. The first display module 901 is for displaying a material editing area of a video clip according to a first script structure, where the material editing area is divided into a plurality of sub-areas, each sub-area corresponds to a script node in the first script structure, the first script structure is for indicating a content paragraph structure of a target video, and each script node is for indicating a content paragraph of the target video. The second display module 902 is for displaying target multimedia material in a target sub-region within the plurality of sub-regions according to a timeline track, where the target multimedia material is multimedia material selected for a target script node, and the target script node is a script node in the first script structure corresponding to the target sub-region. The generating module 903 generates the target video according to the multimedia material displayed in the material editing area, where the target content paragraph of the target video is filled with the target multimedia material, and the target content paragraph corresponds to the target script node.
[0084] In an alternative embodiment, the sub-areas within the material editing area are arranged in vertical alignment in the interface layout.
[0085] In an alternative embodiment, the apparatus further comprises a first determining module and an editing module. The first determination module is used to determine a multimedia material corresponding to the first script node in the material editing area according to an adjustment operation on a target text content of a first script node in the first script structure, and to determine a multimedia segment in the multimedia material corresponding to the target text content. A clip module is used to clip the multimedia segment within the multimedia material according to the adjustment operation.
[0086] In an alternative embodiment, the apparatus further comprises a second determination module and an additional module. A second determination module is used to determine a multimedia material corresponding to the second script node in the material editing area in response to an operation of adding text content to a target text position of a second script node in the first script structure, and to determine a timeline position within the multimedia material corresponding to the target text content. An appending module is used to append a multimedia segment corresponding to the text content to the timeline position within the multimedia material according to the operation of appending the text content.
[0087] In an alternative embodiment, the apparatus further comprises a third determination module and a first adjustment module. The third determination module is used to determine a script node corresponding to the multimedia material according to a clip operation on a target multimedia segment of the multimedia material in the material editing area, and to determine text content corresponding to the target multimedia segment in the script node. A first adjustment module is used to make adjustments to the text content in the script node according to the clip operation.
[0088] In an alternative embodiment, the apparatus further comprises a fourth determination module and a second adjustment module. A fourth determination module is used to determine sub-areas corresponding to the second script node and the third script node, respectively, within the material editing area, according to an adjustment operation on the order of the second script node and the third script node in the first script structure. The second adjustment module is used to perform a ranking adjustment on the multimedia materials in the sub-areas corresponding to the second script node and the third script node in the material editing area according to the adjustment operation on the ranking.
[0089] In an alternative embodiment, the target multimedia material includes candidate multimedia material and the apparatus further comprises a substitution module. A replacement module is used to replace the target multimedia material displayed in the target sub-region with the candidate multimedia material in response to a replacement operation on the target multimedia material and the candidate multimedia material in the target sub-region.
[0090] A video processing device according to an embodiment of the present disclosure displays a material editing area for a video clip according to a first script structure, thereby corresponding sub-areas in the material editing area to script nodes in the first script structure. The device also displays selected multimedia material for a target script node corresponding to the target sub-area according to a timeline track in a target sub-area of the material editing area, and generates a target video according to the multimedia material displayed in the material editing area. The embodiment of the present disclosure can realize a clip for a video based on a material editing area including multiple sub-areas corresponding to script nodes. This enriches video processing methods and enables the provision of video clips that meet diverse user needs.
[0091] In addition to the above methods and devices, embodiments of the present disclosure further provide a computer-readable storage medium having instructions stored therein, the instructions being configured to cause the terminal device to implement the video processing method described in the embodiments of the present disclosure when executed on the terminal device.
[0092] An embodiment of the present disclosure further provides a computer program product, the computer program product including computer programs / instructions, which, when executed by a processor, implement the image processing method according to the embodiment of the present disclosure.
[0093] Furthermore, an embodiment of the present disclosure further provides a video processing device. As shown in FIG. 10 , the device may include a processor 1001, a memory 1002, an input device 1003, and an output device 1004. The number of processors 1001 in the video processing device may be one or more. FIG. 10 illustrates an example in which there is one processor. In some embodiments of the present disclosure, the processor 1001, the memory 1002, the input device 1003, and the output device 1004 may be connected by a bus or other methods. FIG. 10 illustrates an example in which they are connected by a bus.
[0094] The memory 1002 may be used to store software programs and modules. The processor 1001 executes the software programs and modules stored in the memory 1002 to perform various functional applications and data processing of the video processing device. The memory 1002 may mainly include a program storage area and a data storage area. Here, the program storage area may store an operating system, an application program required for at least one function, etc. The memory 1002 may also include a high-speed random access memory, and may further include at least one non-volatile memory such as a disk memory, a flash memory, or other volatile solid-state memory. The input device 1003 may be used to receive input digital or character information and generate signal input related to user settings and function control of the video processing device.
[0095] Specifically, in this embodiment, the processor 1001 uploads executable files corresponding to the processing of one or more application programs to the memory 1002 in accordance with the following instructions: In addition, the processor 1001 executes the application programs stored in the memory 1002 to realize each function of the video processing device described above.
[0096] It should be noted that the use of relational terms such as "first" and "second" herein is intended to distinguish one entity or operation from another, and does not necessarily or imply that these entities or operations are in this relationship or order. Furthermore, the terms "comprises," "comprises," and other variations imply a non-exclusive inclusion. Thus, a list of elements, processes, methods, articles, or devices includes not only those elements but also other elements not explicitly listed, or includes the inherent elements of those processes, methods, articles, or devices. Absent further qualification, an element defined by the phrase "comprises" does not exclude the presence of other identical elements in the process, method, article, or device that includes that element.
[0097] The above are merely specific embodiments of the present disclosure to enable those skilled in the art to understand or practice the disclosure. Various modifications to these examples will be apparent to those skilled in the art. The general principles defined herein may be implemented in other examples without departing from the spirit or scope of the present disclosure. Therefore, the scope of the present disclosure is not limited to the examples described herein, but is accorded the widest scope consistent with the principles and novel features disclosed herein.
[0098] [Incorporated by reference] This application is based on and claims priority from a Chinese patent application bearing application number 20211181785.6 and entitled "Image Processing Method, Apparatus, Device, and Storage Medium," filed on September 15, 2021, the entire contents of which are incorporated herein by reference.
Claims
1. Displaying a material editing area of a video clip according to a first script structure, the material editing area being divided into a plurality of sub-areas, each sub-area corresponding to a script node in the first script structure, the first script structure being for indicating a content paragraph structure of a target video, and each script node being for indicating a content paragraph of the target video; displaying target multimedia material in a target sub-region within the plurality of sub-regions according to a timeline track, the target multimedia material being multimedia material selected for a target script node, the target script node being a script node in the first script structure corresponding to the target sub-region; generating the target video according to the multimedia material displayed in the material editing area, the target content paragraph of the target video being filled with the target multimedia material, and the target content paragraph and the target script node being associated with each other; A video processing method characterized in that the target multimedia material is determined as the multimedia material that has the highest degree of match between the image recognition and / or voice recognition results of the multimedia material and the target script node.
2. The method of claim 1 , wherein the sub-areas within the material editing area are arranged in a vertically aligned interface layout.
3. before generating the target video according to the multimedia material displayed in the material editing area; In response to an adjustment operation on a target text content of a first script node in the first script structure, determine a multimedia material corresponding to the first script node in the material editing area, and determine a multimedia segment in the multimedia material corresponding to the target text content; 2. The method of claim 1, further comprising: clipping the multimedia segments within the multimedia material according to the adjustment operation.
4. before generating the target video according to the multimedia material displayed in the material editing area; In response to an operation of adding text content to a target text position of a second script node in the first script structure, determining a multimedia material corresponding to the second script node in the material editing area, and determining a timeline position in the multimedia material corresponding to the target text position; 2. The method of claim 1, further comprising: adding a multimedia segment corresponding to the text content to the timeline position within the multimedia material in accordance with the operation of adding the text content.
5. before generating the target video according to the multimedia material displayed in the material editing area; determining a script node corresponding to the multimedia material in accordance with a clip operation on a target multimedia segment of the multimedia material in the material editing area, and determining text content in the script node corresponding to the target multimedia segment; 2. The method of claim 1, further comprising: adjusting the text content in the script node according to the clip operation.
6. before generating the target video according to the multimedia material displayed in the material editing area; determining, in the material editing area, sub-areas corresponding to the second script node and the third script node, respectively, in accordance with an adjustment operation on the order of the second script node and the third script node in the first script structure; 2. The method of claim 1, further comprising: adjusting the order of multimedia materials in sub-areas in the material editing area corresponding to the second script node and the third script node, respectively, according to the adjustment operation on the order.
7. the target multimedia material comprises candidate multimedia material; before generating the target video according to the multimedia material displayed in the material editing area; 2. The method of claim 1, further comprising: replacing the target multimedia material displayed in the target sub-region with the candidate multimedia material in response to a replacement operation on the target multimedia material and the candidate multimedia material in the target sub-region.
8. a first display module, a second display module, and a generating module; the first display module is for displaying a material editing area of a video clip according to a first script structure, the material editing area being divided into a plurality of sub-areas, each sub-area corresponding to a script node in the first script structure, the first script structure being for indicating a content paragraph structure of a target video, and each script node being for indicating a content paragraph of the target video; the second display module is for displaying target multimedia material in a target sub-region within the plurality of sub-regions according to a timeline track, the target multimedia material being multimedia material selected for a target script node, the target script node being a script node in the first script structure corresponding to the target sub-region; The generating module is for generating the target video according to the multimedia material displayed in the material editing area, and the target content paragraph of the target video is filled with the target multimedia material, and the target content paragraph corresponds to the target script node; A video processing device in which the target multimedia material is determined as the multimedia material that has the highest degree of match between the recognition results of image recognition and / or voice recognition on the multimedia material and the target script node.
9. A computer readable storage medium having stored thereon instructions which, when executed in a terminal device, cause said terminal device to implement the method of any one of claims 1 to 7.
10. a memory; a processor; and a computer program stored in the memory and executable by the processor; The apparatus, wherein the processor executes the computer program to implement the method according to any one of claims 1 to 7.
11. A computer program which, when executed by a processor, implements the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Editting system for audiovisual work for television news and corresponding text
JP2005318583A
Information presenting apparatus, method, and program
JP2006054517A
Metadata input device and content processor
JP2007052626A
Information processor, processing method and program
JP2007336283A
Time Estimation of Text Position in Video Editing Method and Apparatus
JP2009507453A