Multimedia editing method and device, equipment and medium

By downsampling user materials to generate videos, the problem of long processing times in existing multimedia editing software is solved, resulting in a faster editing process and a better user experience.

CN122073632APending Publication Date: 2026-05-22BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-22
Publication Date
2026-05-22

AI Technical Summary

Technical Problem

Existing multimedia editing software is time-consuming in the intelligent packaging of user materials, resulting in a poor user experience. In particular, the time spent synthesizing the complete video in the preprocessing stage accounts for the largest proportion of the time, leading to a high user cancellation rate.

Method used

Instead of synthesizing a complete video, the user-uploaded material is downsampled first to generate the video. Instead, the downsampled video is directly provided to the server for analysis. The server's editorial recommendations are received, and multimedia editing drafts are generated based on this information.

Benefits of technology

It effectively shortens the total time spent on multimedia editing, saves users' waiting time, improves user experience, and reduces the cancellation rate caused by long editing times.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122073632A_ABST
    Figure CN122073632A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a multimedia editing method and device, equipment and a medium. The method comprises the steps of obtaining a first material; wherein the first material is a material in a first multimedia editing draft; performing down-sampling processing on the first material, generating a first video based on a down-sampling processing result of the first material, providing the first video to a server, and receiving editing recommendation information returned by the server based on the first video; wherein the editing recommendation information at least comprises multimedia data recommended by the server; and obtaining a second multimedia editing draft based on multimedia data in the editing recommendation information, the first material and the first multimedia editing draft. According to the embodiment of the invention, the multimedia editing time can be effectively shortened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of multimedia processing technology, and in particular to a multimedia editing method, apparatus, device and medium. Background Technology

[0002] To enable users to easily and quickly edit videos, some multimedia editing software or websites offer intelligent packaging functions, which intelligently package user-uploaded audio and video materials into engaging videos, thereby significantly reducing users' video editing costs. However, the inventors have found that existing intelligent packaging methods are not ideal. For example, the intelligent multimedia editing process based on user-uploaded materials is time-consuming, resulting in long waiting times and a poor user experience. Summary of the Invention

[0003] In order to solve the above-mentioned technical problems, or at least partially solve the above-mentioned technical problems, this disclosure provides a multimedia editing method, apparatus, device and medium.

[0004] This disclosure provides a multimedia editing method, the method comprising: acquiring a first material; wherein the first material is material in a first multimedia editing draft; performing downsampling processing on the first material, generating a first video based on the downsampling processing result of the first material, and providing the first video to a server, receiving editing recommendation information returned by the server based on the first video; wherein the editing recommendation information includes at least multimedia data recommended by the server; and obtaining a second multimedia editing draft based on the multimedia data in the editing recommendation information, the first material, and the first multimedia editing draft.

[0005] Optionally, after providing the first video to the server, the method further includes: deleting the first video.

[0006] Optionally, the multimedia data includes information about the second material; obtaining the second multimedia editing draft based on the multimedia data in the editor's recommendation information, the first material, and the first multimedia editing draft includes: obtaining the second material based on the multimedia data in the editor's recommendation information; and obtaining the second multimedia editing draft based on the second material, the first material, and the first multimedia editing draft.

[0007] Optionally, the multimedia data includes information about the target template; obtaining the second multimedia editing draft based on the multimedia data in the editor recommendation information, the first material, and the first multimedia editing draft includes: obtaining the target template based on the multimedia data in the editor recommendation information; obtaining a second video clip based on the first material; wherein the second video clip is a video clip used to fill the target template; filling the target template based on the second video clip to obtain a third material; wherein the target template is used to process the second video clip according to specified editing operations; and obtaining the second multimedia editing draft based on the third material and the first multimedia editing draft.

[0008] Optionally, obtaining the second video segment based on the first material includes: determining a first video segment in the first video; wherein the first video segment is at least one segment specified by the server or client in the first video; determining a target segment from the first material based on the time information of the first video segment in the first video; and obtaining the second video segment based on the target segment.

[0009] Optionally, determining the target segment from the first material based on the time information of the first video segment in the first video includes: determining the time information of the target segment in the first material based on the time information of the first video segment in the first video; wherein the target segment includes: a segment obtained by merging at least two first video segments with time intersection, and a first video segment without time intersection; and determining the target segment from the first material based on the time information of the target segment when the first material meets a first preset condition; wherein the first preset condition includes: the first material has not been modified by the target, and the target modification is a modification that changes the playback duration of the material.

[0010] Optionally, determining the target segment from the first material based on the time information of the target segment includes: determining the segment type of the target segment based on the time information of the target segment; wherein the segment type includes a first type and a second type, the first type indicating that the target segment simultaneously corresponds to at least two different materials in the first material, and the second type indicating that the target segment corresponds only to a single material in the first material; when the segment type is the first type, determining the segment content in the at least two different materials corresponding to the time information of the target segment, and synthesizing the segment content in the at least two different materials to obtain the target segment; when the segment type is the second type, determining the segment content in the single material corresponding to the time information of the target segment, and exporting the segment content in the single material to obtain the target segment.

[0011] Optionally, obtaining the second video segment based on the target segment includes: when the number of target segments is at least two, splicing the at least two target segments in chronological order based on the time information of the at least two target segments to obtain a spliced ​​video; and obtaining the second video segment based on the spliced ​​video.

[0012] Optionally, obtaining the second video segment based on the first material includes: when the first material meets a second preset condition, synthesizing all the materials in the first material to obtain a complete video segment; wherein the second preset condition includes: at least some materials in the first material have been modified by a target, the target modification being a modification that changes the playback duration of the materials; and obtaining the second video segment based on the complete video segment.

[0013] This disclosure also provides a multimedia editing device, including: a material acquisition module for acquiring first material; wherein the first material is material in a first multimedia editing draft; a recommendation information receiving module for downsampling the first material, generating a first video based on the downsampling result of the first material, providing the first video to a server, and receiving editing recommendation information returned by the server based on the first video; wherein the editing recommendation information includes at least multimedia data recommended by the server; and a second draft acquisition module for obtaining a second multimedia editing draft based on the multimedia data in the editing recommendation information, the first material, and the first multimedia editing draft.

[0014] This disclosure also provides an electronic device, the electronic device comprising: a processor; a memory for storing executable instructions of the processor; the processor being configured to read the executable instructions from the memory and execute the instructions to implement the multimedia editing method provided in this disclosure.

[0015] This disclosure also provides a computer-readable storage medium storing a computer program for performing the multimedia editing method provided in this disclosure.

[0016] The technical solution provided in this disclosure can provide a first video generated based on the downsampling processing result of the first material (material in the first multimedia editing draft) to the server, and receive editing recommendation information (including at least the multimedia data recommended by the server) returned by the server based on the first video; then, a second multimedia editing draft can be obtained based on the multimedia data in the editing recommendation information, the first material, and the first multimedia editing draft. In the above method, it is not necessary to first synthesize the first material, but to first downsample the first material, thereby generating the first video based on the downsampling processing result and providing the first video to the server. This can effectively shorten the synthesis time of the first material, help to further shorten the multimedia editing time, save user waiting time, and thus improve the user experience.

[0017] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0018] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0019] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 A flowchart illustrating a multimedia editing method for related technologies;

[0021] Figure 2 A flowchart illustrating a multimedia editing method provided in this embodiment of the disclosure;

[0022] Figure 3A schematic diagram of video synthesis provided in this embodiment of the present disclosure;

[0023] Figure 4 A flowchart illustrating a multimedia editing method provided in this embodiment of the disclosure;

[0024] Figure 5 This is a schematic diagram of the structure of a multimedia editing device provided in an embodiment of the present disclosure;

[0025] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure. Detailed Implementation

[0026] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0027] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.

[0028] The inventors discovered through research that the intelligent multimedia editing process (i.e., intelligent packaging of user-uploaded materials) in related technologies is typically time-consuming, resulting in a poor user experience. Some users may even cancel intelligent packaging due to the long waiting time. For easier understanding, please refer to the relevant technologies provided, such as... Figure 1The diagram illustrates a multimedia editing method. The user client first acquires materials located on the audio / video tracks in the draft, including imported videos, images, and audio. Then, based on these materials, the user client sequentially performs operations such as full video compositing, video downsampling, and uploading the video file. The downsampled video file is then provided to the server for processing, with progress polling possible. The server returns the processing result to the user client, representing recommended packaging information based on the video file, such as included materials and available templates. The user client then parses the processing result according to the protocol, downloading the materials indicated by the result and using the pre-processed full video to composite templates, filling template slots with the full video. The resulting template video can also be considered as material. Finally, the materials obtained from the processing result are applied to the draft, and a video can be generated from the draft after applying the materials. This video is the result of intelligent packaging of the user's materials. However, the inventors found that the above process is generally time-consuming, especially in the preprocessing stage, where the time spent synthesizing the complete video using existing materials on the audio and video tracks accounts for the largest proportion, resulting in the highest user cancellation rate during this period. To improve these problems, embodiments of this disclosure provide a multimedia editing method, apparatus, device, and medium, which are described in detail below.

[0029] Figure 2 This is a flowchart illustrating a multimedia editing method provided in an embodiment of the present disclosure. The method can be executed by a multimedia editing device, which can be implemented using software and / or hardware, and is generally integrated into an electronic device. Figure 2 As shown, the method mainly includes the following steps S202 to S208:

[0030] Step S202: Obtain the first material; wherein, the first material is the material in the first multimedia editing draft; for example, the first material can be all the materials in the first multimedia editing draft, or it can include materials of a specified type in the first multimedia editing draft, such as video type, audio type, image type, etc. In specific implementation, the material located on the target track in the first multimedia editing draft can be used as the first material, and the target track includes the video track and the audio track. The multimedia editing draft can contain currently existing materials and editing operation information for the materials.

[0031] Step S204: Downsample the first material, generate a first video based on the downsampling result, and provide the first video to the server. Receive editorial recommendation information returned by the server based on the first video. The editorial recommendation information includes at least multimedia data recommended by the server. For example, the multimedia data may include information about the second material recommended by the server and / or information about the target template. The information about the second material may be, for example, the identification information of the second material or the second material itself. The second material may be any material recommended by the server, such as subtitles, stickers, on-screen text, sound effects, etc., without limitation. The information about the target template may be, for example, the identification of the target template or the target template itself. Video clips from the first material can be filled into the target template, meaning that the target template can be used to process the clips in the first material according to specified editing operations to generate a third material.

[0032] In some implementations, a preset downsampling algorithm can be used to downsample the first source material to obtain a downsampled result. A first video can then be generated based on this downsampled result, such as by synthesizing the first video using the downsampled first source material. The first source material can be one or more pieces, while the first video can be a single video; that is, the first video can be a composite of multiple downsampled first source materials (i.e., downsampled results). In practical applications, if the first source material is a single video and has not undergone any modifications affecting its playback duration, such as speed adjustments or cropping, a downsampling algorithm can be directly used to downsample the first source material (e.g., reduce resolution, modify size) to obtain the first video. If the first source material is not a single video, but includes multiple video segments or audio, or if the first source material requires the aforementioned modifications, a downsampled first video can be directly synthesized from the first source material using a downsampling algorithm. Compared to the above methods... Figure 1 The biggest difference between the existing technology shown is that it does not require synthesizing a complete video based on the first source material. Figure 1 The method described above requires first synthesizing a complete video, then downsampling the complete video, and finally providing the downsampled video file to the server while retaining the complete video for post-processing such as template filling. In this embodiment, however, the downsampled video can be directly synthesized based on the first source material. This can be understood as combining video synthesis and downsampling into a single operation, eliminating the need to synthesize and retain the complete video from the first source material at this stage, thus significantly saving the time required for synthesizing the complete video.

[0033] Providing the downsampled first video to the server helps reduce bandwidth consumption and facilitates server-side analysis and processing. Furthermore, after providing the first video to the server, the method also includes deleting the first video. This method can effectively reduce memory usage.

[0034] The server analyzes the first video using a preset algorithm and can recommend editorial recommendations that match the first video. These editorial recommendations are the processing results generated by the server based on the first video. The editorial recommendations include, but are not limited to, second materials and / or target templates recommended by the server that match the first material, as detailed in the aforementioned content. In addition, the server can also recommend templates that can be filled in the first material or editorial information corresponding to the first material and / or the second material, without any restrictions.

[0035] Step S206: Based on the multimedia data in the editor's recommendation information, the first material, and the first multimedia editing draft, a second multimedia editing draft is obtained.

[0036] In practical applications, the multimedia data and the first material from the aforementioned editorial recommendation information can be applied together to the first multimedia editing draft to obtain the second multimedia editing draft. Specifically, based on the first multimedia editing draft, the multimedia data and the first material can be comprehensively edited to obtain the second multimedia editing draft. The second multimedia editing draft can be used to generate the target video, which can be referred to as the intelligently packaged video.

[0037] In summary, in the above-mentioned intelligent multimedia editing process of user-uploaded materials, the embodiments of this disclosure do not require the first material to be synthesized first. Instead, the first material is downsampled first, and a first video is generated based on the downsampling result. The first video is then provided to the server. This can effectively shorten the synthesis time of the first material, which helps to further shorten the multimedia editing time, save users' waiting time, and thus improve the user experience.

[0038] In some implementations, the multimedia data includes information about the second material. Based on this, step S206, the step of obtaining the second multimedia editing draft based on the multimedia data in the editorial recommendation information, the first material, and the first multimedia editing draft, includes: obtaining the second material based on the multimedia data in the editorial recommendation information; and obtaining the second multimedia editing draft based on the second material, the first material, and the first multimedia editing draft. Wherein, if the information of the second material is an identifier for the second material, the client can download the second material based on the identifier; if the information of the second material is the second material itself, the client directly obtains the second material upon receiving the multimedia data. The specific configuration can be flexibly set and is not limited here. As mentioned earlier, the second material can be any material recommended by the server, such as subtitles, stickers, decorative text, sound effects, etc. The second material can be applied to the first multimedia editing draft to obtain the second multimedia editing draft.

[0039] In some implementations, the multimedia data includes information about the target template. Based on this, step S206, which is the step of obtaining the second multimedia editing draft based on the multimedia data in the editor's recommendation information, the first material, and the first multimedia editing draft, can be referred to as steps one through four below:

[0040] Step 1: Obtain the target template based on the multimedia data in the editor's recommendation information. If the target template information is the target template identifier, the client can download the target template based on the target template identifier. If the target template information is the target template itself, the client directly obtains the target template when receiving the multimedia data. The specific settings can be flexibly configured and are not restricted here.

[0041] The target template mentioned in this disclosure is used to process fill materials according to specified editing operations. Specifically, the target template can be a video template. For example, it includes template materials and a sequence of editing operations. The editing operation sequence indicates the specific type of editing operation (such as adding effects, transitions, adding text, etc.) and the order of these operations. By editing the template materials according to the editing operation sequence, the corresponding video effect can be obtained. Templates typically have one or more slots, which can be filled with other materials such as user-generated materials. By replacing the template material corresponding to the slot with the material filled in that slot, the fill materials can be quickly and conveniently edited based on the template's editing operation sequence, thereby presenting the editing effect corresponding to the video template.

[0042] In practical applications, if a user confirms adding a template, such as by selecting the option to add an intro / chapter template, the server can recommend target templates that match the first video clip. This information may include template identification information. The target template may contain slots. By parsing the content of the first video, the server can also provide the time information of a suitable video clip to fill the slot in the target template. Specifically, this could be the timestamp information of the first clip on the target timeline. The time information of the first video clip corresponds to the template slot. The first video clip can also be a highlight segment of the first video, such as containing exciting scenes or key content. If the user confirms not to add a template, or if the server does not find a template that matches the first video, the editorial recommendation information returned by the server may no longer include information related to the target template. In some specific implementation examples, if the user confirms the addition of a template, and the server finds a target template that matches the first video, and the server also recommends a second material, then the target template's identification information, the timestamp of the first video clip (used to fill the slot in the target template), and the information of the second material mentioned above can be sent to the client as editorial recommendation information.

[0043] Step 2: Obtain the second video segment based on the first material; wherein, the second video segment is the video segment in the first material used to fill the target template.

[0044] In some specific implementation examples, step two can be performed as follows: Steps A through C:

[0045] Step A: Determine the first video segment in the first video; wherein the first video segment is at least one segment specified by the server or client in the first video.

[0046] In practical applications, the server can analyze the first video and recommend suitable video segments to fill the target template. For example, the recommendation information can include not only the target template but also the time information of the first video segment within the first video, allowing the client to directly locate the first video segment. Alternatively, the client can specify the first video segment within the first video. For instance, the user can specify the first video segment to fill the template through the client, or the client can determine the first video segment from the first video based on a specific algorithm or target template. The specific settings are flexible and not limited here. For example, the first video segment can be a highlight segment or a segment containing specific information, and the number can be one or more, such as matching the number of slots in the target template.

[0047] Step B: Based on the time information of the first video segment in the first video, determine the target segment from the first material.

[0048] It should be noted that since the first video is a downsampled video, it is only used for transmission to the server for analysis and will not directly generate the target segment based on the first video. This embodiment of the disclosure determines the target segment based on the time information of the first video segment within the first video, using the first source material (i.e., the original source material without downsampling), in order to obtain the subsequent second video segment. In other words, the second video segment obtained based on the target segment is a video synthesized from the original source material uploaded by the user, thereby ensuring the quality of the second video segment.

[0049] In some specific implementation examples, step B can be performed by referring to steps B1 and B2 as follows:

[0050] Step B1: Based on the time information of the first video segment in the first video, determine the time information of the target segment in the first material; wherein, the target segment includes: a segment obtained by merging at least two first video segments that have time intersection, and a first video segment that does not have time intersection.

[0051] For example, assuming the time information of the first video segment is 5-25s, 10-20s, 30-45s, 35-50s, 52-63s, 70-85s, then the time information of the target segment is 5-25s, 30-50s, 52-63s, 70-85s. The first video segments 5-25s and 10-20s have a temporal overlap, which are merged to obtain the target segment 5-25s. Similarly, the first video segments 30-45s and 35-50s have a temporal overlap, which are merged to obtain the target segment 30-50s.

[0052] Step B2: If the first material meets the first preset condition, the target segment is determined from the first material based on the time information of the target segment. The first preset condition includes: the first material has not been modified by the target, and the target modification is a modification that changes the playback duration of the material. Modifications such as speed adjustment and cropping can be considered as target modifications. Furthermore, the first preset condition may include the first material containing more than one video segment, which can be flexibly set. That is, under the above conditions, the target segment can be obtained by directly synthesizing part of the content in the first material based on the time information of the target segment. Specifically, the above steps of determining the target segment from the first material based on the time information of the target segment can be performed by referring to the following steps (1) to (3):

[0053] (1) Determine the segment type of the target segment based on the time information of the target segment; wherein, the segment type includes a first type and a second type, the first type indicates that the target segment corresponds to at least two different materials in the first material at the same time, and the second type indicates that the target segment corresponds to only a single material in the first material.

[0054] (2) When the fragment type is the first type, determine the fragment content in at least two different materials that correspond to the time information of the target fragment, and synthesize the fragment content in the at least two different materials to obtain the target fragment.

[0055] (3) When the fragment type is the second type, determine the fragment content in the individual material corresponding to the time information of the target fragment, and export the fragment content in the individual material to obtain the target fragment.

[0056] The above methods can efficiently and conveniently obtain the desired target fragments.

[0057] Step C: Obtain the second video segment based on the target segment. In some embodiments, if there is only one target segment, it can be directly used as the second video segment. In other embodiments, if there are multiple target segments, they can be spliced ​​together to facilitate subsequent processing, thereby obtaining the second video segment. For example, step C can be obtained by referring to steps C1 and C2 as follows:

[0058] Step C1: When there are at least two target segments, based on the time information of the at least two target segments, splice the at least two target segments in chronological order to obtain a spliced ​​video.

[0059] Step C2: Obtain the second video segment based on the spliced ​​video.

[0060] In some implementations, the spliced ​​video can be directly used as the second video segment. Subsequently, based on the time period mapping information between the first and second video segments and the time information of the first video segment, the segment content in the second video segment corresponding to the slot of the target template can be determined, thereby filling the template slot. In other implementations, the video segment in the spliced ​​video corresponding to the first video segment can be determined based on the time information of the first video segment; then, the video segment in the spliced ​​video corresponding to the first video segment can be exported to obtain the second video segment.

[0061] In addition, the aforementioned second video clip can be an exported video or it can be presented in a form such as a draft nesting, without any restrictions.

[0062] Step 3: Fill the target template with the second video clip to obtain the third material; the target template is used to process the second video clip according to the specified editing operations.

[0063] For example, template slot filling can be performed based on the correspondence between the second video clip and the slots of the target template. In some embodiments, the aforementioned steps generate one or more second video clips, each corresponding to one or more slots of the target template, and the second video clips can be directly filled into the corresponding slots. In other embodiments, the second video clips generated in the aforementioned steps are video clips synthesized based on partial content of the first material. Based on the time information of the first video clips, the segment content that matches the slots of the target template can be determined from the second video clips, and the corresponding segment content can be filled into the slots of the target template. The target template after slot filling is the third material.

[0064] Step four: Based on the third material and the first multimedia editing draft, obtain the second multimedia editing draft.

[0065] The third material is placed into the first multimedia editing draft to obtain the second multimedia editing draft. In practical applications, if the multimedia data also includes information from the second material, the second multimedia editing draft can be obtained based on the second material, the third material, and the first multimedia editing draft.

[0066] In the above method, there is no need to directly export the complete video synthesized from the first material for template filling. Instead, the second video segment can be obtained based on the time information of the first video segment in the first video and the first material. Template filling can then be performed based on the second video segment. Since the required duration of the second video segment is usually shorter than the duration of the complete video synthesized from the first material, the multimedia editing time can be effectively shortened, saving users' waiting time and thus improving the user experience.

[0067] To facilitate understanding of the above content, please refer to... Figure 3 The diagram illustrates a video compositing process. Assume the draft contains four video clips, all considered as the first source material. These four clips are located on the draft's video track axis at times 0-15s, 15-40s, 40-60s, and 60-90s, respectively. The user can combine these four clips into a downsampled first video and send it to the server. The server can then send the time information of the first video clip, corresponding to times 5-25s, 10-20s, 30-45s, 35-50s, 52-63s, and 70-85s in the draft. For ease of viewing, [further details are needed]. Figure 3The diagram illustrates the distribution of six first video segments based on their timing information. These first video segments can be highlighted or key segments, which can be used for subsequent template filling. Merging these six first video segments yields target segments 1 through 4, which are the four target segments to be exported as drafts: 5-25s, 30-50s, 52-63s, and 70-85s. These four target segments can be stitched together into a complete 66s video. For example, each exported target segment can be used as a second video segment. For easier subsequent processing, the entire stitched video of the four target segments can also be used as the second video segment. In practical applications, the aforementioned second video segments can also be exported as the second video, or presented in a nested draft format; the specific configuration is flexible.

[0068] by Figure 3 For example, in related technologies, the four segments in the draft are usually combined into a complete 90-second video and exported directly. However, in this embodiment, only a 66-second video needs to be exported, thus reducing the time consumption by 24 seconds. It should also be noted that... Figure 3 For example only, in practical applications, for an original video with a total length of 10 minutes, the total duration of the first video segment sent by the server is less than 1 minute. That is, the video that needs to be synthesized from 10 minutes can be optimized to synthesize only 1 minute. This will greatly demonstrate the superiority of the technical solution provided by the embodiments of this disclosure, and can shorten the video synthesis time in the multimedia editing process by a factor of two.

[0069] Furthermore, considering that the first source material may not meet the aforementioned first preset condition in some cases, for example, in some embodiments, step two above, that is, the step of obtaining the second video segment based on the first source material, can be specifically executed with reference to steps a and b as follows:

[0070] Step a: If the first material meets the second preset condition, all materials in the first material are synthesized to obtain a complete video clip. The second preset condition includes: at least some materials in the first material have been modified by the target, that is, at least some materials in the first material have been modified in a way that affects the duration, such as speed adjustment or cropping. Since the server processes the first video generated from the first material modified by the target to obtain the time information of the first video clip, in order to accurately correspond to the time, regardless of whether the first material is a single video, it is necessary to synthesize all the first materials modified by the target to obtain a complete video clip.

[0071] Step b: Obtain the second video segment based on the complete video segment. In practical applications, the complete video segment can be directly used as the second video segment. Subsequently, template filling can be performed based on the URL of the complete video segment, that is, filling according to the time interval of the template slot. Alternatively, the second video segment corresponding to each first video segment can be exported from the complete video segment based on the time information of the first video segment, and the template slot filling can be performed directly on the second video segment.

[0072] It should be noted that steps a and b above are merely fallback solutions for situations where the first source material does not meet the first preset condition, ensuring the reliability of the final multimedia editing based on the editorial recommendations sent by the server. Additionally, for another scenario where the first and second preset conditions are not met—that is, the first source material contains only a single video segment and has not been modified by the target—there is no need to export from the first source material again. Instead, the URL of the first source material can be directly used to fill the template, i.e., filling it according to the time interval of the template slots.

[0073] Based on the foregoing, please refer to the embodiments provided in this disclosure, such as... Figure 4 The diagram shown is a flowchart of a multimedia editing method. Figure 4 and Figure 1 The main difference lies in the following: The preprocessing stage does not require a complete video compositing operation; instead, video compositing is performed in the post-processing stage. Furthermore, video compositing only occurs under specific conditions. For example, based on the server-side processing results, if it is determined that a template needs to be inserted and that the video to be composited (e.g., the material on the audio / video track is not only a single video, and the material has not been modified such as speed adjustment), a highlight video clip is composited. This highlight video clip belongs to the aforementioned second type of video clip. The compositing time for highlight video clips is usually less than the compositing time for a complete video. If the specific conditions are not met, video compositing cannot be performed; instead, the materials recommended by the server can be downloaded and applied directly. It should be noted that... Figure 4 This only briefly illustrates a few key operations, and does not mean that all of them need to be performed. For example, if no template is inserted, the template composition operation does not need to be performed. And for... Figure 1 In any case, a complete video compositing process needs to be performed during the preprocessing stage. Therefore, compared to other methods, the above-described approach provided by this disclosure can greatly save the time spent on video compositing or exporting, thereby effectively shortening the total multimedia editing time, which also shortens the time spent on intelligent packaging and reduces the user cancellation rate caused by the long processing time.

[0074] Furthermore, the inventors compared the above-mentioned technical solutions provided by related technologies and embodiments of this disclosure. Taking the total video duration (also known as the original video duration) of the first material in the draft as 8 minutes and 12 seconds as an example, the related technologies require 58 seconds for complete video compositing in the preprocessing stage, 13 seconds for downsampling based on the complete video, 51 seconds for uploading the video file to the server, 56 seconds for server processing, 13 seconds for downloading and compositing based on the materials recommended by the server, and 2 seconds for material application. However, if the above-mentioned technical solution provided by embodiments of this disclosure is applied, complete video compositing is not required in the preprocessing stage; only video downsampling based on the first material takes 12 seconds. Uploading the video file to the server takes 51 seconds, server processing takes 56 seconds, and then highlight segment compositing based on the server's processing results takes only 2 seconds. Subsequent material downloading and template compositing take 13 seconds, and material application takes 2 seconds. As can be seen from the above, in the related technologies, the local processing on the user end requires 58 + 13 = 71 seconds. In the technical solution provided by the embodiments of this disclosure, the local processing on the user end requires 12 + 2 = 14 seconds. Therefore, the technical solution provided by the embodiments of this disclosure reduces the processing time by 80% compared with the related technologies.

[0075] In summary, the multimedia editing method provided in this disclosure can effectively shorten the multimedia editing time, save user waiting time, and thus improve user experience.

[0076] Corresponding to the aforementioned multimedia editing method, this disclosure further provides a multimedia editing device. Figure 5 This is a schematic diagram of a multimedia editing device provided in an embodiment of the present disclosure. The device can be implemented by software and / or hardware, and is generally integrated into an electronic device, such as... Figure 5 As shown, the multimedia editing device includes:

[0077] The material acquisition module 502 is used to acquire the first material; wherein, the first material is the material in the first multimedia editing draft;

[0078] The recommendation information receiving module 504 is used to perform downsampling processing on the first material, generate a first video based on the downsampling processing result of the first material, provide the first video to the server, and receive editor recommendation information returned by the server based on the first video; wherein, the editor recommendation information contains at least multimedia data recommended by the server;

[0079] The second draft acquisition module 506 is used to obtain a second multimedia editing draft based on the multimedia data in the editor's recommendation information, the first material, and the first multimedia editing draft.

[0080] In the aforementioned device, instead of first compositing the first material, the first material is downsampled first, and the first video is generated based on the downsampling result. The first video is then provided to the server, which effectively shortens the compositing time of the first material, helps to further reduce the multimedia editing time, saves user waiting time, and thus improves the user experience.

[0081] In some implementations, after providing the first video to the server, the apparatus further includes a video deletion module for deleting the first video.

[0082] In some implementations, the multimedia data includes information about the second material; the second draft acquisition module 506 is specifically used to: obtain the second material based on the multimedia data in the editor recommendation information; and obtain the second multimedia editing draft based on the second material, the first material, and the first multimedia editing draft.

[0083] In some implementations, the multimedia data includes information about a target template; the second draft acquisition module 506 is specifically used to: obtain the target template based on the multimedia data in the editor recommendation information; obtain a second video clip based on the first material; wherein the second video clip is a video clip used to fill the target template; fill the target template based on the second video clip to obtain a third material; wherein the target template is used to process the second video clip according to a specified editing operation; and obtain a second multimedia editing draft based on the third material and the first multimedia editing draft.

[0084] In some implementations, the second draft acquisition module 506 is specifically used to: determine a first video segment in the first video; wherein the first video segment is at least one segment specified by the server or client in the first video; determine a target segment from the first material based on the time information of the first video segment in the first video; and obtain a second video segment based on the target segment.

[0085] In some implementations, the second draft acquisition module 506 is specifically used to: determine the time information of a target segment in the first material based on the time information of the first video segment in the first video; wherein the target segment includes: a segment obtained by merging at least two first video segments with time intersection, and a first video segment without time intersection; and, if the first material meets a first preset condition, determine the target segment from the first material based on the time information of the target segment; wherein the first preset condition includes: the first material has not been modified by the target, and the target modification is a modification that changes the playback duration of the material.

[0086] In some implementations, the second draft acquisition module 506 is specifically used to: determine the segment type of the target segment based on the time information of the target segment; wherein the segment type includes a first type and a second type, the first type indicating that the target segment simultaneously corresponds to at least two different materials in the first material, and the second type indicating that the target segment corresponds only to a single material in the first material; when the segment type is the first type, determine the segment content in the at least two different materials corresponding to the time information of the target segment, and synthesize the segment content in the at least two different materials to obtain the target segment; when the segment type is the second type, determine the segment content in the single material corresponding to the time information of the target segment, and export the segment content in the single material to obtain the target segment.

[0087] In some embodiments, the second draft acquisition module 506 is specifically used to: when the number of target segments is at least two, splice the at least two target segments in chronological order based on the time information of the at least two target segments to obtain a spliced ​​video; and obtain a second video segment based on the spliced ​​video.

[0088] In some implementations, the second draft acquisition module 506 is specifically used to: synthesize all the materials in the first material to obtain a complete video clip when the first material meets the second preset conditions; wherein, the second preset conditions include: at least some materials in the first material have been modified by a target, the target modification being a modification that changes the playback duration of the material; and obtain a second video clip based on the complete video clip.

[0089] The multimedia editing device provided in this disclosure can execute the multimedia editing method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of executing the method.

[0090] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device embodiments can be referred to the corresponding process in the method embodiments, and will not be repeated here.

[0091] This disclosure provides an electronic device, which includes: a storage device storing a computer program thereon; and a processing device for executing the computer program in the storage device to implement the steps of any method of this disclosure.

[0092] The following is for reference. Figure 6This diagram illustrates a structural schematic of an electronic device 600 suitable for implementing embodiments of the present disclosure. The terminal devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0093] like Figure 6 As shown, electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from storage device 608 into random access memory (RAM) 603. RAM 603 also stores various programs and data required for the operation of electronic device 600. Processing device 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.

[0094] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 An electronic device 600 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0095] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, it performs the functions defined in the methods of embodiments of this disclosure.

[0096] In addition to the methods and devices described above, embodiments of this disclosure can also be computer program products, comprising computer program instructions that, when executed by a processor, cause the processor to perform the methods provided in the embodiments of this disclosure. The computer program product can be written in any combination of one or more programming languages ​​to perform the operations of the embodiments of this disclosure. These programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on a user computing device, partially on a user device, as a standalone software package, partially on a user computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0097] Furthermore, embodiments of this disclosure may also be computer-readable storage media storing computer program instructions that, when executed by a processor, cause the processor to perform the multimedia editing method provided in embodiments of this disclosure.

[0098] The computer-readable storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may, for example, include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0099] This disclosure also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the multimedia editing method described in this disclosure.

[0100] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0101] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0102] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0103] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0104] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0105] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A multimedia editing method, characterized in that, include: Obtain the first material; wherein, the first material is the material in the first multimedia editing draft; The first material is downsampled, a first video is generated based on the downsampling result, and the first video is provided to the server. The server returns editorial recommendation information based on the first video. The editorial recommendation information includes at least multimedia data recommended by the server. Based on the multimedia data in the editorial recommendation information, the first material, and the first multimedia editing draft, a second multimedia editing draft is obtained.

2. The method according to claim 1, characterized in that, After providing the first video to the server, the method further includes: deleting the first video.

3. The method according to claim 1, characterized in that, The multimedia data includes information about the second material; the process of obtaining the second multimedia editing draft based on the multimedia data in the editorial recommendation information, the first material, and the first multimedia editing draft includes: The second material is obtained based on the multimedia data in the editorial recommendation information; Based on the second material, the first material, and the first multimedia editing draft, a second multimedia editing draft is obtained.

4. The method according to claim 1, characterized in that, The multimedia data includes information about the target template; The process of obtaining a second multimedia editing draft based on the multimedia data in the editorial recommendation information, the first material, and the first multimedia editing draft includes: The target template is obtained based on the multimedia data in the editorial recommendation information; A second video segment is obtained based on the first material; wherein, the second video segment is a video segment used to fill the target template; The target template is filled based on the second video segment to obtain the third material; wherein, the target template is used to process the second video segment according to the specified editing operations; Based on the third material and the first multimedia editing draft, a second multimedia editing draft is obtained.

5. The method according to claim 4, characterized in that, The step of obtaining the second video segment based on the first material includes: Determine a first video segment in the first video; wherein the first video segment is at least one segment specified by the server or client in the first video; Based on the time information of the first video segment in the first video, the target segment is determined from the first material; A second video segment is obtained based on the target segment.

6. The method according to claim 5, characterized in that, The step of determining the target segment from the first material based on the time information of the first video segment in the first video includes: Based on the time information of the first video segment in the first video, the time information of the target segment in the first material is determined; wherein, the target segment includes: a segment obtained by merging at least two first video segments that have time intersection, and a first video segment that does not have time intersection. If the first material meets the first preset condition, the target segment is determined from the first material based on the time information of the target segment; wherein, the first preset condition includes: the first material has not been modified by the target, and the target modification is a modification that changes the playback duration of the material.

7. The method according to claim 6, characterized in that, Determining the target segment from the first material based on the time information of the target segment includes: The segment type of the target segment is determined based on the time information of the target segment; wherein, the segment type includes a first type and a second type, the first type indicating that the target segment simultaneously corresponds to at least two different clips in the first material, and the second type indicating that the target segment corresponds only to a single clip in the first material; When the segment type is the first type, the segment content in at least two different materials corresponding to the time information of the target segment is determined, and the segment content in the at least two different materials is synthesized to obtain the target segment; When the segment type is the second type, the segment content in the individual material corresponding to the time information of the target segment is determined, and the segment content in the individual material is exported to obtain the target segment.

8. The method according to claim 5, characterized in that, The process of obtaining the second video segment based on the target segment includes: When there are at least two target segments, the at least two target segments are spliced ​​together in chronological order based on the time information of the at least two target segments to obtain a spliced ​​video. The second video segment is obtained based on the spliced ​​video.

9. The method according to claim 4, characterized in that, The step of obtaining the second video segment based on the first material includes: If the first material meets the second preset condition, all materials in the first material are synthesized to obtain a complete video clip; wherein, the second preset condition includes: at least some materials in the first material have been modified by a target, and the target modification is a modification that changes the playback duration of the material; The second video segment is obtained based on the complete video segment.

10. A multimedia editing device, characterized in that, include: The material acquisition module is used to acquire the first material; wherein, the first material is the material in the first multimedia editing draft; The recommendation information receiving module is used to downsample the first material, generate a first video based on the downsampling result of the first material, provide the first video to the server, and receive the editor recommendation information returned by the server based on the first video; wherein, the editor recommendation information includes at least the multimedia data recommended by the server; The second draft acquisition module is used to obtain a second multimedia editing draft based on the multimedia data in the editor recommendation information, the first material, and the first multimedia editing draft.

11. An electronic device, characterized in that, The electronic device includes: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the multimedia editing method according to any one of claims 1-9.

12. A computer-readable storage medium, characterized in that, The storage medium stores a computer program for executing the multimedia editing method described in any one of claims 1-9.

13. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the multimedia editing method according to any one of claims 1-9.