Method and apparatus for processing media item, device, and medium

By acquiring multiple templates and candidate configurations for the first media item, and using a machine learning model to select the target template to generate the second media item, the inefficiency problem in the existing technology is solved, achieving more efficient media item processing and generation that better meets user needs.

WO2025251217A1PCT designated stage Publication Date: 2025-12-11BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/097543
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-05
Publication Date
2025-12-11

AI Technical Summary

Technical Problem

The process of processing media items requires a lot of manual operation to obtain media items that meet user needs, resulting in low efficiency and wasted computing resources.

Method used

By acquiring multiple templates and the first media item, candidate configurations and their effectiveness evaluation are determined, and a machine learning model is used to select the target template that best matches the user's needs to generate the second media item.

Benefits of technology

It reduces the computational resource overhead of processing media items, improves processing efficiency, and generates media items that better meet user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024097543_11122025_PF_FP_ABST
    Figure CN2024097543_11122025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a method and apparatus for processing a media item, a device, and a medium. The method comprises: acquiring a first media item and multiple templates, wherein the multiple templates are respectively used for generating multiple second media items from the first media item; on the basis of the multiple templates and the first media item, determining multiple candidate configurations used for respectively generating the multiple second media items, wherein each candidate configuration among the multiple candidate configurations indicates a corresponding template among the multiple templates and the first media item; determining multiple effect evaluations respectively associated with the multiple candidate configurations, wherein each effect evaluation among the multiple effect evaluations indicates the effect evaluation for a corresponding second media item to be generated from the first media item using the template indicated by a corresponding candidate configuration; and on the basis of the multiple effect evaluations, selecting from among the multiple templates a target template used for generating a second media item from the first media item. In this way, multiple effect evaluations can be used to indicate whether multiple candidate second media items to be generated meet user requirements, improving the efficiency of processing media items.
Need to check novelty before this filing date? Find Prior Art

Description

Method, device, apparatus and medium for processing media item TECHNICAL FIELD

[0001] Implementations of the present disclosure generally relate to the field of computers, and particularly relate to a method, device, apparatus and computer-readable storage medium for processing media item. BACKGROUND

[0002] Currently, there are various technical solutions for generating media items, for example, a user can manually create a media item, edit an existing media item, or invoke a machine learning model to generate a media item, etc. However, a large amount of manual operation is required in the process of processing media items in order to obtain a media item that meets the user's needs. At this time, it is desirable to process media items in a simpler and more effective manner, and then obtain a media item that meets expectations.

[0003] SUMMARY

[0004] In a first aspect of the present disclosure, a method for processing a media item is provided. In the method, a first media item and a plurality of templates are obtained, and the plurality of templates are respectively used to generate a plurality of second media items from the first media item. A plurality of candidate configurations for respectively generating the plurality of second media items are determined based on the plurality of templates and the first media item, and a candidate configuration in the plurality of candidate configurations indicates a template in the plurality of templates and the first media item. A plurality of effect evaluations respectively associated with the plurality of candidate configurations are determined, and an effect evaluation in the plurality of effect evaluations represents an effect evaluation of a second media item to be generated from the first media item using the template indicated by the candidate configuration. Based on the plurality of effect evaluations, a target template is selected from the plurality of templates for generating the second media item from the first media item.

[0005] In a second aspect of the present disclosure, an apparatus for processing a media item is provided. The apparatus comprises: an obtaining module configured to obtain a first media item and a plurality of templates, and the plurality of templates are respectively used to generate a plurality of second media items from the first media item; a generating module configured to determine a plurality of candidate configurations for respectively generating the plurality of second media items based on the plurality of templates and the first media item, and a candidate configuration in the plurality of candidate configurations indicates a template in the plurality of templates and the first media item; an evaluation module configured to determine a plurality of effect evaluations respectively associated with the plurality of candidate configurations, and an effect evaluation in the plurality of effect evaluations represents an effect evaluation of a second media item to be generated from the first media item using the template indicated by the candidate configuration; and a selection module configured to select a target template from the plurality of templates for generating the second media item from the first media item based on the plurality of effect evaluations.

[0006] In a third aspect of the disclosure, an electronic device is provided. The electronic device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method according to the first aspect of the disclosure.

[0007] In a fourth aspect of the disclosure, a computer-readable storage medium is provided, having stored thereon a computer program which, when executed by a processor, causes the processor to implement the method according to the first aspect of the disclosure.

[0008] In a fifth aspect of the disclosure, a computer program product is provided, comprising a computer program, wherein the computer program, when executed by a processor, implements the method according to the first aspect of the disclosure.

[0009] It is to be understood that the details set forth herein are not intended to define key or critical features of implementations of the present disclosure, and are not intended to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0010] The above and other features, advantages and aspects of implementations of the present disclosure will become more apparent by describing in detail some implementations thereof with reference to the accompanying drawings. In the drawings:

[0011] FIG. 1 illustrates a block diagram of a media processing process according to one implementation of the present disclosure;

[0012] FIG. 2 illustrates a block diagram for processing a media item according to some implementations of the present disclosure;

[0013] FIG. 3 illustrates a block diagram for generating a template of a media item according to some implementations of the present disclosure;

[0014] FIG. 4 illustrates a block diagram for determining an effect score of a candidate configuration according to some implementations of the present disclosure;

[0015] FIG. 5 illustrates a block diagram for generating a media item according to some implementations of the present disclosure;

[0016] FIG. 6 illustrates a block diagram for generating a second media item using a template according to some implementations of the present disclosure;

[0017] FIG. 7 illustrates a flowchart of a method for processing a media item according to some implementations of the present disclosure;

[0018] FIG. 8 illustrates a block diagram of an apparatus for processing a media item according to some implementations of the present disclosure; and

[0019] FIG. 9 illustrates a block diagram of a device that can implement a plurality of implementations of the present disclosure. DETAILED DESCRIPTION

[0020] Implementations of the present disclosure will be described in detail below with reference to the drawings. Although certain implementations of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the implementations set forth herein; rather, these implementations are provided so that the present disclosure can be more thoroughly and completely understood. It should be understood that the drawings and implementations of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0021] In the description of implementations of the present disclosure, the term "includes" and its conjugates are open-ended, meaning "including but not limited to". The term "based on" is intended to mean "based, at least in part, on" unless explicitly stated otherwise. The term "one implementation" or "implementation" is understood to mean "at least one implementation". The term "some implementations" is understood to mean "at least some implementations". Other explicitly and implicitly recited definitions can be found below. As used herein, the term "model" can represent an association between various data. For example, the association can be obtained based on various technical solutions that are currently known and / or will be developed in the future.

[0022] It can be understood that the data involved in the technical solutions of the present disclosure (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the relevant laws and regulations and the relevant provisions.

[0023] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type of personal information involved in the present disclosure, the scope of use, the use scenario, etc. should be informed to the user and the authorization of the user should be obtained through appropriate means according to the relevant laws and regulations.

[0024] For example, in response to receiving the active request of the user, the user is sent prompt information to explicitly prompt the user that the operation requested to be executed will require the acquisition and use of the personal information of the user. Thus, the user can voluntarily choose whether to provide the personal information to the software or hardware such as the electronic device, the application program, the server or the storage medium, etc. that executes the operation of the technical solutions of the present disclosure according to the prompt information.

[0025] As an optional but non-limiting implementation, in response to receiving the active request of the user, the way of sending prompt information to the user, for example, can be the way of pop-up window, and the prompt information can be presented in the form of text in the pop-up window. In addition, the pop-up window can also carry selection controls for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0026] It can be appreciated that the above-mentioned notification and obtaining user authorization process is only illustrative and does not limit the implementation of the present disclosure, and other ways that meet relevant laws and regulations can also be applied to the implementation of the present disclosure.

[0027] The term "in response to" as used herein refers to a state in which a corresponding event occurs or a condition is met. It will be understood that the timing of the execution of the subsequent action in response to the event or condition is not necessarily strongly associated with the time at which the event occurs or the condition is met. For example, in some cases, the subsequent action can be performed immediately when the event occurs or the condition is met; in other cases, the subsequent action can be performed after a period of time after the event occurs or the condition is met.

[0028] Example Environment

[0029] A variety of technical solutions for processing media items have been proposed. However, a large amount of manual operation is required in the process of processing media items in order to obtain media items that meet expectations. Referring to FIG. 1, an application environment according to some implementations of the present disclosure is described, which shows a block diagram 100 of a media processing process according to an implementation of the present disclosure. In the context of the present disclosure, the specific process of processing media items will be described with video as an example of a media item. Alternatively and / or additionally, the media item can include other formats, such as but not limited to images, documents including text and images, and / or other formats of rich text data.

[0030] As shown in FIG. 1, a first media item 110 can be processed using a template 140 in order to obtain a second media item 120 from the first media item 110. For example, the template 140 can specify one or more media elements to be included in the second media item 120. At this time, the generated second media item 120 can include more rich visual content, such as visual content 130, 132, 134, 136, etc., thereby carrying more information.

[0031] Although a plurality of templates have been provided and can be used to generate a plurality of second media items, respectively. However, the generation process requires a large amount of computing resources and needs to confirm one by one whether the plurality of second media items meet the user's needs. Further, the generated second media item 120 can not meet the user's needs, which leads to the necessity of manually editing the second media item 120 using a media editing tool, thereby being unable to generate a large-scale media item. At this time, it is desirable to process media items in a simpler and more efficient manner, thereby obtaining media items that meet expectations.

[0032] Summary of Processing Media Items

[0033] To at least partially address the deficiencies in the prior art, according to one implementation of the present disclosure, a method for processing a media item is proposed. Generally, a plurality of effect evaluations of processing a first media item with a plurality of templates to generate a second media item can be determined respectively. The plurality of effect evaluations can be compared and an effect evaluation matching a user demand can be selected, and then the first media item can be processed using a template corresponding to the selected effect evaluation.

[0034] An overview according to one implementation of the present disclosure is described with reference to FIG. 2, which shows a block diagram 200 for processing a media item according to some implementations of the present disclosure. As shown in FIG. 2, a first media item can be obtained, as well as a plurality of templates 212, 214. Here, the plurality of templates can be from a predetermined template library 210, and the plurality of templates 212, 214 are respectively for generating a plurality of second media items from the first media item.

[0035] A plurality of candidate configurations 220, 222 for respectively generating the plurality of second media items can be determined based on the plurality of templates 212, 214 and the first media item. A candidate configuration in the plurality of candidate configurations can indicate a template in the plurality of templates and the first media item. For example, the candidate configuration 220 can indicate the template 212 and the first media item 110, and the candidate configuration 222 can indicate the template 214 and the first media item 110. At this time, a candidate configuration can indicate data required for generating a second media item. For example, the first media item 110 can be processed with the template 212 to generate a second media item; for another example, the first media item 110 can be processed with the template 214 to generate a second media item.

[0036] Further, a plurality of effect evaluations 230, 232 respectively associated with the plurality of candidate configurations can be determined. An effect evaluation in the plurality of effect evaluations represents an effect evaluation of a second media item to be generated from the first media item with a template indicated by a candidate configuration. For example, the effect evaluation 230 represents an effect evaluation of a second media item to be generated from the first media item 110 with the template 212 indicated by the candidate configuration 220, and the effect evaluation 232 represents an effect evaluation of a second media item to be generated from the first media item 110 with the template 214 indicated by the candidate configuration 222.

[0037] An effect evaluation can be represented in a variety of ways, for example, a continuous numerical value in a range of 0 to 1 (or other range) can be used to represent an effect evaluation, the larger the value is, the higher the matching degree with a user demand is, and the smaller the value is, the lower the matching degree with a user demand is. Alternatively and / or additionally, a discrete format (e.g., high, medium, low) can be used to represent an effect evaluation.

[0038] The target template for generating the second media item from the first media item can be selected from the plurality of templates 212,..., and 214 based on the plurality of effect evaluations 230,..., and 232. The plurality of effect evaluations 230,..., and 232 can be compared to determine the effect evaluation 232 (e.g., with the largest value) that matches the user demand more, and then the corresponding target template (e.g., target 214) is selected and the second media item 240 is generated with the corresponding candidate configuration 222.

[0039] With the implementation of the present disclosure, instead of actually generating a plurality of second media items with a plurality of templates, a plurality of effect evaluations can be used to represent whether a plurality of candidate second media items to be generated match the user demand, and then a template that matches the user demand more is selected to generate the corresponding second media item. In this way, the computational resource overhead for processing media items can be greatly reduced, and thus the efficiency of processing media items is improved.

[0040] Detailed process of processing media items

[0041] Having described the overview of processing media items, more details about processing media items are described below with reference to the accompanying drawings. According to some implementations of the present disclosure, the first media item can be obtained in various ways. Specifically, the first media item can be obtained from a media sharing application (e.g., the first media sharing application). It should be appreciated that a media sharing application can include a large number of original media items published by a large number of users and provide a rich data source. The first media item can be obtained from a plurality of original media items published by a plurality of users of the first media sharing application, thereby improving the efficiency of obtaining media items.

[0042] According to some implementations of the present disclosure, the first media item is a media clip extracted from a plurality of original media items. For example, in a data promotion scenario, assuming that it is desired to promote the first media sharing application and / or media items in the application, all or a part of the original media items can be used as the first media item. Generally, a video published by a user can be long (e.g., 5 minutes, etc.), and a key part (e.g., 10 seconds, etc.) can be extracted from the video. For example, a machine learning model can be used to analyze the content of the original media item, and then division and selection processing is performed on the original media item based on user demand, thereby finding a media clip that matches the user demand more (e.g., attracts more user interest, etc.).

[0043] According to some implementations of the present disclosure, a selected template can be utilized to generate a second media item from a first media item, and the template represents a layout of a plurality of media elements to be added into the second media item. More information about the template is described with reference to FIG. 3, which illustrates a block diagram 300 of a template for generating a media item according to some implementations of the present disclosure. As shown in FIG. 3, the template 214 can include a plurality of media elements 310, 320, 330, 340, and 350, etc.

[0044] The template 214 can be represented in various ways, for example, the template can be represented using an image, and a plurality of regions can be defined in the image to represent the plurality of media elements respectively. For another example, the template can be represented using a custom format, for example, can be represented in an array manner according to pixel coordinates of positions where the media elements are located, etc.

[0045] According to some implementations of the present disclosure, in the process of determining the related effect evaluation of a certain candidate configuration, the features associated with the candidate configuration can be determined based on the template specified in the candidate configuration and the first media item. Specifically, the template-related features and the related features of the first media item are determined respectively, and then the features of the candidate configuration are determined. More details about determining the features and then determining the effect score are described with reference to FIG. 4, which illustrates a block diagram 400 for determining the effect score of a candidate configuration according to some implementations of the present disclosure.

[0046] As shown in FIG. 4, in the case where the first media item is a video 410, the video features 414 can be extracted using an encoder 412 for extracting video features. In the case where the template is represented using an image, the template features 424 can be extracted using an encoder 422 for extracting image features. With the maturity of neural network technology, the original picture, video data can be calculated using neural network technology, the features are generated and used for subsequent strategy model analysis, clustering, classification, etc. Since the video includes image frames, pre-trained neural networks can be used to extract feature information of each image frame. Specifically, a residual network and / or other network can be used. Through this step, the original video can be processed into an N*D feature vector, where N represents the number of frames of the video, and D represents the dimension of the feature vector of the video. Each image frame corresponds to a feature vector.

[0047] Further, the video features 414 and the template features 424 can be input to the neural network 450, and the corresponding effectiveness score 452 can be determined. For example, the video features 414 and the template features 424 can be combined to obtain features of the candidate configuration, and the features can be input to the neural network 450. The neural network 450 herein can be a pre-obtained machine learning model. In this way, the powerful processing capability of the machine learning model can be utilized to determine the effectiveness score 452 in a more accurate manner.

[0048] According to some implementations of the present disclosure, the machine learning model can be obtained based on a generation target of generating the second media item from the first media item. Continuing the example above, in the scenario of data promotion, it is assumed that the target of generating the second media item from the first media item is to let more users access the first media item. At this time, the training data of the machine learning model can be obtained based on the target, for example, the first reference media item and the second reference media item with more user access can be selected. For another example, it is assumed that the target of generating the second media item from the first media item is to let more users download the application recommended in the first media item, the first reference media item and the second reference media item with more user download can be selected, and so on. With the example implementations of the present disclosure, the effectiveness score output by the machine learning model can be more in line with the generation target, so that the second media item generated by the selected template is more in line with the user demand.

[0049] According to some implementations of the present disclosure, more factors can be considered in the process of generating the features. For example, the candidate configuration can further include background audio (e.g., music) for generating the second media item. It should be understood that the background audio herein can represent audio for replacing the original background audio in the first media item, that is, the second media item generated in this way can have a brand new background audio, so that the second media item is more in line with the demand.

[0050] Specifically, the background audio can be selected from an audio library including a plurality of background audios. At this time, the features of the candidate configuration can be updated by the features of the background audio. Continuing to refer to FIG. 4, the music 430 can be selected, and the music features 434 can be extracted by using the encoder 432 for extracting music features. Further, the video features 414, the template features 424 and the music features 434 can be input to the neural network 450, and the corresponding effectiveness score 452 can be obtained. In this way, the audio information can be considered in the process of determining the effectiveness score, so that the determined effectiveness score is more accurate.

[0051] According to some implementations of the present disclosure, the candidate configuration can further include an attribute of the first media item. Here, the attribute can include at least any of a status of the first media item, a category of the first media item, a content template of the first media item, and an audio of the first media item, etc. With continued reference to FIG. 4, a status numerical feature 442 can be extracted from the status of the first media item. For the first media item published in the first media sharing application, the status can represent a play status, a like status, a follow status, etc. of the first media item. A category numerical feature 444 can be extracted from the category of the first media item, e.g., a category to which the presented content of the first media item belongs, a food category, a scenery category, a song category, etc.

[0052] Alternatively and / or additionally, a template numerical feature 446 can be extracted from a content template of the first media item. Here, the content template refers to a template used in the process of making the first media item (e.g., a configuration defining a style, a time length, a shot, etc. in the first media item), which is different from the template 214 used to generate the second media item. A music numerical feature 448 can be extracted from an audio of the first media item. Here, the audio of the first media item refers to an audio (e.g., a background music) used by the first media item itself, which is different from the music 430 used to generate the second media item.

[0053] According to some implementations of the present disclosure, the status numerical feature 442, the category numerical feature 444, the template numerical feature 446, and the music numerical feature 448 can be obtained, and then input to the neural network 440 (e.g., for combination and / or dimension of individual features, etc.), to obtain an attribute feature of the first media item. Further, the attribute feature obtained in the above manner is used to update the feature of the candidate configuration. Specifically, the video feature 414, the template feature 424, the music feature 434, and the attribute feature can be input to the neural network 450, to obtain a corresponding effect score 452. In this way, more abundant information can be considered in the process of determining the effect score, so that the determined effect score is more accurate.

[0054] It should be appreciated that although FIG. 4 only illustrates the process of determining the effect score 452 based on the template 420 and the music 430. Alternatively and / or additionally, a predetermined template library and a music library can be provided. Each template in the template library and each music in the music library can be traversed, and then a plurality of effect scores of a plurality of candidate configurations can be determined by way of combination. Assuming that the template library includes K templates and the music library includes L musics, then K*L candidate configurations and corresponding K*L effect scores can be obtained. The template and the music corresponding to the highest effect score can be selected, and then the second media item is generated from the first media item. In this way, it is not necessary to occupy a large amount of computing resources to actually generate K*L second media items, but the template and the music that can obtain a better effect score can be determined by the machine learning model.

[0055] In the case where the best effect score has been determined, the corresponding second media item can be generated by using the template and the music that can obtain a better effect score. More details are described with reference to FIG. 5, which illustrates a block diagram 500 for generating a media item according to some implementations of the present disclosure. According to some implementations of the present disclosure, the template library 530 can provide a large number of templates, and the music library 510 can provide a large number of musics. The original video in the video library 520 can be processed into different segments for material delivery performance, and the delivery performance of the video can be estimated by using the machine learning model described above.

[0056] As shown in FIG. 5, the music 512 can be selected from the music library 510, the video segment 524 can be extracted from the video library 520 by using the video understanding model 522, the template 532 can be selected from the template library 530, and then the video 540 can be generated. Here, the video segment 524 can be a complete video in the video library 520, or a segment in the complete video. In this way, the video 540 that is more in line with the user's demand can be generated. According to some implementations of the present disclosure, the corresponding script 514 (for example, the description of the video segment 524, etc.) can be generated by using the video understanding model 522, and then the script 514 is added to the video 540.

[0057] According to some implementations of the present disclosure, the template can further represent a plurality of association relationships between a plurality of media elements and a plurality of attributes of the first media item. Further, based on the plurality of association relationships, a plurality of attributes can be respectively added to a plurality of positions of the plurality of media elements corresponding to the target template to generate the second media item. More information is described with reference to FIG. 6, which illustrates a block diagram 600 for generating a second media item by using a template according to some implementations of the present disclosure.

[0058] As shown in FIG. 6, the plurality of attributes include at least any one of the following of the first media item: content of the first media item, a description of the first media item, an access address of an application for accessing the first media item, and an identification of the application. Specifically, the respective attributes can be added to positions in the template 214 corresponding to the respective media elements.

[0059] The media element 310 in the template 214 can correspond to the identification of the application for accessing the first media item, i.e., the application identification 612 can be added to the position of the media element 310. The media element 320 can correspond to the content of the first media item, i.e., the content 614 can be added to the position of the media element 320. The media element 330 can correspond to the description of the first media item, i.e., the description 616 can be added to the position of the media element 330. The media elements 340 and 350 can correspond to the access addresses of the application, i.e., the access addresses 620 and 622 can be added to the positions of the media elements 340 and 350, respectively (the access address 620 can be used to download the application for installation in one operating system, and the access address 622 can be used to download the application for installation in one operating system). At this time, the second media item includes the address for accessing the first media sharing application.

[0060] With the example implementation of the present disclosure, the attributes presented in the second media item and the positions at which the attributes are presented can be specified in a more accurate and effective manner, so that the generated second media item is more in line with the user demand.

[0061] According to some implementations of the present disclosure, the description of the first media item is extracted from the first media item. Specifically, the script extracted by the video understanding model 522 shown in FIG. 5 can be used as the description. With the example implementation of the present disclosure, the presented description will be more in line with the original content of the first media item.

[0062] According to some implementations of the present disclosure, the second media item can be provided in a second media sharing application different from the first media sharing application. For example, the first media sharing application can be an application for publishing short videos, and the second media sharing application can be an application for publishing multimedia data. In this way, more users can be attracted to watch the generated second media item across different applications, thereby improving the efficiency of data promotion.

[0063] In the process of data promotion, high-quality creative materials play a crucial role. High-quality materials can attract users and enable users to obtain more information. According to some implementations of the present disclosure, the content suitable for making materials can be found in the media sharing application for material processing, and finally for delivery.

[0064] According to some implementations of the present disclosure, high-quality materials can be produced by intelligently editing original media content with multi-modal techniques and content generation capabilities. Specifically, a video understanding process can apply multi-modal techniques to understand a video and provide a basis for subsequent content extraction, music recommendation, template recommendation. A segment extraction process can split and select the original video to find suitable video segments. A script generation process can generate suitable scripts for video content with language models. A music recommendation process can recommend suitable music for each video content. A template recommendation process can recommend suitable templates for each video content. A finished product effect estimation process can build a machine learning model and estimate the effect of the finished product after processing, and select suitable material processing methods according to the effect.

[0065] With the maturity of multi-modal and large model technologies, a computer vision model can be used to recognize and understand original user-generated content to obtain video understanding information. On the basis of video understanding, multi-modal techniques can be used to split the video to obtain atomized content, and language models can be used to process suitable script information. Machine learning can be used to select the most suitable music and templates for the video. Finally, multi-modal techniques can be used to splice the content into finished materials.

[0066] Example process

[0067] FIG. 7 illustrates a flowchart of a method 700 for processing a media item according to some implementations of the present disclosure. At block 710, a first media item and a plurality of templates are obtained, the plurality of templates being respectively for generating a plurality of second media items from the first media item. At block 720, a plurality of candidate configurations for respectively generating the plurality of second media items are determined based on the plurality of templates and the first media item, a candidate configuration of the plurality of candidate configurations indicating a template of the plurality of templates and the first media item. At block 730, a plurality of effect evaluations respectively associated with the plurality of candidate configurations are determined, an effect evaluation of the plurality of effect evaluations representing an effect evaluation of a second media item to be generated from the first media item with the template indicated by the candidate configuration. At block 740, a target template for generating a second media item from the first media item is selected from the plurality of templates based on the plurality of effect evaluations.

[0068] According to some implementations of the present disclosure, determining an effect evaluation of the plurality of effect evaluations includes determining a feature associated with the candidate configuration based on the template and the first media item, and determining the effect evaluation with a machine learning model based on the feature.

[0069] According to some implementations of the present disclosure, the template represents a layout of a plurality of media elements to be added into the second media item, and determining the feature includes determining the feature based on a feature of the template and a feature of the first media item.

[0070] According to some implementations of the present disclosure, the candidate configuration further includes background audio for generating the second media item, the background audio being selected from a plurality of background audios, and determining the features further includes updating the features with features of the background audio.

[0071] According to some implementations of the present disclosure, the candidate configuration further includes attributes of the first media item, and determining the features further includes updating the features with features of the attributes, the attributes including at least any of: a status of the first media item, a category of the first media item, a content template of the first media item, and audio of the first media item.

[0072] According to some implementations of the present disclosure, the machine learning model is acquired based on a generation target of generating the second media item from the first media item.

[0073] According to some implementations of the present disclosure, the template further indicates a plurality of association relationships between a plurality of media elements and a plurality of attributes of the first media item, and the method further includes adding the plurality of attributes to a plurality of locations of the plurality of media elements corresponding to the target template respectively based on the plurality of association relationships to generate the second media item.

[0074] According to some implementations of the present disclosure, the plurality of attributes include at least any of: content of the first media item, a description of the first media item, an access address of an application for accessing the first media item, and an identification of the application.

[0075] According to some implementations of the present disclosure, the description of the first media item is extracted from the first media item.

[0076] According to some implementations of the present disclosure, acquiring the first media item includes acquiring the first media item from a plurality of original media items published by a plurality of users of a first media sharing application, and the method further includes providing the second media item in a second media sharing application.

[0077] According to some implementations of the present disclosure, the first media item is a media clip extracted from the plurality of original media items.

[0078] Example apparatus and devices

[0079] FIG. 8 shows a block diagram of an apparatus 800 for processing a media item, according to some implementations of the present disclosure. The apparatus 800 includes an obtaining module 810 configured to obtain a first media item and a plurality of templates, the plurality of templates being respectively for generating a plurality of second media items from the first media item; a generating module 820 configured to determine, based on the plurality of templates and the first media item, a plurality of candidate configurations for respectively generating the plurality of second media items, a candidate configuration of the plurality of candidate configurations indicating a template of the plurality of templates and the first media item; an evaluating module 830 configured to determine a plurality of effect evaluations respectively associated with the plurality of candidate configurations, an effect evaluation of the plurality of effect evaluations representing an effect evaluation of a second media item to be generated from the first media item with the template indicated by the candidate configuration; and a selecting module 840 configured to select, based on the plurality of effect evaluations, a target template of the plurality of templates for generating a second media item from the first media item.

[0080] According to some implementations of the present disclosure, the evaluating module includes a feature determining module configured to determine, based on the template and the first media item, a feature associated with the candidate configuration, and a calling module configured to determine, based on the feature, the effect evaluation with the machine learning model.

[0081] According to some implementations of the present disclosure, the template represents a layout of a plurality of media elements to be added into the second media item, and the feature determining module includes a combining module configured to determine the feature based on a feature of the template and a feature of the first media item.

[0082] According to some implementations of the present disclosure, the candidate configuration further includes a background audio for generating the second media item, the background audio being selected from a plurality of background audios, and the feature determining module further includes an updating module configured to update the feature with a feature of the background audio.

[0083] According to some implementations of the present disclosure, the candidate configuration further includes an attribute of the first media item, and the feature determining module further includes updating the feature with a feature of the attribute, the attribute including at least any of a state of the first media item, a category of the first media item, a content template of the first media item, and an audio of the first media item.

[0084] According to some implementations of the present disclosure, the machine learning model is obtained based on a generation target of generating the second media item from the first media item.

[0085] According to some implementations of the present disclosure, the template further represents a plurality of association relationships between the plurality of media elements and a plurality of attributes of the first media item, and the apparatus further includes an adding module configured to add, based on the plurality of association relationships, the plurality of attributes to a plurality of positions respectively of the plurality of media elements corresponding to the target template to generate the second media item.

[0086] According to some implementations of the present disclosure, the plurality of attributes comprises at least any one of the following of the first media item: content of the first media item, a description of the first media item, an access address of an application for accessing the first media item, and an identification of the application.

[0087] According to some implementations of the present disclosure, the description of the first media item is extracted from the first media item.

[0088] According to some implementations of the present disclosure, the obtaining module comprises an extracting module configured to obtain the first media item from a plurality of original media items published by a plurality of users of the first media sharing application, and the apparatus further comprises a providing module configured to provide the second media item in the second media sharing application.

[0089] According to some implementations of the present disclosure, the first media item is a media clip extracted from the plurality of original media items.

[0090] FIG. 9 illustrates a block diagram of a device 900 capable of implementing a number of implementations of the present disclosure. It should be understood that the computing device 900 illustrated by FIG. 9 is merely an example and should not be construed to limit the functionality and scope of the implementations described herein. The computing device 900 illustrated by FIG. 9 can be used to implement the methods described above.

[0091] As illustrated by FIG. 9, the computing device 900 is in the form of a general- purpose computing device. Components of the computing device 900 can include, but are not limited to, one or more processors or processing units 910, a memory 920, a storage device 930, one or more communication units 940, one or more input devices 950, and one or more output devices 960. The processing unit 910 can be a real or virtual processor and is capable of executing various processing in accordance with programs stored in the memory 920. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of the computing device 900.

[0092] The computing device 900 typically includes a plurality of computer storage media. Such media can be volatile, nonvolatile, removable, and / or non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Storage 920 can be volatile (such as RAM) non-volatile (such as ROM, EEPROM, flash memory, etc.) or some combination of the two. Storage 930 can be removable or non-removable and can include machine readable media such as flash drives, magnetic disks (e.g., "hard" disks), or any other media capable of storing information and / or data (e.g., training data for training) and accessible by the computing device 900.

[0093] The computing device 900 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 9, a disk drive or other computer readable media drive can provide for reading from and writing to a removable, non-removable, volatile, or non-volatile computer readable medium. In these instances, each drive can be connected to the bus by one or more data media interfaces. The memory 920 can include a computer program product 925 having one or more program modules configured to carry out the various methods or actions of the various implementations of the present disclosure.

[0094] The communication unit 940 enables communications with other computing devices over a communication media. Additionally, the functionality of the components of the computing device 900 can be implemented in a single computing cluster or a plurality of computer machines that are capable of communicating over a communication connection. Thus, the computing device 900 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network nodes in the networking environment.

[0095] The input device 950 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 960 can be one or more output devices, such as a display, a speaker, a printer, etc. The computing device 900 can also communicate with one or more external devices (not shown) such as a storage device, a display device, etc. through the communication unit 940, as needed, one or more devices that enable a user to interact with the computing device 900, or any devices (e.g., a network card, a modem, etc.) that enable the computing device 900 to communicate with one or more other computing devices. Such communication can be carried out through an input / output (I / O) interface (not shown).

[0096] According to implementations of the present disclosure, a computer-readable storage medium having computer-executable instructions stored thereon is provided, where the computer-executable instructions are executed by a processor to implement the method described above. According to implementations of the present disclosure, a computer program product tangibly stored on a non-transitory computer-readable medium and comprising computer-executable instructions is also provided, where the computer-executable instructions are executed by a processor to implement the method described above. According to implementations of the present disclosure, a computer program product having a computer program stored thereon is provided, where the program is executed by a processor to implement the method described above.

[0097] Various aspects of the disclosure are now described with reference to the drawings. In general, the drawings described below are diagrammatic and schematic representations of actual or conceptual structures and processes, and are not limiting of the scope of the present disclosure. In the drawings, the same reference numerals are used to represent similar or like items.

[0098] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can be coupled to a computer or other programmable data processing apparatus, such that the computer readable storage medium can provide instructions to the computer or other programmable data processing apparatus, which execute the instructions to produce a computer implemented process, such that the instructions stored in the computer readable storage medium can produce a means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0099] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process, such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0100] The computer program product of the present disclosure can have a signal including said computer program. This signal can be electronic, electromagnetic, optical, or any other suitable type of signal. Such a signal can be provided through a communication connection, such as electrical wiring, optical fiber, wireless interface, etc. Examples of computer program products include computer program implemented on a personal computer, server, or other networked device. A non-transitory computer readable medium, such as a floppy disk, CD-ROM, DVD-ROM, Blu-ray Disc, hard disk drive, or any other suitable non-transitory computer readable medium can store the computer program product.

[0101] Having described several implementations of the present disclosure, it will be clear to those skilled in the art that many modifications, additions, and substitutions are possible without departing from the scope and spirit of the described implementations. Many modifications and variations of the present disclosure are possible in light of the above teachings. It is, therefore, to be understood that within the scope of the appended claims, the present disclosure can be practiced otherwise than as specifically described. While the present disclosure has been described with reference to the implementation figures, it will be understood by those skilled in the art that various changes can be made and equivalents can be substituted for elements thereof without departing from the scope of the present disclosure. In addition, many modifications can be made to adapt a particular situation or material to the teachings of the disclosure without departing from its scope. Therefore, it is contemplated to cover any and all adaptations and modifications falling within the scope of the appended claims. It should also be understood that the terminology used herein is for the purpose of describing particular implementations only and is not intended to be limiting.

Claims

1. A method for processing a media item, comprising: obtaining a first media item and a plurality of templates, the plurality of templates being respectively used to generate a plurality of second media items from the first media item; determining a plurality of candidate configurations for respectively generating the plurality of second media items based on the plurality of templates and the first media item, a candidate configuration of the plurality of candidate configurations indicating a template of the plurality of templates and the first media item; determining a plurality of effectiveness evaluations respectively associated with the plurality of candidate configurations, an effectiveness evaluation of the plurality of effectiveness evaluations representing an effectiveness evaluation of a second media item to be generated from the first media item with the template indicated by the candidate configuration; and selecting a target template for generating the second media item from the first media item based on the plurality of effectiveness evaluations from the plurality of templates. 2.The method of claim 1, wherein determining the effectiveness evaluation of the plurality of effectiveness evaluations comprises: determining a feature associated with the candidate configuration based on the template and the first media item; and determining the effectiveness evaluation with a machine learning model based on the feature. determining the feature based on a feature of the template and a feature of the first media item.

3. The method of claim 2, wherein the template representation is of a layout of a plurality of media elements to be added to the second media item, and determining the feature comprises: updating the feature with a feature of the background audio.

4. The method of claim 2, wherein the candidate configuration further comprises background audio for generating the second media item, the background audio being selected from a plurality of background audios, and determining the feature further comprises: updating the feature with a feature of the attribute, the attribute comprising at least any of a state of the first media item, a category of the first media item, a content template of the first media item, and an audio of the first media item.

5. The method of claim 2, wherein the candidate configuration further comprises attributes of the first media item, and determining the feature further comprises: 6.The method of claim 2, wherein the machine learning model is obtained based on a generation target of generating the second media item from the first media item. adding the plurality of attributes to a plurality of positions of a plurality of media elements corresponding to the target template respectively based on the plurality of association relationships to generate the second media item.

7. The method of claim 3, wherein the template further represents a plurality of association relationships between the plurality of media elements and a plurality of attributes of the first media item, and the method further comprises: 8.The method of claim 7, wherein the plurality of attributes comprise at least any of the following of the first media item: a content of the first media item, a description of the first media item, an access address of an application used to access the first media item, and an identification of the application. 9.The method of claim 8, wherein the description of the first media item is extracted from the first media item. obtaining the first media item from a plurality of original media items published by a plurality of users of a first media sharing application, and the method further comprises: providing the second media item in a second media sharing application.

10. The method of claim 7, wherein obtaining the first media item comprises: 11.The method of claim 10, wherein the first media item is a media clip extracted from the plurality of original media items. 12.An apparatus for processing a media item, comprising: an obtaining module configured to obtain a first media item and a plurality of templates, the plurality of templates being respectively used to generate a plurality of second media items from the first media item; ​ a generating module configured to determine, based on the plurality of templates and the first media item, a plurality of candidate configurations for respectively generating the plurality of second media items, a candidate configuration in the plurality of candidate configurations indicating a template in the plurality of templates and the first media item; an evaluating module configured to determine a plurality of effect evaluations respectively associated with the plurality of candidate configurations, an effect evaluation in the plurality of effect evaluations representing an effect evaluation of a second media item to be generated from the first media item with the template indicated by the candidate configuration; and a selecting module configured to select, based on the plurality of effect evaluations, a target template in the plurality of templates for generating the second media item from the first media item.

13. An electronic device, comprising: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions which when executed by the at least one processing unit cause the electronic device to perform the method according to any one of claims 1-11.

14. A computer readable storage medium having stored thereon a computer program which, when executed by a processor, causes the processor to implement the method according to any one of claims 1-11.

15. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-11. ​ ​

Citation Information

Patent Citations

  • Image-text typesetting method and image-text typesetting device based on artificial intelligence and electronic equipment

    CN110795925A

  • Recommendation information display method and device

    CN116186412A

  • Multimedia resource delivery method and device, electronic equipment and storage medium

    CN116932787A

  • Multimedia data processing method and device, electronic equipment and storage medium

    CN117726716A

  • Dynamically generating customized media effects

    US20180191797A1