Method for processing media item, electronic device and computer readable storage media
Patent Information
- Application Number
- BR112025028856
- Authority / Receiving Office
- BR · BR
- Patent Type
- Applications
- Publication Date
- 2026-08-11
Smart Images

Figure 00000000_0000_ABST
Description
1 / 44 METHOD FOR PROCESSING MEDIA ITEM, ELECTRONIC DEVICE AND COMPUTER-READABLE STORAGE MEDIUM FIELD
[001] Implementations of the present disclosure relate generally to the field of computers and, in particular, to computer-readable methods, apparatus, devices and storage media for processing media items. BACKGROUND
[002] Several technical solutions have been proposed for generating media items; for example, a user can manually create a media item, edit an existing media item, or invoke a machine learning model to generate a media item, and so on. However, a large number of manual operations are required in the media item processing to obtain media items that meet the user's needs. At this point, it may be desirable to process media items in a simpler and more efficient way to obtain the desired media items. SUMMARY
[003] In a first aspect of the present disclosure, a method is provided for processing a media item. In the method, a first media item and a plurality of Petition 870260064563, dated 01 / 07 / 2026, page 15 / 116 Two out of four models are obtained, with the plurality of models being used to generate a plurality of second media items from the first media item, respectively. Based on the plurality of models and the first media item, a plurality of candidate configurations is generated to generate the plurality of second media items, respectively, with a candidate configuration from the plurality of candidate configurations indicating a model from the plurality of models and the first media item. A plurality of effect evaluations, respectively associated with the plurality of candidate configurations, is determined, with an effect evaluation from the plurality of effect evaluations representing an effect evaluation of a second media item that should be generated from the first media item using the model indicated by the candidate configuration.Based on the plurality of effect evaluations, a target model is selected from the plurality of models to generate the second media item from the first media item.
[004] In a second aspect of the present disclosure, an apparatus for processing a media item is provided. The apparatus includes: a procurement module configured to procure a first media item and a plurality of templates, the plurality of templates being used to generate a plurality of second media items from the first item. Petition 870260064563, dated 01 / 07 / 2026, page 16 / 116 3 / 44 media, respectively; a generation module configured to determine, based on the plurality of models and the first media item, a plurality of candidate configurations to generate the plurality of second media items, respectively, a candidate configuration from the plurality of candidate configurations indicating a model from the plurality of models and the first media item; an evaluation module configured to determine a plurality of effect evaluations respectively associated with the plurality of candidate configurations, an effect evaluation from the plurality of effect evaluations representing an effect evaluation of a second media item that should be generated from the first media item using the model indicated by the candidate configuration;and a selection module configured to select, based on the plurality of effect evaluations, from the plurality of models, a target model to generate the second media item from the first media item.
[005] In a third aspect of the present disclosure, an electronic device is provided. The electronic device includes: at least one processing unit; and at least one memory, coupled to at least one processing unit and storing instructions executed by at least one processing unit, the instructions, Petition 870260064563, dated 01 / 07 / 2026, page 17 / 116 4 / 44 when executed by at least one processing unit, causing the electronic device to execute the method according to the first aspect of the present disclosure.
[006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, storing a computer program, the computer program, when executed by a processor, causing the processor to implement the method according to the first aspect of the present disclosure.
[007] In a fifth aspect of the present disclosure, a computer program product is provided, including a computer program, wherein the computer program, when executed by a processor, implements the method according to the first aspect of the present disclosure.
[008] It should be understood that what is described in this summary is not intended to identify key features or essential characteristics of the implementations of this disclosure, nor to limit the scope of this disclosure. Other features disclosed in this document will become easily understandable through the description below. BRIEF DESCRIPTION OF THE DRAWINGS
[009] The characteristics, advantages and aspects Petition 870260064563, dated 01 / 07 / 2026, page 18 / 116 5 / 44 above and other various implementations of the present disclosure will become more evident from the detailed description that follows, together with the accompanying drawings. In the drawings, the same reference numbers or similar reference numbers refer to the same or similar elements, wherein: FIG. 1 illustrates a block diagram of a media processing process according to an implementation of the present disclosure; FIG. 2 illustrates a block diagram for processing media items according to some implementations of the present disclosure; FIG. 3 illustrates a block diagram of a model for generating media items according to some implementations of the present disclosure; FIG. 4 illustrates a block diagram for determining an effect score of a candidate configuration according to some implementations of the present disclosure; FIG. 5 illustrates a block diagram for generating a media item according to some implementations of the present disclosure; FIG. 6 illustrates a block diagram for generating a second media item with a template according to some implementations of the present disclosure; Petition 870260064563, dated 01 / 07 / 2026, page 19 / 116 6 / 44 FIG. 7 illustrates a flowchart of a method for processing a media item according to some implementations of the present disclosure; FIG. 8 illustrates a block diagram of an apparatus for processing media items according to some implementations of the present disclosure; and FIG. 9 illustrates a block diagram of a device capable of implementing various implementations of the present disclosure. DETAILED DESCRIPTION
[0010] Implementations of this disclosure will be described in more detail with reference to the accompanying drawings, in which some implementations of this disclosure have been illustrated. However, it should be understood that this disclosure can be implemented in a variety of ways and, therefore, should not be interpreted as limited to the implementations disclosed in this document. Rather, these implementations are provided for a detailed and complete understanding of this disclosure. It should be understood that the drawings and implementations of this disclosure are used for illustrative purposes only and do not limit the scope of protection of this disclosure.
[0011] As used in this document, the term include and its variants should be read as terms Petition 870260064563, dated 01 / 07 / 2026, page 20 / 116 7 / 44 open terms meaning “include, but not be limited to”. The term “based on” should be read as “based at least in part on”. The term “an implementation” or “the implementation” should be read as “at least one implementation”. The term “some implementations” should be read as “at least some implementations”. Other definitions, explicit and implicit, may be included below. As used in this document, the term “model” may represent associations between respective data. For example, the above association may be obtained based on various technical solutions currently known and / or to be developed in the future.
[0012] It should be understood that the data involved in this technical solution (including, but not limited to, the data itself, the acquisition or use of data) must comply with the requirements of the relevant laws and regulations and pertinent provisions.
[0013] It should be understood that, before applying the technical solutions disclosed in the respective modalities of this disclosure, the user must be informed about the type, scope of use and scenario of use of the personal information involved in this disclosure in an appropriate manner, in accordance with the relevant laws and regulations, and the user's authorization must be obtained. Petition 870260064563, dated 01 / 07 / 2026, page 21 / 116 8 / 44
[0014] For example, in response to receiving an active user request, prompt information is sent to the user to explicitly inform them that the requested operation will acquire and use their personal information. Therefore, based on the prompt information, the user can decide for themselves whether to provide personal information to software or hardware, such as electronic devices, applications, servers, or storage media that perform operations of the technical solutions of this disclosure.
[0015] As an optional, but not limiting, implementation in response to receiving an active user request, the method of sending prompt information to the user may, for example, include a pop-up window, and the prompt information may be presented in text form within the pop-up window. Additionally, the pop-up window may also contain a selection control for the user to choose to agree or disagree with providing personal information to the electronic device.
[0016] It should be understood that the above process of notification and obtaining user authorization is merely illustrative and does not limit the implementations of this disclosure. Other methods that comply with relevant laws and regulations are also applicable to Petition 870260064563, dated 01 / 07 / 2026, page 22 / 116 9 / 44 implementations of the present disclosure.
[0017] As used in this document, the term “in response to” indicates a state in which a corresponding event occurs or a condition is met. It should be understood that the timing of the execution of a subsequent action performed in response to the event or condition is not necessarily strongly correlated with the timing of when the event or condition occurs or is established. For example, in some cases, the subsequent action may be performed immediately after the event occurs or after the condition is met. In other cases, the subsequent action may be performed only after a period of time has elapsed since the event occurred or the condition was established. Example Environment
[0018] Several technical solutions for media item processing have been proposed; however, a large number of manual operations are required in the media item processing process to obtain the desired media items. FIG. 1 is a block diagram of a media processing process according to some implementations of the present disclosure. In the context of the present disclosure, a specific media item processing process will be described using a video as an example of a media item. Alternatively and / or additionally, the Petition 870260064563, dated 01 / 07 / 2026, page 23 / 116 10 / 44 media items may include other formats, including, but not limited to, images, documents including text and images, and / or other enriched text data formats.
[0019] As shown in FIG. 1, the first media item 110 can be processed using the model 140 to obtain the second media item 120 from the first media item 110. For example, the model 140 can specify one or more media elements to be included in the second media item 120. At this point, the generated second media item 120 can include richer visual content, such as visual contents 130, 132, 134, 136, etc., to carry more information.
[0020] Although several models are currently provided and several models can be used to generate multiple second media items, respectively, the generation process requires a large amount of computational resources, and it is necessary to determine whether a plurality of second media items meets the user's requirements, one by one. Furthermore, the generated second media item may not meet the user's requirements, resulting in the need for manual editing of the second media item using the media editing tool, making large-scale media item generation impossible. At this point, it may be desirable to process the media items in a more... Petition 870260064563, dated 01 / 07 / 2026, page 24 / 116 11 / 44 simple and efficient way to obtain the desired media items. Summary of Media Item Processing
[0021] In order to at least partially solve the deficiencies of the prior art, according to an implementation of the present disclosure, a method for processing a media item is presented. Generally, multiple models can be used to process the first media item and generate a plurality of effect evaluations of the second media item, respectively. The plurality of effect evaluations can be compared and an effect evaluation that meets the user's requirements can be selected to process the first media item using a model corresponding to the selected effect evaluation.
[0022] Referring to FIG. 2, a summary is described according to an implementation of the present disclosure, and FIG. 2 illustrates a block diagram 200 for processing a media item according to some implementations of the present disclosure. As shown in FIG. 2, a first media item and a plurality of templates 212, ···, 214 can be obtained. Here, the plurality of templates can be from a predetermined template library 210, and the plurality of templates 212, ···, 214 is used to generate a plurality of second media items from the first media item, respectively. Petition 870260064563, dated 01 / 07 / 2026, page 25 / 116 12 / 44
[0023] A plurality of candidate configurations 220, ··· and 222 to generate, respectively, a plurality of second media items can be determined based on the plurality of models 212, ··· and 214 and the first media item. A candidate configuration from the plurality of candidate configurations can indicate a model from the plurality of models and the first media item. For example, candidate configuration 220 can indicate model 212 and the first media item 110, ···, and candidate configuration 222 can indicate model 214 and the first media item 110. At this point, the candidate configuration can indicate the data needed to generate the second media item. For example, the first media item 110 can be processed with model 212 to generate a second media item; And, in another example, the first media item 110 can be processed with the model 214 to generate a second media item.
[0024] In addition, multiple effect evaluations 230, ··· and 232 associated with multiple candidate configurations can be determined, respectively. An effect evaluation of the plurality of effect evaluations represents an effect evaluation of a second media item to be generated from the first media item using the model indicated by the candidate configuration. For example, Petition 870260064563, dated 01 / 07 / 2026, page 26 / 116 13 / 44 effect evaluation 230 represents an effect evaluation of a second media item to be generated from the first media item 110 using the model 212 indicated by the candidate configuration 220, and effect evaluation 232 represents an effect evaluation of a second media item to be generated from the first media item 110 using the model 214 indicated by the candidate configuration 222.
[0025] The effect rating can be represented in several ways, for example, a continuous numerical value from 0 to 1 (or other range) can be used to represent the effect rating; a higher numerical value indicates a greater degree of correspondence with the user's need, and a lower numerical value indicates a lower degree of correspondence with the user's need. Alternatively and / or additionally, the effect rating can be represented using a discrete format (e.g., high, medium, and low).
[0026] The target model, for generating the second media item from the first media item, can be selected from the plurality of models 212, ··· and 214 based on the plurality of effect evaluations 230, ··· and 232. The plurality of effect evaluations 230, ··· and 232 can be compared to determine the effect evaluation 232. Petition 870260064563, dated 01 / 07 / 2026, page 27 / 116 14 / 44 (for example, having a maximum value) that best matches a user requirement. Then, a corresponding target model (for example, model 214) is selected and a second media item 240 is generated using the corresponding candidate configuration 222.
[0027] With the implementation of the present disclosure, it is not necessary to generate multiple second media items using multiple models, but multiple effect evaluations can be used to indicate whether multiple candidate second media items to be generated meet the user's requirement, and then a model that best matches the user's need is selected to generate a corresponding second media item. In this way, the overhead of computational resources for media item processing can be greatly reduced, thus improving the efficiency of media item processing. Detailed Procedure for Media Item Processing
[0028] Having described a summary of the media item processing, further details about the media item processing are described below with reference to the drawings. According to some implementations of the present disclosure, the first media item can be obtained in several ways. Specifically, the first media item can be Petition 870260064563, dated 01 / 07 / 2026, page 28 / 116 15 / 44 obtained from a media sharing application (e.g., the first media sharing application). It should be noted that the media sharing application may include a large number of original media items published by a large number of users and provide rich data sources. The first media item may be obtained from the plurality of original media items published by the plurality of users of the first media sharing application, thus improving the efficiency of media item acquisition.
[0029] According to some implementations of the present disclosure, the first media item is a media segment extracted from the plurality of original media items. For example, in a data promotion scenario, it is assumed that the first media sharing application and / or the media item within the application is promoted, and all or part of the original media item can be used as the first media item. Generally, the user-published video can be long (e.g., 5 minutes, etc.), and the main part (e.g., 10 seconds, etc.) can be extracted from the video. For example, the content of the original media item can be analyzed using a machine learning model and then split processing and Petition 870260064563, dated 01 / 07 / 2026, page 29 / 116 16 / 44 selection can be performed on the original media item based on the user's need, thus finding a media segment that is more consistent with the user's needs (e.g., attracting more user interest, etc.).
[0030] According to some implementations of the present disclosure, a selected model can be used to generate the second media item from the first media item, and the model represents a configuration of the plurality of media elements to be added to the second media item. In FIG. 3, more information is described about a model, which shows a block diagram 300 of a model for generating media items according to some implementations of the present disclosure. As shown in FIG. 3, the model 214 can include a plurality of media elements 310, 320, 330, 340 and 350, and the like.
[0031] Model 214 can be represented in several ways, for example, an image can be used to represent the model, and multiple regions can be defined in the image to represent multiple media elements, respectively. In another example, the model can be represented using a self-defined format, for example, it can be represented in matrix form according to the pixel coordinates of a position where the media element is located. Petition 870260064563, dated 01 / 07 / 2026, page 30 / 116 17 / 44 is located, and similar.
[0032] According to some implementations of the present disclosure, in a process of determining a related effect rating of a candidate configuration, a feature associated with the candidate configuration can be determined based on a model and the first media item specified in the candidate configuration. Specifically, features related to the model and features related to the first media item are determined, respectively, and then features of the candidate configuration are determined. Further details on the determination of the features and then the determination of the effect scores are described with reference to FIG. 4, which illustrates a 400 block diagram for determining effect scores for candidate configurations according to some implementations of the present disclosure.
[0033] As shown in FIG. 4, where the first media item is a video 410, the video feature 414 can be extracted with an encoder 412 to extract video features. In a situation where an image is used to represent a model, the model feature 424 can be extracted with an encoder 422 to extract image features. With the maturity of Petition 870260064563, dated 01 / 07 / 2026, page 31 / 116 18 / 44 Neural network technologies can be used to implement calculations on original images and video data, generate features, and use those features in scenarios for analysis, clustering, classification, and similar subsequent policy modeling. Since the video includes an image frame, the feature information from each image frame can be extracted using a pre-trained neural network. In particular, a residual network and / or another network can be implemented. Through this step, the original video can be processed into a feature vector represented by N * D, where N represents the video frame number, D represents the dimension of the video feature vector, and each image frame corresponds to a feature vector.
[0034] Furthermore, video feature 414 and model feature 424 can be fed into neural network 450 to determine a corresponding effect score 452. For example, video features 414 and model features 424 can be combined to obtain the candidate configuration feature, and then the feature is fed into neural network 450. The neural network 450 described in this document can be a pre-acquired machine learning model. In this way, the powerful processing capacity of the machine learning model can be Petition 870260064563, dated 01 / 07 / 2026, page 32 / 116 19 / 44 used to determine the effect score 452 more accurately.
[0035] According to some implementations of the present disclosure, the machine learning model can be obtained based on the purpose of generating the second media item from the first media item. Continuing with the example above, in the data promotion scenario, it is assumed that the purpose of generating the second media item from the first media item is to allow more users to access the first media item. In this case, the training data for the machine learning model can be obtained based on purpose, for example, a first reference media item and a second reference media item can be selected, leading to a greater number of user accesses.In another example, it is assumed that the purpose of generating the second media item from the first media item is to get more users to download the application recommended in the first media item; then, the first reference media item and the second reference media item can be selected, leading to a greater number of user downloads, and so on. According to exemplary implementations of the present disclosure, the effect score generated by the machine learning model can be more consistent with the... Petition 870260064563, dated 01 / 07 / 2026, page 33 / 116 20 / 44 generation purpose, so that the second media item generated by the selected model best meets the user's requirements.
[0036] According to some implementations of this disclosure, more factors may be considered in generating the resources. For example, the candidate configuration may also include background audio (e.g., music) to generate the second media item. It should be understood that the background audio described in this document may represent audio used to replace the original background audio in the first media item. That is, the second media item thus generated may have new background audio, so that the second media item is more consistent with the requirement.
[0037] Specifically, background audio can be selected from an audio library that includes a plurality of background audios. At this point, the candidate configuration feature can be updated using the background audio feature. With continuous reference to FIG. 4, music 430 can be selected and music feature 434 can be extracted by an encoder 432 to extract music features. Furthermore, video feature 414, model feature 424, and music feature 434 can be fed into the neural network 450 to obtain a corresponding effect score 452. In this way, the information of Petition 870260064563, dated 01 / 07 / 2026, page 34 / 116 21 / 44 audio tracks can be considered in the process of determining the effect score, so that the determined effect score is more accurate.
[0038] According to some implementations of the present disclosure, the candidate configuration may further include an attribute of the first media item. Here, the attribute may include at least one of the following: a state of the first media item, a category of the first media item, a content model of the first media item and audio of the first media item, and so forth. Referring again to FIG. 4, the state value resource 442 may be extracted from the state of the first media item. For the first media item published in the first media sharing application, the state may represent the playback state, the similar state, the next state of the first media item, and similar states.The category value resource 444 can be extracted from the category of the first media item, for example, a category to which the presented content of the first media item belongs, such as a food category, a landscape category, a music category, and the like.
[0039] Alternatively and / or additionally, the 446 model value resource can be extracted from the content model of the first media item. Here, the model of Petition 870260064563, dated 01 / 07 / 2026, page 35 / 116 22 / 44 content refers to a template used in the creation process of the first media item (e.g., defining settings such as style, duration, take, and similar in the first media item). Here, the content template is different from the 214 template for generating the second media item. The 448 music value resource can be extracted from the audio of the first media item. Here, the audio of the first media item refers to the audio (e.g., background music) used by the first media item itself, which is different from the 430 music that generated the second media item.
[0040] According to some implementations of the present disclosure, the state value feature 442, the category value feature 444, the model value feature 446, and the music value feature 448 can be obtained, and then these features are fed into the neural network 440 (e.g., to combine and / or scale multiple features, etc.) to obtain the attribute feature of the first media item. Furthermore, the candidate configuration feature is updated using the attribute feature obtained in the previous manner. Specifically, the video feature 414, the model feature 424, the music feature 434, and the attribute feature can be fed into the neural network 450 to obtain a corresponding effect score 452. In this way, Petition 870260064563, dated 01 / 07 / 2026, page 36 / 116 23 / 44 richer information can be considered in the process of determining the effect score, so that the determined effect score is more accurate.
[0041] It should be understood that, although FIG. 4 shows only the process of determining the effect score 452 based on the model 420 and the music 430, alternatively and / or additionally, a predetermined model library and music library can be provided. Each model in the model library can be traversed, and each piece of music in the music library is traversed, thus determining a plurality of effect scores for the plurality of candidate configurations in a combined manner. Assuming the model library includes K models and the music library includes L musical pieces, K * L candidate configurations and K * L corresponding effect scores can be obtained. The model and music corresponding to the highest effect score can be selected to generate the second media item from the first media item.In this way, it is not necessary to spend a large amount of computational resources to generate the K*L seconds of media items, but a model and a song that lead to a better effect score can be determined by the machine learning model.
[0042] In the case where the ideal effect score Petition 870260064563, dated 01 / 07 / 2026, page 37 / 116 Once 24 / 44 has been determined, the second corresponding media item can be generated using the model and music that lead to the best effect score. More details are described with reference to FIG. 5, which illustrates a 500-block diagram for generating media items according to some implementations of the present disclosure. According to some implementations of the present disclosure, the 530-model library can provide a large number of models, and the 510-music library can provide a large number of songs. The delivery performance for different segments of the original videos in the 520-video library may be different, and the video delivery performance can be estimated using the machine learning model above.
[0043] As shown in FIG. 5, music 512 can be selected from music library 510, video segment 524 can be extracted from video library 520 using video comprehension model 522, model 532 can be selected from model library 530, and then video 540 can be generated. Here, video segment 524 can be a complete video in video library 520 or it can be a segment of the complete video. In this way, video 540 most consistent with the user's requirements can be generated. According to some implementations of the present disclosure, the model of Petition 870260064563, dated 01 / 07 / 2026, page 38 / 116 25 / 44 video comprehension 522 can be used to generate a corresponding redaction 514 (e.g., a description of video segment 524, etc.), thus adding redaction 514 to video 540.
[0044] According to some implementations of the present disclosure, the model may further represent a plurality of association relations between the plurality of media elements and the plurality of attributes of the first media item. Furthermore, a plurality of attributes may be added, respectively, to the plurality of positions of the plurality of media elements corresponding to the target model, based on the plurality of association relations, to generate the second media item. More information is described with reference to FIG. 6, which illustrates a 600 block diagram of the generation of a second media item with a model, according to some implementations of the present disclosure.
[0045] As shown in FIG. 6, the plurality of attributes includes at least one of the following: a content of the first media item, a description of the first media item, an access address of an application used to access the first media item, and an application identifier. Specifically, a corresponding attribute can be added to a corresponding position for each element. Petition 870260064563, dated 01 / 07 / 2026, page 39 / 116 26 / 44 media specified in model 214.
[0046] Media element 310 in model 214 can correspond to the application ID for accessing the first media item, i.e., application ID 612 can be added to the position of media element 310. Media element 320 can correspond to the content of the first media item, i.e., content 614 can be added to the position of media element 320. Media element 330 can correspond to a description of the first media item, i.e., description 616 can be added to the position of media element 330. Media elements 340 and 350 can correspond to the application access address, i.e., access addresses 620 and 622 can be added to the positions of media elements 340 and 350, respectively (access address 620 can be used to download an application installed on one operating system, and access address 622 can be used to download an application installed on another operating system).In this case, the second media item includes an address used to access the first media sharing application.
[0047] With exemplary implementations of the present disclosure, the attributes presented in the second media item and the position in which the attribute is presented Petition 870260064563, dated 01 / 07 / 2026, page 40 / 116 27 / 44 can be specified more precisely and efficiently, making the second generated media item more consistent with user requirements.
[0048] According to some implementations of the present disclosure, the description of the first media item is extracted from the first media item. Specifically, the text extracted by the 522 video comprehension model shown in FIG. 5 can be used as the description. With exemplary implementations of the present disclosure, the presented description will be more consistent with the original content of the first media item.
[0049] According to some implementations of the present disclosure, the second media item may be provided in a second media sharing application different from the first media sharing application. For example, the first media sharing application may be an application for publishing short videos and the second media sharing application may be an application for publishing multimedia data. In this way, more users may be attracted to watch the second media items generated in different applications, thus improving the efficiency of data promotion.
[0050] In data promotion, creative materials of Petition 870260064563, dated 01 / 07 / 2026, page 41 / 116 28 / 44 High quality plays a crucial role. High-quality material can attract the user, allowing them to obtain richer information. According to some implementations of this disclosure, suitable content for creating the material can be searched on the media sharing application for material processing and finally provided to users.
[0051] According to some implementations of the present disclosure, multimodal technologies and content generation capabilities can be used, and high-quality materials can be produced by intelligently editing the original media content. Specifically, the video comprehension process can apply multimodal technology to understand the video and provide a basis for subsequent content extraction, music recommendation, and model recommendation. The segment extraction process can split and select the original video to find a suitable video segment. The text generation process can use a language model to generate appropriate text to match the video content. The music recommendation process can recommend appropriate music for each video content. The model recommendation process can recommend a suitable model for each video content. The estimation process of Petition 870260064563, dated 01 / 07 / 2026, page 42 / 116 29 / 44 final effect can build a machine learning model, estimate a final effect for the media item, and select an appropriate material processing mode according to the effect.
[0052] With the maturity of multimodal technologies and large-scale models, a computer vision model can be used to recognize and understand original user-generated content in order to obtain video comprehension information. Based on video comprehension, videos can be segmented using multimodal technology to obtain atomized content. Meanwhile, appropriate text information is processed using a language model. Music and models that best match the video can be selected by machine learning. Finally, the content is merged into a final product using multimodal technology. Example of a Process
[0053] FIG. 7 shows a flowchart of a method 700 for processing a media item according to some implementations of the present disclosure. In block 710, a first media item and a plurality of models are obtained, the plurality of models being used to generate a plurality of second media items from the first media item, respectively. In block 720, based on Petition 870260064563, dated 01 / 07 / 2026, page 43 / 116 In block 30 / 44, a plurality of models and, in the first media item, a plurality of candidate configurations is generated to generate the plurality of second media items, respectively, a candidate configuration from the plurality of candidate configurations indicating a model from the plurality of models and the first media item. In block 730, a plurality of effect evaluations, respectively associated with the plurality of candidate configurations, is determined, an effect evaluation from the plurality of effect evaluations representing an effect evaluation of a second media item to be generated from the first media item using the model indicated by the candidate configuration. In block 740, based on the plurality of effect evaluations, a target model is selected from the plurality of models to generate the second media item from the first media item.
[0054] According to some implementations of the present disclosure, determining the effect evaluation of the plurality of evaluation effects includes: determining a feature associated with the candidate configuration based on the model and the first media item; and determining the effect evaluation using a machine learning model based on the feature.
[0055] According to some implementations of Petition 870260064563, dated 01 / 07 / 2026, page 44 / 116 31 / 44 present revelation, the model represents a configuration of a plurality of media elements that must be added to the second media item, and the determination of the characteristic includes: determining the characteristic based on a characteristic of the model and a characteristic of the first media item.
[0056] According to some implementations of the present disclosure, the candidate configuration further includes background audio to generate the second media item, the background audio being selected from a plurality of background audios, and the feature determination further includes updating the feature with a background audio feature.
[0057] According to some implementations of the present disclosure, the candidate configuration further includes an attribute of the first media item, and the determination of the characteristic further includes: updating the characteristic with a characteristic of the attribute, wherein the attribute includes at least one of the following: a state of the first media item, a category of the first media item, a content model of the first media item, and the audio of the first media item.
[0058] According to some implementations of the present disclosure, the machine learning model is Petition 870260064563, dated 01 / 07 / 2026, page 45 / 116 32 / 44 obtained based on a generation purpose to generate the second media item from the first media item.
[0059] According to some implementations of the present disclosure, the model further represents a plurality of association relations between the plurality of media elements and a plurality of attributes of the first media item, and the method further includes: generating the second media item by adding, based on the plurality of association relations, the plurality of attributes in a plurality of positions corresponding to a plurality of media elements of the target model, respectively.
[0060] According to some implementations of the present disclosure, the plurality of attributes includes at least one of the following: the first media item: content of the first media item, a description of the first media item, an access address from an application to access the first media item, and an identification of the application.
[0061] According to some implementations of the present disclosure, the description of the first media item is extracted from the first media item.
[0062] According to some implementations of the present disclosure, obtaining the first media item includes: obtaining the first media item from a plurality of original media items published by a Petition 870260064563, dated 01 / 07 / 2026, page 46 / 116 33 / 44 plurality of users of a first media sharing application, and the method also includes: providing the second media item in a second media sharing application.
[0063] According to some implementations of the present disclosure, the first media item is a media segment drawn from the plurality of original media items. Example of Apparatus and Device
[0064] FIG. 8 shows a block diagram of an apparatus 800 for processing a media item according to some implementations of the present disclosure. The apparatus 800 includes: an acquisition module 810 configured to obtain a first media item and a plurality of models, the plurality of models being used to generate a plurality of second media items from the first media item, respectively; a generation module 820 configured to determine, based on the plurality of models and the first media item, a plurality of candidate configurations to generate the plurality of second media items, respectively, a candidate configuration from the plurality of candidate configurations indicating a model from the plurality of models and the first media item; an evaluation module 830 configured to determine a plurality of effect evaluations, respectively Petition 870260064563, dated 01 / 07 / 2026, page 47 / 116 34 / 44 associated with the plurality of candidate configurations, an effect evaluation of the plurality of effect evaluations representing an effect evaluation of a second media item to be generated from the first media item using the model indicated by the candidate configuration; and a selection module 840 configured to select, based on the plurality of effect evaluations, from the plurality of models, a target model to generate the second media item from the first media item.
[0065] According to some implementations of the present disclosure, the evaluation module includes: a feature determination module configured to determine a feature associated with the candidate configuration based on the model and the first media item; and a call module configured to determine the effect evaluation using a feature-based machine learning model.
[0066] According to some implementations of the present disclosure, the model represents a configuration of a plurality of media elements that must be added to the second media item, and the feature determination module includes: a combination module configured to determine the feature based on a feature of the model and a feature of the Petition 870260064563, dated 01 / 07 / 2026, page 48 / 116 35 / 44 first media item.
[0067] According to some implementations of the present disclosure, the candidate configuration further includes background audio to generate the second media item, the background audio being selected from a plurality of background audios, and the feature determination module further includes: an update module configured to update the feature with a feature from the background audio.
[0068] According to some implementations of the present disclosure, the candidate configuration further includes an attribute of the first media item, and the feature determination module further includes: an update module configured to update the feature with a feature of the attribute, the attribute including at least one of the following: a state of the first media item, a category of the first media item, a content model of the first media item, and audio of the first media item.
[0069] According to some implementations of the present disclosure, the machine learning model is obtained based on a generation purpose to generate the second media item from the first media item.
[0070] According to some implementations of Petition 870260064563, dated 01 / 07 / 2026, page 49 / 116 36 / 44 present disclosure, the model also represents a plurality of association relations between the plurality of media elements and a plurality of attributes of the first media item, and the device also includes: an addition module configured to generate the second media item by adding, based on the plurality of association relations, the plurality of attributes in a plurality of positions corresponding to a plurality of media elements of the target model, respectively.
[0071] According to some implementations of the present disclosure, the plurality of attributes includes at least one of the following: the first media item: the content of the first media item, a description of the first media item, an access address from an application to access the first media item, and an identification of the application.
[0072] According to some implementations of the present disclosure, the description of the first media item is extracted from the first media item.
[0073] According to some implementations of the present disclosure, the retrieval module includes: an extraction module configured to retrieve the first media item from a plurality of original media items published by a plurality of users of a first media sharing application, and the device includes Petition 870260064563, dated 01 / 07 / 2026, page 50 / 116 37 / 44 still: a provisioning module configured to provide the second media item in a second media sharing application.
[0074] According to some implementations of the present disclosure, the first media item is a media segment extracted from the plurality of original media items.
[0075] FIG. 9 illustrates a block diagram of a 900 device that can implement a plurality of implementations of the present disclosure. It should be understood that the 900 computing device shown in FIG. 9 is only an example and does not constitute any limitation to the functions and scope of the implementations described in this document. The 900 computing device shown in FIG. 9 can be used to implement the method described above.
[0076] As shown in FIG. 9, the computing device 900 is in the form of a general-purpose computing device. The components of the computing device 900 may include, but are not limited to, one or more processors or processing units 910, a memory 920, a storage device 930, one or more communication units 940, one or more input devices 950, and one or more output devices 960. The processing unit 910 may be a physical or virtual processor and may perform various processing based on the Petition 870260064563, dated 01 / 07 / 2026, page 51 / 116 38 / 44 programs stored in memory 920. In a multiprocessor system, a plurality of processing units execute computer-executable instructions in parallel to enhance the parallel processing capability of the computing device 900.
[0077] The computing device 900 generally includes a plurality of computer storage media. Such media may be any media accessible by the computing device 900, including, but not limited to, volatile and non-volatile media, removable and non-removable media. Memory 920 may be volatile memory (e.g., a register, a cache, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash) or any combination thereof. Storage device 930 may be removable or non-removable media and may include machine-readable media (e.g., memory, a pen drive, a magnetic disk) or any other media that may be used to store information and / or data (e.g., training data for training) and be accessed within the computing device 900.
[0078] The 900 computing device may also include additional removable / non-removable storage media. Petition 870260064563, dated 01 / 07 / 2026, page 52 / 116 39 / 44 removable, volatile / non-volatile. Although not shown in FIG. 10, a disk drive for reading or writing to a removable, non-volatile disk (e.g., floppy disk) and an optical disk drive for reading or writing to a removable, non-volatile optical disk may be provided. In such cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. Memory 920 may include a computer program product 925 with one or more program modules, and these program modules are configured to perform various methods or actions of various implementations of the present disclosure.
[0079] The 940 communication unit implements communication with another computing device through a communication medium. Furthermore, the functions of the 900 computing device components can be performed by a single computing group or by a plurality of computing machines, and these computing machines can communicate through communication connections. Therefore, the 900 computing device can operate in a network environment using a logical connection with one or more servers, a personal computer (PC), or another general network node.
[0080] The 950 input device can be either Petition 870260064563, dated 01 / 07 / 2026, page 53 / 116 40 / 44 plus various input devices, such as a mouse, a keyboard, a trackball, a voice input device, and the like. The output device 960 can be one or more output devices, for example, a monitor, a speaker, a printer, and so on. The computing device 900 can also communicate via the communication unit 940 with one or more external devices (not shown), as needed, where the external device, for example, a storage device, a display device, and so on, communicates with one or more devices that allow users to interact with the computing device 900, or with any device (such as a network card, a modem, and the like) that allows the computing device 900 to communicate with one or more other computing devices. Such communication can be performed via an Input / Output (I / O) interface (not shown).
[0081] According to the exemplary implementations of the present disclosure, a computer-readable storage medium is provided in which computer-executable instructions are stored, and the computer-executable instructions are executed by a processor to implement the method described above. According to the exemplary implementations of the present disclosure Petition 870260064563, dated 01 / 07 / 2026, page 54 / 116 41 / 44 disclosure, a computer program product is further provided, which is tangibly stored on a non-transient, computer-readable medium and includes computer-executable instructions that are executed by a processor to implement the method described above. According to the exemplary implementations of the present disclosure, a computer program product is provided that stores a computer program, and that program, when executed by a processor, implements the method described above.
[0082] Aspects of the present disclosure are described in this document with reference to illustrations of flowcharts and / or block diagrams of methods, apparatus (systems) and computer program products according to the implementations of the invention. It is understood that each block of the illustrations of the flowchart and / or block diagrams, and combinations of blocks in the illustrations of the flowchart and / or block diagrams, can be implemented by computer-readable program instructions.
[0083] These computer-readable program instructions can be given to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing device to produce a machine, so that the instructions, which are executed Petition 870260064563, dated 01 / 07 / 2026, page 55 / 116 42 / 44 through the computer processor or other programmable data processing device, create means to implement the functions / acts specified in the block or blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that may direct a computer, a programmable data processing device and / or other devices to operate in a specific manner, such that the computer-readable storage medium containing instructions stored therein includes a manufacturing article including instructions that implement aspects of the function / act specified in the block or blocks of the flowchart and / or block diagram.
[0084] Computer-readable program instructions may also be loaded into a computer, other programmable data processing device, or other device to cause a series of operational steps to be executed on the computer, other programmable device, or other device to produce a computer-implemented process, such that the instructions executed on the computer, other programmable device, or other device implement the functions / acts specified in the flowchart and / or block(s) of the block diagram. Petition 870260064563, dated 01 / 07 / 2026, page 56 / 116 43 / 44
[0085] Flowcharts and block diagrams in The figures illustrate the architecture, functionality, and operation of possible implementations of computer program systems, methods, and products according to various implementations of the present disclosure. In this sense, each block in the flowchart or block diagrams may represent a module, segment, or portion of code, which includes one or more executable instructions to implement the specified logical function(s). It should also be noted that, in some alternative implementations, the functions indicated in the block may occur out of the order shown in the figures. For example, two blocks shown in succession may, in fact, be executed substantially simultaneously, or the blocks may sometimes be executed in reverse order, depending on the functionality involved.It should also be noted that each block in the block diagrams and / or flowchart illustration, and combinations of blocks in the block diagrams and / or flowchart illustration, can be implemented by systems based on purpose-built hardware that perform the specified functions or actions, or combinations of purpose-built hardware and computer instructions.
[0086] The descriptions of the various implementations of the present disclosure have been presented for purposes Petition 870260064563, dated 01 / 07 / 2026, page 57 / 116 44 / 44 are illustrative, but are not intended to be exhaustive or limited to the implementations disclosed. Many modifications and variations will be evident to those skilled in the art, without departing from the scope and essence of the implementations described. The terminology used in this document has been chosen to better explain the principles of the implementations, the practical application or technical improvement in relation to technologies found on the market, or to allow others skilled in the art to understand the implementations disclosed in this document. Petition 870260064563, dated 01 / 07 / 2026, page 58 / 116
Claims
1 / 4 CLAIMS 1. A method for processing a media item characterized in that it comprises: obtaining a first media item and a plurality of models, the plurality of models being used to generate a plurality of second media items from the first media item, respectively; determining, based on the plurality of models and the first media item, a plurality of candidate configurations to generate the plurality of second media items, respectively, a candidate configuration from the plurality of candidate configurations indicating a model from the plurality of models and the first media item; determining a plurality of effect evaluations, respectively associated with the plurality of candidate configurations, an effect evaluation from the plurality of effect evaluations representing an effect evaluation of a second media item to be generated from the first media item using the model indicated by the candidate configuration;and select, based on the plurality of effect evaluations, from the plurality of models, a target model to generate the second media item from the first media item.
2. Method according to claim 1, characterized in that the determination of the effect evaluation of the plurality of evaluation effects comprises: determining a feature associated with the candidate configuration based on the model and the first media item; and determining the effect evaluation using a machine learning model based on the feature. Petition 870260064563, dated 01 / 07 / 2026, p. 59 / 116 2 / 4 3. Method, according to claim 2, characterized in that the model represents a configuration of a plurality of media elements that must be added to the second media item, and the determination of the characteristic comprises: determining the characteristic based on a characteristic of the model and a characteristic of the first media item.
4. Method according to claim 2, characterized in that the candidate configuration further comprises background audio to generate the second media item, the background audio being selected from a plurality of background audios, and the feature determination further comprises updating the feature with a feature from the background audio.
5. Method according to claim 2, characterized in that the candidate configuration additionally comprises an attribute of the first media item, and the determination of the feature further comprises: updating the feature with a feature of the attribute, the attribute comprising at least one of: a state of the first media item, a category of the first media item, a content model of the first media item, and audio of the first media item.
6. Method, according to claim 2, characterized in that the machine learning model is obtained based on a generation purpose to generate the second media item from the first media item.
7. Method, according to claim 3, characterized in that the model additionally represents a plurality of association relations between the plurality of media elements and a plurality of attributes of the first media item, and the method additionally comprises: Petition 870260064563, dated 01 / 07 / 2026, page 60 / 116 3 / 4 generating the second media item by adding, based on the plurality of association relations, the plurality of attributes in a plurality of positions corresponding to a plurality of media elements of the target model, respectively.
8. A method according to claim 7, characterized in that the plurality of attributes comprises at least one of the following: the content of the first media item, a description of the first media item, an access address from an application to access the first media item, and an identification of the application.
9. Method according to claim 8, characterized in that the description of the first media item is extracted from the first media item.
10. A method according to claim 7, characterized in that obtaining the first media item comprises: obtaining the first media item from a plurality of original media items published by a plurality of users of a first media sharing application, and the method further comprises: providing the second media item in a second media sharing application.
11. Method, according to claim 10, characterized in that the first media item is a media segment extracted from the plurality of original media items.
12. Electronic device characterized in that it comprises: at least one processing unit; and at least one memory coupled to at least one processing unit and storing instructions executed by at least one processing unit, the instructions, when executed by at least one processing unit, causing the electronic device to execute a method as defined in any one of claims 1 to 11.
13. A computer-readable storage medium characterized in that it stores instructions therein, the instructions, when executed by a processor, cause the processor to implement a method as defined in any one of claims 1 to 11. Petition 870260064563, dated 01 / 07 / 2026, pp. 62 / 116